AIDC
DC AI 热点

全部 AI 动态

9月19日2026-09-19
AWS 机器学习✦ 精选AI 评分 75/10004:52

亚马逊云科技回顾 SageMaker AI 2026 年多项推理功能更新

亚马逊云科技(AWS)汇总了 Amazon SageMaker AI 在 2026 年初至今推出的 13 项推理相关更新,涵盖全托管端点与 Amazon SageMaker HyperPod Inference 两大部署路径。重点功能包括推理配置推荐、容量感知实例池、分层 KV 缓存机制,以及预填充与解码解耦(disaggregated prefill and decode)等大模型推理优化技术,旨在提升云端 AI 部署效率并降低算力成本。

阅读原文 ↗推荐理由:系统盘点了 AWS SageMaker 在大模型推理加速与算力资源池化管理方面的前沿落地特性。# Amazon# 推理# 算力# 大模型
NVIDIA Developer Blog✦ 精选AI 评分 68/10003:04

使用 AIPerf 对大规模大语言模型推理进行基准评测

在系统上部署大语言模型后,评估其实际推理速度与性能表现是关键难题。本文介绍了用于大规模大模型推理评测的工具 AIPerf,旨在帮助开发者和工程师摆脱直觉猜测,建立系统化、规模化的性能基准测试方法,准确衡量模型在不同负载下的响应速度与吞吐能力。

阅读原文 ↗推荐理由:介绍了大语言模型规模化推理基准测试工具 AIPerf,对模型部署与性能调优具备实用参考价值。# 大模型# 推理# 评测# 算力
9月17日2026-09-17
IT168 服务器存储 · AI与算力(网页)✦ 精选规则精选20:20

华为发布全球首个采用NPO的超节点——昇腾960超节点

今日,华为全联接大会2026在上海启幕,华为副董事长、轮值董事长汪涛发表题为“智启新未来,打造智能世界的硅基黑土地”的主题演讲,发布全球首个采用NPO技术的超节点——昇腾960超节点,加速十万亿规模的大模型训练和推理。

9月16日2026-09-16
NVIDIA · AI 筛选✦ 精选规则精选23:00

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics. Higher system performance means more tokens generated, resulting in higher revenue. Efficient scaling means throughput grows proportionally as hardware gets added, requiring fewer resources to serve users at scale. Continuous optimization means generating more value from infrastructure investments. […]

NVIDIA · AI 筛选✦ 精选规则精选06:24

‘Now We Can Know Everything and Do Anything,’ Jensen Huang Says at Dreamforce

Know everything. Do anything. That was the message NVIDIA founder and CEO Jensen Huang brought to Salesforce Dreamforce Tuesday, joining CEO Marc Benioff onstage in an appearance that coincided with the announcement of Koa — Salesforce’s first CRM reasoning model, built on NVIDIA Nemotron 3 Super. Huang didn’t just take the stage. He walked into […]

9月10日2026-09-10
9月9日2026-09-09
9月7日2026-09-07
AI Roundup · X AI coding圈日报规则精选14:00

OpenAI Publishes Its RSI Numbers, Astra on Low Beats Sol on High & a PR Review Toolkit

OpenAI spent Sunday talking about recursive self-improvement. Chief Scientist Jakub Pachocki's essay An Alien Mind says internal results give him a strong expectation that progress can be sustained into RSI, that chain-of-thought monitoring is becoming progressively less reliable on the Astra class, and that OpenAI will unilaterally withhold scaling if needed while calling for mandated safety bars. The companion data post is the more concrete document: the median OpenAI researcher now burns over $600 a day of inference at API prices, the 90th percentile over $7,000, the research org runs 3.1 agent-workdays per human workday, and a July 20 inf

9月6日2026-09-06
AI Roundup · X AI coding圈日报规则精选14:00

Astra's Quiet Wins: Cached Reasoning Swaps, Cross-Window Notes & a Twitter Clone in Minecraft

The first full weekend with GPT-6 Astra in everyone's hands, and the interesting findings are the unglamorous ones nobody put in a launch video. You can now change reasoning effort mid-conversation without invalidating the prompt cache , because the effort change is appended to the end of the context instead of rewriting the top. Codex has an experimental compaction mode where Astra keeps notes across context windows and can search earlier windows including tool calls, off by default and buried in a TOML flag. The hallucination-rate drop is on pages 20 and 21 of the system card and OpenAI barely mentioned it. Third parties are filling in the

9月5日2026-09-05
9月4日2026-09-04
AI Roundup · X AI coding圈日报规则精选14:00

GPT-6 Astra Lands, 99.9% on ARC With the Right Harness & the Model That Hides Its Thoughts

OpenAI shipped GPT-6 Astra , priced exactly like Fable at $10 in and $50 out, and the day split three ways. The capability story is real: 99.9% on ARC-AGI-3, two Lean-verified Erdős problems no model had touched, a prime-gap bound improved for the first time since the 1930s, and Latent Space's writeup after 20 billion tokens calling it an AI engineer you can hire for under six dollars an hour. The benchmark story is messier: the ARC score needs OpenAI's own harness that preserves hidden reasoning state (the standard harness gets 62.7%), Artificial Analysis has Astra level with GPT-5.6 Sol and five points behind Fable 5.1 on general intelligen

9月3日2026-09-03
9月2日2026-09-02
9月1日2026-09-01
8月21日2026-08-21
8月19日2026-08-19
8月13日2026-08-13
Microsoft Research✦ 精选规则精选00:00

MindTopo reveals VLMs’ spatial reasoning abilities

A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research .

8月12日2026-08-12
Microsoft Research✦ 精选规则精选00:00

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research .

8月8日2026-08-08
7月18日2026-07-18
6月25日2026-06-25