DCAI
DC AI 热点

全部 AI 动态

9月24日2026-09-24
Data Center Dynamics✦ 精选AI 评分 72/10015:00

华为发布全新UnifiedBus计算架构,助力SuperPoD及超算集群互连

为应对快速增长的算力需求,华为正式推出了一套全新的计算架构。该架构采用UnifiedBus互连技术,专门用于支持SuperPoD与大规模算力集群之间的高效互连。该设计旨在突破计算节点间的连接瓶颈,优化集群整体性能与扩展能力,为大规模AI训练及高性能计算基础设施提供更强的底层算力支持。

阅读原文 ↗推荐理由:展示了华为在超大规模算力集群互连技术上的新突破,对AI算力基础设施建设具有参考价值。# 华为# UnifiedBus# SuperPoD# 算力集群# 芯片互联
Data Center DynamicsAI 评分 40/10015:00

随着数据中心电力技术演进,相关设备物流挑战日益复杂

随着数据中心开始采用更多样化的发电与电力供应技术,相关设备的运输、仓储和安装等物流流程正变得愈发复杂。报道指出,伴随计算能耗剧增,新型供电解决方案的引入虽然缓解了能源压力,但也对数据中心上下游的供应链物流管理提出了更高要求,部署重型及特种电力设备已成为基础设施建设中的关键考验。

阅读原文 ↗推荐理由:探讨了数据中心电力技术演进对底层供应链与物流部署带来的实际挑战。# 数据中心# 电力供应# 物流管理# 基础设施
arXiv 人工智能✦ 精选AI 评分 75/10012:00

Orthrus:面向迭代检索的异构批处理高效推理服务系统

现代信息检索系统通常同时结合嵌入模型与生成模型处理复杂查询,但现有系统往往将二者隔离部署,粗粒度分配硬件易引发计算气泡,导致 GPU 利用率低和吞吐不足。为此,研究团队提出服务系统 Orthrus。该系统在统一推理循环中引入异构批处理机制,旨在有效整合具有资源冲突特性的嵌入与生成工作负载,提升动态查询下的硬件利用效率。(注:素材摘要文末存在截断)

阅读原文 ↗推荐理由:针对检索增强等场景中嵌入与生成模型混部低效的痛点,提出了创新的异构推理批处理系统设计。# AI推理# 异构计算# GPU利用率# 信息检索# 系统架构
量子位AI 评分 45/10008:48

成立九年,中科类脑将积累注入“Token工厂”

报道聚焦于成立九年的中科类脑公司,指出其将过往的技术与业务积累整合并投入到“Token工厂”的建设与布局中,暗示其在大模型时代围绕算力资源与Token生成服务进行的转型与探索。由于输入素材未提供详细正文或摘要内容,关于该“Token工厂”的具体技术架构、产品形态、商业落地细节及最新进展等信息尚不明确,有待后续报道补充。

阅读原文 ↗推荐理由:展示了类脑AI企业向大模型算力与Token基础设施转型的动态,但缺乏正文细节。# 中科类脑# Token工厂# 算力基础设施# 大模型# 商业动态
Data Center Dynamics✦ 精选AI 评分 60/10008:00

澳大利亚发布数据中心指南,推动高耗能企业融入电网建设

该内容探讨了大型能源消耗主体如何更好地为能源系统及其服务的社区做出贡献。澳大利亚为此出台了相关数据中心指导准则,旨在推动数据中心等高耗能设施成为电网的“良好公民”,探索其在支撑能源系统稳定与可持续发展方面的协作平台与基础。(注:素材信息有限,主要呈现了准则的愿景与探讨方向)

阅读原文 ↗推荐理由:体现了海外针对数据中心高能耗挑战、推动算力基础设施与国家电网协同治理的政策导向。# 数据中心# 算力基础设施# 绿色计算# 电网协同# 澳大利亚
InfoQ · 架构与云计算AI 评分 58/10005:57

4096张卡如何成为“一台计算机”?解析华为超节点布局

该内容探讨华为在AI算力基础设施领域的超节点技术布局,重点关注如何通过先进的互联与系统架构,将4096张计算卡高效协同、聚合成具备统一计算能力的“单台计算机”形态,以满足大模型时代的超大规模并行训练需求。由于提供的原文摘要信息不足,关于具体的互联协议、拓扑结构及软件栈支持等详细技术细节有待进一步查阅。

阅读原文 ↗推荐理由:聚焦华为在万卡级/千卡级超节点集群互联架构上的工程探索,对国产算力基础设施演进具有参考价值。# 华为# 超节点# 算力基础设施# 集群互联# AI芯片
Data Center Knowledge✦ 精选AI 评分 68/10005:42

数据中心用水挑战:从效率指标转向现实韧性考量

水资源正从数据中心的后台配套公用设施转变为核心瓶颈制约。随着行业发展,数据中心建设不能仅依赖平均效率指标,而需根据具体场地开展韧性规划,重点考量峰值用水压力、当地水文状况以及对周边社区的影响,以应对日益严峻的现实资源挑战。

阅读原文 ↗推荐理由:探讨了算力扩张背景下数据中心所面临的关键水资源制约与韧性规划转向。# 数据中心# 算力基础设施# 水资源# 绿色计算# 可持续发展
Google DeepMindAI 评分 55/10000:00

推进私有AI计算:引入安全服务器端内存技术

该动态宣布为面向个人AI的私有AI计算(Private AI Compute)引入安全的服务器端内存技术。由于目前提供的发布信息较为简短,文章主要聚焦于在云端服务器环境中强化内存隐私与数据安全,旨在兼顾大规模算力扩展与个人AI数据的机密性。具体技术细节及完整实现方案尚需进一步披露。

阅读原文 ↗推荐理由:聚焦通过安全服务器端内存提升私有AI计算的安全性,关乎个人AI在云端处理时的数据隐私保障。# 私有AI计算# 服务器端内存# 数据隐私# 云计算基础设施# 个人AI
9月23日2026-09-23
InfoQ · 架构与云计算AI 评分 45/10023:47

Dropbox 升级 Riviera 平台以支持 AI 工作负载

据消息披露,云存储与协作服务商 Dropbox 对其内部 Riviera 平台进行了升级,旨在增强对各类人工智能工作负载的支持能力。由于原始输入素材仅包含简略标题且缺乏具体正文内容,关于此次 Riviera 平台升级的技术架构细节、支持的具体 AI 任务类型、计算资源配置以及上线时间等详细信息,目前均无法进一步确认。

阅读原文 ↗推荐理由:展示了老牌云存储厂商针对 AI 算力与基础设施架构的技术演进,但由于原文缺少正文细节,参考价值有限。# Dropbox# Riviera# AI基础设施# 工作负载# 云计算
9月22日2026-09-22
Simon Willison规则精选06:25

Cloudflare Python Workers are now generally available

Cloudflare Python Workers are now generally available After a two year preview, Cloudflare's support for running Python code in their server-side Workers platform is now stable: "Python is now a first-class, fully supported language on the Cloudflare Developer Platform". A neat thing about this is how it works. Cloudflare are running Python compiled to WebAssembly via Pyodide in their V8-based workerd runtime. This comes with some limitations, documented here - most notably both multiprocessing and threading are non-functional in the WebAssembly VM. One particularly interesting detail of this is the local development environment story - their

NVIDIA · AI 筛选✦ 精选规则精选02:00

NVIDIA Launches DSX Ready to Qualify Power and Cooling Products for AI Factories

Every AI factory needs power and cooling that fit its computing architecture. As AI infrastructure expands, power, cooling, water, site and grid constraints are shaping what builders can deploy. Choosing products that fit the complete factory design helps builders turn computing capacity into useful AI output. To help builders make those decisions, NVIDIA is introducing […]

AWS 机器学习✦ 精选规则精选00:36

How BMW Group detects cost anomalies across 14,000 cloud accounts

BMW Group operates CLEA, a FinOps platform monitoring more than 14,000 cloud accounts. This post shows how BMW added automated daily cost anomaly detection, moving from reactive dashboards to proactive alerts using Prophet forecasting, AWS Step Functions, and a serverless pipeline that processes every account for about $50 per month.

9月19日2026-09-19
Ollama 更新✦ 精选规则精选08:15

v0.34.3-rc1

server: allow registry cross-host redirects among allowlisted hosts (…

vLLM 更新✦ 精选规则精选05:15

v0.30.0rc2

[Bugfix][NIXL] Avoid receive reports for notification-only requests (…

AWS 机器学习✦ 精选规则精选04:52

Amazon SageMaker Inference: 2026 year-to-date launches in review

Amazon SageMaker AI shipped 13 inference launches in year-to-date across two deployment paths: fully managed endpoints and Amazon SageMaker HyperPod Inference. This post reviews each launch, from inference recommendations and capacity-aware instance pools to tiered KV caching and disaggregated prefill and decode.

AWS 机器学习✦ 精选规则精选00:52

Introducing Kimi K3 on Amazon Bedrock

Kimi K3 from Moonshot AI is now available on Amazon Bedrock, giving you a powerful new open-weight option for coding and knowledge work. It offers native vision, a 1-million-token context window, and explicit prompt caching to reduce latency and input costs.

9月18日2026-09-18
Ollama 更新✦ 精选规则精选00:45

v0.34.2-rc2: mlxrunner: Release freed KV buffers during speculative decode

The decode loop releases MLX's pool of freed buffers every 256 generated tokens, which is also how often the KV cache grows and drops its previous, smaller buffers. The check fires only when the token count lands exactly on a multiple of 256. Speculative decoding emits several tokens per round, so most rounds step over the boundary and the pool is never released. Each growth at a long context leaves several GB of buffers that no later allocation can reuse, so the runner's footprint keeps climbing over a long generation until the system runs out of memory. We now release the pool whenever a round crosses a multiple of 256 tokens, which is what

9月17日2026-09-17
Ollama 更新✦ 精选规则精选05:06

v0.34.2-rc1: mlxrunner: lay out model by contract, checkpoint and construction

model is one package with three jobs: the contract between the runner and the architectures, the opened checkpoint, and building nn layers from checkpoint tensors. Its files did not say which was which. base.go carried the folded package's name over the interfaces and the registry, root.go held the safetensors header scan next to Root, and quant.go mixed the nvfp4 global-scale helpers with quant parameter resolution. base.go becomes model.go, named for what it holds. root.go keeps Root and Open; TensorQuantInfo and the header scan join quant.go, so everything the checkpoint says about quantization is read and resolved in one file. The global-

9月16日2026-09-16
NVIDIA · AI 筛选✦ 精选规则精选23:00

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics. Higher system performance means more tokens generated, resulting in higher revenue. Efficient scaling means throughput grows proportionally as hardware gets added, requiring fewer resources to serve users at scale. Continuous optimization means generating more value from infrastructure investments. […]

NVIDIA · AI 筛选✦ 精选规则精选21:00

Emerald AI, Google and NVIDIA Launch Alliance to Advance Flexible AI Data Centers

AI factories are the infrastructure of the intelligence era. Scaling them responsibly will depend as much on innovation across the grid as inside the data center. Today, Emerald AI, Google and NVIDIA announced the launch of the AI Energy Management Alliance (AEMA), a first-of-its-kind coalition advancing data centers that can dynamically manage their electricity use […]

MIT Technology Review AI规则精选20:47

Building the materials foundation for AI

The AI boom is becoming a materials challenge. As AI pushes computing into new territory, the materials behind that infrastructure are becoming just as crucial as the algorithms running on it. Semiconductors and data centers are approaching physical limits around performance, thermal management, electrical efficiency, and reliability, creating new demands for materials that can do…

Kubernetes Blog✦ 精选规则精选02:30

Kubernetes v1.37: Pod-Level Resource Managers graduated to Beta

With the release of Kubernetes v1.37, the Pod-Level Resource Managers feature has graduated to Beta status (disabled by default)! First introduced as an Alpha feature in Kubernetes v1.36 , this enhancement builds on Pod-Level Resources by equipping Kubelet's Topology Manager, CPU Manager, and Memory Manager to use Pod-level resource declarations ( .spec.resources ) directly when making hardware placement decisions. Bringing pod-level resources to node managers Before this feature, obtaining exclusive NUMA-aligned CPU cores or memory for latency-critical applications forced cluster operators into an all-or-nothing choice: assign integer resour

NVIDIA · AI 筛选✦ 精选规则精选00:55

From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production

On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked, Silicon Valley Power sent a signal to an AI factory to adjust its power consumption. Varun Sivaram was watching on Zoom with about forty others — his team at Emerald AI in their San Francisco conference room, engineers […]