华为发布全新UnifiedBus计算架构,助力SuperPoD及超算集群互连
为应对快速增长的算力需求,华为正式推出了一套全新的计算架构。该架构采用UnifiedBus互连技术,专门用于支持SuperPoD与大规模算力集群之间的高效互连。该设计旨在突破计算节点间的连接瓶颈,优化集群整体性能与扩展能力,为大规模AI训练及高性能计算基础设施提供更强的底层算力支持。
为应对快速增长的算力需求,华为正式推出了一套全新的计算架构。该架构采用UnifiedBus互连技术,专门用于支持SuperPoD与大规模算力集群之间的高效互连。该设计旨在突破计算节点间的连接瓶颈,优化集群整体性能与扩展能力,为大规模AI训练及高性能计算基础设施提供更强的底层算力支持。
随着数据中心开始采用更多样化的发电与电力供应技术,相关设备的运输、仓储和安装等物流流程正变得愈发复杂。报道指出,伴随计算能耗剧增,新型供电解决方案的引入虽然缓解了能源压力,但也对数据中心上下游的供应链物流管理提出了更高要求,部署重型及特种电力设备已成为基础设施建设中的关键考验。
现代信息检索系统通常同时结合嵌入模型与生成模型处理复杂查询,但现有系统往往将二者隔离部署,粗粒度分配硬件易引发计算气泡,导致 GPU 利用率低和吞吐不足。为此,研究团队提出服务系统 Orthrus。该系统在统一推理循环中引入异构批处理机制,旨在有效整合具有资源冲突特性的嵌入与生成工作负载,提升动态查询下的硬件利用效率。(注:素材摘要文末存在截断)
报道聚焦于成立九年的中科类脑公司,指出其将过往的技术与业务积累整合并投入到“Token工厂”的建设与布局中,暗示其在大模型时代围绕算力资源与Token生成服务进行的转型与探索。由于输入素材未提供详细正文或摘要内容,关于该“Token工厂”的具体技术架构、产品形态、商业落地细节及最新进展等信息尚不明确,有待后续报道补充。
该内容探讨了大型能源消耗主体如何更好地为能源系统及其服务的社区做出贡献。澳大利亚为此出台了相关数据中心指导准则,旨在推动数据中心等高耗能设施成为电网的“良好公民”,探索其在支撑能源系统稳定与可持续发展方面的协作平台与基础。(注:素材信息有限,主要呈现了准则的愿景与探讨方向)
该内容探讨华为在AI算力基础设施领域的超节点技术布局,重点关注如何通过先进的互联与系统架构,将4096张计算卡高效协同、聚合成具备统一计算能力的“单台计算机”形态,以满足大模型时代的超大规模并行训练需求。由于提供的原文摘要信息不足,关于具体的互联协议、拓扑结构及软件栈支持等详细技术细节有待进一步查阅。
水资源正从数据中心的后台配套公用设施转变为核心瓶颈制约。随着行业发展,数据中心建设不能仅依赖平均效率指标,而需根据具体场地开展韧性规划,重点考量峰值用水压力、当地水文状况以及对周边社区的影响,以应对日益严峻的现实资源挑战。
该动态宣布为面向个人AI的私有AI计算(Private AI Compute)引入安全的服务器端内存技术。由于目前提供的发布信息较为简短,文章主要聚焦于在云端服务器环境中强化内存隐私与数据安全,旨在兼顾大规模算力扩展与个人AI数据的机密性。具体技术细节及完整实现方案尚需进一步披露。
据消息披露,云存储与协作服务商 Dropbox 对其内部 Riviera 平台进行了升级,旨在增强对各类人工智能工作负载的支持能力。由于原始输入素材仅包含简略标题且缺乏具体正文内容,关于此次 Riviera 平台升级的技术架构细节、支持的具体 AI 任务类型、计算资源配置以及上线时间等详细信息,目前均无法进一步确认。
Signed-off-by: Andreas Karatzas akaratza@amd.com Co-authored-by: OpenAI Codex noreply@openai.com
mlx: speed up Qwen 3.8 prompt processing Use MLX's gated-delta kernel for long scans and fold dense MLP global scales into SwiGLU. address comments
Concurrency sweeps help you right-size a generative AI endpoint on Amazon SageMaker AI by systematically benchmarking it at increasing load levels. This post walks through deploying a model, running automated concurrency sweeps with the CreateAIBenchmarkJob API, and using the results to make data-driven capacity decisions about fleet size.
Cloudflare Python Workers are now generally available After a two year preview, Cloudflare's support for running Python code in their server-side Workers platform is now stable: "Python is now a first-class, fully supported language on the Cloudflare Developer Platform". A neat thing about this is how it works. Cloudflare are running Python compiled to WebAssembly via Pyodide in their V8-based workerd runtime. This comes with some limitations, documented here - most notably both multiprocessing and threading are non-functional in the WebAssembly VM. One particularly interesting detail of this is the local development environment story - their
Every AI factory needs power and cooling that fit its computing architecture. As AI infrastructure expands, power, cooling, water, site and grid constraints are shaping what builders can deploy. Choosing products that fit the complete factory design helps builders turn computing capacity into useful AI output. To help builders make those decisions, NVIDIA is introducing […]
BMW Group operates CLEA, a FinOps platform monitoring more than 14,000 cloud accounts. This post shows how BMW added automated daily cost anomaly detection, moving from reactive dashboards to proactive alerts using Prophet forecasting, AWS Step Functions, and a serverless pipeline that processes every account for about $50 per month.
server: allow registry cross-host redirects among allowlisted hosts (…
[Bugfix][NIXL] Avoid receive reports for notification-only requests (…
Amazon SageMaker AI shipped 13 inference launches in year-to-date across two deployment paths: fully managed endpoints and Amazon SageMaker HyperPod Inference. This post reviews each launch, from inference recommendations and capacity-aware instance pools to tiered KV caching and disaggregated prefill and decode.
Kimi K3 from Moonshot AI is now available on Amazon Bedrock, giving you a powerful new open-weight option for coding and knowledge work. It offers native vision, a 1-million-token context window, and explicit prompt caching to reduce latency and input costs.
The decode loop releases MLX's pool of freed buffers every 256 generated tokens, which is also how often the KV cache grows and drops its previous, smaller buffers. The check fires only when the token count lands exactly on a multiple of 256. Speculative decoding emits several tokens per round, so most rounds step over the boundary and the pool is never released. Each growth at a long context leaves several GB of buffers that no later allocation can reuse, so the runner's footprint keeps climbing over a long generation until the system runs out of memory. We now release the pool whenever a round crosses a multiple of 256 tokens, which is what
Signed-off-by: jiahanc 173873397+jiahanc@users.noreply.github.com Co-authored-by: OpenAI Codex codex@openai.com
Abandoned jobs, instances, and volumes can run indefinitely. FinOps tools help, but GPU-era AI demands new zombie-hunting methods.
Release vllm-proto 0.3.0
model is one package with three jobs: the contract between the runner and the architectures, the opened checkpoint, and building nn layers from checkpoint tensors. Its files did not say which was which. base.go carried the folded package's name over the interfaces and the registry, root.go held the safetensors header scan next to Root, and quant.go mixed the nvfp4 global-scale helpers with quant parameter resolution. base.go becomes model.go, named for what it holds. root.go keeps Root and Open; TensorQuantInfo and the header scan join quant.go, so everything the checkpoint says about quantization is read and resolved in one file. The global-
System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics. Higher system performance means more tokens generated, resulting in higher revenue. Efficient scaling means throughput grows proportionally as hardware gets added, requiring fewer resources to serve users at scale. Continuous optimization means generating more value from infrastructure investments. […]
AI factories are the infrastructure of the intelligence era. Scaling them responsibly will depend as much on innovation across the grid as inside the data center. Today, Emerald AI, Google and NVIDIA announced the launch of the AI Energy Management Alliance (AEMA), a first-of-its-kind coalition advancing data centers that can dynamically manage their electricity use […]
The AI boom is becoming a materials challenge. As AI pushes computing into new territory, the materials behind that infrastructure are becoming just as crucial as the algorithms running on it. Semiconductors and data centers are approaching physical limits around performance, thermal management, electrical efficiency, and reliability, creating new demands for materials that can do…
Validated by PR #56538 CI at fa2a26f .
With the release of Kubernetes v1.37, the Pod-Level Resource Managers feature has graduated to Beta status (disabled by default)! First introduced as an Alpha feature in Kubernetes v1.36 , this enhancement builds on Pod-Level Resources by equipping Kubelet's Topology Manager, CPU Manager, and Memory Manager to use Pod-level resource declarations ( .spec.resources ) directly when making hardware placement decisions. Bringing pod-level resources to node managers Before this feature, obtaining exclusive NUMA-aligned CPU cores or memory for latency-critical applications forced cluster operators into an all-or-nothing choice: assign integer resour
On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked, Silicon Valley Power sent a signal to an AI factory to adjust its power consumption. Varun Sivaram was watching on Zoom with about forty others — his team at Emerald AI in their San Francisco conference room, engineers […]