DCAI
DC AI 热点

全部 AI 动态

8月18日2026-08-18
8月7日2026-08-07
8月5日2026-08-05
AWS HPC✦ 精选规则精选08:00

Resilient HPC and ML on AWS: Running Tightly Coupled Workloads on Spot Instances

This post references AWS ParallelCluster. Check out AWS Parallel Computing Service (AWS PCS), our new managed Slurm service for running HPC and AI workloads on AWS. This post was contributed by Santosh Kumar, Bhagyaraju Kasina, Dr. Sandeep Sovani and Dr. Max Starr Researchers and engineering teams running High Performance Computing (HPC) jobs face a constant […]

7月23日2026-07-23
7月21日2026-07-21
7月16日2026-07-16
Transformers 更新✦ 精选规则精选17:41

Patch release: v5.14.1

Patch release v5.14.1 This patch solves a few issues which appeared when integrating Inkling model, most notably an issue affecting models using EncoderDecoderCache during assisted generation. It also fixes an issue that could appear during prefill with StaticCache and sdpa without padding for Inkling which uses a position_bias. It contains the following commits: Fix sdpa prefill with position_bias ( #47359 ) by @Cyrilvallez Fix assisted decoding for models with EncoderDecoder cache & OlmoHybrid ( #47361 ) by @Cyrilvallez [FP8] Bump kernels version ( #47344 ) by @vasqu Fix deepgemm on multiple devices ( #47323 ) by @IlyasMoutawwakil

7月15日2026-07-15
Kubernetes Blog✦ 精选规则精选02:00

Building a Custom Metrics Exporter for Kubernetes

Kubernetes ships with built-in awareness of CPU and memory, but most real-world scaling decisions depend on signals that live entirely outside that narrow window: how many messages are waiting in a queue, how long the last batch job took, how many active WebSocket connections a pod is holding. When the built-in metrics are not enough, a metrics exporter bridges that gap. This post walks through writing one from scratch, packaging it as a container, and wiring it into a cluster so that Prometheus — and ultimately the HorizontalPodAutoscaler — can consume it. What a metrics exporter actually does An exporter is a small HTTP server with a single

7月11日2026-07-11
Transformers 更新✦ 精选规则精选17:15

Patch release v5.13.1

Patch release v5.13.1 This patch is focused on enabling transformers for the latest release of vllm! Be more defensive with remap_legacy_layer_types for custom models ( #47245 ) from @hmellor Fix custom code which doesn't know about the new linear layer type names ( #47174 ) from @hmellor Fix case where _LazyAutoMapping.register is passed a str key ( #47148 ) from @hmellor

6月26日2026-06-26
AWS HPC✦ 精选规则精选08:15

Transforming HPC Operations with Intelligent Workload Orchestration on AWS

This post was contributed by Manu Pillai, Gloria Macia and Natalia Jimenez, PhD Organizations running high-performance computing (HPC) workloads today operate largely as they have for decades: users manually specify required compute specifications for each of their jobs. Users spend valuable time analyzing workload requirements, selecting instance types, and troubleshooting infrastructure issues – time that […]

6月25日2026-06-25
Kubernetes Blog✦ 精选规则精选02:00

Spotlight on WG Device Management

The rising popularity of AI, Edge, and Telecommunications workloads on Kubernetes has led to new requirements for hardware management. We now need hardware specification beyond CPU time and memory allocations. This includes allocating GPUs, TPUs, network interfaces, and other hardware, sometimes after pod start and occasionally through time-sharing. Efficiently managing this specialized hardware is the mission of the Device Management Working Group . Their cornerstone project, Dynamic Resource Allocation (DRA) , recently graduated to GA, marking a fundamental shift in how the project handles hardware-intensive workloads at scale. In this spot

6月16日2026-06-16
Transformers 更新✦ 精选规则精选01:29

Patch release v5.12.1

Patch release v5.12.1 Updated the lower bound for PEFT and a fix for auto tokenizer to properly resolve the mistral tokenizer (when mistral-common is installed). This is similar to v.5.10.3 minus the fixes that were already included in the main release - vLLM will first target 5.10.3 🤗 Fix peft lower bound #46605 by @hmellor ( #46605 ) mistral common backend fix #46667 by @itazap ( #46667 ) Full Changelog : v5.12.0...v5.12.1

6月13日2026-06-13
6月12日2026-06-12
6月11日2026-06-11
6月8日2026-06-08
AWS HPC✦ 精选规则精选21:51

Reducing costs by 50% while processing population-scale genomics with Mountpoint for Amazon S3 and AWS Batch

This post was contributed by Kambiz Shahim, Ankit Kalyani, and Chris Wright. Oxford Nanopore Technologies used Mountpoint for Amazon S3, AWS Batch, and Nextflow to build EPI2ME Cloud, a managed compute environment for processing human genomes at population-scale reliably and securely while reducing computational costs by 50%. EPI2ME Cloud forms Oxford Nanopore’s suite of local […]

6月4日2026-06-04
6月3日2026-06-03
AWS HPC✦ 精选规则精选00:29

Monitoring AWS Parallel Computing Service

This post was contributed by Ronald Hudson and Nate Haynes High Performance Computing (HPC) on AWS demands precise monitoring, like the racing telemetry used by Formula 1 teams to deliver results. Like race engineers tracking car performance, AWS Parallel Computing Service (AWS PCS) administrators must monitor computing metrics in real-time. This vigilance is critical because […]

6月2日2026-06-02