全部 AI 动态
Watermarking in vLLM
How vLLM implements distribution-preserving Gumbel-max text watermarking with efficient GPU kernels, statistical detection, speculative decoding, and repeated-c
Announcing vllm-metal: Concurrent Serving on Apple Silicon
vllm-metal brings vLLM's paged, continuously batched serving stack to Apple Silicon, with flatter TTFT under concurrent agent load, batched MTP, and automatic M
PD Serving of Qwen3.8-2.4T
How vLLM reaches 5K throughput and 180 interactivity on Qwen3.8-2.4T with GB300 NVL72 PD serving and how to reproduce results yourself.
Scaling Multi-GPU Video Captioning with PyNvVideoCodec and vLLM
How to leverage NVIDIA Hardware Video Decoders to Achieve Multi-GPU Scaling in Video Captioning and Description tasks.
How we trained the fastest DSpark for Kimi-K3 using GB300 NVL72
How Speculators and Mooncake enabled multi-node DSpark training for Kimi K3.
vLLM x Novita AI: Chord, Faster INT4 MoE for Kimi K2.x. Up to 1.3x on H200, 2.15x on Untuned B300
Novita AI has open-sourced Chord, a high-performance W4A16 MoE CUDA kernel for Kimi K2.x serving shapes, with a Humming-compatible indexed path and grouped SM90
vime × RL-Kernel × AMD: Bitwise Train–Rollout Consistency on ROCm
vime and RL-Kernel align selected-token logprobs bit for bit across Megatron training and vLLM rollout on AMD Instinct MI300X, with zero mismatches across 200 G
Kimi K3 Performance Optimizations in vLLM: The Road to 2.8× Throughput
Kimi K3 serving optimizations across scheduling, KDA prefix caching, ReplaySSM state recovery, PD disaggregation and state offload, parallelism, MoE, and GPU ke
Following the Bottleneck: Optimizing MiniMax M3 on AMD Instinct MI355X
A performance model for LLM serving: inspect local shapes, remove repeated work, verify data movement and dispatch, then follow the queue.
Tiered KV Cache Offloading in vLLM
A host-centric framework for scaling KV cache across host memory, filesystems, object stores, and remote peers — reducing recomputation and increasing serving c
GLM 5.3 Optimizations, Part 1: Hybrid HiSparse Offloading in vLLM
vLLM integrates HiSparse as a pressure-driven memory tier that composes with the Hybrid Memory Allocator and KV offloading, letting GLM 5.3 requests keep decodi
vLLM x AgentX: Optimizing for Real-World Agentic Serving
How vLLM optimizes KV cache management, parallelism, scheduling, and P/D disaggregation for agentic workloads, validated on SemiAnalysis AgentX with up to 130K
Serving LLMs on Tenstorrent Hardware: Inside the vLLM TT Plugin
Tenstorrent accelerators join vLLM as an out-of-tree platform plugin, driven by mesh-architecture choices: phase-based scheduling, single-process data paralleli
MiniMax H3 on vLLM-Omni: From System-Wide Optimization to Real-Time Serving with FastVideo’s FastH3
How vLLM-Omni optimizes and scales the complete MiniMax H3 stack, then integrates FastVideo’s four-step FastH3 for generation faster than playback.
Large-Scale Sharded Weight Transfer with Ray Direct Transport (RDT) in vLLM
We implement a native sharded weight transfer engine in vLLM utilizing Ray Direct Transport (RDT), achieving weight transfer for the Kimi K2 model in BF16 on 48
IsoExec: Unified Execution to Eliminate Trainer-Inference Mismatch in SkyRL
IsoExec unifies numerical execution across SkyRL's vLLM and Megatron runtimes, reducing the average rollout-versus-training logprob difference below 1e-6 on Qwe
VeRL-Omni v0.2.0: Faster Diffusion RL and Stable Omni Training
A release focused on higher-throughput diffusion rollout, reusable omni adapters, and broader recipe coverage.
Distributed Layerwise Offload: Scaling Toward 200B+ DiT Models Efficiently in vLLM-Omni
Distributed Layerwise Offload shards and streams DiT weights across devices, serving a measured 124 GB Cosmos3 model on 64 GB HBM and estimating a path toward 2