DCAI
← 返回全部动态
vLLM 官方博客(网页)规则精选09月10日 00:00

Tiered KV Cache Offloading in vLLM

A host-centric framework for scaling KV cache across host memory, filesystems, object stores, and remote peers — reducing recomputation and increasing serving c

阅读 vLLM 官方博客(网页) 原文 ↗