A translation table plus asynchronous live-token repacking lets token-granular KV eviction reclaim underutilized physical blocks in PagedAttention-style serving runtimes, improving memory efficiency and throughput in paired vLLM experiments.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
vToken: Token-Level Virtualization for Reclaimable KV Caches
A translation table plus asynchronous live-token repacking lets token-granular KV eviction reclaim underutilized physical blocks in PagedAttention-style serving runtimes, improving memory efficiency and throughput in paired vLLM experiments.