By replaying Mooncake production traces, the authors show that KVC metadata workloads have high reuse, 86.8% sequential access, and mixed random lookups, and that existing key-value stores deliver poor, variable p99 latency.
vLLM vs TensorRT- LLM 12, Automatic Prefix Caching
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.ET 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference
By replaying Mooncake production traces, the authors show that KVC metadata workloads have high reuse, 86.8% sequential access, and mixed random lookups, and that existing key-value stores deliver poor, variable p99 latency.