DepthWeave-KV achieves 8.3x KV cache memory reduction with near-full-cache task quality by factorizing key-value states across transformer layers using shared bases and token-adaptive residuals.
Title resolution pending
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
VaSE improves KV cache eviction accuracy for reasoning models by over 4% versus prior eviction methods at 4x compression through value-magnitude protection and stochastic diversity.
A survey organizing serving-time KV cache optimization techniques into temporal, spatial, and structural system behaviors, analyzing cross-behavior co-design patterns and open challenges.
citing papers explorer
-
DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression
DepthWeave-KV achieves 8.3x KV cache memory reduction with near-full-cache task quality by factorizing key-value states across transformer layers using shared bases and token-adaptive residuals.
-
Value-Aware Stochastic KV Cache Eviction for Reasoning Models
VaSE improves KV cache eviction accuracy for reasoning models by over 4% versus prior eviction methods at 4x compression through value-magnitude protection and stochastic diversity.
-
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization
A survey organizing serving-time KV cache optimization techniques into temporal, spatial, and structural system behaviors, analyzing cross-behavior co-design patterns and open challenges.