S4R compresses the KV cache up to 5x by learning low-rank subspaces from sampled prompt tokens and reconstructing only a small set of relevant keys and values during decoding.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
S$^4$R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching
S4R compresses the KV cache up to 5x by learning low-rank subspaces from sampled prompt tokens and reconstructing only a small set of relevant keys and values during decoding.