Pith. sign in

hub

Expected attention: KV cache compression by estimating attention from future queries distribution

26 Pith papers cite this work. Polarity classification is still indexing.

26 Pith papers citing it

hub tools

citation-role summary

background 3

citation-polarity summary

years

2026 25 2024 1

roles

background 3

polarities

background 3

representative citing papers

CodeComp: Structural KV Cache Compression for Agentic Coding

cs.CL · 2026-04-11 · unverdicted · novelty 7.0

CodeComp uses Joern-extracted Code Property Graph priors for training-free structural KV cache compression, outperforming attention-only baselines on bug localization and code generation while matching full-context patch quality.

Express Language Modeling

cs.LG · 2026-06-09 · unverdicted · novelty 6.0

Express converts non-causal attention approximations to causal versions, achieving log^{3/2}(n)/s error with O(s) memory and O(s^2 log^2(n)) overhead when combined with Thinformer.

End-to-End Context Compression at Scale

cs.CL · 2026-06-08 · unverdicted · novelty 6.0

LCLMs are scaled 0.6B-encoder 4B-decoder compressors pre-trained on over 350B tokens that improve the Pareto frontier for general-task performance, compression speed, and peak memory in long-context language model inference.

Selectivity Estimation for Semantic Filters on Image Data

cs.DB · 2026-06-03 · unverdicted · novelty 6.0

Semantic Histograms treat semantic image filters as implicit range queries in embedding space and use two specificity estimators whose ensemble reduces end-to-end query optimization and execution overhead by up to 86%.

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference

cs.AR · 2026-05-17 · unverdicted · novelty 6.0

VeriCache turns lossy KV cache compression into lossless LLM inference by drafting with compressed cache and verifying drafts with full cache, achieving up to 4x throughput with identical outputs.

Rethinking LoRA Memory Through the Lens of KV Cache Compression

cs.CL · 2026-06-04 · unverdicted · novelty 5.0

Document LoRA acts as decoding-time parametric memory that recovers 13-21 ROUGE-L points under heavy KV cache compression in QA, performing best when the base model encodes the document and the adapter is used only at generation with QA supervision.

NestedKV: Nested Memory Routing for Long-Context KV Cache Compression

cs.CL · 2026-05-26 · unverdicted · novelty 5.0

NestedKV is a new training-free KV cache compression technique using nested memory anchors and multi-time-scale anomaly scoring that outperforms prior methods like KeyDiff on long-context benchmarks when the retained cache fraction is small.

citing papers explorer

Showing 26 of 26 citing papers.