Pith. sign in

Paper Citation Record · LEDGER

Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2402.09398.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.09398 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:40:37.634455Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T07:04:20.871745Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c935123c-84ea-458a-80dd-3a8a694ba6a4 · inbound

PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling cites this paper.

PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:58:29.127993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T09:58:29.057357Z digest=sha256:af924adace5770581ce61463dc994c9fe04fd37141330a53fa62aa973a01b979

Observation 4b48bf19-1d0a-4999-b8fc-deceae71d9ed · inbound

MARM: Unlocking the Future of Recommendation Systems through Memory Augmentation and Scalable Complexity cites this paper.

MARM: Unlocking the Future of Recommendation Systems through Memory Augmentation and Scalable Complexity Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T20:44:06.061422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:44:06.061422Z digest=sha256:850a6a38019cd6185c1bfd6b5c89c6c5e376b8f22e07d415a3fd44e4f6bef503

Observation e1c26979-def8-487e-93b5-e8de2d577094 · inbound

InputSnatch: Stealing Input in LLM Services via Timing Side-Channel Attacks cites this paper.

InputSnatch: Stealing Input in LLM Services via Timing Side-Channel Attacks Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:29:29.338559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:29:29.338559Z digest=sha256:ccaab8bae6e72bd8c59eea4a52ab25c0a910f170b8071f1b97bf850d21d7de56

Observation f6a282a6-7ad4-4513-aa8e-66ae7775740d · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:46.940015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:46.940015Z digest=sha256:bc59f7ca7aaeab23f45b75e6bbff2010d2d15a5e5b3da5b7ffcaf08c9896131a

Observation efcfaefd-a247-4353-a861-efdbe90e7350 · inbound

Climber: Toward Efficient Scaling Laws for Large Recommendation Models cites this paper.

Climber: Toward Efficient Scaling Laws for Large Recommendation Models Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T20:15:13.933003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:15:13.933003Z digest=sha256:dbae473e2035fa73261bdf8c16245b7de97f25d820f7435112a5574cb5994ca0

Observation 49442074-052e-4d9c-a885-1bd1d94332c4 · inbound

Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation cites this paper.

Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T05:40:21.420376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:40:21.420376Z digest=sha256:30548fcc4ed58a31ac39a78bb69bdcd482d30b5bd05256f5d20695bf93edb794

Observation 7fcf1d28-fe4a-4a9b-8672-6af3e98962f2 · inbound

RCStat: A Statistical Framework for using Relative Contextualization in Transformers cites this paper.

RCStat: A Statistical Framework for using Relative Contextualization in Transformers Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:40:37.634455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:40:37.634455Z digest=sha256:b8e8b7b3fb5d395bf73ffbaaad4c3ba76caf14a290215ab65035b8d8729343f1

Observation 1dc4b474-fa34-4e3b-b537-3d0469b20443 · inbound

Tensor Cache: Eviction-conditioned Associative Memory for Transformers cites this paper.

Tensor Cache: Eviction-conditioned Associative Memory for Transformers Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:50:23.550340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-25T05:49:43.160574Z digest=sha256:1e7c064259b5881444cff3bf7791dbea93f04eeca1814c3a91f7591e333137d4

Observation c7bf44d1-23f0-4280-a864-204f71d92004 · inbound

Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding cites this paper.

Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:04:20.873934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T07:03:08.617257Z digest=sha256:1c1cf138bd58e5af03adf9991dac2263e9fcd91208d46cadb975212c4db323fa

Observation b59e9c71-c4f4-41f5-a174-fe493555aa17 · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:6f7c4a59c8d4a680b054a5b4aedb06e496e094c96ed7eb5db5a441cc0c58a2b4

Observation 77917668-f16d-4abe-b609-fe7f6bb63bec · inbound

GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference cites this paper.

GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T09:52:44.106186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:52:44.106186Z digest=sha256:3ec44373fc15268c116b3b794d37c28cebcd54d5c9c1f159db9ebf55acd0450c