Pith. sign in

Paper Citation Record · LEDGER

Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2403.09636.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.09636 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:37:07.614689Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.094915Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8c34dbbd-0ade-4422-bffa-c6884b0912b5 · inbound

A Survey of Mamba cites this paper.

A Survey of Mamba Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 139

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:13:30.580273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-23T22:09:19.917854Z digest=sha256:2a1e24a85471428e5f45c829affcb2c66afc75d5fc022e90e455d652486691e0

Observation ffe77f9b-36cd-461f-8c2e-dbb838f22c99 · inbound

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation cites this paper.

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.343855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-23T18:31:35.391674Z digest=sha256:854f421df48c60e3665506c46819d0a3e059796341cf2e0f62fa68ca43230102

Observation e45af902-9bf3-44bc-bf9b-cd7f903b4e55 · inbound

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache cites this paper.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.614689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.614689Z digest=sha256:ffa1e23f35f1f0c7405b8c2ff0352c9cd890c9b00227aa4fa3a9c27a37361c8c

Observation deeff9c6-4613-428d-b87c-589864345c4d · inbound

DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs cites this paper.

DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:57:17.963561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:57:17.963561Z digest=sha256:c3e54c452f8b275e6e70a08c4b7ffad7d37742fd44d743bf250311891d348791

Observation f69b1452-feb0-4910-b4e3-6dc32ed2a7bb · inbound

Provence: efficient and robust context pruning for retrieval-augmented generation cites this paper.

Provence: efficient and robust context pruning for retrieval-augmented generation Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T13:43:43.814854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:43:43.814854Z digest=sha256:815972adad1edd13f077e9dfee5d1d5aa586f3ed7c059c76025e93def5386633

Observation 4db5cccd-a4d0-42f5-bcf4-8a0e0964a462 · inbound

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents cites this paper.

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T14:13:38.251701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:13:38.251701Z digest=sha256:a1ef23e95a9440425eebf423bb16533848978157d6fe8e6cc4f7dc662d92b532

Observation 9058b287-13a9-472f-8692-f8d0b39a65b5 · inbound

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache cites this paper.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.950468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.950468Z digest=sha256:9cc40ae11ce9a3c536dff6a19b7289978113d15f281ff646d1201ee4aa89db28

Observation cb57c20c-15ac-49ce-a4b9-65c607e9ffb3 · inbound

PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding cites this paper.

PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:37:14.222217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T09:33:56.616828Z digest=sha256:b6d444053f7c9fb0a5a7dfcf8234e3d90b91c2c8f58d1a2e2b0729a030d1afb3

Observation 65147fb4-8ae0-4ecd-9133-6c47dfc388b1 · inbound

Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures cites this paper.

Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T20:15:52.776309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:15:52.776309Z digest=sha256:6ec298a1a9dbd3ea620e4c34512b1a3575739ab130c064b286509c93a6fb649c

Observation b6763228-711e-4bcd-8240-3fc1b58fd1d5 · inbound

Controllably Efficient Language Models cites this paper.

Controllably Efficient Language Models Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:58.921784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:58.921784Z digest=sha256:1909f4bc9c7eeae347e324981144b2a90b4a0503c43a68879a2fb9991205ef49

Observation def493b9-6fda-471b-9d74-9ccbd470179a · inbound

eOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and Quantization cites this paper.

eOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and Quantization Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:00:50.840025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T20:18:04.392331Z digest=sha256:7ef332936a937a50b32598f323569b58e1d886292499316fddd8087653e8079c

Observation 2b02129f-4394-4a54-9464-2cf8a86f6c50 · inbound

Sparse Prefix Caching for Hybrid and Recurrent LLM Serving cites this paper.

Sparse Prefix Caching for Hybrid and Recurrent LLM Serving Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:53:04.425034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T08:49:24.880528Z digest=sha256:9d8e629914faf8fef944347927cd4d81ceed3e8911bb60376ad7df8bcd9660d0

Observation 546364dd-325b-4ad1-a1f1-8c5b47a9126e · inbound

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction cites this paper.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:26.589479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:3e34a3e1b99c53d01e34b2d0e91a1224feaffa0ab6d553ebab7f7ad4c12940a4

Observation 0307c91a-251b-4556-a284-eac48f39b2fd · inbound

MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference cites this paper.

MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:56:15.415430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T16:07:22.196982Z digest=sha256:a98f4d73282cdb060bd812615243cee14a097fbefce5589ed03e21987c78c450

Observation b714451a-be66-4a3f-b7c6-4554e09617e3 · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:9b883e0a7a7676104043d721b14fa48e82578aa16baf1b025f61cbb06ee70355

Observation 4ff14cb3-c98e-4347-b95b-9a310c60e865 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.096075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:46dab5079e1179a23d317e9870ab931724238d5b7cf13c0ec357a5257a8b084e

Observation a613a1ae-93f0-4f72-9b28-f777bcef99f4 · inbound

Hierarchical Domain Generalization cites this paper.

Hierarchical Domain Generalization Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T20:54:08.847193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T20:54:08.847193Z digest=sha256:ab009fd6f571aece3f9e733423a821e3c0ebb1c66f29484ab1049eebca6a98cc

Observation dc986122-34a1-4295-aa6b-8deca33979f4 · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:52.665974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:43:52.665974Z digest=sha256:ed1238f4ea246782bfabede96b3cf8a74a1a0fcf3bae411fe6fe3872a031176d

Observation 8f4fd67f-3406-436a-b83f-69e1650e6e9c · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:11:48.853872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:11:48.853872Z digest=sha256:a279792e383fc9448493f4de315c4f6928313e2d11f41f0340cecbfb59b78a7c