Pith. sign in

Paper Citation Record · LEDGER

MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2405.14366.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.14366 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:20:32.412901Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a54f002f-c183-4e53-88e0-19641b7e1091 · inbound

Hymba: A Hybrid-head Architecture for Small Language Models cites this paper.

Hymba: A Hybrid-head Architecture for Small Language Models MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.412901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.412901Z digest=sha256:4c1fe3c176ed6a648ae7a36becd36f44fe1b45a4aae9a5ce904c9cef9e46e5a3

Observation 39ffc95a-dca4-42ef-829d-d6bc784359bd · inbound

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache cites this paper.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.595756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.595756Z digest=sha256:3a9d3d2453abe70d84f315f968890ecc413fae5f069e6152267964e7ba377a51

Observation ab5fa997-8eea-4f97-b365-6962333cce6e · inbound

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity cites this paper.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.056231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.056231Z digest=sha256:408e8c48d892bca58e7885d9c2bf6a4b7f5325bb0bc2323983cf73d4bdf3ff1d

Observation b35376c6-1d3e-4277-9e74-9f3fa2bdc621 · inbound

DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs cites this paper.

DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:57:17.950846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:57:17.950846Z digest=sha256:d44dfbbb74f3018f5c55995da94821a5ab4cd94fc875e0cdf79a71093f7fa131

Observation a09a16a1-07cc-4ba0-a737-ae241f0805eb · inbound

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing cites this paper.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.131384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.131384Z digest=sha256:1b39656198f5cd00936689216a45859fe6048ee5acec9c96ac2b46f35a7fbd63

Observation 51faa30f-39bb-4c62-ba21-5637acb71c24 · inbound

A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression cites this paper.

A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:29.447672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:29.447672Z digest=sha256:7db9483995173bd14fcc7666e753109f20c67b4e49bff5ac151a2fe88a949432

Observation 437891ab-6e45-49eb-b1ab-69b9315df6f1 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:47.643051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:47.643051Z digest=sha256:2f57096ca9befc8d68fa794fcaac73d1b08a23fe0f5f2bc1da8c13b12765332e

Observation de681a8d-578e-4497-80dd-db5200ce5d35 · inbound

Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads cites this paper.

Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:41:05.013818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:41:05.013818Z digest=sha256:711febc3ec1a64c386face508806de044c42d87eb7e835317b8c12d328efd506

Observation 8100fd37-f4e5-42b6-8a23-f126bcce6ec9 · inbound

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models cites this paper.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.274336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.274336Z digest=sha256:3e8329f72c4c9a25e0e06b2ea3eda4454022191671d825cb355a2b650463c7cd

Observation 0c27c169-0ae3-49b1-9f7d-caf7525136ff · inbound

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression cites this paper.

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:31.255973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T04:15:36.906263Z digest=sha256:40a4c9c948fc3a712df74e8cefcf982e666c6bb6d988f14416026fe4f523776b

Observation cff4f956-d691-4937-a2a9-a02e8aa7f41c · inbound

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents cites this paper.

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T14:13:38.173419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:13:38.173419Z digest=sha256:b1ad232e1722b45e661b5c6cc77b50211bb91dde7de39d85790c9bc1c4a5d049

Observation 10b2764a-bae7-4d44-83c7-958a0cdd247b · inbound

TransMLA: Multi-Head Latent Attention Is All You Need cites this paper.

TransMLA: Multi-Head Latent Attention Is All You Need MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T11:48:01.204132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:48:01.204132Z digest=sha256:ffa28449c323a9ab45f08b08704c446cbb182dfc20d7010075d53e0c9c0e64f0

Observation f5509a30-74a1-4411-b871-35b83934b4b2 · inbound

LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation cites this paper.

LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T22:12:11.625691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T22:07:37.906916Z digest=sha256:85876e3fbcf68178408f57e16a5c1fee8910c92f37f5b9bd84657dc0be6a00fd

Observation 46aec037-8f7c-40db-bf10-3ec3a2194891 · inbound

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression cites this paper.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.805734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.805734Z digest=sha256:ed25f01d73b2e63cb155f14e5099d00e971948b8896cf0ce676e2dfeb7e1dd88

Observation efa4715b-447e-4500-b65e-e60faca4728b · inbound

Semantic Scheduling for LLM Inference cites this paper.

Semantic Scheduling for LLM Inference MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T01:09:20.984252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:09:20.984252Z digest=sha256:0abf1ed9a71540790bf29b7244d6a17b307ae233289d917e3c75cc47aeabc8fc

Observation 338b3950-8251-4c0f-99be-28abbd3377f5 · inbound

SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers cites this paper.

SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:01.211946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:10:01.211946Z digest=sha256:f45b457753ed67b000ba7756a1b55f84fb7c36fa3909abeb193f1f1c34978853

Observation 958f7d1e-9fa0-4a05-91c7-7d69f4b8c3dd · inbound

S2O: Early Stopping for Sparse Attention via Online Permutation cites this paper.

S2O: Early Stopping for Sparse Attention via Online Permutation MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:36:32.930897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T19:32:52.948154Z digest=sha256:b2c92e32196101199c213b12c54cc3cdb2b6840824f77ff68d0a40091d9ca5fb

Observation 20a2a60e-c8b6-45a4-b02f-46fc08f9e349 · inbound

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling cites this paper.

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:27:24.426091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T19:46:43.514413Z digest=sha256:756a53a8573726b7a8f19f70aad96073fc537c8fcd11796d2faf8d237a7fec05

Observation 6aa01486-410a-4af5-afd6-ddaead8d8ea1 · inbound

RoPE-Aware Bit Allocation for KV-Cache Quantization cites this paper.

RoPE-Aware Bit Allocation for KV-Cache Quantization MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:59:57.378678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T01:05:13.790381Z digest=sha256:eccf644e95bef1f85e2924848eaee237a35ef38cb4852b322dd8e6821ab17c36

Observation 90c45ba2-3c61-48ec-823f-22d5f61ed95d · inbound

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference cites this paper.

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-07-11T00:17:45.321384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-11T00:12:13.830917Z digest=sha256:847214159f8abcefc1e3e2f697790fe9a294a3e2ca94a85dbfd4cbd85ce43b1d

Observation 1912ef61-47c5-439e-be22-b44e35444151 · inbound

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression cites this paper.

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 103

Resolution
verified exact
local_arxiv, observed 2026-07-08T03:14:31.504835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-08T03:07:23.648382Z digest=sha256:d25d4dc51925b39befc3d484faed36fe3e53358c6b6a3858781aaf544be7134f

Observation 8c1ab5a5-eceb-4b81-b1dc-7e73b170419a · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.024412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:7806bbc910e8be8eb6a3c34861a9e68b5800e3fb52f30fbdb67b68059e443822

Observation 9f43cf07-6576-47ee-b798-22ec9e9db6d5 · inbound

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers cites this paper.

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T23:22:31.121825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:22:31.121825Z digest=sha256:b201540917e4534e151217595991549e43100466d2b527ad5668cf3d3bea3b87

Observation cb39eb81-8397-4bdb-985c-5218d34daf3d · inbound

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling cites this paper.

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T09:51:03.345462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:51:03.345462Z digest=sha256:d240a512be58cc76ca61e1d4a9cf6cc59696e666c3399e3d11ade7aa0a378411