Pith. sign in

Paper Citation Record · LEDGER

Effectively Compress KV Heads for LLM

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2406.07056.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.07056 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:32:34.524396Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:12:34.646797Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ca559c28-8d52-4068-a408-0d456a2caf2c · inbound

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation cites this paper.

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation Effectively Compress KV Heads for LLM

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.338559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T18:31:35.391674Z digest=sha256:85810ea39014d27c9a99462c7bb54d137d145caa18c8b15448d0ccdd968074f8

Observation d5467938-fa24-4ea2-884c-4e003de5fc59 · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding Effectively Compress KV Heads for LLM

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:34.524396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:34.524396Z digest=sha256:0b287dfe9bda9e06c1fb60da3a180bea211585e8d7ddb71ab14c03b81c071c3c

Observation 0f93231f-9814-4dc5-85c4-ecb43a2b7e43 · inbound

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering cites this paper.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Effectively Compress KV Heads for LLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.947823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.947823Z digest=sha256:545cbd5780486bfa3d2d9206c9e24d148fd993c13b208ec29d9307ba1ba72777

Observation 1ce2cfb8-7c12-45f3-b120-633a5b4dde42 · inbound

Cartridges: Lightweight and general-purpose long context representations via self-study cites this paper.

Cartridges: Lightweight and general-purpose long context representations via self-study Effectively Compress KV Heads for LLM

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:35.180125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:35.180125Z digest=sha256:03c6eb3f786cbd566be75e93351bee33a4e624a268ce0bdacf5942a83cb00c1e

Observation 59e4af24-f7c0-4ad3-8846-51f8fbccd7e4 · inbound

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding cites this paper.

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding Effectively Compress KV Heads for LLM

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:20:28.754354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:20:28.754354Z digest=sha256:fd518aad5fa633d2e2a9ec9920c42af50ff4349db0d23271f7dad93b71746621

Observation 2ccd4bf6-e3fa-41e6-823b-ffa68218497f · inbound

LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models cites this paper.

LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models Effectively Compress KV Heads for LLM

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:33:21.617755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:33:21.617755Z digest=sha256:5307878821c85b95620fa9bd177059cfd32562be9ecf341bb889a9eb8489425a

Observation 3d124d0f-0688-4e16-8aea-961b5b574e41 · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration Effectively Compress KV Heads for LLM

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.817194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.817194Z digest=sha256:cf1e5faec8571de79db0f765552c94be4f9f29ff2ec9b955c93e1157d79cbd56

Observation 129d056b-aa76-4449-8ef0-fd01fe67112e · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Effectively Compress KV Heads for LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:09.928919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:09.928919Z digest=sha256:6a0255618a91e927e51b424077df181c8817f8628fb25322ae75ec1fc3678150

Observation 6d4fd0d3-0437-4586-b14c-22365ba354aa · inbound

WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models cites this paper.

WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models Effectively Compress KV Heads for LLM

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:18.045245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T21:05:09.254086Z digest=sha256:41883a47109b94ea7941468a44d1f6748a18f5a887c9c17cf405305bb7bea691

Observation 88fd4ed8-96dd-4c18-b2c2-ce95482be7a3 · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation Effectively Compress KV Heads for LLM

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:02:34.497786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:69d64ced34ca024f34da25fbbc111bd8d18fee4f054160f42c2d7cac9e6e8263

Observation adf25fc3-b10d-4806-a842-42e91bdda4b0 · inbound

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models cites this paper.

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models Effectively Compress KV Heads for LLM

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.648373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T19:09:29.347284Z digest=sha256:a71229586c5c7dba389466dece13e726835ec95ee036a60f1dfd9ceaacf37893

Observation 906ecff6-6869-4d8a-9b1c-114a7c9ddbb5 · inbound

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling cites this paper.

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling Effectively Compress KV Heads for LLM

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T09:51:03.388419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:51:03.388419Z digest=sha256:637dff644be60a2a0466f50d89f56004f7fe7c6d7d47fae95e94793cecf0ea93