Pith. sign in

Paper Citation Record · LEDGER

Layer-Condensed KV Cache for Efficient Inference of Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2405.10637.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.10637 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:57:15.793887Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T23:16:23.532286Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8facd824-f426-4216-a2f3-95d1e0c5a236 · inbound

When Attention Sink Emerges in Language Models: An Empirical View cites this paper.

When Attention Sink Emerges in Language Models: An Empirical View Layer-Condensed KV Cache for Efficient Inference of Large Language Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:41:03.764666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T17:41:03.674759Z digest=sha256:db7a830a86c41e0a9e89154515fcd692c702332c77572787fa3eb4fd18e808ed

Observation c79d223c-ceef-4c26-8a6d-3c70c232ac82 · inbound

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression cites this paper.

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression Layer-Condensed KV Cache for Efficient Inference of Large Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:31.312626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T04:15:36.906263Z digest=sha256:38c0b9052a857e45a107f2ccddb72edcc8a6371fa32c974e6dfc4340df74c3c8

Observation 7b6f0485-9572-4977-8526-ce7db81918e2 · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration Layer-Condensed KV Cache for Efficient Inference of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.793887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.793887Z digest=sha256:f27e54fecd3e7049337e8d2561726baefd02922ece278c52fe297fa3e6aec927

Observation 60be5ed8-16ce-4a7e-941f-1cd3df3493bf · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Layer-Condensed KV Cache for Efficient Inference of Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:09.408451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:09.408451Z digest=sha256:2fc34c9d6429eb79b1e712cea2960a031060c5607a307436f7324f992f7263e6

Observation ec10e0a1-1afa-4399-bb3f-7967d322121b · inbound

CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs cites this paper.

CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs Layer-Condensed KV Cache for Efficient Inference of Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:14:00.165021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T04:13:13.231293Z digest=sha256:fac4c56aac57c38640a5b0ec15d71ab7b909ca51afe0e08908ff80633f408c3a

Observation 78a08397-e062-44bb-aad8-398ec5914b0f · inbound

ClusterFusion++: Expanding Cluster-Level Fusion to Full Transformer-Block Decoding cites this paper.

ClusterFusion++: Expanding Cluster-Level Fusion to Full Transformer-Block Decoding Layer-Condensed KV Cache for Efficient Inference of Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:26:17.311258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T05:34:08.596351Z digest=sha256:9edaad0e1088c47b12b06000e1bc0adca4aa4f1e44d3663ea60d133ae69383d8

Observation 66c5a0d8-90cb-4d5a-b40c-6265b75558e4 · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? Layer-Condensed KV Cache for Efficient Inference of Large Language Models

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:16:23.534411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T14:35:48.292081Z digest=sha256:89687219448f170447cd10a422d12ab52831d1e76ea916ea154adf8f6b8d0db5

Observation a5ab00fb-e0d4-41e5-a9d3-70dc5eccf3e1 · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? Layer-Condensed KV Cache for Efficient Inference of Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T12:41:23.956019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T12:41:23.956019Z digest=sha256:d4c7ac9618cc4ac2e769673e82359bd76ec39e449501ba34dc2528bebb01a015

Observation ccc29096-3f45-49e3-85f1-84b97ec275aa · inbound

Structured Thoughts For Improved Reasoning And Context Pruning cites this paper.

Structured Thoughts For Improved Reasoning And Context Pruning Layer-Condensed KV Cache for Efficient Inference of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T12:08:05.502310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T12:08:05.502310Z digest=sha256:7d4f8088bd0217627c26cffc18193481bff7e09c4c7725b5e49640ab5ee12611

Observation dc2c4eee-5961-41b9-800a-1342369c9d8b · inbound

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory cites this paper.

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory Layer-Condensed KV Cache for Efficient Inference of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T12:56:46.292098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T12:56:46.292098Z digest=sha256:4e807172e15957d9011986f83c13124d8786a41c7214413a2bba93b9e3533596