Pith. sign in

Paper Citation Record · LEDGER

You Only Cache Once: Decoder-Decoder Architectures for Language Models

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2405.05254.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.05254 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:59:11.548817Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.070431Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 76089141-4f48-4e51-8561-9c31cf68f0cb · inbound

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation cites this paper.

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.415000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T18:31:35.391674Z digest=sha256:d8ab9356ccc253e3305e61c65d495d597ce5ceeb1eb9344ed4834350c6fa7d8b

Observation 03d9c865-d3ed-4065-b639-a14223410df7 · inbound

Large Language Models Can Self-Improve in Long-context Reasoning cites this paper.

Large Language Models Can Self-Improve in Long-context Reasoning You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T21:59:11.548817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:59:11.548817Z digest=sha256:3c2f398e9fcb295479a23c3734fc575bf7ecc8d99b5edacfa460d1a818598e37

Observation 5b7f8d72-611b-43a4-9635-6f890c34e254 · inbound

Star Attention: Efficient LLM Inference over Long Sequences cites this paper.

Star Attention: Efficient LLM Inference over Long Sequences You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.334337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.334337Z digest=sha256:a4e292727d9e6fd603065569a8b2121818a492da06f8757c161b098c08171a67

Observation f6ace2cf-c3aa-4b23-8867-ff849a4d6a98 · inbound

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning cites this paper.

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:55.284100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:55.284100Z digest=sha256:f968f92a44817ea59823c4e7f07dd5361cb92b67b59f8c8e7bcf47c9b1c2ae72

Observation 46a7c86d-7356-41e5-a6ff-ba62edc91568 · inbound

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference cites this paper.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:18.263580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:18.263580Z digest=sha256:7f200f8a9480097e68c4ad37ece1bbc2173eec4073d892f8da54cda2b030eb97

Observation 0082e2e0-a6fc-43d5-b057-138f82f37d4d · inbound

Gated Delta Networks: Improving Mamba2 with Delta Rule cites this paper.

Gated Delta Networks: Improving Mamba2 with Delta Rule You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 167

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:50:24.320504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T14:50:23.991809Z digest=sha256:c9942c096d571c9b7a78878f081fec3b89c5028ffe10a2d3710a59644e3893c5

Observation 9c290bd5-84e9-4430-ac12-2e7c84472eb9 · inbound

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries cites this paper.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.973242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.973242Z digest=sha256:b939255a1c5c0bb6ea09972a4ac1e4b633732ea208258c8af692908f56dd2d9c

Observation f46d77db-056b-4eba-8efa-464c6fefa035 · inbound

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression cites this paper.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.780704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.780704Z digest=sha256:1f784793f5298a7c79ddfae6589deaa9aca09ce759292bde90642d4b22483360

Observation c692c58b-428a-4649-92b0-2dd0221feab9 · inbound

DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs cites this paper.

DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T11:57:17.968732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:57:17.968732Z digest=sha256:eee0abc1a5139b670745723be8320f4ea92818d7db68a16ac33a8425b80d6341

Observation a4bb8a40-fea6-4c74-b8f3-10187ff911ab · inbound

A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression cites this paper.

A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:29.489118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:29.489118Z digest=sha256:5a2d6ec37cffff24cf88adb7091ebcb37507976b6a4760950774d382d1728554

Observation 050fda77-e494-49c3-b987-3809c52a804e · inbound

Bootstrap Your Own Context Length cites this paper.

Bootstrap Your Own Context Length You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T04:28:39.127780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:28:39.127780Z digest=sha256:51821472a2191dd12f2526c205ae3e8de2ba19c21222d30008fd2feda79b0123

Observation 3b185fd5-546f-49f4-a807-6c99dcf8cd9f · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 191

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:48.698761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:48.698761Z digest=sha256:f3c434d97649c0394211991abd6a2e65d1e3f6dc471549e69f20dcc151ff8864

Observation 39082ba9-aa4c-41be-bb5e-00b5b1b30af0 · inbound

Parallel Key-Value Cache Fusion for Position Invariant RAG cites this paper.

Parallel Key-Value Cache Fusion for Position Invariant RAG You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:44:49.140730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:44:49.140730Z digest=sha256:d33c76b2b49da60d3789877e79792df243ed95e440a3fbdfae20ef39af13bc92

Observation 87415344-5f4d-4e26-837e-257f1bcaa594 · inbound

Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models cites this paper.

Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T15:50:21.978594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:50:21.978594Z digest=sha256:5316a49d403d53181db735b6f87098c976212ec4614bf98f965e463bbd9b1303

Observation e9629a70-76bc-479b-b9ba-0bfaaaeab813 · inbound

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression cites this paper.

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:31.196543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T04:15:36.906263Z digest=sha256:8758eb9cde104b014bbc3fd72315a3a7e771595620e66a316486e4438e8843ba

Observation 7eb8ef10-7404-4e97-b97d-a9d1fd8d5eeb · inbound

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents cites this paper.

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T14:13:38.329919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:13:38.329919Z digest=sha256:0aaf3865cdf67c50838d2d0645b2a3dc74ca8a6fa532d6c576aca4e042fea078

Observation ab040785-c0d6-42d3-94d1-0e5dd2b9d466 · inbound

ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training cites this paper.

ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:22.006655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:51:22.006655Z digest=sha256:fe1757b654c8876f1afac27c198456e36c5cf27bc36b6e93d6bcb605c21f875e

Observation ca85e05f-e502-4809-b022-bace30f01456 · inbound

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding cites this paper.

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:20:27.714146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:20:27.714146Z digest=sha256:eccb22d002c02e49ad4af5b522e66fb010b79a130016e4885f4b1d31e668429c

Observation a5db910f-b2cd-4b5f-be99-8f1a334ccd22 · inbound

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression cites this paper.

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:37.363103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:37.363103Z digest=sha256:2bfd6a166137b43a271539ca1bf75457a0d135a314e1b60e81916d9426f30836

Observation ef02b930-ad70-4fd8-8ef7-3bc981c325c4 · inbound

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs cites this paper.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.234433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.234433Z digest=sha256:d86fbdd402aaf18cab5c32685d10f09adeb4703892365c23613b21e806becb6c

Observation 13304490-bca6-4a82-90c0-81c7b3e72b6d · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.771372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.771372Z digest=sha256:155cdb26b1201dc01567c4a54b02c16832baf267ba763572360948d6db15069a

Observation 6e83b91d-a9ee-496d-9d4f-db0645586b0f · inbound

Kimi Linear: An Expressive, Efficient Attention Architecture cites this paper.

Kimi Linear: An Expressive, Efficient Attention Architecture You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:49:10.824263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T23:49:10.555255Z digest=sha256:d52f4067abc2ccfd5fe85fb8e172cca87e315a69db3a00fc284969e3b36d0efd

Observation a2aa1d70-5671-4611-9f6c-ec2aff9c48e9 · inbound

Block-Based Double Decoders cites this paper.

Block-Based Double Decoders You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:07.458682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T22:05:25.465669Z digest=sha256:b249e0b255262bc75c58d82acfc01fc18760d7b6ba39082c3edae9d1c8be85aa

Observation 3cecfd45-a9ed-41cc-a1b3-4df3fd28b6de · inbound

Block-Based Double Decoders cites this paper.

Block-Based Double Decoders You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:47.083968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T22:10:23.224012Z digest=sha256:3037da57f0d70f3d34e1f903d3990405dda5960902342753ffd2ccc69e43821a

Observation 14cd498d-80d4-44dd-8f87-f02c818b831f · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:16:23.419010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T14:35:48.292081Z digest=sha256:f2ec83bfd70a8ffae341c4dc724e3fb02e57e6d31680552fa20209fa7e89aa85

Observation 4f3f47b6-d345-4f36-9191-c7a620b9f0d3 · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T12:41:23.547953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T12:41:23.547953Z digest=sha256:214ff90386b4f0ed7ef11824617586dfd552ee00d980919627964fd18db85702

Observation 0a9fdbd2-e375-4bad-bf49-76840635598a · inbound

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing cites this paper.

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.824157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T01:06:04.896501Z digest=sha256:39314ba1d5fb2ed37e56bbc9c9412aa5851247610b3264e46916cd22b18c8082

Observation 57cf74bd-3e17-483f-acee-f328c61f57d7 · inbound

Q-Delta: Beyond Key-Value Associative State Evolution cites this paper.

Q-Delta: Beyond Key-Value Associative State Evolution You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:07:26.548538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T18:30:51.523567Z digest=sha256:a4857ef2322a69f7bf77c406e8a8497b23ca5e5709444d97ab18d53d0ad118e4

Observation 93396da0-cd30-4d86-b9da-056eb7a98255 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 111

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.071688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:84bb69e294a7fa2b80103dfd737ef528e9097440ffd7e08c3cd9faa8ac87cd8c

Observation 2348b8b4-20be-4ce9-ac7f-1ae1c7839e96 · inbound

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers cites this paper.

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T23:22:31.614891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:22:31.614891Z digest=sha256:6a56915d49051b88d455f5adb13573e035a5a65f51b3d1372bd32f0fca7f389b

Observation 72683e7f-bb28-4eb5-b906-475de5930321 · inbound

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling cites this paper.

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T09:51:03.340861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:51:03.340861Z digest=sha256:3b1bcaa467d68458fa0525366623df08da211af94f80c90696448cf4dac412b2

Observation eb06958c-9289-4e84-b333-3ac60a66a525 · inbound

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory cites this paper.

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T12:56:45.924478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T12:56:45.924478Z digest=sha256:06050eb1c35e22bba969c8c8d8b847961c2159d87ac7db338879a3a3f74c32f0