Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2405.05254.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T21:59:11.548817Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T01:36:44.070431Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 76089141-4f48-4e51-8561-9c31cf68f0cb · inbound
LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 03d9c865-d3ed-4065-b639-a14223410df7 · inbound
Large Language Models Can Self-Improve in Long-context Reasoning You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b7f8d72-611b-43a4-9635-6f890c34e254 · inbound
Star Attention: Efficient LLM Inference over Long Sequences You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6ace2cf-c3aa-4b23-8867-ff849a4d6a98 · inbound
CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46a7c86d-7356-41e5-a6ff-ba62edc91568 · inbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0082e2e0-a6fc-43d5-b057-138f82f37d4d · inbound
Gated Delta Networks: Improving Mamba2 with Delta Rule You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 167
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9c290bd5-84e9-4430-ac12-2e7c84472eb9 · inbound
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f46d77db-056b-4eba-8efa-464c6fefa035 · inbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c692c58b-428a-4649-92b0-2dd0221feab9 · inbound
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4bb8a40-fea6-4c74-b8f3-10187ff911ab · inbound
A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 050fda77-e494-49c3-b987-3809c52a804e · inbound
Bootstrap Your Own Context Length You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b185fd5-546f-49f4-a807-6c99dcf8cd9f · inbound
A Survey on Large Language Model Acceleration based on KV Cache Management You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 191
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39082ba9-aa4c-41be-bb5e-00b5b1b30af0 · inbound
Parallel Key-Value Cache Fusion for Position Invariant RAG You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87415344-5f4d-4e26-837e-257f1bcaa594 · inbound
Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9629a70-76bc-479b-b9ba-0bfaaaeab813 · inbound
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7eb8ef10-7404-4e97-b97d-a9d1fd8d5eeb · inbound
Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab040785-c0d6-42d3-94d1-0e5dd2b9d466 · inbound
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca85e05f-e502-4809-b022-bace30f01456 · inbound
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5db910f-b2cd-4b5f-be99-8f1a334ccd22 · inbound
DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef02b930-ad70-4fd8-8ef7-3bc981c325c4 · inbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13304490-bca6-4a82-90c0-81c7b3e72b6d · inbound
CaliDrop: KV Cache Compression with Calibration You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e83b91d-a9ee-496d-9d4f-db0645586b0f · inbound
Kimi Linear: An Expressive, Efficient Attention Architecture You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a2aa1d70-5671-4611-9f6c-ec2aff9c48e9 · inbound
Block-Based Double Decoders You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3cecfd45-a9ed-41cc-a1b3-4df3fd28b6de · inbound
Block-Based Double Decoders You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 14cd498d-80d4-44dd-8f87-f02c818b831f · inbound
Do Value Vectors in Deep Layers Need Context from the Residual Stream? You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4f3f47b6-d345-4f36-9191-c7a620b9f0d3 · inbound
Do Value Vectors in Deep Layers Need Context from the Residual Stream? You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a9fdbd2-e375-4bad-bf49-76840635598a · inbound
You Only Index Once: Cross-Layer Sparse Attention with Shared Routing You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 57cf74bd-3e17-483f-acee-f328c61f57d7 · inbound
Q-Delta: Beyond Key-Value Associative State Evolution You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 93396da0-cd30-4d86-b9da-056eb7a98255 · inbound
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2348b8b4-20be-4ce9-ac7f-1ae1c7839e96 · inbound
Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72683e7f-bb28-4eb5-b906-475de5930321 · inbound
SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb06958c-9289-4e84-b333-3ac60a66a525 · inbound
Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.