Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2405.12981.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T17:22:46.999943Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation c00b63c8-c9a6-4544-978d-b68ed14a0472 · inbound
LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ec66c560-a6b2-4182-a552-453dd9e2a2e0 · inbound
PoM: Efficient Image and Video Generation with the Polynomial Mixer Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baa5f6ec-1d50-4f0b-9471-2fe7120524d9 · inbound
Hymba: A Hybrid-head Architecture for Small Language Models Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24e3e211-2a95-4ad7-ae8a-add9f701ff2f · inbound
CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 592c3a29-2acc-4656-848d-50800257f37c · inbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87ff2fb3-099b-47c4-8c1a-56e45b31dc06 · inbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0785f33-eacf-4f05-ab0a-3c3e6e476661 · inbound
UniForm: A Reuse Attention Mechanism Optimized for Efficient Vision Transformers on Edge Devices Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3572522b-9059-4164-93e1-a936006ebe83 · inbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5c07440-d1ed-432c-a602-7b2ea45b4c59 · inbound
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 357356ce-20ee-44fa-b58b-cab024cf6828 · inbound
A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6d12d0e-1a79-4b0e-8f87-f82d9fe68267 · inbound
Multi-matrix Factorization Attention Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc63f34-a742-404f-b05d-86d8dcaa63a1 · inbound
A Survey on Large Language Model Acceleration based on KV Cache Management Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 205
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0982b95-36e1-4cb1-8f07-7655903b600f · inbound
Foundations of Large Language Models Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a8512ca-d088-4f2a-b192-7d6577b5871b · inbound
Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60eeb619-40af-4ecb-9862-e62d815b8235 · inbound
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1db8d200-bea1-4583-a4f8-07dca902121d · inbound
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d4937893-d0a2-45b9-8c82-11ad5be1b371 · inbound
Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30d4b816-ebab-44fb-b757-ee0633875c68 · inbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bbe0a22-fc85-4ac7-81a7-82c15fcbdcf7 · inbound
MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd97bfa7-81dd-4579-b112-b259b2a85479 · inbound
Hardware-Efficient Attention for Fast Decoding Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6a8c350-03bf-426b-891a-a5d23d81a1fd · inbound
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 345bb7a4-d061-4786-9478-af7175703411 · inbound
CaliDrop: KV Cache Compression with Calibration Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d5022ab-4df1-4ddb-b02e-098ef4e89e67 · inbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0afe267d-674f-4f95-8d75-5f91593e0c76 · inbound
Stochastic KV Routing: Enabling Adaptive Depth-Wise Cache Sharing Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 75f6d43b-16fc-4490-bdf2-28c20b0612d6 · inbound
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3673d03d-8509-48e2-be5b-673b600ca091 · inbound
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3a38211c-62ab-4de3-a429-0a1f48999f38 · inbound
Do Value Vectors in Deep Layers Need Context from the Residual Stream? Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 306faa8c-84e0-49b9-808f-58f778b2b301 · inbound
Do Value Vectors in Deep Layers Need Context from the Residual Stream? Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94bdc96c-9f3f-447c-be23-74ce6c3bbf63 · inbound
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 347be885-1d25-4ac8-91f9-9c10bb0815f5 · inbound
SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.