Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T00:11:57.763089Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2606.24467.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T00:11:57.763089Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T12:56:45.291959Z
A source-named dated measurement, never combined with another source.
Source: cited_works
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ed88dafd-6b39-4038-90b5-adccec79cdad · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1da52462-0a7a-4175-a718-a484a8f40fc7 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 161d8000-00ce-4020-bce7-9a781805fa75 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 082c550d-f1f2-4e4b-9086-4876c57c6822 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference The Llama 3 Herd of Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ba327b48-32b6-4ba4-a725-b7086922418b · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Qwen2.5 Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 840454fa-54ab-4dfd-8a6c-e695e68eabcb · outbound
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db03baaf-a8e7-4ab6-ad70-a8f34aa04018 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95bb6d88-1f3e-47bd-8e44-4d6174afc6c8 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 176e4426-a23b-4b4d-96b4-7f985a40e75f · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e45f06e1-7dfe-43e0-bd94-2fe77032d648 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59c55722-3d3e-4f8b-b4c8-15fa7c2c5b6e · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57f7f1be-1426-4318-968a-5fbbcaeee64d · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a59fb7dd-2e04-460d-8b4d-4e58d22fb382 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b0c13995-5a76-4ded-9274-2fb9a43db4c8 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae98eed7-e0a9-484e-b8d9-a1fe2adcd90c · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a4aec25-f8a9-4260-be27-fa2d8910f6e0 · outbound
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ca51b5c-94c1-4a51-bab3-4691d41f8f00 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f0c9e18-9979-4692-9235-4dffb5f47504 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Not all heads matter: A head-level KV cache compression method with integrated retrieval and reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cdcca5e3-b201-4475-84b2-8babd76b69a8 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2d926237-75ef-4cc0-9cd3-eaec92c1b960 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d5d0be6-e42c-4357-9e61-444de75be77d · outbound
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33567cdf-0ff6-4191-9970-d799ef3abda9 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81338119-39bd-47fc-b16b-1fce231afc7f · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13e049c5-6e4f-4053-8fa5-69be06fac3b9 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24e356cb-8a28-4b84-8ee9-47151dc07cd3 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference In-context Learning and Induction Heads
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7c08a395-bf1a-49ed-838e-735220d663a0 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc7873b9-9a69-4ffb-8bf2-7ca248dcc20f · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Attention Heads of Large Language Models: A Survey
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a08b16b2-a9a9-4867-a995-00fcfbe4a800 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8729ed96-5f0a-4d9b-a027-454555b9498e · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3443cbb6-d9b8-46e5-923a-70d22af1dae9 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f9276c2-1721-4d88-bea3-3c653ebbe4fe · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Which Attention Heads Matter for In-Context Learning?
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation da331ebc-b840-425a-af55-3348867563e3 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a31b896d-ed4e-4428-8bb3-3714b9e5800b · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 83395ff6-9e1a-478e-ad13-8686005f2d84 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18816ac6-2b77-4c89-9f59-bcd679dc6ff6 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Kamradt, Needleinahaystack, https://github.com/gkamradt/LLMTest_NeedleInAHaystack, 2023
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5787d02d-c279-4a55-9d16-fb257ad763a3 · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Dao, Flashattention-2: Faster attention with better parallelism and work partitioning, in: The Twelfth International Conference on Learning Representations, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f655f36-06c5-4b15-895a-e09ed15f658f · outbound
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38333a1e-4bf2-4285-8410-eb6961e9fada · outbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22bdc162-20b6-4d40-8af9-6cef3be9f150 · inbound
Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.