Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T17:54:40.241144Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2607.01237.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T17:54:40.241144Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 80d8136a-b205-4969-8543-214a27830729 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fec53c5-7593-415f-8f71-ab3edcf77ef2 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Chain-of-thought prompting elicits reasoning in large language models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e18fe63-e85f-441f-9519-2d6d0825cebf · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Qwen3 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b339e05-75aa-4724-ab84-c3047cb81653 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression CRANE: Reasoning with constrained LLM generation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d897440-0980-48d5-b75b-aac80246a55d · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression A Survey on Large Language Model Acceleration based on KV Cache Management
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 335dbaa4-29f2-468a-a406-8b2e36a447af · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression LLM Inference Unveiled: Survey and Roofline Model Insights
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d93f3c0a-8555-47ff-9b4b-af02c91d07da · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression R-KV: Redundancy-aware KV cache compression for reasoning models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2a3af9d-b160-416e-868d-1f48e120884f · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Ada-KV: Optimizing KV cache eviction by adaptive budget allocation for efficient LLM inference
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed5bee5c-3c90-41b8-8291-083765bb628b · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression SnapKV: LLM knows what you are looking for before generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63c4c100-f793-4e9f-af04-8c88ccf4e9cc · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression KV cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3345f60f-9dc5-4f5d-b6ab-8de28f5de93c · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Efficient streaming language models with attention sinks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e478993-cd24-4aeb-8541-ed3bd8c5c5a2 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5116eaeb-bf7c-4fb6-a92e-60434d4dc6db · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Inference-time hyper-scaling with KV cache compression
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15b06e5f-d426-4701-94ad-eeb15ebfaef3 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Lee, Sangdoo Yun, and Hyun Oh Song
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82ceb0a5-06db-4e0d-ac3e-37b2dc3541f5 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Expected attention: Kv cache compres- sion by estimating attention from future queries distribution.arXiv preprint arXiv:2510.00636, 2025
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbbab5dc-b0d4-4e5d-b221-7b9f76b6d62d · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Gonzalez, Hao Zhang, and Ion Stoica
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bfc3dce-0fec-4c44-8810-353f7a78a7c8 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Criticbench: Benchmarking llms for critique-correct reasoning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d8bc600-dc96-494d-a2ee-bfcfe26830aa · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression The Llama 3 Herd of Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb4f415-7e74-4c7d-a47f-539e634a7b93 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression ChunkKV: Semantic-preserving KV cache compression for efficient long-context LLM inference
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e30c151d-eb63-4b81-ac4c-b5860959317a · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Efficient many- shot in-context learning with dynamic block-sparse attention
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9e6893b-1cfb-4058-aa25-7a2a7c835586 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression ThinKV: Thought-adaptive KV cache compression for efficient reasoning models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 750fa320-e1a8-478d-8bcc-4bb9ce9bc88a · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression QuoKA: Query-oriented KV selection for efficient LLM prefill
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79c718dc-7f54-45a4-9361-af01cab2e35c · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Icecache: Memory-efficient KV-cache management for long-sequence LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5a24779-c8b7-4b78-82f7-9a24bd5cb16e · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Cache what lasts: Token retention for memory-bounded KV cache in LLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 999036f9-037c-43e6-8e96-5112f799346d · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Keydiff: Key similarity-based KV cache eviction for long-context LLM inference in resource-constrained environments
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e383c10-4093-4d2f-8011-0cbb4b6de5a3 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Spargeattention: Accurate and training-free sparse attention accelerating any model inference
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 677ab63c-0793-4954-88c1-c2af34493308 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression MoBA: Mixture of block attention for long-context LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0eeae7f-99e4-49bf-9d91-04e44b6701e9 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Twilight: Adaptive attention sparsity with hierarchical top-$p$ pruning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 118e17f8-1133-4b77-82f2-038bbc8cab5a · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Xattention: Block sparse attention with antidiagonal scoring
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee3d6372-d41e-40d7-9753-1b6458120aa0 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Efficient attention mechanisms for large language models: A survey.arXiv preprint arXiv:2507.19595, 2025
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76cad784-ca42-43f8-9b26-23cd6d04ee87 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Deepseek-v4: Towards highly efficient million-token context intelligence, 2026
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 028e36d6-ed91-4bc3-bfb8-bbc1187b1239 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Kwai Summary Attention Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b312c64b-aaae-4a59-baf1-cb57768ac2b7 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Where does in-context learning \\ happen in large language models? InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10ac2138-1f68-4805-a52e-61ab2333a2f0 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression GPT-4 Technical Report
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d48e8749-ea35-4244-969d-1251a3e85bb6 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression A survey on large language model acceleration based on KV cache management.Transactions on Machine Learning Research, 2025
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7cc8594-c288-4b44-9823-d51196e2b317 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Gonzalez, Clark Barrett, and Ying Sheng
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd657f61-0eb9-4b77-b823-fa4fdc01e8bc · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d94c1bf0-fc2c-4c66-9078-7f105845093d · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression CAKE: Cascading and adaptive KV cache eviction with layer preferences
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c5d13d9-9675-4c0a-b820-c72e1fc49d62 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression American invitational mathematics examination (aime) 2024, 2024
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a37132e-3a93-4810-8d07-c4668696b1d5 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2edd8f80-9b6d-4e67-86f7-278ed53eb693 · outbound
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Data engineering for scaling language models to 128k context
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.