Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T13:57:15.844310Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2507.19906.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T13:57:15.844310Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7202c61d-e046-47d4-a4a6-dd9616e5842a · outbound
CaliDrop: KV Cache Compression with Calibration GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2676b63f-3f03-4c4b-8fed-6fa6fcb6f796 · outbound
CaliDrop: KV Cache Compression with Calibration Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f6b2c56f-f772-4ace-89cc-1e45982ccf17 · outbound
CaliDrop: KV Cache Compression with Calibration Longbench: A bilingual, multitask benchmark for long context understanding
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 345bb7a4-d061-4786-9478-af7175703411 · outbound
CaliDrop: KV Cache Compression with Calibration Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71be8909-196f-489c-9806-77b6bc72c888 · outbound
CaliDrop: KV Cache Compression with Calibration PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72c9b27c-2c3b-4a5d-87d3-bc691b0d4124 · outbound
CaliDrop: KV Cache Compression with Calibration Palu: Compressing KV-Cache with Low-Rank Projection
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57454bc2-d061-44aa-ae76-7f3e17e17fe3 · outbound
CaliDrop: KV Cache Compression with Calibration SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7806b74f-bbdc-40fc-8703-d73765ce6bdd · outbound
CaliDrop: KV Cache Compression with Calibration A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11d233ba-8593-4977-af05-59f648ad1ee3 · outbound
CaliDrop: KV Cache Compression with Calibration QAQ: Quality Adaptive Quantization for LLM KV Cache
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df3fa1f2-e40d-4e1b-a806-412d58cc275e · outbound
CaliDrop: KV Cache Compression with Calibration The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96d90404-72c7-4a92-a26e-07c29d737a99 · outbound
CaliDrop: KV Cache Compression with Calibration Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf847818-5763-493b-a472-5af7c4a4b078 · outbound
CaliDrop: KV Cache Compression with Calibration Not all heads matter: A head-level kv cache compression method with integrated retrieval and reasoning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f42b0cec-832c-44e0-ba4a-1db6ace0f765 · outbound
CaliDrop: KV Cache Compression with Calibration Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83187b8a-959e-40c0-a79a-35d4d6a8652f · outbound
CaliDrop: KV Cache Compression with Calibration ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffc9ce03-1f66-45f5-a56c-51d090804ebd · outbound
CaliDrop: KV Cache Compression with Calibration KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01f02dc7-00ff-4b0d-9a53-8034531de720 · outbound
CaliDrop: KV Cache Compression with Calibration RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f8d88bd-33bb-4e47-b3c8-d8853acda7ea · outbound
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f92c6351-36f5-467d-8658-d8e1baa35229 · outbound
CaliDrop: KV Cache Compression with Calibration Hydragen: High-Throughput LLM Inference with Shared Prefixes
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5310eae0-5bb4-4afe-ac14-c9176c428768 · outbound
CaliDrop: KV Cache Compression with Calibration GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30bf4c5f-1431-4f2c-b1cd-cc5c284d9e16 · outbound
CaliDrop: KV Cache Compression with Calibration SnapKV: LLM Knows What You are Looking for Before Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3d33bfc-f572-4d21-b961-be136cd0003b · outbound
CaliDrop: KV Cache Compression with Calibration DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb05e49-b401-4fdc-b3a4-2db964c44424 · outbound
CaliDrop: KV Cache Compression with Calibration DeepSeek-V3 Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcdcf966-4e0f-41c1-b5b8-effde72e5654 · outbound
CaliDrop: KV Cache Compression with Calibration Lost in the middle: How language models use long contexts
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3a2c4d2-f9f9-4352-8645-e4e7df5738a0 · outbound
CaliDrop: KV Cache Compression with Calibration Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 365ca0fe-707b-42f2-9793-ab56062917e7 · outbound
CaliDrop: KV Cache Compression with Calibration KIVI: A tuning-free asymmetric 2bit quantization for KV cache
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 597ffd37-bc5a-4e92-844c-e9e60014ebe3 · outbound
CaliDrop: KV Cache Compression with Calibration CAKE: Cascading and adaptive KV cache eviction with layer preferences
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ba1b3470-96d1-46fd-a945-120164147543 · outbound
CaliDrop: KV Cache Compression with Calibration Fast Transformer Decoding: One Write-Head is All You Need
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00e1e89f-cb37-4df4-a152-1064e9263afb · outbound
CaliDrop: KV Cache Compression with Calibration Flexgen: High-throughput generative inference of large language models with a single gpu
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13304490-bca6-4a82-90c0-81c7b3e72b6d · outbound
CaliDrop: KV Cache Compression with Calibration You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfe709df-f236-40e4-8208-1325c1a77790 · outbound
CaliDrop: KV Cache Compression with Calibration Quest: Query-aware sparsity for efficient long-context llm inference
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d9e996b0-3d34-438f-8007-6825fdca0878 · outbound
CaliDrop: KV Cache Compression with Calibration Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0d8488a-6186-46fc-9703-dec6a9253b90 · outbound
CaliDrop: KV Cache Compression with Calibration With Greater Text Comes Greater Necessity: Inference-Time Training Helps Long Text Generation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f55d36b7-9840-4ccd-a473-1fbad5488290 · outbound
CaliDrop: KV Cache Compression with Calibration Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b6f0485-9572-4977-8526-ce7db81918e2 · outbound
CaliDrop: KV Cache Compression with Calibration Layer-Condensed KV Cache for Efficient Inference of Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 475fbee9-8551-418f-8854-3e867a0151ed · outbound
CaliDrop: KV Cache Compression with Calibration Efficient streaming language models with attention sinks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 06bfe2da-1a58-40bf-9640-c9831ed69550 · outbound
CaliDrop: KV Cache Compression with Calibration RefreshKV: Updating Small KV Cache During Long-form Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1280fc6c-c0e7-48ed-90fa-0b80692b91d2 · outbound
CaliDrop: KV Cache Compression with Calibration PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94bf7996-0ad1-4918-bb21-0b53445cdb60 · outbound
CaliDrop: KV Cache Compression with Calibration No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d124d0f-0688-4e16-8aea-961b5b574e41 · outbound
CaliDrop: KV Cache Compression with Calibration Effectively Compress KV Heads for LLM
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aef9934-a407-4c49-8162-428eb3976373 · outbound
CaliDrop: KV Cache Compression with Calibration KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0600c39f-b547-4039-b16f-492813fa9ea3 · outbound
CaliDrop: KV Cache Compression with Calibration DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a025ba81-5030-40a3-aec2-5800b2d8dab2 · outbound
CaliDrop: KV Cache Compression with Calibration Cam: Cache merging for memory-efficient llms inference
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f312100d-bf84-4f50-a921-f44980329ab3 · outbound
CaliDrop: KV Cache Compression with Calibration Barrett, Zhangyang Wang, and Beidi Chen
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 20efee29-b8eb-44cc-a975-33186125f2c7 · outbound
CaliDrop: KV Cache Compression with Calibration DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8161cfc0-3de8-4c17-bc42-c96431256f5e · outbound
CaliDrop: KV Cache Compression with Calibration RelayAttention for Efficient Large Language Model Serving with Long System Prompts
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.