Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:37:07.697300Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 3 inbound Pith citation observations for arXiv:2411.18077.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:37:07.697300Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:38:47.483457Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T20:59:01.556666Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e7ad4eaf-d1d5-4d47-a479-334cd7017ad5 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5af67f7f-5cea-4d68-bef3-4751fcbf1add · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 592c3a29-2acc-4656-848d-50800257f37c · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d456d26-f44c-45d2-870e-b89f54a6f0c5 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03199717-1dcd-425a-9376-503a74745195 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 199d81b7-0328-4935-ada3-998fede18f89 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8433fc1-ed3c-4e87-9e95-4dd9418229ce · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32394cf1-ad8c-4554-8e12-3cdc54a95116 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cf2ff58-02f5-4c4c-91b4-2c10cc73130d · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6ce3c48-7d7e-4ca0-840a-3a7eb8132d43 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Mistral 7B
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95b0495f-cc05-4fd9-b454-a6d9643c4add · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 904d9b76-48d9-40e5-8e95-fd183426b773 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache SnapKV: LLM Knows What You are Looking for Before Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adbad6c2-b9d1-468a-b410-838668088066 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39ffc95a-dca4-42ef-829d-d6bc784359bd · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f7676d9-5061-4d50-a2a4-ea539a3a263e · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Multi-head or Single-head? An Empirical Comparison for Transformer Training
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1734be7-5d10-467c-9a4d-755149e86b5a · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ae1c708-e961-4a3a-b887-930d580ce0d4 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1f830ab9-1b00-4ec0-95da-5b50c75c7476 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e45af902-9bf3-44bc-bf9b-cd7f903b4e55 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8b8127c-4618-4611-bc2b-9dbea5525753 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7057606f-fac1-4ad3-aeb0-15f972b7e1ae · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57c805ba-a686-4df7-9460-27b12a7190c3 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3fa7a7c7-6b28-4d72-bb1f-86ea6ca9b1c5 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd394d09-30f7-4f6c-bfe9-99f30fa42aff · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 80ae3419-7cf7-48bd-860e-22fbb5e45893 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3788dc32-33fe-4100-9964-5c88c6e2a9c1 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 97a6cc9b-743d-4e2d-8581-1a9877b43c1f · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Do Large Language Model Benchmarks Test Reliability?
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e18dec0-c176-4357-b125-3a597f77447c · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 04249263-2fab-41bd-a950-dabe643a7754 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c5ca8fff-152a-451d-aea9-2bebe4f1518c · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6564002-5de1-46ec-94b7-1698ddbae20a · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Retrieval Head Mechanistically Explains Long-Context Factuality
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd0916af-31f0-4431-ac51-2fc67cc52b04 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 20e63b59-4d5b-4e35-bafc-1cbcebf0873c · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c22d63e-a693-4633-bfd8-eb0c8e03e13b · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Efficient Streaming Language Models with Attention Sinks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ced9945-c15e-4e04-939c-711b8dbb8902 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d88fdc-0c2f-4b57-87f3-7a4dfaadb303 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7267b29a-8ed7-424e-b943-25a38825d5d3 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9340b39b-65a6-4abf-adff-3f0066468f2c · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 248e51a3-32a2-4a0c-b655-ab701a21bfa8 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Barrett, Zhangyang Wang, and Beidi Chen
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93f5e1f6-0042-405c-9339-19fa85f70aa3 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache online" 'onlinestring :=
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b66d035f-d673-4e4c-92c0-6329e84a19c7 · outbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache write newline
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1956160-20f3-4901-b2c6-5b51d48b213a · inbound
A Survey on Large Language Model Acceleration based on KV Cache Management MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c752e517-becb-4267-8201-5ca35c10a045 · inbound
Minimal-Intervention KV Retention via Set-Conditioned Diversity MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bbcf9321-28d6-400b-946f-c5dddb42a72e · inbound
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.