Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2401.18079.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T17:10:08.496622Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T06:15:00.866473Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 49358396-5a24-42bf-a26e-8110a97c2a08 · inbound
ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f0ce2c05-fd48-4cd7-8700-7480987707f1 · inbound
SGLang: Efficient Execution of Structured Language Model Programs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1d83d1f5-4133-4531-ba96-1b93e15fbeb0 · inbound
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8d622d52-b7d1-44e9-8a56-00fc999812ca · inbound
A Survey on Efficient Inference for Large Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 219
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b6d19c84-6bc0-4145-8471-c17e03e40217 · inbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b8ca3e16-9a01-4171-917a-ef22f8c98e96 · inbound
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7b6d9095-11a4-4ee2-aac3-bd77a67b6b77 · inbound
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 983a32bc-107b-44c9-a3e0-41e1c4621837 · inbound
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 15760c83-790d-4e7a-b65e-b4d4e8dcae70 · inbound
Token Sample Complexity of Attention KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 1963
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8e2c0b5-458f-4baf-b766-d3d418fc5841 · inbound
SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d4f3d5-68a8-4f87-9c2e-6f24aedf0a2d · inbound
Sequential KV Cache Compression via Probabilistic Language Tries: Beyond the Per-Vector Shannon Limit KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation cde0251b-0d4b-4c31-bb72-819a7691390c · inbound
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c7fc36aa-0ad4-44f9-93ab-ccb34322ac62 · inbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d57de22e-4e01-405f-8866-348250b34b03 · inbound
HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0cd80202-ee07-453b-b303-5406d11e96aa · inbound
HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6413a979-1151-4d0a-b089-2bb3db893e42 · inbound
VORT: Adaptive Power-Law Memory for NLP Transformers KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e94c01d3-5644-4303-8353-e0656cc4ef32 · inbound
Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f06e10bb-9277-48af-bb23-f84e9d816afa · inbound
SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation cf1e2c40-e430-4817-aeb3-4d4d392feff0 · inbound
Runtime-Certified Bounded-Error Quantized Attention KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4bfb4e11-c9c5-404c-906c-7a5c0666f1d5 · inbound
Adaptive Mass-Segmented KV Compression for Long-Context Reasoning KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e1012cfb-6667-488f-bf37-162f4e42edaf · inbound
Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e4e72868-c583-4321-a84e-da0be5a57f9e · inbound
Do Value Vectors in Deep Layers Need Context from the Residual Stream? KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2c5b01e8-7ed2-4d15-b55c-7efbd6e447c3 · inbound
Do Value Vectors in Deep Layers Need Context from the Residual Stream? KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66f98172-5a50-4385-9606-d3916ec2c9c9 · inbound
STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank Control KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0d00ab5b-bcb2-46ef-b136-993c3c1a8dbf · inbound
From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 19ab3f4d-16ff-4062-b6f0-4af4d02abc83 · inbound
Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2bd10106-03ec-442b-b763-c59f9b220815 · inbound
GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8e5f9032-d7cd-4df2-aa58-ef1b88649739 · inbound
Edge-Deployable LLM Fine-Tuning on a Single GPU for Telecom Network Troubleshooting KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 318bfd26-f9f5-4a53-a10a-d232fb84b2f9 · inbound
Fractal KV-Cache Archives: Lossless Symbolic Storage with In-Place Retrieval for Long-Context LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 21137c33-95f5-443c-af0c-65450a67d571 · inbound
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 079a02e3-8f49-4470-85f5-7c3e64a2be09 · inbound
Stateful Worlds, Stateless Elasticity: Exact-State Serving for Interactive World Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 370400af-5b24-4d45-9b3e-a0fe5069a658 · inbound
Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c399909b-00ee-449b-aa0e-cdb766908015 · inbound
Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad5553b7-0e5a-4377-b8c2-422f584943ce · inbound
High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27aab2f4-0d4c-4593-9b3e-55914237ada8 · inbound
SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ac5d3aa-ed7d-4947-bbca-b783de3c136c · inbound
Where Facts Go Missing: A Layerwise Taxonomy and Per-Layer Attribution of Information Omission in Air-Gapped LLMAgent Pipelines KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 270ca459-9ac7-4564-81ca-0a1b40cfb920 · inbound
A Photonic-CXL Memory Appliance for Scalable KV Cache Management in LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c113f63-5339-43b9-95e9-a9b797220e11 · inbound
WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.