Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 47 inbound Pith citation observations for arXiv:2403.05527.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:38:47.324395Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T01:36:44.078109Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 4c274a29-48b0-4191-a303-036c206199ef · inbound
SGLang: Efficient Execution of Structured Language Model Programs GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1c1458c6-6319-4acf-bd58-3c52b8d758d3 · inbound
Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 183
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 564100b0-98ef-440b-bb33-3416f19c907b · inbound
A Survey on Large Language Model Acceleration based on KV Cache Management GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d80ec495-591c-4b2b-98d8-d3be9eca364f · inbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1c3a8f3-05a1-48fe-98ae-f71f82e1b2af · inbound
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6284fb8-9ee7-4f61-9079-8df0022ccc90 · inbound
PolarQuant: Quantizing KV Caches with Polar Transformation GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a312f7a9-8693-4c31-a9e8-20d62d99656f · inbound
HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0872bdf5-f05e-4ce2-b24d-f984a37244e5 · inbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d3def96-2c07-4025-8407-26f66ab15109 · inbound
Inference-time sparse attention with asymmetric indexing GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d31dcdd4-0ac6-4064-8f01-5ef9ccc71d1c · inbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ccfd40f-ed8b-4b96-b152-e504110dea8f · inbound
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 85e7572d-1ce2-4964-828f-7d62707c5a7a · inbound
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 704c6af7-e828-4af3-a14c-e0b581fc656f · inbound
Win Fast or Lose Slow: Balancing Speed and Accuracy in Latency-Sensitive Decisions of LLMs GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e68ccd2-51c8-47b1-a411-fa6b04f267a7 · inbound
TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38eff18d-581c-469e-8f75-d9b88ac6bfaf · inbound
Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d693ca11-8206-4880-b0da-41d86b5d3cf7 · inbound
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5310eae0-5bb4-4afe-ac14-c9176c428768 · inbound
CaliDrop: KV Cache Compression with Calibration GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c0fae0e-b197-4574-816e-8b75130645d1 · inbound
CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5270e236-184d-421b-947a-e5617649799a · inbound
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e588fc88-901c-4328-a74b-94078c24fee2 · inbound
PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17518d99-034c-4515-9726-e2b0bf2e9a12 · inbound
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 26180e94-c9d8-4f74-b93b-983fce023816 · inbound
AdaHOP: Fast and Accurate Low-Precision Training via Outlier-Pattern-Aware Rotation GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ec38aa1a-be53-4a07-baac-d56e0703140e · inbound
Quantization Dominates Rank Reduction for KV-Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e33bd435-4711-41ab-9309-1743ef658143 · inbound
Open-TQ-Metal: Fused Compressed-Domain Attention for Long-Context LLM Inference on Apple Silicon GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8196466f-24c9-400c-b563-0b79058b682c · inbound
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 853184f9-946d-43a3-97c6-c185c304cb92 · inbound
eOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 797c33ad-141e-49f6-8239-64969751bc78 · inbound
HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation df78a131-79cd-47f3-88b3-b5c660288077 · inbound
HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 441450c7-0f1c-42e7-9131-5233b47e043f · inbound
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 13fa2f57-062c-406d-be67-dd173170a9b3 · inbound
Search Your Block Floating Point Scales! GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c1a85a81-d71c-4f62-829f-0e32b90b9bd7 · inbound
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ca781ce1-a450-4219-a64e-c587edc7e076 · inbound
SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2a04797a-b0e8-4365-b0dc-8c53e5360920 · inbound
SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f7c3e792-1262-4343-8413-d6b058d1550a · inbound
A Simple Plug-in for Improving Eviction-Based KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9303a006-593e-45b9-a8f5-a8099fd3cf29 · inbound
SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b0381d89-0d97-4b00-8633-f384ba79211f · inbound
Cartridges at Scale: Training Modular KV Caches over Large Document Collections GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 50e46d3f-fb28-403a-8037-c3b8c563c7ef · inbound
HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 17ab5027-5332-4216-9e0e-d8146a8dcd68 · inbound
IntentKV: Cross-Turn Intent-Aware KV Cache Pruning for Agent Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2783ee1a-32f5-46ff-a527-f7e05aed2797 · inbound
Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6e81c4dd-1205-4d79-a1a9-75556b9e8042 · inbound
RoPE-Aware Bit Allocation for KV-Cache Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 95bb6d88-1f3e-47bd-8e44-4d6174afc6c8 · inbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0ed7ecab-b701-4395-bc41-f5a5b61cc369 · inbound
MosaicKV: Serving Long-Context LLM with Dynamic Two-D KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7cdd29e4-e91d-495a-807c-9983e269bd1a · inbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afa6ae33-80fd-418f-97e1-0fb83778e755 · inbound
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c5c2f566-6573-44f2-83fe-7b5390964c35 · inbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b22813b0-b0d5-4967-9874-23a3fff43484 · inbound
Practical Online KV Cache Compaction for LLM Agents: An Empirical Study GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f2f53e8-1e16-4af1-b737-44aa270e9204 · inbound
Runtime Observability for Heterogeneous Attention Memory GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.