Pith. sign in

Paper Citation Record · LEDGER

GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 59 inbound Pith citation observations for arXiv:2403.05527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.05527 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 59 of 59 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:05:38.937281Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.078109Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4c274a29-48b0-4191-a303-036c206199ef · inbound

SGLang: Efficient Execution of Structured Language Model Programs cites this paper.

SGLang: Efficient Execution of Structured Language Model Programs GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.181115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:4e03a60c3eb3d01b252fcc4cdd21038b5b293a07cf3da19bb23eec1a62a33c19

Observation 1c1458c6-6319-4acf-bd58-3c52b8d758d3 · inbound

Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey cites this paper.

Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 183

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:32:36.921300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T11:32:36.738536Z digest=sha256:d53f76ff623168e773e7376413f2a2f17ec57febb6e68eb958a017775ad7e259

Observation 787d8d4b-0a63-4247-99d0-0fe6d7293239 · inbound

Multi-Bin Batching for Increasing LLM Inference Throughput cites this paper.

Multi-Bin Batching for Increasing LLM Inference Throughput GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:56:23.192461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:56:23.192461Z digest=sha256:88fb3a4188e9800eb4a6b2f55bf7f072a76d25571b861adf1d74145fd86cae71

Observation 01665fc2-6ef2-4415-bf6a-6d14cac1887e · inbound

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference cites this paper.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:17.868031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:17.868031Z digest=sha256:047558d14944121b61113d1d0b230ef2543e1d6f15aeb082c625427f6b5610f4

Observation 4bfc02b8-2457-4587-bdeb-4c0b0f90e1c1 · inbound

Low-Rank Correction for Quantized LLMs cites this paper.

Low-Rank Correction for Quantized LLMs GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:26.804829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:26.804829Z digest=sha256:8413305d1c51810208cc83fcc600066fcc4c6dd23e15436342def277a26e5ecc

Observation 44398d71-fd32-41e3-987c-c6fb764c9fef · inbound

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries cites this paper.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.949962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.949962Z digest=sha256:1c2dcf7bbbca20379c0231a65af4e7a01ef4f18dfe7cb039869092089f8d2184

Observation e6f96a3d-b421-445a-84a6-69172711e76a · inbound

CSR:Achieving 1 Bit Key-Value Cache via Sparse Representation cites this paper.

CSR:Achieving 1 Bit Key-Value Cache via Sparse Representation GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:53.743169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:53.743169Z digest=sha256:ca1ec6342282f4b2142a5e1a994b2c05d57d4b6294aa1601380f00fa83f89aff

Observation ae71be05-5aad-457e-b094-8edc2005b616 · inbound

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression cites this paper.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.712247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.712247Z digest=sha256:15c74db8204d64c5a4340c93c7ecc9005518a4b77ad8dc02f862230050a2bd35

Observation 8a50b5f7-d468-4e8b-bfbc-e2ed5fb7edd7 · inbound

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals cites this paper.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.849655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.849655Z digest=sha256:ddd63292697e4be693cb82c4d93abbf02397607737350a8cc16805577ed27f11

Observation 8108889f-c38f-4a3f-b0bb-daabff8c9a4b · inbound

DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs cites this paper.

DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:57:17.934982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:57:17.934982Z digest=sha256:1e8c8367d9244398a86ed348e11a3e1c62ce5dd77d6b2228b36079cd096ebe2a

Observation 564100b0-98ef-440b-bb33-3416f19c907b · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:47.324395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:47.324395Z digest=sha256:37e9219c51e388be37b2e9d90651e581b9e623310b1e454db73aa75351fa6a2b

Observation 2377e117-9098-4867-96d4-78956e29de2a · inbound

Efficiently Serving Large Multimodal Models Using EPD Disaggregation cites this paper.

Efficiently Serving Large Multimodal Models Using EPD Disaggregation GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:29:34.728199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:29:34.728199Z digest=sha256:e5ce8f206d903a77a1546a122e932abd9924a2e407bc6b6d982cc7af43353cef

Observation d80ec495-591c-4b2b-98d8-d3be9eca364f · inbound

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models cites this paper.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.250786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.250786Z digest=sha256:ea07b074fe114a9dbe5522a665e1324afdfea3241c22786fa4bfc698a50afa9f

Observation b1c3a8f3-05a1-48fe-98ae-f71f82e1b2af · inbound

PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration cites this paper.

PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T18:46:51.073129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:46:51.073129Z digest=sha256:d8cddcce49c119946088f7bfbaacd1850503867d6e0c1705e1050192cf710ce0

Observation e6284fb8-9ee7-4f61-9079-8df0022ccc90 · inbound

PolarQuant: Quantizing KV Caches with Polar Transformation cites this paper.

PolarQuant: Quantizing KV Caches with Polar Transformation GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T13:26:55.388294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:26:55.388294Z digest=sha256:14f03fc5e8dbd45e6055c799c9e3acfc18201901113353a86139f9a589664d1e

Observation a312f7a9-8693-4c31-a9e8-20d62d99656f · inbound

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference cites this paper.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.224806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.224806Z digest=sha256:5b412fbb28bc999c8678f00cea09cd278da61bff8ff5a1576a02662f426f04d1

Observation 0872bdf5-f05e-4ce2-b24d-f984a37244e5 · inbound

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters cites this paper.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.538291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.538291Z digest=sha256:3d5d05247330f62c8161c5b53a7c9102a83faef738ed16b316341db51e3a6cbf

Observation 5d3def96-2c07-4025-8407-26f66ab15109 · inbound

Inference-time sparse attention with asymmetric indexing cites this paper.

Inference-time sparse attention with asymmetric indexing GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T05:56:55.130242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:56:55.130242Z digest=sha256:d2afd82dfa55994f2f4200e8deb7a7116755bd704ecac18ab8fa2f767bd0ab56

Observation d31dcdd4-0ac6-4064-8f01-5ef9ccc71d1c · inbound

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache cites this paper.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.935822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.935822Z digest=sha256:497feb241fa7b6385920a163df5a8e8a46edae4834825eb9fe238537fc40ce5e

Observation 3ccfd40f-ed8b-4b96-b152-e504110dea8f · inbound

LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation cites this paper.

LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:12:11.639029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T22:07:37.906916Z digest=sha256:2ff0817881f3d40d04d54b793e3f42c3c2eae5c127e389170e1e9e878cd44579

Observation 85e7572d-1ce2-4964-828f-7d62707c5a7a · inbound

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate cites this paper.

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:09:22.423835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T08:09:22.226608Z digest=sha256:b3b495968bac13ae571b6b96255b92cc849e76647c748397977974fd423f2475

Observation 314a7d1b-a983-4a50-a744-b0f41a356f60 · inbound

Accurate KV Cache Quantization with Outlier Tokens Tracing cites this paper.

Accurate KV Cache Quantization with Outlier Tokens Tracing GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:38.937281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:05:38.937281Z digest=sha256:ca4156f1d01a288f5b3473028031905c71c7706443da812df390c5521b2065b6

Observation 704c6af7-e828-4af3-a14c-e0b581fc656f · inbound

Win Fast or Lose Slow: Balancing Speed and Accuracy in Latency-Sensitive Decisions of LLMs cites this paper.

Win Fast or Lose Slow: Balancing Speed and Accuracy in Latency-Sensitive Decisions of LLMs GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:45.550536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:45.550536Z digest=sha256:1f98e168a33a1cb4850a8ce4eb608841fe522dd8aace228f157aac42dc95f34a

Observation 7e68ccd2-51c8-47b1-a411-fa6b04f267a7 · inbound

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization cites this paper.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.300614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.300614Z digest=sha256:ffa20d79134c987aaf3e6788a770ece18511db93501a080c8ba482b1dad94ddf

Observation 38eff18d-581c-469e-8f75-d9b88ac6bfaf · inbound

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression cites this paper.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.282281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.282281Z digest=sha256:5ef5a81b1fdd3be86e2a09c5ec7f6419123c0a7f785004586e8f047b148ced71

Observation d693ca11-8206-4880-b0da-41d86b5d3cf7 · inbound

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering cites this paper.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.879572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.879572Z digest=sha256:67dad49cb8daa4e76bbbdf01566ea65881bf84e2eab957513132d2d916d1108e

Observation 5310eae0-5bb4-4afe-ac14-c9176c428768 · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.727858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.727858Z digest=sha256:b437522f974c81a9caf31b22d26e1080a4a0a4e5fa329ad8497fd3faf01079d9

Observation 8c0fae0e-b197-4574-816e-8b75130645d1 · inbound

CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation cites this paper.

CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T05:00:04.860070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:00:04.860070Z digest=sha256:5f787a5511d426d7eaf34f96b9e91cfebec7b94161131ea4807dfa5d60cf0298

Observation 5270e236-184d-421b-947a-e5617649799a · inbound

LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind cites this paper.

LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:20:42.120063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T22:16:07.638869Z digest=sha256:f5854c64d1ff5cee116c925947417e51e293403cf493c9925f9d785ab7560d96

Observation e588fc88-901c-4328-a74b-94078c24fee2 · inbound

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference cites this paper.

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:13.082094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:16:13.082094Z digest=sha256:85602299eff7dcfcb745ed171b516879012daf5bf78bf01bcf23b8d8abd37be1

Observation 17518d99-034c-4515-9726-e2b0bf2e9a12 · inbound

Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization cites this paper.

Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:00:44.636266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T07:58:24.456859Z digest=sha256:9f1d0935f01661179a73580b796c152cacf5f06efefd4eedcce9456bca52e714

Observation 26180e94-c9d8-4f74-b93b-983fce023816 · inbound

AdaHOP: Fast and Accurate Low-Precision Training via Outlier-Pattern-Aware Rotation cites this paper.

AdaHOP: Fast and Accurate Low-Precision Training via Outlier-Pattern-Aware Rotation GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:15.737296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T20:51:39.641700Z digest=sha256:40e500cb8f7b019b798e9ab3ad379df5c7e1abc41d1b56ba060e9f6b30ed7093

Observation ec38aa1a-be53-4a07-baac-d56e0703140e · inbound

Quantization Dominates Rank Reduction for KV-Cache Compression cites this paper.

Quantization Dominates Rank Reduction for KV-Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:04.560722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T15:43:48.986845Z digest=sha256:61ca534dfbead25c29177608d4b4c703b40f70bb05833148aa96051351c94a4b

Observation e33bd435-4711-41ab-9309-1743ef658143 · inbound

Open-TQ-Metal: Fused Compressed-Domain Attention for Long-Context LLM Inference on Apple Silicon cites this paper.

Open-TQ-Metal: Fused Compressed-Domain Attention for Long-Context LLM Inference on Apple Silicon GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:52:14.192689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T07:18:26.936652Z digest=sha256:dc66692935d43522a2f67414a7c95c8b789cc617dd9ad7bb3d4e34d7cc69d5ad

Observation 8196466f-24c9-400c-b563-0b79058b682c · inbound

SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving cites this paper.

SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:14:08.289804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T03:09:51.453839Z digest=sha256:0598ef2a94e86ac4334ab268e9d515ef5b26f03e759cc112d43044f7aab1c30a

Observation 853184f9-946d-43a3-97c6-c185c304cb92 · inbound

eOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and Quantization cites this paper.

eOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:50.797121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T20:18:04.392331Z digest=sha256:8262ad3996462ad1be43ab0bf62877d72eb3b608e6190dce2f81c32b00fca260

Observation 797c33ad-141e-49f6-8239-64969751bc78 · inbound

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization cites this paper.

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:29.722007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T17:07:28.163304Z digest=sha256:7c0aee7e4e493d4f06ccdf942a71b01a0cafaf7199bb04bed66193486e0b18d1

Observation df78a131-79cd-47f3-88b3-b5c660288077 · inbound

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization cites this paper.

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:42.525075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T18:33:19.999707Z digest=sha256:8137d80ad41ac17a48bb9d2b9b49727589389a9f97a59f7d66a177b93a953c43

Observation 441450c7-0f1c-42e7-9131-5233b47e043f · inbound

FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression cites this paper.

FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.506719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T01:42:45.446358Z digest=sha256:1a097fe5e21c1934fb62ed792f03393d70a67b5a345253353cf0b8223aa14c30

Observation 13fa2f57-062c-406d-be67-dd173170a9b3 · inbound

Search Your Block Floating Point Scales! cites this paper.

Search Your Block Floating Point Scales! GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:24.096602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T05:52:26.558984Z digest=sha256:d59a73a0810161b5fe11bba6d4f25d5849c98c02eb7ebe29c641ee846d141a58

Observation c1a85a81-d71c-4f62-829f-0e32b90b9bd7 · inbound

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization cites this paper.

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:18:16.457796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:16:07.797702Z digest=sha256:663fb27d707250a928cf87bc0ddb29f5357307ae5d5f9181f601f08723305fba

Observation ca781ce1-a450-4219-a64e-c587edc7e076 · inbound

SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference cites this paper.

SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:33:43.288306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T20:30:33.208386Z digest=sha256:5995a301ed6db45669422b3809355d26900c2fdd7cb59c9dccdc0cbd616d24ef

Observation 2a04797a-b0e8-4365-b0dc-8c53e5360920 · inbound

SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference cites this paper.

SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.844560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T21:27:19.945653Z digest=sha256:0174679923d11ecfdb78f366980aa5bdff5fe35ee76c34c3005a3cf4eb49a509

Observation f7c3e792-1262-4343-8413-d6b058d1550a · inbound

A Simple Plug-in for Improving Eviction-Based KV Cache Compression cites this paper.

A Simple Plug-in for Improving Eviction-Based KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:55:24.195044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-25T04:51:04.068354Z digest=sha256:240ce7ab8978161f983d606e23c8fd2483551e3c347c1f0d8e9459dacfe9ba14

Observation 9303a006-593e-45b9-a8f5-a8099fd3cf29 · inbound

SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference cites this paper.

SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:32:34.748513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T19:31:28.053090Z digest=sha256:d0750615e092323389df8cfc08da96909dae315c4a5157c1fb4e871aa76aebd8

Observation b0381d89-0d97-4b00-8633-f384ba79211f · inbound

Cartridges at Scale: Training Modular KV Caches over Large Document Collections cites this paper.

Cartridges at Scale: Training Modular KV Caches over Large Document Collections GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:16:48.165901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T06:11:18.942446Z digest=sha256:e3e6368a8f748dc199e2fabf9f3b486b3cb379d8ad79ddd28fc21e198045e59c

Observation 50e46d3f-fb28-403a-8037-c3b8c563c7ef · inbound

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling cites this paper.

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:27:24.340384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T19:46:43.514413Z digest=sha256:9777ea8c182995a471edb638502bac23191c2f5a4ca6927ed70e4d640b85e7e2

Observation 17ab5027-5332-4216-9e0e-d8146a8dcd68 · inbound

IntentKV: Cross-Turn Intent-Aware KV Cache Pruning for Agent Inference cites this paper.

IntentKV: Cross-Turn Intent-Aware KV Cache Pruning for Agent Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:22.957997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T20:04:44.703782Z digest=sha256:af288361a2fd6a9af084e05bf1e3a404676e7bc08349521ac066a96526bcf1ab

Observation 2783ee1a-32f5-46ff-a527-f7e05aed2797 · inbound

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse cites this paper.

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:09:48.918409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T07:13:19.593093Z digest=sha256:8e3bce58ff691a6cc3d34b4fe3abd7951e2ca71575af41f4845a289795724d66

Observation 6e81c4dd-1205-4d79-a1a9-75556b9e8042 · inbound

RoPE-Aware Bit Allocation for KV-Cache Quantization cites this paper.

RoPE-Aware Bit Allocation for KV-Cache Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:59:57.355432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T01:05:13.790381Z digest=sha256:c0178143fb463d48bce3f785cecc8d493c860fb64859e2083f741c72d2fe32bd

Observation 95bb6d88-1f3e-47bd-8e44-4d6174afc6c8 · inbound

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference cites this paper.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:58.047166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:c3b8e156acaf3a201e583e0c36383a67ff61d689dce12fc41899e328e133c277

Observation 0ed7ecab-b701-4395-bc41-f5a5b61cc369 · inbound

MosaicKV: Serving Long-Context LLM with Dynamic Two-D KV Cache Compression cites this paper.

MosaicKV: Serving Long-Context LLM with Dynamic Two-D KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:08.402573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-02T16:12:12.043127Z digest=sha256:4da262373d6f060ec5260c5c23e971bff77e9168b8388c755d4401dc1067702e

Observation 7cdd29e4-e91d-495a-807c-9983e269bd1a · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:9d563a13a6630858c006a034e84e2804cc6bfe184f2411774a25f4b0139fded3

Observation afa6ae33-80fd-418f-97e1-0fb83778e755 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.079355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:220433d8dfe37ae60ebb5c6457755130d6ba179374169c07fe778997f83c67d4

Observation c5c2f566-6573-44f2-83fe-7b5390964c35 · inbound

Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization cites this paper.

Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:56:40.965166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T00:55:52.215700Z digest=sha256:f98dfb34fa8a4edf635e7383e735ea6183001ef56b81a9892dd1b65d938bb5ba

Observation b22813b0-b0d5-4967-9874-23a3fff43484 · inbound

Practical Online KV Cache Compaction for LLM Agents: An Empirical Study cites this paper.

Practical Online KV Cache Compaction for LLM Agents: An Empirical Study GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:48:33.732453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:48:33.732453Z digest=sha256:2619a7dc0b43d18b0dc9e573ed9424cac77093e271cd64941bf2cd29e3b599cd

Observation 23bd3a26-47ec-430c-8306-8dc3e62b3e96 · inbound

AnchorKV: Anchor-Residual KV Cache Compression cites this paper.

AnchorKV: Anchor-Residual KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:52.889978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:02:52.889978Z digest=sha256:372caab27c51760b21a4ab9730fb453d356d7ab7542a9c4980207455b2b1efe1

Observation 9f2f53e8-1e16-4af1-b737-44aa270e9204 · inbound

Runtime Observability for Heterogeneous Attention Memory cites this paper.

Runtime Observability for Heterogeneous Attention Memory GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T22:11:58.745454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:11:58.745454Z digest=sha256:727e837d9ef9a34e69ba8e533215279add5dd802128be675e62153a18e7db917

Observation 71aa1f3d-9747-494d-ae24-a2759075741e · inbound

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents cites this paper.

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:51.138707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:50:51.138707Z digest=sha256:9262e6b2923b296b32da9ecd0820e312145f4d197118a4e08bc7d66354bf878b