Pith. sign in

Paper Citation Record · LEDGER

GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 47 inbound Pith citation observations for arXiv:2403.05527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.05527 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 47 of 47 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:38:47.324395Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.078109Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4c274a29-48b0-4191-a303-036c206199ef · inbound

SGLang: Efficient Execution of Structured Language Model Programs cites this paper.

SGLang: Efficient Execution of Structured Language Model Programs GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.181115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:ccb24ed59d49c8da8f5b92b8a2030f8d09ed02a2e0ca73c3daf1026df331296b

Observation 1c1458c6-6319-4acf-bd58-3c52b8d758d3 · inbound

Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey cites this paper.

Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 183

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:32:36.921300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T11:32:36.738536Z digest=sha256:7db2318cf3c8968a97207f57a4c3c21fb369a39c83db1d9f563b4a5e29e5a57c

Observation 564100b0-98ef-440b-bb33-3416f19c907b · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:47.324395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:47.324395Z digest=sha256:0d18b0218c17964e3d3a04367149e6f04bedb533cc0c7264c1bae4288b7e748a

Observation d80ec495-591c-4b2b-98d8-d3be9eca364f · inbound

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models cites this paper.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.250786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.250786Z digest=sha256:7dd39753672532bf11a51290c11990f38500fbd23119f74ca7d4fb4751262884

Observation b1c3a8f3-05a1-48fe-98ae-f71f82e1b2af · inbound

PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration cites this paper.

PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T18:46:51.073129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:46:51.073129Z digest=sha256:969b920c2decd5fa4aff1418b377c3e7c40ee62fd21e60b11332bb64b5c701a4

Observation e6284fb8-9ee7-4f61-9079-8df0022ccc90 · inbound

PolarQuant: Quantizing KV Caches with Polar Transformation cites this paper.

PolarQuant: Quantizing KV Caches with Polar Transformation GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T13:26:55.388294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:26:55.388294Z digest=sha256:fe25b7cbfe73d4c1e9b6e70516b7a6365be0d13b92d51aaef3fdcd74274b9271

Observation a312f7a9-8693-4c31-a9e8-20d62d99656f · inbound

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference cites this paper.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.224806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.224806Z digest=sha256:4a9cef0b664ea3e8f3f913236c6cf0cd58a748ba28cf52365d2ee1e243d64b6c

Observation 0872bdf5-f05e-4ce2-b24d-f984a37244e5 · inbound

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters cites this paper.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.538291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.538291Z digest=sha256:13e3cc5e169f7ca0e8f116ff3f166239cb04809022df8d4b24efdc599cec614b

Observation 5d3def96-2c07-4025-8407-26f66ab15109 · inbound

Inference-time sparse attention with asymmetric indexing cites this paper.

Inference-time sparse attention with asymmetric indexing GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T05:56:55.130242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:56:55.130242Z digest=sha256:50bc5b210183b2c3d0e8477d5bc0aaed2d5631d9f23db06f82edce9fb9f58099

Observation d31dcdd4-0ac6-4064-8f01-5ef9ccc71d1c · inbound

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache cites this paper.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.935822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.935822Z digest=sha256:b5c5cadbdd21159791a86274e6e22c667f7128b6d1a73670c99bebc7b1c0a039

Observation 3ccfd40f-ed8b-4b96-b152-e504110dea8f · inbound

LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation cites this paper.

LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:12:11.639029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T22:07:37.906916Z digest=sha256:da56426015201d4d546bde38b343522a29b0ea196023655f481cf03685b6579b

Observation 85e7572d-1ce2-4964-828f-7d62707c5a7a · inbound

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate cites this paper.

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:09:22.423835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T08:09:22.226608Z digest=sha256:833c608e620852333ad6e7e6de5e588187e3849b8b4f1c1ca5c7dd8e9855b78a

Observation 704c6af7-e828-4af3-a14c-e0b581fc656f · inbound

Win Fast or Lose Slow: Balancing Speed and Accuracy in Latency-Sensitive Decisions of LLMs cites this paper.

Win Fast or Lose Slow: Balancing Speed and Accuracy in Latency-Sensitive Decisions of LLMs GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:45.550536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:45.550536Z digest=sha256:aa535780e6b80db95c6a78b3e511c31b66466d1691d2e453549ff002825a8227

Observation 7e68ccd2-51c8-47b1-a411-fa6b04f267a7 · inbound

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization cites this paper.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.300614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.300614Z digest=sha256:57db6af50b51a2747c428152ac251e742cd64b2bb517f0685529ab425a0cf4e7

Observation 38eff18d-581c-469e-8f75-d9b88ac6bfaf · inbound

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression cites this paper.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.282281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.282281Z digest=sha256:4c0c14ab55348d6c44854f1190c4f7a1622ec723f2708bb866eb662bdd4beef2

Observation d693ca11-8206-4880-b0da-41d86b5d3cf7 · inbound

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering cites this paper.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.879572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.879572Z digest=sha256:e5894f1fac7ac20ccbb5db29fd46a4f63f5f804c75eebb5736aa6c6fbd11c9d8

Observation 5310eae0-5bb4-4afe-ac14-c9176c428768 · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.727858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.727858Z digest=sha256:da10a67e98be16b9eb37f2fdf1d52b31098f01ac4793c1b51fcea34cca7a490c

Observation 8c0fae0e-b197-4574-816e-8b75130645d1 · inbound

CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation cites this paper.

CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T05:00:04.860070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:00:04.860070Z digest=sha256:955fa9d57189718aeb5f602155828e6ec8af831eda7321180993af4a0f1f8582

Observation 5270e236-184d-421b-947a-e5617649799a · inbound

LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind cites this paper.

LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:20:42.120063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T22:16:07.638869Z digest=sha256:8895527edb698fb7426c6fb3a1b0ed3d5dce70f8d601bf1d609aa0851376d3fd

Observation e588fc88-901c-4328-a74b-94078c24fee2 · inbound

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference cites this paper.

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:13.082094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:16:13.082094Z digest=sha256:4534d412e8c327f1d51a08dba10b8582c5ab85fd46c3b65001228b85416c4dca

Observation 17518d99-034c-4515-9726-e2b0bf2e9a12 · inbound

Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization cites this paper.

Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:00:44.636266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T07:58:24.456859Z digest=sha256:4a0c74bdbded4bd1605546b1fe23b9c2e43ab6b54cb7c1540c35a345058666f9

Observation 26180e94-c9d8-4f74-b93b-983fce023816 · inbound

AdaHOP: Fast and Accurate Low-Precision Training via Outlier-Pattern-Aware Rotation cites this paper.

AdaHOP: Fast and Accurate Low-Precision Training via Outlier-Pattern-Aware Rotation GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:15.737296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:51:39.641700Z digest=sha256:a8b685499566c39d4d76633bd58a3383a6d598ee2842c2b4a602565bc53a6315

Observation ec38aa1a-be53-4a07-baac-d56e0703140e · inbound

Quantization Dominates Rank Reduction for KV-Cache Compression cites this paper.

Quantization Dominates Rank Reduction for KV-Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:04.560722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T15:43:48.986845Z digest=sha256:d76ace3f8876c53b9254432ad09a54a816dc9876b3a3b68d7caa09169543f2d9

Observation e33bd435-4711-41ab-9309-1743ef658143 · inbound

Open-TQ-Metal: Fused Compressed-Domain Attention for Long-Context LLM Inference on Apple Silicon cites this paper.

Open-TQ-Metal: Fused Compressed-Domain Attention for Long-Context LLM Inference on Apple Silicon GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:52:14.192689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T07:18:26.936652Z digest=sha256:62620aa7c4813fbfa8b55d409afdce07c09aa5658041685f5254b49487163f75

Observation 8196466f-24c9-400c-b563-0b79058b682c · inbound

SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving cites this paper.

SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:14:08.289804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T03:09:51.453839Z digest=sha256:54af30de7e73a93710cd4bb3fa7ead165e1ee13df13b2dbf6b1429d6838b75b4

Observation 853184f9-946d-43a3-97c6-c185c304cb92 · inbound

eOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and Quantization cites this paper.

eOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:50.797121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T20:18:04.392331Z digest=sha256:08799d355ce748772eedeb90f7040fae448b4b69075f718982863d81317bd4d1

Observation 797c33ad-141e-49f6-8239-64969751bc78 · inbound

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization cites this paper.

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:29.722007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T17:07:28.163304Z digest=sha256:2758a2ea55acaff75abcb6631cf6cec0876c95c98c46ce776611dedd4d1a6228

Observation df78a131-79cd-47f3-88b3-b5c660288077 · inbound

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization cites this paper.

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:42.525075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T18:33:19.999707Z digest=sha256:98fea2155e0f1f0e162a742ae1e67f4b73dd0bb308d4a80612defcbc637209b9

Observation 441450c7-0f1c-42e7-9131-5233b47e043f · inbound

FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression cites this paper.

FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.506719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:42:45.446358Z digest=sha256:cfd19a8ef1bcb45bdc84463bd3e149a1decbefb463b6193d8f1e90dcbe94be0a

Observation 13fa2f57-062c-406d-be67-dd173170a9b3 · inbound

Search Your Block Floating Point Scales! cites this paper.

Search Your Block Floating Point Scales! GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:24.096602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T05:52:26.558984Z digest=sha256:f9b2ee8e3868ea6585bae3a555bdbeb5db47258c48306a84eae5b8bb849ebdfc

Observation c1a85a81-d71c-4f62-829f-0e32b90b9bd7 · inbound

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization cites this paper.

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:18:16.457796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:16:07.797702Z digest=sha256:667965fcacec9962f1716b8a511eeb9702bb7a9da46a7d4c8cea3c5fb14d778f

Observation ca781ce1-a450-4219-a64e-c587edc7e076 · inbound

SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference cites this paper.

SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:33:43.288306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T20:30:33.208386Z digest=sha256:fc874616d90f75df223b74965fd59168f9a048405cba136e19b713e267758c5d

Observation 2a04797a-b0e8-4365-b0dc-8c53e5360920 · inbound

SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference cites this paper.

SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.844560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T21:27:19.945653Z digest=sha256:a85124865d3b3cb10b5e76385c6337541c7b53272f9a03e1d472efc16dd93c4a

Observation f7c3e792-1262-4343-8413-d6b058d1550a · inbound

A Simple Plug-in for Improving Eviction-Based KV Cache Compression cites this paper.

A Simple Plug-in for Improving Eviction-Based KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:55:24.195044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T04:51:04.068354Z digest=sha256:5654536485008b08aaf80fe2ea8ef0d811951e511a7f3289bf225d74c316b298

Observation 9303a006-593e-45b9-a8f5-a8099fd3cf29 · inbound

SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference cites this paper.

SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:32:34.748513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T19:31:28.053090Z digest=sha256:a36860e7331a02c2e88e75129000305e6078367002e389ef426837805eb0672c

Observation b0381d89-0d97-4b00-8633-f384ba79211f · inbound

Cartridges at Scale: Training Modular KV Caches over Large Document Collections cites this paper.

Cartridges at Scale: Training Modular KV Caches over Large Document Collections GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:16:48.165901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T06:11:18.942446Z digest=sha256:d9167e876a12befe1d8671ca4f9d97784677caecb629aa00cbdf40594c5a5df6

Observation 50e46d3f-fb28-403a-8037-c3b8c563c7ef · inbound

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling cites this paper.

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:27:24.340384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T19:46:43.514413Z digest=sha256:485e918ca6851f4336c3c8578c7e36c2aeec7bf3b5531859d521033326d64018

Observation 17ab5027-5332-4216-9e0e-d8146a8dcd68 · inbound

IntentKV: Cross-Turn Intent-Aware KV Cache Pruning for Agent Inference cites this paper.

IntentKV: Cross-Turn Intent-Aware KV Cache Pruning for Agent Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:22.957997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T20:04:44.703782Z digest=sha256:accfb1a1569e1cceae1f161bacebdc17098f604bdf6631eeefaa68b9db75d4a6

Observation 2783ee1a-32f5-46ff-a527-f7e05aed2797 · inbound

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse cites this paper.

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:09:48.918409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T07:13:19.593093Z digest=sha256:692bafa8bc2e88ddad59aa0727a4d96359e2a6f66fb4110834ff3f1109f8ba2c

Observation 6e81c4dd-1205-4d79-a1a9-75556b9e8042 · inbound

RoPE-Aware Bit Allocation for KV-Cache Quantization cites this paper.

RoPE-Aware Bit Allocation for KV-Cache Quantization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:59:57.355432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T01:05:13.790381Z digest=sha256:5dd4337e805d49c80a9afcd2058957952df179e32516b242214fd519275d2b87

Observation 95bb6d88-1f3e-47bd-8e44-4d6174afc6c8 · inbound

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference cites this paper.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:58.047166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:85396a5cb49c0fcbab2e085ec735697aa2e9d956e6a4d142d8c0e07a7f8f4b6b

Observation 0ed7ecab-b701-4395-bc41-f5a5b61cc369 · inbound

MosaicKV: Serving Long-Context LLM with Dynamic Two-D KV Cache Compression cites this paper.

MosaicKV: Serving Long-Context LLM with Dynamic Two-D KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:08.402573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T16:12:12.043127Z digest=sha256:ab65bf8669600e708f02fc83aeb9584b90d31489ef6de47a3fb347fd8f170695

Observation 7cdd29e4-e91d-495a-807c-9983e269bd1a · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:bd366638e450c8a1ea3be4f8a147ff82403829e6f823061a9ff866e93fbe4307

Observation afa6ae33-80fd-418f-97e1-0fb83778e755 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.079355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:9b6ca6251162ef79bdae4278c70ab3692873c6b3bd91288b30fdb9b9bb561f66

Observation c5c2f566-6573-44f2-83fe-7b5390964c35 · inbound

Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization cites this paper.

Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:56:40.965166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:55:52.215700Z digest=sha256:5fdfea7839ae279187c16b1edba23d37aa2f88ca92c9dcd1520bdd0651e9e549

Observation b22813b0-b0d5-4967-9874-23a3fff43484 · inbound

Practical Online KV Cache Compaction for LLM Agents: An Empirical Study cites this paper.

Practical Online KV Cache Compaction for LLM Agents: An Empirical Study GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:48:33.732453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:48:33.732453Z digest=sha256:f7ee821f649b8aa189881e5a7a1e9edcaae451dbf2934bf258af996d876f7387

Observation 9f2f53e8-1e16-4af1-b737-44aa270e9204 · inbound

Runtime Observability for Heterogeneous Attention Memory cites this paper.

Runtime Observability for Heterogeneous Attention Memory GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T22:11:58.745454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:11:58.745454Z digest=sha256:60d4c995b09a6c213477ee730343476f645a1ade20ea3ec3e11b1751099156b5