Pith. sign in

Paper Citation Record · LEDGER

KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2401.18079.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.18079 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:10:08.496622Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 49358396-5a24-42bf-a26e-8110a97c2a08 · inbound

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models cites this paper.

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T13:49:33.837276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T13:49:33.747672Z digest=sha256:43f4876daf349de39a9258b273652f08422835cf59a3a88bb2b7573debdda5e9

Observation f0ce2c05-fd48-4cd7-8700-7480987707f1 · inbound

SGLang: Efficient Execution of Structured Language Model Programs cites this paper.

SGLang: Efficient Execution of Structured Language Model Programs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.106533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:66570e16d0c9f2d7ea625a4167c433d9e298dc285a4ff37a806df214a8d98d6f

Observation 1d83d1f5-4133-4531-ba96-1b93e15fbeb0 · inbound

KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache cites this paper.

KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:53:12.320304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T08:53:12.253243Z digest=sha256:834366c3d805c3639304c8c1f80c29353456ca179ee0598a157d99c5a679b16c

Observation 8d622d52-b7d1-44e9-8a56-00fc999812ca · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 219

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:39:33.280738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:04fcee0d175f7971102e48f66a7a33e1432a3b5a8af37751bccf07ca2e245bf9

Observation b6d19c84-6bc0-4145-8471-c17e03e40217 · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:26.480168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:5ee61abe63d560e54fdae8f3a5e885b93a26295394d00cd5b78e2a10484ba3b0

Observation b8ca3e16-9a01-4171-917a-ef22f8c98e96 · inbound

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision cites this paper.

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:45:36.452789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T19:45:36.337956Z digest=sha256:ee36e2a19add621d72fe7b69f439942b4cf36aba54b94f90e14845afc5466ed6

Observation 7b6d9095-11a4-4ee2-aac3-bd77a67b6b77 · inbound

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate cites this paper.

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:09:22.397439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T08:09:22.226608Z digest=sha256:a167983a6b3bc2a9dc42f4dc169ab4e87f16e1eb9388f8e18387f1b63204ea36

Observation 983a32bc-107b-44c9-a3e0-41e1c4621837 · inbound

EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments cites this paper.

EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T21:54:22.821265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-21T21:54:05.442769Z digest=sha256:bb007df2e018aaf3a3816754d0c8394127c92e64307448cc1d4de2ce48f38418

Observation 15760c83-790d-4e7a-b65e-b4d4e8dcae70 · inbound

Token Sample Complexity of Attention cites this paper.

Token Sample Complexity of Attention KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 1963

Resolution
unresolved
no resolver link, observed 2026-08-03T17:10:08.496622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:10:08.496622Z digest=sha256:86bbc01d5e129f1018b66e44ba53aa7bc0d143581a5424e7a5ffc64ff63fca53

Observation d8e2c0b5-458f-4baf-b766-d3d418fc5841 · inbound

SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training cites this paper.

SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T20:43:37.256381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:43:37.256381Z digest=sha256:16e269e09b9b340925f7c666cc85fe058f5e865144bdadcdf0d9c8a0d9865ae5

Observation d5d4f3d5-68a8-4f87-9c2e-6f24aedf0a2d · inbound

Sequential KV Cache Compression via Probabilistic Language Tries: Beyond the Per-Vector Shannon Limit cites this paper.

Sequential KV Cache Compression via Probabilistic Language Tries: Beyond the Per-Vector Shannon Limit KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:05.119893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:46:14.866351Z digest=sha256:7e551039a3cc8349b49b051d10b894e3bbd9fa0479f76fdc92b35b29e2945c48

Observation cde0251b-0d4b-4c31-bb72-819a7691390c · inbound

PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference cites this paper.

PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:34.356303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T03:55:32.585495Z digest=sha256:72783389b251e8ea58cbe786b857656b8a60f84e0cb0e6b6c0bbb5c8b8eef344

Observation c7fc36aa-0ad4-44f9-93ab-ccb34322ac62 · inbound

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization cites this paper.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.002040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:931749e730bd581622954d48553dbc0e48bd3f01d711550e43ab264b69f7b315

Observation d57de22e-4e01-405f-8866-348250b34b03 · inbound

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization cites this paper.

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:29.776641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T17:07:28.163304Z digest=sha256:2412f97409fdb23398041cb498f59c3ac2e8cb3d4449896a516e9a8a4794a9a8

Observation 0cd80202-ee07-453b-b303-5406d11e96aa · inbound

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization cites this paper.

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:42.515117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T18:33:19.999707Z digest=sha256:4257a6da565a1835c9ab031d181675ee4d189f8ff00ce7c5cfada25932fd5510

Observation 6413a979-1151-4d0a-b089-2bb3db893e42 · inbound

VORT: Adaptive Power-Law Memory for NLP Transformers cites this paper.

VORT: Adaptive Power-Law Memory for NLP Transformers KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:26.259266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T02:01:53.000938Z digest=sha256:64a7054b89743447c10a1e610c2696ab8af4c9a5072631f5200fa4dbd3f0b0fd

Observation e94c01d3-5644-4303-8353-e0656cc4ef32 · inbound

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction cites this paper.

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:13:15.987749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T12:13:14.295659Z digest=sha256:ff6e601eea027aa81bd3d606d30b377324e9a64acf059ada1ac7d362d89adccf

Observation f06e10bb-9277-48af-bb23-f84e9d816afa · inbound

SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference cites this paper.

SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:33:43.345093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T20:30:33.208386Z digest=sha256:cf519eff12449d7b337d110e423894f7c96bb548290820502117ab5099d63210

Observation cf1e2c40-e430-4817-aeb3-4d4d392feff0 · inbound

Runtime-Certified Bounded-Error Quantized Attention cites this paper.

Runtime-Certified Bounded-Error Quantized Attention KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:49:41.109858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T05:45:42.295527Z digest=sha256:3a13917eae18b3146ab26687f32b6b74fed420f65a1f180ae6abc9fd5848b55f

Observation 4bfb4e11-c9c5-404c-906c-7a5c0666f1d5 · inbound

Adaptive Mass-Segmented KV Compression for Long-Context Reasoning cites this paper.

Adaptive Mass-Segmented KV Compression for Long-Context Reasoning KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:25:23.488548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-25T05:22:47.198185Z digest=sha256:ef342739f80e455dbe12d4fb56edaaab063ba704ddf381204efdcc1b721a8a5d

Observation e1012cfb-6667-488f-bf37-162f4e42edaf · inbound

Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization cites this paper.

Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:14:39.072739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T12:13:18.805668Z digest=sha256:8414f3b2e8b5824ef0a2c5f4ac105dab34b6479cdf31b8146f413123beadd5f2

Observation e4e72868-c583-4321-a84e-da0be5a57f9e · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:16:23.448886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T14:35:48.292081Z digest=sha256:0e3a8b8f4e097c304fc19eedcedfe1053b1ebe74949b2829023be96ee4981ce0

Observation 2c5b01e8-7ed2-4d15-b55c-7efbd6e447c3 · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T12:41:20.552739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T12:41:20.552739Z digest=sha256:7d58a33d04414055cb7ef9a68d4ec2aac9f1fa75edce15f790156eba26202ebf

Observation 66f98172-5a50-4385-9606-d3916ec2c9c9 · inbound

STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank Control cites this paper.

STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank Control KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:57:26.297219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T18:34:20.645677Z digest=sha256:5e58bbd6f7473150f85e7aa6cbc09f00c3855fe9eee792432f7555b6f20c3987

Observation 0d00ab5b-bcb2-46ef-b136-993c3c1a8dbf · inbound

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs cites this paper.

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:01:07.733642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-27T16:57:02.709967Z digest=sha256:bab750b7fa2f089a1c5b10646257392dc2f48bf62b0e2e14e68619a2eee3f3c0

Observation 19ab3f4d-16ff-4062-b6f0-4af4d02abc83 · inbound

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation cites this paper.

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:16.033365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T15:37:34.129339Z digest=sha256:9cf62ae5984e86d560172e6c438c175a8560d3fa4118fb8c2ea335dd7b92b214

Observation 2bd10106-03ec-442b-b763-c59f9b220815 · inbound

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache cites this paper.

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:57:06.412465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-07-02T15:55:40.177742Z digest=sha256:c96717ac995b754463a20de1602d1cb3737b83f8a57a3545393e952950c71377

Observation 8e5f9032-d7cd-4df2-aa58-ef1b88649739 · inbound

Edge-Deployable LLM Fine-Tuning on a Single GPU for Telecom Network Troubleshooting cites this paper.

Edge-Deployable LLM Fine-Tuning on a Single GPU for Telecom Network Troubleshooting KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T17:31:15.972423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:31:15.972423Z digest=sha256:aef07d059ad160302b15ea77fad3b73678cd3410da158e77cd983fdfa1f0ef6a

Observation 318bfd26-f9f5-4a53-a10a-d232fb84b2f9 · inbound

Fractal KV-Cache Archives: Lossless Symbolic Storage with In-Place Retrieval for Long-Context LLM Inference cites this paper.

Fractal KV-Cache Archives: Lossless Symbolic Storage with In-Place Retrieval for Long-Context LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-09T19:16:27.748453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-09T19:13:15.765652Z digest=sha256:9e6934a8b08afb57ca402efef5487cf3e8dfb1d1d2b47bf812d7563fdafa9685

Observation 21137c33-95f5-443c-af0c-65450a67d571 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.040808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:fe87a8cba902f6f667ca43bca4564c7d35ec7bc1a9b58be54fe8bebebf60a899

Observation 079a02e3-8f49-4470-85f5-7c3e64a2be09 · inbound

Stateful Worlds, Stateless Elasticity: Exact-State Serving for Interactive World Models cites this paper.

Stateful Worlds, Stateless Elasticity: Exact-State Serving for Interactive World Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T12:07:13.700787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:07:13.700787Z digest=sha256:dcdd2da40cafb6c6331dbb91680b0a97cbd56ed6cf4a93640a6bc0524f2a3d30

Observation 370400af-5b24-4d45-9b3e-a0fe5069a658 · inbound

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems cites this paper.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.433775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.433775Z digest=sha256:7a5149ddbe574374454ed55356efdb7d0463b3170bee088e11195e083ad8c626

Observation c399909b-00ee-449b-aa0e-cdb766908015 · inbound

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers cites this paper.

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T23:22:30.957456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:22:30.957456Z digest=sha256:05b8d5d6ee3243e235d596b39372b96e88a4378f3f62f16bc7a1deaec199e84a

Observation ad5553b7-0e5a-4377-b8c2-422f584943ce · inbound

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration cites this paper.

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T09:50:32.327649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:50:32.327649Z digest=sha256:686cc243c88a4c1c85f17ac294e150e6c861ffe4998eb48f3895f4d0f7a4df16

Observation 27aab2f4-0d4c-4593-9b3e-55914237ada8 · inbound

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling cites this paper.

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T09:51:03.359796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:51:03.359796Z digest=sha256:6658dc71f7050bd9577fcc572af7751d85c1dd6727d1d0d72a2a147c7f2b22e0

Observation 6ac5d3aa-ed7d-4947-bbca-b783de3c136c · inbound

Where Facts Go Missing: A Layerwise Taxonomy and Per-Layer Attribution of Information Omission in Air-Gapped LLMAgent Pipelines cites this paper.

Where Facts Go Missing: A Layerwise Taxonomy and Per-Layer Attribution of Information Omission in Air-Gapped LLMAgent Pipelines KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T04:48:25.926385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:48:25.926385Z digest=sha256:d40e14579fe97a98403519b56733a8afea74edc6b91504adfac47bf5d6373d82

Observation 270ca459-9ac7-4564-81ca-0a1b40cfb920 · inbound

A Photonic-CXL Memory Appliance for Scalable KV Cache Management in LLM Inference cites this paper.

A Photonic-CXL Memory Appliance for Scalable KV Cache Management in LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-30T11:12:05.608333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:12:05.608333Z digest=sha256:74c825eca271b6562e122c96953041a2c88ca9384a3836e398cb16aaf2f2a006

Observation 4c113f63-5339-43b9-95e9-a9b797220e11 · inbound

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization cites this paper.

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T00:43:49.304692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:43:49.304692Z digest=sha256:18648de1bd2575e89e186bb93b758ac61d0d0c7a4bc400559c96bd2f04699c75