Pith. sign in

Paper Citation Record · LEDGER

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models

As of 23 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 6 inbound Pith citation observations for arXiv:2501.19392.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.19392 v4

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:34:01.382301Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-25T04:51:04.068354Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T04:55:24.203465Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact2
  • verified fuzzy3
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ad5d31d-d34d-4027-8f5b-c90498d5144b · outbound

This paper cites write newline.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.163378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.163378Z digest=sha256:db1df80dabcf49cd624e82e51919e0f8f66b76550ba6fa3062891c2f6f14bf7a

Observation 251e4d38-cb7c-48ca-8464-a9042a9ff396 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.168712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.168712Z digest=sha256:c62188d24167e77bca7001ddce81d2faa9e97e594d671c4b8c0cad41b8077193

Observation af7d1eb7-9eae-4398-a32d-2dc231825daa · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Understanding intermediate layers using linear classifier probes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.172410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.172410Z digest=sha256:e2af8e9f650653584d78b2aee0e1ab55f6a1e096d03ae4ef7bc8538cb1ca19e5

Observation 3d464322-1ec4-4d62-add5-38a920dc34e3 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.176784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.176784Z digest=sha256:26dcfcecfd2d407079c20adff1a43e3d78796b7762d0259babd036fe7e999d38

Observation 3ab4ce44-4aa3-45d0-8dbf-615a79164ca4 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.181193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.181193Z digest=sha256:59854b16ee1e490309aae36730ed58b669d20112e13f62e6a3975107f04a6923

Observation 4fae5ebc-01a5-44ce-8995-db483bb3304d · outbound

This paper cites M., Gebru, T., McMillan-Major, A., and Shmitchell, S.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models M., Gebru, T., McMillan-Major, A., and Shmitchell, S

Reference 6

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T20:34:02.075389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.185806Z digest=sha256:a997a52f2bfa95f560aceb93a172c4363ef396fa5a4537af8df6202931b7b748

Observation 9e72fe70-0f02-408e-ad53-19ad07a7a089 · outbound

This paper cites Palu: Compressing KV-Cache with Low-Rank Projection.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Palu: Compressing KV-Cache with Low-Rank Projection

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.189891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.189891Z digest=sha256:663927e8ed2428926333cdd74e78a9e3ff9d17d799b15c2c7bf4801e4dbdb05e

Observation db045350-6224-45c4-a0ee-2771b8660ea2 · outbound

This paper cites PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.194069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.194069Z digest=sha256:b416a8d0d3e0580a1fb20219ae71ab965fde85c76f8d2811e42fcd16ae161531

Observation ed214d9c-c945-4b69-b3b3-c80e9c68c2e9 · outbound

This paper cites Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.197891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.197891Z digest=sha256:6829e2e022b4a7d543bcfe0738d82c486855af6c9d527349f69be82fc966f1b1

Observation 7230ed0d-e4f3-4763-8368-4242e1735827 · outbound

This paper cites QAQ: Quality Adaptive Quantization for LLM KV Cache.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models QAQ: Quality Adaptive Quantization for LLM KV Cache

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.201633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.201633Z digest=sha256:edd1750d1bd15e7927aea66ceafefaf5c25f219b0e770a2fe247edb13aa38188

Observation 9ed2fe7f-cb97-493d-871f-444192003f39 · outbound

This paper cites SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.205850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.205850Z digest=sha256:dbb9bb62c84ecc5f71a3337a236e935aa589068c0065f31265c1d20a5c70318c

Observation 4a74aea1-516b-4ebb-8de6-e9e59c0f7091 · outbound

This paper cites The Llama 3 Herd of Models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.209552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.209552Z digest=sha256:b8a033415be5d97517518ba56cc904920bed3b362228b1764cb82e845d02050c

Observation d792f3f2-6b06-45a3-a052-5d75ce161ae3 · outbound

This paper cites Towards Measuring the Representation of Subjective Global Opinions in Language Models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Towards Measuring the Representation of Subjective Global Opinions in Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.212965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.212965Z digest=sha256:4017058fee966790d4975463e58dbaa110c13914556e8c07f7aef333e9eb4b5e

Observation 43925cb0-b1e0-4fa5-9897-df956ab7e938 · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Extreme Compression of Large Language Models via Additive Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.216140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.216140Z digest=sha256:02a507ac4ed608bb7cdfd583060031a8315998ee666d0370da73eafc8c6874ea

Observation 8a374d1a-a87f-403a-b8ae-0f658c2b9086 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.220117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.220117Z digest=sha256:e499060ab6d80fa272f8260cbf7313f673dde4534004a56c7ee0b5d9a86b20f4

Observation 8c7041fe-7771-41b7-8d3c-ab6a53ac0f1e · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.224437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.224437Z digest=sha256:24422014a61df656eda565ce8526788500fc39cad8d68d0a2cb2a306dad2c395

Observation a2276e2c-5d4c-46cc-81dd-afd48dae68ad · outbound

This paper cites an unresolved cited work.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:34:02.220412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.228274Z digest=sha256:7d4aeb4f2e07854b13e8d1d6124194eeb0282235fb412534e0d6ab7dd81a810a

Observation 95a1ce10-7f8c-4dca-8d26-f4291b42853f · outbound

This paper cites Fast matrix multiplications for lookup table-quantized llms.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Fast matrix multiplications for lookup table-quantized llms

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:02.211651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.231283Z digest=sha256:efc6ee25822af6aff289bb28ad28c7a10f52617c953219790c94e760b2c6e4fd

Observation b8b6f30e-69ab-4186-804d-ceb69798c826 · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.234695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.234695Z digest=sha256:0aa5c8b6224d55bcee630b52a031232fa6a6e0fe1079a82c36eb7da88b1c3d17

Observation 63704b6d-d0da-4d2f-9966-e36b374e5998 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.238595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.238595Z digest=sha256:d9de7a03588af9805f4e8442f0129910d5ef91e8b6d9f398bb52a4a92d07883c

Observation 82fe702c-a446-44c3-8773-9407f9878293 · outbound

This paper cites Stochastic distributed learning with gradient quantization and double-variance reduction.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Stochastic distributed learning with gradient quantization and double-variance reduction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.242633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.242633Z digest=sha256:4c704e2bde8a65f3d52641407a71ce76dbda6df9aea76b1490ed5eb27e80fab9

Observation 4fa84ad4-54f3-4f07-8dc1-78f7ea39e1a7 · outbound

This paper cites Optimum-quanto: A pytorch quantization backend for optimum.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Optimum-quanto: A pytorch quantization backend for optimum

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:02.196286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.246441Z digest=sha256:a4c3d2a9aa3c85f81385b2204707904d0203e2bdaeae6f6f149dba91e41c1ade

Observation d80ec495-591c-4b2b-98d8-d3be9eca364f · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.250786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.250786Z digest=sha256:ea07b074fe114a9dbe5522a665e1324afdfea3241c22786fa4bfc698a50afa9f

Observation ddf7a799-c6fc-4f2c-95bb-f12f480a8f7f · outbound

This paper cites H., Gonzalez, J., Zhang, H., and Stoica, I.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models H., Gonzalez, J., Zhang, H., and Stoica, I

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.254750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.254750Z digest=sha256:1a4f1a2cf287283cb880badab55a795d2845b20d80f57f328a0e769f34fa7237

Observation 363d5d4e-f8ef-4346-a9f0-3e4ee7b35342 · outbound

This paper cites A Survey on Large Language Model Acceleration based on KV Cache Management.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models A Survey on Large Language Model Acceleration based on KV Cache Management

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.257427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.257427Z digest=sha256:28f2fc27b00abd25dbe81db0a48ce53e6ba010697405e812a767004d448a105f

Observation 3bf02429-cf6d-4568-9e3c-7552eecc6d51 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models SnapKV: LLM Knows What You are Looking for Before Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.261682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.261682Z digest=sha256:33027f12a386b25b2042c5e389f92bbc41afa28e8b2d06a4640a05d6098efd86

Observation ef36e9a9-22d8-4947-a948-1d7fb234a758 · outbound

This paper cites SCBench: A KV Cache-Centric Analysis of Long-Context Methods.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.265626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.265626Z digest=sha256:f258e9e334e515666c6f14e88a21de650eb7668130120ad4d9cf16c94f4058b7

Observation f4d0dc2c-42d4-46ac-9b48-cf2b1d6fb4cd · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.269681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.269681Z digest=sha256:a2534ee25817d844d8087ecafb782ff56ff4a5c47839a938286d9a54b4586cee

Observation 8100fd37-f4e5-42b6-8a23-f126bcce6ec9 · outbound

This paper cites MiniCache: KV Cache Compression in Depth Dimension for Large Language Models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.274336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.274336Z digest=sha256:755abb74c8723fa1e0a5c061f33d0233d1cf34800a07dabfc4ed1f6bf669a673

Observation a3c491b0-3ab9-45ff-a0f2-11633560713e · outbound

This paper cites IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.278582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.278582Z digest=sha256:63db1f012dbb7d4e2cab728ae3749e49b3b1ccf0d780109b766008102236e77c

Observation 70c22dd8-2852-4c04-93f9-ee6de6058fd1 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.281832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.281832Z digest=sha256:31193fe151bb5d46a6ecbcd8e521162bf22cc1cdc2ba803c28a1198b377af2d0

Observation be5b2b39-17d7-4e12-baa0-d2ac0ad6d211 · outbound

This paper cites PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.285840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.285840Z digest=sha256:d6e332a4b92bbd7d072b916bea63e4d3b5de2dbb947da18fb214ba8aa6ff3705

Observation f0f3d9b0-5619-4ccf-a814-6d412e9edcfc · outbound

This paper cites Pushing the Limits of Large Language Model Quantization via the Linearity Theorem.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Pushing the Limits of Large Language Model Quantization via the Linearity Theorem

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.289119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.289119Z digest=sha256:2ac790f53b013fc3bb6f98d535589c994bb1bb62f2f0f0e482c477384d2c77d5

Observation ec4bf091-6217-4858-9150-f8330403e6bc · outbound

This paper cites Pointer Sentinel Mixture Models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Pointer Sentinel Mixture Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.293602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.293602Z digest=sha256:7a9a13c87072b5a2d0a93557903e3e4b68e41ce2a5264af1cb78b610da1a09ee

Observation d05ec6d2-877e-4df3-a98f-56f099406ef1 · outbound

This paper cites PyTorch : An imperative style, high-performance deep learning library.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models PyTorch : An imperative style, high-performance deep learning library

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:02.180789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.297355Z digest=sha256:83852fe10cd52eae8eb32905c762d6826c5e0845f8f30e6a2f5b3640c4c2e7fa

Observation d5100856-34e6-483f-8383-9210abf52b3d · outbound

This paper cites Your Transformer is Secretly Linear.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Your Transformer is Secretly Linear

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.300035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.300035Z digest=sha256:06dc82a84a01a8b6873e952f565df6c7cf3b82ff0844846138ee4d9ad9eeafb9

Observation f446cae9-71a6-4b8b-99eb-2b2ff23675c6 · outbound

This paper cites an unresolved cited work.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.303327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.303327Z digest=sha256:e043840e7d834b8e32430341b4fb95789df39b386b7899e33a63f109d7eee800

Observation e5ecc08a-e6ac-4eec-84ef-bda8409605dc · outbound

This paper cites Societal Biases in Language Generation: Progress and Challenges.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Societal Biases in Language Generation: Progress and Challenges

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.306566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.306566Z digest=sha256:f3e658b980921749c48ed874c7052696f1f9ddaf8bcdfaa9268a44ffac6c480b

Observation a73bffc1-c291-4e1b-b4cc-510722dcd859 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.309979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.309979Z digest=sha256:46af02cfbced246f7084cbd3a57076c9bb12dac6276e41bc91d22d269643ffe7

Observation a9b5555b-cc54-48a1-87a2-84fa96ffc1e3 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Qwen2.5: A party of foundation models, September 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.313944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.313944Z digest=sha256:e595c59c537e8b4b84b2b5fd736f0f4a1a82155b609adbbd0fafddde04b81c30

Observation fd55ea0e-1543-4e06-87e7-f2ab5767f297 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.316831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.316831Z digest=sha256:e307dcf8236b01722ed324c0f2a6293368f5081baf65a41b335453b59369ffb2

Observation 9926d273-81d8-46b8-9d2d-605d88f75458 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.320318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.320318Z digest=sha256:fd27e3d05973994c0529e5acdd00c7976732748b4c28f7e446f7b93416636f25

Observation 5c3e592f-d489-49e3-825a-8c282473eae6 · outbound

This paper cites an unresolved cited work.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:34:02.162672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.323403Z digest=sha256:3d39bac5085e4e4e24cd5e11eff78a1b4b39b4c014eef582674904d0799d1aff

Observation f89d223d-34de-4a17-bc3c-b7f9018f125b · outbound

This paper cites an unresolved cited work.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:34:02.150601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.326506Z digest=sha256:de9b66999e8faf233049ddbecad5df3345ea2afbdac6109b5b77718a8cc6b60c

Observation 9cefe2c7-1865-4dcd-b5e4-bdf3cb2700a8 · outbound

This paper cites GPTVQ: The Blessing of Dimensionality for LLM Quantization.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models GPTVQ: The Blessing of Dimensionality for LLM Quantization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.330531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.330531Z digest=sha256:a1eefb2aa644db94c98b7faab5c86711a88da3daaee4a6ff57a93111f866567f

Observation b1a1c0d2-a727-469b-95c0-61804011d507 · outbound

This paper cites Attention is all you need.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Attention is all you need

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.335869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.335869Z digest=sha256:239cca122db52decba449090354638f1da9790756e4b5562d2fe96d8777b5c4d

Observation c20bd02e-0937-454f-a476-5e06d12ab2bd · outbound

This paper cites an unresolved cited work.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.339571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.339571Z digest=sha256:4b1278f4d60be2e48551d0d608edd1e476c29f246818590373518738807ea785

Observation 35ce49bf-5326-4b1f-8bab-7cc0fa564058 · outbound

This paper cites Ethical and social risks of harm from Language Models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Ethical and social risks of harm from Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.342390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.342390Z digest=sha256:5ad8289343caa0be64c16eade58db0c72418dc18067cbf9e5b9831b74f392c68

Observation 27820dda-1aaa-41e4-b023-1620dc7e8ae4 · outbound

This paper cites an unresolved cited work.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:34:02.128852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.346803Z digest=sha256:d304bd51be5ec1d8d6db0350b6b68ff1e314074f1de620915189827c440f8676

Observation f5ec474a-361b-46dd-98f6-707024ab02b4 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.350988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.350988Z digest=sha256:b35619bd84533f158a22b1a8d668dc4a726d65495e7059846e4bf6856454d951

Observation 2645100c-35ab-44b7-a89d-6b2d880f895f · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Efficient Streaming Language Models with Attention Sinks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.355022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.355022Z digest=sha256:5fada58607aa2cb6065aea7f1d3e7de4ed1e1da6712382b62a418564dae13363

Observation 685d429d-b916-4c2d-b21e-3e38f7f79ca4 · outbound

This paper cites Qwen2 Technical Report.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Qwen2 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.358141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.358141Z digest=sha256:c39a0b5a42d2686cb88384dfa4d8281995a025df292f29ce36ae79515219c00a

Observation e36af4c4-b7d4-4940-8b8d-a614cfbfadfe · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.361362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.361362Z digest=sha256:673abbb037a8da87ddfa656bbae0fabc9ab209e5ea840a5a3323170f7eceabd8

Observation e75cced3-0e74-4d9b-a114-f262531871d2 · outbound

This paper cites KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.364320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.364320Z digest=sha256:735cd73ac020d4efbc155dbaebcc93daf9970087edde8e863e2832d4a06e0f20

Observation 7761c196-453c-4558-84f1-8e046a244107 · outbound

This paper cites WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.368235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.368235Z digest=sha256:b296bff00f2cf62f5c9a4652432b651f58f331f2fd384a9096eeadd8c5dedfe1

Observation 09da41e6-ac57-436c-8e5e-4482df3496cc · outbound

This paper cites QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.371452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.371452Z digest=sha256:b42db61c766cb820261bd78f9fb519d7cf00ec0c7ec573b2f27acdb5c9af9ab4

Observation e6291d10-d012-4508-af3a-b78590459ccf · outbound

This paper cites FDC: Fast KV Dimensionality Compression for Efficient LLM Inference.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models FDC: Fast KV Dimensionality Compression for Efficient LLM Inference

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-09T20:34:01.426422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.375678Z digest=sha256:da02ae94d6da097fbeca0e8c97aea0295e9454af4d62d2f1099495a45b24f56f

Observation b458603b-e590-4201-bbfc-0d58d785b0d5 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.379174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.379174Z digest=sha256:a6f5204255a31010b1290afb90ade93260999249e013cb5f23dbb742a98b4efd

Observation b11442c7-e928-4a77-9eee-1621558020bd · outbound

This paper cites Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.382301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.382301Z digest=sha256:914677ee6a66cea0344a43c9b08b8bc07e9732508f7187113fa390ba09d94ac1

Pith citing papers

Observation 0c036e79-b643-41f2-aef1-56bf2295a7f7 · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:06:01.037853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:49:13.353700Z digest=sha256:e7860841569c8b3a7bfaa226929dcd2d198ef7efe156b3282a699e092b8d9a28

Observation 1c7f9c93-c852-478e-a569-06591cefc921 · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:56.308454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T01:48:21.341462Z digest=sha256:d128e2c7a83eb38aa13e7f551d9b4e6941eac4667794fcf507baec337720e69e

Observation 7779274f-0729-4706-a7d2-54b82190afe9 · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:37:29.392954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T07:35:30.155592Z digest=sha256:ad817a188e9a5d145d646119f8ca20d3fbb1761acecb0b45d6b4f0de23de2c4e

Observation 34b564f1-c187-439d-8388-da7bf1f7fea5 · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:42:39.675232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T16:41:30.021302Z digest=sha256:2c9b2718d416d3b5c595e3de1354c5593e926b9dd50a5cb5985130b0770fc348

Observation 37f7136f-56d8-40a3-b40a-157028fe8598 · inbound

RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache cites this paper.

RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:33.526717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T01:16:15.869198Z digest=sha256:e573b1e55bf9ce020c24b02aff0eb24fe8d3ad6b2bdde0bbdb5d2a446228a327

Observation 074dd5fb-7213-4fb4-8bac-8cce632be527 · inbound

A Simple Plug-in for Improving Eviction-Based KV Cache Compression cites this paper.

A Simple Plug-in for Improving Eviction-Based KV Cache Compression Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:55:24.206757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-25T04:51:04.068354Z digest=sha256:0f2218f995978773c8cc26e86850461e8e4d2a2b5aca3904117ae00eb0712aaa