Pith. sign in

Paper Citation Record · LEDGER

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

As of 14 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 7 inbound Pith citation observations for arXiv:2505.18610.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18610 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:27.538376Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:51:05.406333Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-05T05:50:43.580203Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2b20b506-b216-4e0d-b3f6-033e12c8755b · outbound

This paper cites American invitational mathematics examination, 2025.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs American invitational mathematics examination, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:28.842501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:34:25.149159Z digest=sha256:7f7aba18c00829f416836d31db69afd5b4928801dd552ef48c7b835bace13eb1

Observation 7675a78b-e5df-4b7b-9435-23cd6dc9daf0 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.275482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.275482Z digest=sha256:fc32919ca4e39f0f81a37bfeede8ac40c5c0c3904514f8a6101deaee24131dd7

Observation 4984c2db-1f4b-4200-ab32-71651135d69a · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Extending Context Window of Large Language Models via Positional Interpolation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.333985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.333985Z digest=sha256:3d95418649d65d48f67b5cb3b6152776566421b9f104a81d460549363eb2c79f

Observation 8a6a85a1-b026-4b4d-8961-f89d3ad5f6f7 · outbound

This paper cites Carnegie mellon informatics and mathematics competition, 2025.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Carnegie mellon informatics and mathematics competition, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:28.723653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:34:25.377424Z digest=sha256:e91be777cf86dd12953b77fd0eb510fb3170e2c0765c52058c31e116022bc6de

Observation d4cf4010-ac3c-41ca-b7c0-385bba73c91c · outbound

This paper cites CVXPY: A Python-embedded modeling language for convex optimization.Journal of Machine Learning Research, 17(83):1–5, 2016.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs CVXPY: A Python-embedded modeling language for convex optimization.Journal of Machine Learning Research, 17(83):1–5, 2016

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.407380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.407380Z digest=sha256:3957048afde2048c173f02489bdc51761b611f73f28aea4e4d0dbaea32f0de92

Observation 0bb28117-42f1-45b9-97d9-8d2d9dc4e9b6 · outbound

This paper cites SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.438387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.438387Z digest=sha256:97709d523702caf666951da5cbc37eac4f188cea0e0f9d115ef197331d5b2529

Observation b3f6d26e-29d0-45cf-959b-9bb1f42cc5f3 · outbound

This paper cites Moa: Mixture of sparse attention for automatic large language model compression.arXiv preprint arXiv:2406.14909, 2024.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Moa: Mixture of sparse attention for automatic large language model compression.arXiv preprint arXiv:2406.14909, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.524088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.524088Z digest=sha256:8570116d29dbcf22d90bd258073afa6e09121a548a7484eee1d0a4e3983dc277

Observation 9e8454ea-7e57-4706-a720-2c881b122819 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.573877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.573877Z digest=sha256:b9a65230a5c53229e02b53d22e424fcc350691c9f64bd297dcd8fb3191cdf29a

Observation 6bfc87e1-e9ec-4e0e-a3d0-15c80281a929 · outbound

This paper cites Quantize What Counts: More for Keys, Less for Values.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Quantize What Counts: More for Keys, Less for Values

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.594534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.594534Z digest=sha256:3d49dc7ef9296c894b90495ce12f8f7307a11e22e55c85688ba295024b3d8085

Observation 9240860a-f53e-4064-ac51-83e9689daa9e · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.647588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.647588Z digest=sha256:6b878b354eeb2e43ab78c6f585c4f85719c5575fdacc7a13c261f5d4eb5c1d9f

Observation db4b4924-5a15-422f-b1e1-81dadb0a2338 · outbound

This paper cites Llm-mq: Mixed-precision quantization for efficient llm deployment.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Llm-mq: Mixed-precision quantization for efficient llm deployment

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:28.607938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:34:25.718605Z digest=sha256:071199ca27e880e5a5a84a3aa711530d527f3cb5dfcf7abe11c615afe8a4b818

Observation d2857518-6375-4364-a8f9-6808db1e86c6 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.833313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.833313Z digest=sha256:2a5c9ef300010d42da784c3db93b8f9d03903b0cc72b910fc2c274db737e65cd

Observation 2c239ce3-f9a5-45f6-832d-c4cbfe5c4158 · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.898209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.898209Z digest=sha256:cd78cfe93f71162baa800ffad49dc33b3090e44ebc3cda3c51cb4d59e0b7cbc9

Observation 3701df9d-5b34-4c88-bae0-f5f8e4659a31 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.951695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.951695Z digest=sha256:2bb859903d518fd3d7099e903aebd367cab339688a8b2d861912dfce95f10cdd

Observation 05a9d736-233f-40be-a294-c05402c37922 · outbound

This paper cites Intactkv: Improving large language model quantization by keeping pivot tokens intact.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Intactkv: Improving large language model quantization by keeping pivot tokens intact

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:28.356439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:34:26.021077Z digest=sha256:ac4046e1374b584b6c00cadb046c1d07e0a71a5a68c90b65b183debfb1817a16

Observation 621798a3-de2e-4b6f-a659-4116715d871a · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:26.080933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:26.080933Z digest=sha256:d77d57a183c596b9413d2e55ab78e0687b8c84711852ab27513a133904de4608

Observation a24314fe-0752-48a6-84ec-aac281827bef · outbound

This paper cites Introducing openai o1, September 2024.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Introducing openai o1, September 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:28.139404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:34:26.156677Z digest=sha256:23bfcf30764cdf381eab3d490985024361beb246debcff305d9a470d29ac9055

Observation 1cdbaf85-b8b8-4cf4-a88c-d75c9c3d53f0 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Fast Transformer Decoding: One Write-Head is All You Need

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:26.276661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:26.276661Z digest=sha256:8d2c824b6e5597eb727a5f214726df531f0454c086d5efb572e247921f9ace7a

Observation a617a6db-d533-463d-ab27-f19a9c9e765f · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:26.328850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:26.328850Z digest=sha256:ef1eaf0f312efa5af92bf834c9840af28b3235b53e45d8fa089b85a0308b3738

Observation c805816e-3b54-44e3-9dca-12eead72654f · outbound

This paper cites RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:26.485614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:26.485614Z digest=sha256:06e957c2ed0e7888b6f92fbbd00306d3edcbc1f6d70893cb9f8d32fafc0a5724

Observation 3bb5bc09-014c-4013-9ef9-91deb9bddd52 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:26.603479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:26.603479Z digest=sha256:6bbbb0808c2066a75c1393899130cc3b541b17ff7b6a414e9e2603042258b2cc

Observation 4f59fc53-bcd7-4d96-9798-ce6da8931fcc · outbound

This paper cites an unresolved cited work.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:26.757650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:26.757650Z digest=sha256:ae2d73c3ce1c79940048b4998329e95ea3c74356374ec848fa783a88942e5c02

Observation 3d5a41f4-3681-400c-9ac5-2705c36c194d · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Efficient Streaming Language Models with Attention Sinks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:26.912924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:26.912924Z digest=sha256:c30ee9674b18f0f7a480f094f2cc7cead58eb886a3eb540640e693bf2bf3d287

Observation 9f56688f-352e-42ab-9a88-3b54731b5200 · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:27.041485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:27.041485Z digest=sha256:134613ca6240a74563d1ce3ea3e414d0e446c4e8829f995e46711f9a410b4dc2

Observation c81bd392-b7a7-49c4-974a-27b5949d74d5 · outbound

This paper cites WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:27.199163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:27.199163Z digest=sha256:05b7a007917032336dcb0775e024233fbbba6c5b966ac0bdefef493f6dcdaff4

Observation fc0885bd-c87e-4f3f-a23e-4ebf8585fb8b · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:27.373613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:27.373613Z digest=sha256:7b3a30b1258b2f3518ff3679bc285929d337f2762437e4e2988a3d82772d9b61

Observation 8f54927b-1e40-4966-847c-e6c7f192e5b0 · outbound

This paper cites Mixdq: Memory-efficient few-step text-to-image diffusion models with metric-decoupled mixed precision quantization.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Mixdq: Memory-efficient few-step text-to-image diffusion models with metric-decoupled mixed precision quantization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:27.967366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:34:27.538376Z digest=sha256:9aeac19458fa71a04307d32e6542f120be6fa2bd84f337b07196932fa38df29c

Pith citing papers

Observation df4693bc-bdfd-4a6e-9a8a-910229384ab7 · inbound

PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models cites this paper.

PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:05.406333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:05.406333Z digest=sha256:f600f11b8a2f035f8c4ef51b4963f9e1a3b2f95c7dda1beaae19fcb9acfe1b3c

Observation ddf4ad6e-87ee-4209-91f6-3b3c441c7aa6 · inbound

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory cites this paper.

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-14T01:19:53.637977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-11T01:09:21.811344Z digest=sha256:88c6fdd87b898f60ba70a36425e792f8f250c8024c3c074ba03c92e0d0d33378

Observation 194f9ca4-905b-4a57-a95f-2123d5b4d514 · inbound

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory cites this paper.

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-14T01:19:53.637977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-05T05:41:13.869451Z digest=sha256:3590a17a8b34f0ed9cabcaf902e1f2804a6fd5d84e2b732a3e7c5a1a5b523c7d

Observation 2c0a9ec0-70f8-4829-95a0-c4ed6d93817e · inbound

KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving cites this paper.

KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-14T01:19:53.637977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T17:42:06.870534Z digest=sha256:6593a4229834128efad6c76f601377e3074008aae743f217baa0cb3a28aaf58d

Observation 201f393f-7216-4db9-8702-940eaa3cca96 · inbound

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization cites this paper.

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-14T01:19:53.637977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T12:16:07.797702Z digest=sha256:92d0a5f368b29bd17346052fd2d3b4fa6de05f3892959bb23f213b3fdb4e4283

Observation a57804d3-3d5f-4d26-9817-433780185387 · inbound

RoPE-Aware Bit Allocation for KV-Cache Quantization cites this paper.

RoPE-Aware Bit Allocation for KV-Cache Quantization PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-14T01:19:53.637977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:05:13.790381Z digest=sha256:958e77636ed467c70c4743d73d79bd61f196625dbdb55c34721e35c86c1df877

Observation 2f48ec61-67e9-4b2a-8aa7-4acaa3ffe79d · inbound

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration cites this paper.

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T09:50:32.732340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:50:32.732340Z digest=sha256:33d1e6825d7ef1ab2fd04542fbbbeef4082ca029f147130396e238f57c386fd0