Pith. sign in

Paper Citation Record · LEDGER

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

As of 12 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 7 inbound Pith citation observations for arXiv:2505.18610.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18610 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:27.538376Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:51:05.406333Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-05T05:50:43.580203Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2b20b506-b216-4e0d-b3f6-033e12c8755b · outbound

This paper cites American invitational mathematics examination, 2025.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs American invitational mathematics examination, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:28.842501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:34:25.149159Z digest=sha256:afbaf713dee9c63d86b373c04a117567b407a75f5acde342fee443db1a061877

Observation 7675a78b-e5df-4b7b-9435-23cd6dc9daf0 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.275482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.275482Z digest=sha256:21b2b90afb664c6574a7224b153d8572fb1a40882d706867bfe0325663a56a30

Observation 4984c2db-1f4b-4200-ab32-71651135d69a · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Extending Context Window of Large Language Models via Positional Interpolation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.333985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.333985Z digest=sha256:b582cca1f52ad88fad1558b6252580f1dae6762b35a1318f81b4cabec792486c

Observation 8a6a85a1-b026-4b4d-8961-f89d3ad5f6f7 · outbound

This paper cites Carnegie mellon informatics and mathematics competition, 2025.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Carnegie mellon informatics and mathematics competition, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:28.723653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:34:25.377424Z digest=sha256:c0834e4d4abca0aea65e38ca9da64ece2f3a2bece35ebb19a58dd05a5f999563

Observation d4cf4010-ac3c-41ca-b7c0-385bba73c91c · outbound

This paper cites CVXPY: A Python-embedded modeling language for convex optimization.Journal of Machine Learning Research, 17(83):1–5, 2016.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs CVXPY: A Python-embedded modeling language for convex optimization.Journal of Machine Learning Research, 17(83):1–5, 2016

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.407380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.407380Z digest=sha256:0fd043a7174169b3ad07fe4e99d30fff335ba130604ce58b4e3594ff08d9e9ba

Observation 0bb28117-42f1-45b9-97d9-8d2d9dc4e9b6 · outbound

This paper cites SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.438387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.438387Z digest=sha256:f73665b945fef5c502d4671122f3c1c737a538337e0dbe64a7b5b59d3fa52b0e

Observation b3f6d26e-29d0-45cf-959b-9bb1f42cc5f3 · outbound

This paper cites Moa: Mixture of sparse attention for automatic large language model compression.arXiv preprint arXiv:2406.14909, 2024.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Moa: Mixture of sparse attention for automatic large language model compression.arXiv preprint arXiv:2406.14909, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.524088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.524088Z digest=sha256:4e0b842395f18900cbc0f714b1681f4815a76b221655187e57f7a8a1d8d460e3

Observation 9e8454ea-7e57-4706-a720-2c881b122819 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.573877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.573877Z digest=sha256:3a12246fecdf09daad8359212e427d380ebefe9ff6f419caabee9b71b91825b5

Observation 6bfc87e1-e9ec-4e0e-a3d0-15c80281a929 · outbound

This paper cites Quantize What Counts: More for Keys, Less for Values.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Quantize What Counts: More for Keys, Less for Values

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.594534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.594534Z digest=sha256:781a20b387a50c112e31cbeebbf849abfd990b60cfc996ec854ee1bc9f99876f

Observation 9240860a-f53e-4064-ac51-83e9689daa9e · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.647588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.647588Z digest=sha256:a988a0c57faa5e1ad6eb99c0d77ee600f538e45b143563bf83962d080f1a90e2

Observation db4b4924-5a15-422f-b1e1-81dadb0a2338 · outbound

This paper cites Llm-mq: Mixed-precision quantization for efficient llm deployment.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Llm-mq: Mixed-precision quantization for efficient llm deployment

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:28.607938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:34:25.718605Z digest=sha256:b33d526ebfc301481352e61a10bc9fb4f78ed2e94a32f11e7378092bdae4b21b

Observation d2857518-6375-4364-a8f9-6808db1e86c6 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.833313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.833313Z digest=sha256:78da391e4ad22b8148fbe5dba5799bd95c048e2418c24927094304dd4700bed6

Observation 2c239ce3-f9a5-45f6-832d-c4cbfe5c4158 · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.898209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.898209Z digest=sha256:5d8fe4b368b440d0d2b54fa92ca8d5b30577f7f85cab089d902fa3cfe8753856

Observation 3701df9d-5b34-4c88-bae0-f5f8e4659a31 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:25.951695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:25.951695Z digest=sha256:3a2e37a66682e856ffa759d32ae6daf7a0483199b247a4e187fe6882e5a146a2

Observation 05a9d736-233f-40be-a294-c05402c37922 · outbound

This paper cites Intactkv: Improving large language model quantization by keeping pivot tokens intact.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Intactkv: Improving large language model quantization by keeping pivot tokens intact

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:28.356439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:34:26.021077Z digest=sha256:87a8befb7b8882f2a4aba78dd312d2de331d26c8b1d3702705562c22ec763a94

Observation 621798a3-de2e-4b6f-a659-4116715d871a · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:26.080933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:26.080933Z digest=sha256:fa0a474336b2cf547c7cafc60ab6feffedb4f1c301c38e1b1a6df662bf64833e

Observation a24314fe-0752-48a6-84ec-aac281827bef · outbound

This paper cites Introducing openai o1, September 2024.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Introducing openai o1, September 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:28.139404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:34:26.156677Z digest=sha256:0a9df5c857361bc8d885527c517c0bf5fc8ef232f4ff46d33c4795d6abf670bd

Observation 1cdbaf85-b8b8-4cf4-a88c-d75c9c3d53f0 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Fast Transformer Decoding: One Write-Head is All You Need

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:26.276661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:26.276661Z digest=sha256:9754e18459782b93c4a58405ae8e669752a1953c5b1749a4533c36bfb9504d45

Observation a617a6db-d533-463d-ab27-f19a9c9e765f · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:26.328850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:26.328850Z digest=sha256:85db39399fb4829f71b6d3e03ddb20bfd4196f1a26fa6a792984b70b484ae105

Observation c805816e-3b54-44e3-9dca-12eead72654f · outbound

This paper cites RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:26.485614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:26.485614Z digest=sha256:6a0f996472322d870fa196b96d74ca75d58977a570e17d695f3c1fcd8b0037ae

Observation 3bb5bc09-014c-4013-9ef9-91deb9bddd52 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:26.603479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:26.603479Z digest=sha256:808522ecdeb83634539df501d90529e9247aacd4838e43991b5af98831229136

Observation 4f59fc53-bcd7-4d96-9798-ce6da8931fcc · outbound

This paper cites an unresolved cited work.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:26.757650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:26.757650Z digest=sha256:19b48554b4b0a8d859907779819d6495a01f9ca9caaa7ac9a38082f79150e427

Observation 3d5a41f4-3681-400c-9ac5-2705c36c194d · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Efficient Streaming Language Models with Attention Sinks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:26.912924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:26.912924Z digest=sha256:8cf8e768047ea9656d4ff105e88122c1dbc792aa2f7cfa6c6250d544ae7b8b4a

Observation 9f56688f-352e-42ab-9a88-3b54731b5200 · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:27.041485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:27.041485Z digest=sha256:7f15538ffb12a140392e6cfb99178d713cdbe3b084f1b21dab0261f285fb9804

Observation c81bd392-b7a7-49c4-974a-27b5949d74d5 · outbound

This paper cites WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:27.199163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:27.199163Z digest=sha256:a9911cf5337f6afd13c21721140ee12281ff4a4529fa9907e096407ba43ae12c

Observation fc0885bd-c87e-4f3f-a23e-4ebf8585fb8b · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:27.373613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:27.373613Z digest=sha256:1765b946454ecde8857e0bf2505de3eaef14bc1cb1c768f495961cedcc9471f9

Observation 8f54927b-1e40-4966-847c-e6c7f192e5b0 · outbound

This paper cites Mixdq: Memory-efficient few-step text-to-image diffusion models with metric-decoupled mixed precision quantization.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Mixdq: Memory-efficient few-step text-to-image diffusion models with metric-decoupled mixed precision quantization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:27.967366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:34:27.538376Z digest=sha256:8734b70082ab6fa0e397ceb81e33dd387bf03c1efc017d07a8d441fb1fdc8339

Pith citing papers

Observation df4693bc-bdfd-4a6e-9a8a-910229384ab7 · inbound

PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models cites this paper.

PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:05.406333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:05.406333Z digest=sha256:46db0c4893bede1dfdce2b0f2737ef399005b7d88390072c6e97cc404c8a0d2a

Observation ddf4ad6e-87ee-4209-91f6-3b3c441c7aa6 · inbound

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory cites this paper.

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-14T01:19:53.637977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T01:09:21.811344Z digest=sha256:3e23e873e9113745937df29da585ceb6173bf2b947b7a96191638b547e3cfc34

Observation 194f9ca4-905b-4a57-a95f-2123d5b4d514 · inbound

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory cites this paper.

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-14T01:19:53.637977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-05T05:41:13.869451Z digest=sha256:b2ea8e691375329affec24788c1d248645c2e51fa2fddddfa536f8f591d47be2

Observation 2c0a9ec0-70f8-4829-95a0-c4ed6d93817e · inbound

KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving cites this paper.

KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-14T01:19:53.637977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T17:42:06.870534Z digest=sha256:8814a02e16a75c3660289e7eb73cc59e375b7467604e4f49bc77c5d60a0c85d6

Observation 201f393f-7216-4db9-8702-940eaa3cca96 · inbound

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization cites this paper.

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-14T01:19:53.637977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T12:16:07.797702Z digest=sha256:dcdd0be7a687420ac41ca5e579766002049dd16686861144aef0126dadcb7f80

Observation a57804d3-3d5f-4d26-9817-433780185387 · inbound

RoPE-Aware Bit Allocation for KV-Cache Quantization cites this paper.

RoPE-Aware Bit Allocation for KV-Cache Quantization PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-14T01:19:53.637977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T01:05:13.790381Z digest=sha256:cfbf9b7669fe3fc7d4bf17fc0678310888639c8c490b1ea756054d0aa41aa9e7

Observation 2f48ec61-67e9-4b2a-8aa7-4acaa3ffe79d · inbound

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration cites this paper.

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T09:50:32.732340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:50:32.732340Z digest=sha256:7a1097ad0588793b05660875014fa1345c8277b97a7a5663de809d03375c846c