Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:27.538376Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 7 inbound Pith citation observations for arXiv:2505.18610.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:27.538376Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:51:05.406333Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-05T05:50:43.580203Z
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2b20b506-b216-4e0d-b3f6-033e12c8755b · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs American invitational mathematics examination, 2025
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7675a78b-e5df-4b7b-9435-23cd6dc9daf0 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4984c2db-1f4b-4200-ab32-71651135d69a · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Extending Context Window of Large Language Models via Positional Interpolation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a6a85a1-b026-4b4d-8961-f89d3ad5f6f7 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Carnegie mellon informatics and mathematics competition, 2025
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d4cf4010-ac3c-41ca-b7c0-385bba73c91c · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs CVXPY: A Python-embedded modeling language for convex optimization.Journal of Machine Learning Research, 17(83):1–5, 2016
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bb28117-42f1-45b9-97d9-8d2d9dc4e9b6 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3f6d26e-29d0-45cf-959b-9bb1f42cc5f3 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Moa: Mixture of sparse attention for automatic large language model compression.arXiv preprint arXiv:2406.14909, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e8454ea-7e57-4706-a720-2c881b122819 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bfc87e1-e9ec-4e0e-a3d0-15c80281a929 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Quantize What Counts: More for Keys, Less for Values
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9240860a-f53e-4064-ac51-83e9689daa9e · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db4b4924-5a15-422f-b1e1-81dadb0a2338 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Llm-mq: Mixed-precision quantization for efficient llm deployment
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d2857518-6375-4364-a8f9-6808db1e86c6 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c239ce3-f9a5-45f6-832d-c4cbfe5c4158 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3701df9d-5b34-4c88-bae0-f5f8e4659a31 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05a9d736-233f-40be-a294-c05402c37922 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Intactkv: Improving large language model quantization by keeping pivot tokens intact
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 621798a3-de2e-4b6f-a659-4116715d871a · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a24314fe-0752-48a6-84ec-aac281827bef · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Introducing openai o1, September 2024
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1cdbaf85-b8b8-4cf4-a88c-d75c9c3d53f0 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Fast Transformer Decoding: One Write-Head is All You Need
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a617a6db-d533-463d-ab27-f19a9c9e765f · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c805816e-3b54-44e3-9dca-12eead72654f · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb5bc09-014c-4013-9ef9-91deb9bddd52 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f59fc53-bcd7-4d96-9798-ce6da8931fcc · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d5a41f4-3681-400c-9ac5-2705c36c194d · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Efficient Streaming Language Models with Attention Sinks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f56688f-352e-42ab-9a88-3b54731b5200 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c81bd392-b7a7-49c4-974a-27b5949d74d5 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc0885bd-c87e-4f3f-a23e-4ebf8585fb8b · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f54927b-1e40-4966-847c-e6c7f192e5b0 · outbound
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs Mixdq: Memory-efficient few-step text-to-image diffusion models with metric-decoupled mixed precision quantization
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation df4693bc-bdfd-4a6e-9a8a-910229384ab7 · inbound
PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddf4ad6e-87ee-4209-91f6-3b3c441c7aa6 · inbound
RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 194f9ca4-905b-4a57-a95f-2123d5b4d514 · inbound
RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2c0a9ec0-70f8-4829-95a0-c4ed6d93817e · inbound
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 201f393f-7216-4db9-8702-940eaa3cca96 · inbound
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a57804d3-3d5f-4d26-9817-433780185387 · inbound
RoPE-Aware Bit Allocation for KV-Cache Quantization PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2f48ec61-67e9-4b2a-8aa7-4acaa3ffe79d · inbound
High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.