Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:42:23.114204Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2505.18231.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:42:23.114204Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 04021691-f782-4ee0-911a-c5d942775906 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c064fde5-75fc-4a2e-b708-2129e03f7fad · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Quarot: Outlier-free 4-bit inference in rotated llms.Advances in Neural Information Processing Systems, 37:100213– 100240, 2025
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1864e14-059f-4226-8a0c-6a126f8d17bb · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache LongBench: A bilingual, multitask benchmark for long context understanding
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88c24244-3562-4128-9757-43225bc40b2e · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba5f5ffe-9949-4414-a9f7-cecf98a97fd4 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Palu: Kv- cache compression with low-rank projection
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91388cae-ec1a-4f50-b9c5-1d2c76dc15a4 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be3bddb7-379b-45c4-ad85-3dab61ece3b6 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 589311a8-858a-4cf4-86fd-0d9430dc88b7 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 523759e0-e41f-47d2-b35d-306f8731b9da · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache SDR: Efficient Neural Re-ranking using Succinct Document Representation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6527f5f-df86-49c4-b3d3-88fdb0cc812f · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache QLoRA: Efficient Finetuning of Quantized LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daa42476-cae2-4cc7-a772-afad805be48b · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Extreme Compression of Large Language Models via Additive Quantization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20dde1d7-2727-4da2-8f76-345abae103fd · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b40e5c0-9aa5-4cf1-b909-8b9a77f3007a · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache A framework for few-shot language model evaluation, 07 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b8af31b-88bf-48d7-8a27-2733d47331b5 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e792160-6fcd-460f-ab99-611bae593cf0 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96eb77fc-124f-4476-92df-21cf9a40f843 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions.SIAM review, 53(2):217–288, 2011
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 735a74b6-9c5a-4958-af5b-36956bc96a53 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a3d52a4-87ed-4ce3-9d03-863457c7e3e0 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Measuring Massive Multitask Language Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ca32f2e-10e1-42c3-a799-bdc70c0a2837 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67e9cea5-60b0-4f83-a1c7-19a6a6315861 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ffa03f2-aeb4-406d-82c3-eb69fd9ac9d4 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c33b91b4-7076-4fed-839a-2d4ea3d92494 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache SqueezeLLM: Dense-and-Sparse Quantization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 510a1e5d-fb85-4824-9e6c-b785a122dce8 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Efficient memory management for large language model serving with pagedattention
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93600400-dce6-4a09-8da9-685a63ff55d3 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d7e3402-796f-45cc-8aba-8a6650c27255 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e30a4de3-986a-4d57-bf11-6c0d9001326c · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ff29666-6aef-41c8-840f-82f242a95a2f · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c0a7fbb-d6e8-41a6-9cc4-bf018e610764 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d3692ca-f0bf-41ec-ae6a-a95c3ae19497 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache SpinQuant: LLM quantization with learned rotations
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86282d67-e22d-4b85-a153-824487fbbdd5 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef42813c-e02f-4efe-a7e1-e6948ece603f · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a2d3072-513e-4c5a-8db2-9ed89116cbc4 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Cake: Cascading and adaptive kv cache eviction with layer preferences.arXiv preprint arXiv:2503.12491, 2025
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c80056c-736e-482b-884d-e635ae7bda2a · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Coqa: A conversational question answering challenge.Transactions of the Association for Computational Linguistics, 7:249–266, 2019
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93d8bf7c-4959-4bc8-8d67-58a6147bd157 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Hanson-wright inequality and sub-gaussian concentra- tion
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a99de4a-d23d-4050-850c-d43e52fedb8c · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e15ebc8-9549-4bb6-a683-c712d4cb7235 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0a527a4-b230-468f-bd55-1992104ff61b · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache QuIP#: Even better LLM quantization with hadamard incoherence and lattice codebooks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 07b96d21-7971-43bc-ac17-0773479f1a63 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Attention is all you need.Advances in neural information processing systems, 30, 2017
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50e1376b-fe23-4b48-b5f8-6a7a9f9b03e5 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 487e206a-3751-47e3-9ec8-a14135340265 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Chain-of-thought prompting elicits reasoning in large language models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96754ce-5019-4912-aa00-8454d5ad2387 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87d1acc8-019d-4a99-bedb-7d94516415f0 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Benchmarking the Reliability of Post-training Quantization: a Particular Focus on Worst-case Performance
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96f64d5c-4832-4640-93ca-2772e53b30f5 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Kv cache is 1 bit per channel: Efficient large language model inference with coupled quantization.Advances in Neural Information Processing Systems, 37:3304–3331, 2024
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3f8a2f2-0842-4cc6-93e4-361c804c15e6 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c769c61-a318-4598-9404-46aaa563c585 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Decompose Σ = Cov(X) =D+A , D= diag(Σ),A ii = 0,∥A∥ F ≤Γ
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 005a9e7c-b858-409e-88b6-4d08833ba59f · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ae2d4ba-a8d4-4bcd-a9df-c8af1065cd99 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Then hTDh= 1−¯ε, f(h) :=h TAh= 1 d sTAs, and Var(Yi) = (1−¯ε) +f(h)
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ff43d2c-dac1-43d0-9560-f3cb61434e5a · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Applying the Hanson-Wright inequality [34], for anyu >0 Pr |sTAs|> u ≤2 exp −c u2/Γ2 wherecis a universal constant
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 94c89684-f3fd-4a5a-9186-136648659cb3 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Exponent becomes −ln(2/α) ; hence Pr(|f(h)|> t)≤α
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9145d042-9aea-4cae-9d61-0018c71de493 · outbound
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.