Pith. sign in

Paper Citation Record · LEDGER

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache

As of 8 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2505.18231.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18231 v3

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:42:23.114204Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved34
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 04021691-f782-4ee0-911a-c5d942775906 · outbound

This paper cites GPT-4 Technical Report.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:19.441861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:19.441861Z digest=sha256:805ec22f70a341a187070e5e630fa8d3422c79169b59bf1303bb794d3a9f8a8a

Observation c064fde5-75fc-4a2e-b708-2129e03f7fad · outbound

This paper cites Quarot: Outlier-free 4-bit inference in rotated llms.Advances in Neural Information Processing Systems, 37:100213– 100240, 2025.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Quarot: Outlier-free 4-bit inference in rotated llms.Advances in Neural Information Processing Systems, 37:100213– 100240, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:26.876599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:19.520380Z digest=sha256:c3394ffb2f0dc8051f9d5d60e58ad8bcfc7c7e665181f70ddc3c15644eea61cb

Observation d1864e14-059f-4226-8a0c-6a126f8d17bb · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache LongBench: A bilingual, multitask benchmark for long context understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:26.751190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:19.618799Z digest=sha256:b5dcc672102c3e3b5c2bc660762dff6c0c8be8e81d2afcc14389fc92713d0ac2

Observation 88c24244-3562-4128-9757-43225bc40b2e · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:19.668535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:19.668535Z digest=sha256:0b431fc3590e0487fc584e04f874eb25722a10aaec72be6a62ef18be49515715

Observation ba5f5ffe-9949-4414-a9f7-cecf98a97fd4 · outbound

This paper cites Palu: Kv- cache compression with low-rank projection.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Palu: Kv- cache compression with low-rank projection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:26.582959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:19.709613Z digest=sha256:22e0a6d62aa7b5e311b3c0b280750a2a169b754d8877d1b8efd1d710a181a26a

Observation 91388cae-ec1a-4f50-b9c5-1d2c76dc15a4 · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:19.764346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:19.764346Z digest=sha256:2b2d097e0dc9623b55bbfd09b55cad3703b6b23bc4fcd77de441f644d874d090

Observation be3bddb7-379b-45c4-ad85-3dab61ece3b6 · outbound

This paper cites PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:19.813220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:19.813220Z digest=sha256:df61b5b0cd964278b4c926e4e25fd77a2f71bf3e79225427fd9883162c86bdd8

Observation 589311a8-858a-4cf4-86fd-0d9430dc88b7 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:19.966304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:19.966304Z digest=sha256:700415a2a99856feed7afa7889e805c42e83168e38653529680ed18803648ad8

Observation 523759e0-e41f-47d2-b35d-306f8731b9da · outbound

This paper cites SDR: Efficient Neural Re-ranking using Succinct Document Representation.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache SDR: Efficient Neural Re-ranking using Succinct Document Representation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:42:23.794083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:20.113131Z digest=sha256:c9bfbf9668668d623dd254f318d99684605c1a75e950c17147ddd5c9f5a1f211

Observation e6527f5f-df86-49c4-b3d3-88fdb0cc812f · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache QLoRA: Efficient Finetuning of Quantized LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.170685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.170685Z digest=sha256:ec0ec6161c634ab008af533fafa27b78594037bbfc413136abeb82521e7c0e44

Observation daa42476-cae2-4cc7-a772-afad805be48b · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Extreme Compression of Large Language Models via Additive Quantization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.242101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.242101Z digest=sha256:b32e1bc70861d353fee446c660f2f8a58ab2ac7ca3ed6a05058adc923cafc51e

Observation 20dde1d7-2727-4da2-8f76-345abae103fd · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.321754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.321754Z digest=sha256:5ebcd954373b714d091f6143c43b9fa06104fc2d64f6c0888120f10f18530b78

Observation 7b40e5c0-9aa5-4cf1-b909-8b9a77f3007a · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache A framework for few-shot language model evaluation, 07 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:26.414042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:20.403085Z digest=sha256:e3f8514de70dbe6a1633ad0d3698a08fc736cd73f003684670403ae86e892620

Observation 3b8af31b-88bf-48d7-8a27-2733d47331b5 · outbound

This paper cites The Llama 3 Herd of Models.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.499175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.499175Z digest=sha256:fc0944710e69ba796a803775f9ddad0e56660df26e571bd3541d42a6904d86ab

Observation 0e792160-6fcd-460f-ab99-611bae593cf0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.623145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.623145Z digest=sha256:685b1710a4664b51f04dec45b78412545a18f04db6e8f0c3ae9d952624be8c64

Observation 96eb77fc-124f-4476-92df-21cf9a40f843 · outbound

This paper cites Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions.SIAM review, 53(2):217–288, 2011.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions.SIAM review, 53(2):217–288, 2011

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.725384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.725384Z digest=sha256:f9dbba9ec573fcc8d0a64a4398f6b8edd2460c3e99f8ea9d9d134980424f0738

Observation 735a74b6-9c5a-4958-af5b-36956bc96a53 · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.783045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.783045Z digest=sha256:2417fe163f7632089b6fd6b9c4985daa42dcfbfc2360bb19497839ec71622876

Observation 4a3d52a4-87ed-4ce3-9d03-863457c7e3e0 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Measuring Massive Multitask Language Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.854098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.854098Z digest=sha256:85ccc94c6212f89888f524ad3f49db6b7513b5f4c91ea575b5e93b3d89f711d2

Observation 5ca32f2e-10e1-42c3-a799-bdc70c0a2837 · outbound

This paper cites Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.904104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.904104Z digest=sha256:fa0be169d5172188ddece628f11e5f20ec7031453c1d2e206f45150ddec47317

Observation 67e9cea5-60b0-4f83-a1c7-19a6a6315861 · outbound

This paper cites OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.965959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.965959Z digest=sha256:6ff72bfe6e9950f0b73ad34a68c462c98d4823c19f2f03d30a76acc4b1248a71

Observation 0ffa03f2-aeb4-406d-82c3-eb69fd9ac9d4 · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.019251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.019251Z digest=sha256:fbad6a65ca42d402bbfcbca042659dc07f9d7036da61043a96695d5ca38d7038

Observation c33b91b4-7076-4fed-839a-2d4ea3d92494 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache SqueezeLLM: Dense-and-Sparse Quantization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.069575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.069575Z digest=sha256:5fbcf80f0fc372ece7204a886c27f3313dd9e74b4c10ac6d5ad970b04962aeff

Observation 510a1e5d-fb85-4824-9e6c-b785a122dce8 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Efficient memory management for large language model serving with pagedattention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.150721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.150721Z digest=sha256:b30c358b8032a90e6426e0aa742378430af5eee00b89f939d0493c62a1aef73c

Observation 93600400-dce6-4a09-8da9-685a63ff55d3 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.194134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.194134Z digest=sha256:ccf45a8d076ae188b88cf3a83cc80363c6294485c6a9ed0c9d1f1c43930ce1c0

Observation 1d7e3402-796f-45cc-8aba-8a6650c27255 · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970, 2024.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.275247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.275247Z digest=sha256:efcbac85268ada087849eb0aabed1700927bbee6aa0e0d0593a143984a00d13b

Observation e30a4de3-986a-4d57-bf11-6c0d9001326c · outbound

This paper cites MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.377519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.377519Z digest=sha256:1d7050675efeb8ef644ffc45ed678eb6bbda0f90f5ac55eb39b1286e05479abc

Observation 7ff29666-6aef-41c8-840f-82f242a95a2f · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:26.198825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:21.476258Z digest=sha256:7970b9d5038deab95d9f56acda004bb12b1e6335fc438bdbce084bd71fadfef9

Observation 0c0a7fbb-d6e8-41a6-9cc4-bf018e610764 · outbound

This paper cites VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.528632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.528632Z digest=sha256:7bcd405de94f62105e7232eb7fb6546a54513e320e327aa4a4a62fb9b75d425d

Observation 0d3692ca-f0bf-41ec-ae6a-a95c3ae19497 · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache SpinQuant: LLM quantization with learned rotations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.563050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.563050Z digest=sha256:ee311e0b0ead8fbf00db0e1300dfa8120ff8abf5045c6148d17ce56247251b92

Observation 86282d67-e22d-4b85-a153-824487fbbdd5 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.631333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.631333Z digest=sha256:366a0f59c49d30aa67e64b56f265119f7b0283137e38aff78570f27ffaad8af8

Observation ef42813c-e02f-4efe-a7e1-e6948ece603f · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:25.985348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:21.706702Z digest=sha256:bca215eb2163756b651f8089fab47c76eb360e8d4779b202ab1ac799b1cbc32c

Observation 8a2d3072-513e-4c5a-8db2-9ed89116cbc4 · outbound

This paper cites Cake: Cascading and adaptive kv cache eviction with layer preferences.arXiv preprint arXiv:2503.12491, 2025.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Cake: Cascading and adaptive kv cache eviction with layer preferences.arXiv preprint arXiv:2503.12491, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.756963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.756963Z digest=sha256:0e972f02d000b4db814211f5895dd35271b56eced77a5310876dba37734fc4ad

Observation 4c80056c-736e-482b-884d-e635ae7bda2a · outbound

This paper cites Coqa: A conversational question answering challenge.Transactions of the Association for Computational Linguistics, 7:249–266, 2019.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Coqa: A conversational question answering challenge.Transactions of the Association for Computational Linguistics, 7:249–266, 2019

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.846993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.846993Z digest=sha256:116f11139582e410fc62ae9d73fc7a1db4f571b0ac52289435e61329d795dd69

Observation 93d8bf7c-4959-4bc8-8d67-58a6147bd157 · outbound

This paper cites Hanson-wright inequality and sub-gaussian concentra- tion.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Hanson-wright inequality and sub-gaussian concentra- tion

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:25.738176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:21.943014Z digest=sha256:1c77b3026528d3318ac997dc21e1c8437126107d32f3018c8426bde96896513a

Observation 8a99de4a-d23d-4050-850c-d43e52fedb8c · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:25.616849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:22.054272Z digest=sha256:47aa6eafbabed27ed811f0ce4d11d1636f9ff10d8570999b5d08524cd19fa603

Observation 9e15ebc8-9549-4bb6-a683-c712d4cb7235 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:22.125724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:22.125724Z digest=sha256:5fe457d962b14f664c6d7f9cef35449ea9616c050a4daea16d95ae872721d0d4

Observation b0a527a4-b230-468f-bd55-1992104ff61b · outbound

This paper cites QuIP#: Even better LLM quantization with hadamard incoherence and lattice codebooks.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache QuIP#: Even better LLM quantization with hadamard incoherence and lattice codebooks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:25.408508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:22.217465Z digest=sha256:854db7b7acc2754c2dfe4e5fed20444066910ff8633eeb9b2f4794358200d1b6

Observation 07b96d21-7971-43bc-ac17-0773479f1a63 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:22.302830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:22.302830Z digest=sha256:ea3a4d0c09e7069ba4ea4a7bdd2fc950bf433d52a20dd6888d4f1a754a5fbef0

Observation 50e1376b-fe23-4b48-b5f8-6a7a9f9b03e5 · outbound

This paper cites BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:22.378319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:22.378319Z digest=sha256:d2dc41599f2b7b663dc97790c86097962e46a121a5e8c4601b8f2e91df7233f3

Observation 487e206a-3751-47e3-9ec8-a14135340265 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Chain-of-thought prompting elicits reasoning in large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:22.452938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:22.452938Z digest=sha256:62a54fb2abfe05c2e5f0e2e1e09d4ad331554c3894bdeec1147f545d92e535fe

Observation d96754ce-5019-4912-aa00-8454d5ad2387 · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:22.512307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:22.512307Z digest=sha256:293b23c29512dda2bcf79d6e7e4a7b774c82c0a3d0f40fe13b090d2ce22214b3

Observation 87d1acc8-019d-4a99-bedb-7d94516415f0 · outbound

This paper cites Benchmarking the Reliability of Post-training Quantization: a Particular Focus on Worst-case Performance.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Benchmarking the Reliability of Post-training Quantization: a Particular Focus on Worst-case Performance

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:42:23.446611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:22.560543Z digest=sha256:398eab397c9c7956ed456660239afd73f2ac344671fc1390aa5b2d4d8fdbf1af

Observation 96f64d5c-4832-4640-93ca-2772e53b30f5 · outbound

This paper cites Kv cache is 1 bit per channel: Efficient large language model inference with coupled quantization.Advances in Neural Information Processing Systems, 37:3304–3331, 2024.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Kv cache is 1 bit per channel: Efficient large language model inference with coupled quantization.Advances in Neural Information Processing Systems, 37:3304–3331, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:25.245202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:22.657980Z digest=sha256:51944c1e906d7ab50671cc13dd78774725f7445e97f27717cb9e55816ebe8b18

Observation b3f8a2f2-0842-4cc6-93e4-361c804c15e6 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:25.047251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:22.696914Z digest=sha256:d9ea403f2787cbf5b6e551c77e037fe0d4c157e43c54003843d7edf2317178ee

Observation 4c769c61-a318-4598-9404-46aaa563c585 · outbound

This paper cites Decompose Σ = Cov(X) =D+A , D= diag(Σ),A ii = 0,∥A∥ F ≤Γ.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Decompose Σ = Cov(X) =D+A , D= diag(Σ),A ii = 0,∥A∥ F ≤Γ

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:24.860495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:22.744073Z digest=sha256:c8124178e5369887cc7662f67281f40c18fd059ea4206956d2a0fe06560d5bd3

Observation 005a9e7c-b858-409e-88b6-4d08833ba59f · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:24.627060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:22.835376Z digest=sha256:a6a575a6c909168c42294055f033686a957e4e1ca901c40b00a93d4e20c35243

Observation 9ae2d4ba-a8d4-4bcd-a9df-c8af1065cd99 · outbound

This paper cites Then hTDh= 1−¯ε, f(h) :=h TAh= 1 d sTAs, and Var(Yi) = (1−¯ε) +f(h).

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Then hTDh= 1−¯ε, f(h) :=h TAh= 1 d sTAs, and Var(Yi) = (1−¯ε) +f(h)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:24.481106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:22.891916Z digest=sha256:2ef27539592e98ab8b9d52ad153b6a8d89e14f10b43bc00078aa6190d0dab3bc

Observation 3ff43d2c-dac1-43d0-9560-f3cb61434e5a · outbound

This paper cites Applying the Hanson-Wright inequality [34], for anyu >0 Pr |sTAs|> u ≤2 exp −c u2/Γ2 wherecis a universal constant.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Applying the Hanson-Wright inequality [34], for anyu >0 Pr |sTAs|> u ≤2 exp −c u2/Γ2 wherecis a universal constant

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:24.295084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:22.978513Z digest=sha256:58e4397aa6d7a619f0ae8271062021da8b4c69ad200e9cc8981416c2cf4731f2

Observation 94c89684-f3fd-4a5a-9186-136648659cb3 · outbound

This paper cites Exponent becomes −ln(2/α) ; hence Pr(|f(h)|> t)≤α.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Exponent becomes −ln(2/α) ; hence Pr(|f(h)|> t)≤α

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:24.086439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:23.036088Z digest=sha256:684d1293f5256fc353f91a7f1d691062f61dff9a46c2d0894c4220587d1040da

Observation 9145d042-9aea-4cae-9d61-0018c71de493 · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 50

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:42:23.342602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:23.114204Z digest=sha256:474fd5fe4a9e6ffd7631e5905382fb389148cf250817b003e1271f9f1da830c7

Pith citing papers

No inbound Pith citation observations are available.