Pith. sign in

Paper Citation Record · LEDGER

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache

As of 24 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2505.18231.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18231 v3

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:42:23.114204Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved34
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 04021691-f782-4ee0-911a-c5d942775906 · outbound

This paper cites GPT-4 Technical Report.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:19.441861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:19.441861Z digest=sha256:5acc65472fce79aa5be9956374303baa354d216b43d7240bc307394e463dcbd5

Observation c064fde5-75fc-4a2e-b708-2129e03f7fad · outbound

This paper cites Quarot: Outlier-free 4-bit inference in rotated llms.Advances in Neural Information Processing Systems, 37:100213– 100240, 2025.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Quarot: Outlier-free 4-bit inference in rotated llms.Advances in Neural Information Processing Systems, 37:100213– 100240, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:26.876599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:19.520380Z digest=sha256:c19cbc69f6812561198ff3538e205d47ad54092e6d6eae700d5121405a2d935a

Observation d1864e14-059f-4226-8a0c-6a126f8d17bb · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache LongBench: A bilingual, multitask benchmark for long context understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:26.751190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:19.618799Z digest=sha256:65ac69b1685dee18097f540d973e06d6cb2e87f55a9802225b3d509d0b52a073

Observation 88c24244-3562-4128-9757-43225bc40b2e · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:19.668535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:19.668535Z digest=sha256:b2fb601b2aec758053b4c0789cfae8673472fcef86c2d5a7626118a7eff59846

Observation ba5f5ffe-9949-4414-a9f7-cecf98a97fd4 · outbound

This paper cites Palu: Kv- cache compression with low-rank projection.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Palu: Kv- cache compression with low-rank projection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:26.582959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:19.709613Z digest=sha256:31912028d774af3e67fb702b5923ebafc0aa130a4f17999b61d75bdc8c91341a

Observation 91388cae-ec1a-4f50-b9c5-1d2c76dc15a4 · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:19.764346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:19.764346Z digest=sha256:98dc76dc3a5682a5ce6d4f35a2b86e0991743edeaba8eeb5dde3e9bf25f0db39

Observation be3bddb7-379b-45c4-ad85-3dab61ece3b6 · outbound

This paper cites PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:19.813220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:19.813220Z digest=sha256:3803c80cf4452a0a0bf545cfd79771400b3f0713e32439003654049914cc8509

Observation 589311a8-858a-4cf4-86fd-0d9430dc88b7 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:19.966304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:19.966304Z digest=sha256:2acd4375c9da885ef5bb8cd9852a2fb93033d2b52f6d8ac226f7436165533e4d

Observation 523759e0-e41f-47d2-b35d-306f8731b9da · outbound

This paper cites SDR: Efficient Neural Re-ranking using Succinct Document Representation.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache SDR: Efficient Neural Re-ranking using Succinct Document Representation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:42:23.794083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:20.113131Z digest=sha256:c7c99f498454128db7a01515af8c82ec7563bbebf070dc97a3cd3e9c17c8ef2e

Observation e6527f5f-df86-49c4-b3d3-88fdb0cc812f · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache QLoRA: Efficient Finetuning of Quantized LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.170685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.170685Z digest=sha256:ff3e585eece79abd50d7e6cdfd68ee4f2530794a51d869f6771f314a3a84bdf9

Observation daa42476-cae2-4cc7-a772-afad805be48b · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Extreme Compression of Large Language Models via Additive Quantization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.242101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.242101Z digest=sha256:bfe6091d2f40c22bc92a005e4bf43cd9cf7578e2630016da11af0e8766b81bb7

Observation 20dde1d7-2727-4da2-8f76-345abae103fd · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.321754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.321754Z digest=sha256:7c1d2cff7496b36824ffc03bc75696a001113f507e811cfb01c7d16841cd2f5a

Observation 7b40e5c0-9aa5-4cf1-b909-8b9a77f3007a · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache A framework for few-shot language model evaluation, 07 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:26.414042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:20.403085Z digest=sha256:b9dd36256a0008b5b5a25468a230252747977fd31aea7f16d0b59d3fd0a1b867

Observation 3b8af31b-88bf-48d7-8a27-2733d47331b5 · outbound

This paper cites The Llama 3 Herd of Models.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.499175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.499175Z digest=sha256:50852fee54ce07b170c157444f8ad4c5d1f5f5e79b74f9c9c56efc3923905128

Observation 0e792160-6fcd-460f-ab99-611bae593cf0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.623145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.623145Z digest=sha256:016d1f20b89266bfccb16bc2a97721dd0cd4e68e8b58acc50045ec5392439c0a

Observation 96eb77fc-124f-4476-92df-21cf9a40f843 · outbound

This paper cites Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions.SIAM review, 53(2):217–288, 2011.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions.SIAM review, 53(2):217–288, 2011

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.725384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.725384Z digest=sha256:66772a1ef5a9787a48e58e872cde7d6bb67f822873e12555e991fca09b5134d4

Observation 735a74b6-9c5a-4958-af5b-36956bc96a53 · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.783045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.783045Z digest=sha256:e2e9cded86705519781c1ab5c0e8767eb4ad166722b633f87435c85536c841ae

Observation 4a3d52a4-87ed-4ce3-9d03-863457c7e3e0 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Measuring Massive Multitask Language Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.854098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.854098Z digest=sha256:778662dcc714c8e1045c3f0958522ce3c9887ab292a86307ddbc0298068e2044

Observation 5ca32f2e-10e1-42c3-a799-bdc70c0a2837 · outbound

This paper cites Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.904104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.904104Z digest=sha256:c4203ae96dc74b8a00362fd962c53f4b9c5f505eb5a064b7cd2ce79cdb13484c

Observation 67e9cea5-60b0-4f83-a1c7-19a6a6315861 · outbound

This paper cites OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.965959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.965959Z digest=sha256:01c0c9fd7f0b649e486abcaec27574d554dcd01324981e0fd0cf88adbaa8415d

Observation 0ffa03f2-aeb4-406d-82c3-eb69fd9ac9d4 · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.019251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.019251Z digest=sha256:e658c2992cf5fe5e122fc87d942261ba96ae8134508cc174b34c21489c5ebe4b

Observation c33b91b4-7076-4fed-839a-2d4ea3d92494 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache SqueezeLLM: Dense-and-Sparse Quantization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.069575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.069575Z digest=sha256:44536c618b866387b92dc2dba00162b99d5f3059e6419ba08ff4ebab3b6db4af

Observation 510a1e5d-fb85-4824-9e6c-b785a122dce8 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Efficient memory management for large language model serving with pagedattention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.150721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.150721Z digest=sha256:7ba8579a8af53d52e174533e0b002d052d061604f3b850478fedcab3d116be43

Observation 93600400-dce6-4a09-8da9-685a63ff55d3 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.194134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.194134Z digest=sha256:3d22a43c06eb0facc298d5c5e7b509ae0d43527061267289aa4f3233ac4eb5ed

Observation 1d7e3402-796f-45cc-8aba-8a6650c27255 · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970, 2024.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.275247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.275247Z digest=sha256:79672dd92dcb13c47460aa9751e07e6b551f2d11492ac21471d5941eae6b39f8

Observation e30a4de3-986a-4d57-bf11-6c0d9001326c · outbound

This paper cites MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.377519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.377519Z digest=sha256:d156dba39cbf9f7a657228d5d22b7c964ad358246650f77dea9588a9d4b0d98f

Observation 7ff29666-6aef-41c8-840f-82f242a95a2f · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:26.198825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:21.476258Z digest=sha256:ac586c9752d901e98888e607f0f80e3ce2840ae3d720bf6c695a126f880e99cf

Observation 0c0a7fbb-d6e8-41a6-9cc4-bf018e610764 · outbound

This paper cites VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.528632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.528632Z digest=sha256:7254ca2487fadfe121644aaaae8567c792fced20ffa613a373acfb8f3d2ffbee

Observation 0d3692ca-f0bf-41ec-ae6a-a95c3ae19497 · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache SpinQuant: LLM quantization with learned rotations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.563050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.563050Z digest=sha256:21b9e6f18506f3c0e8e9963dcdc6e1fe6850b6c7ef4c7e69426453dbec30255c

Observation 86282d67-e22d-4b85-a153-824487fbbdd5 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.631333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.631333Z digest=sha256:673ac959cb72eea95f986f4412fbb202bd51f75d9f8ae4e8be7d37f6d97a1719

Observation ef42813c-e02f-4efe-a7e1-e6948ece603f · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:25.985348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:21.706702Z digest=sha256:397355f0bcfdb064e8feb7efe8ce25355cb3870523fc5554d95c705bdcc594c7

Observation 8a2d3072-513e-4c5a-8db2-9ed89116cbc4 · outbound

This paper cites Cake: Cascading and adaptive kv cache eviction with layer preferences.arXiv preprint arXiv:2503.12491, 2025.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Cake: Cascading and adaptive kv cache eviction with layer preferences.arXiv preprint arXiv:2503.12491, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.756963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.756963Z digest=sha256:0ebe6501f81b2e2bd97dfa9e1b2232a13d939accf73a02e86833f5fee235c24d

Observation 4c80056c-736e-482b-884d-e635ae7bda2a · outbound

This paper cites Coqa: A conversational question answering challenge.Transactions of the Association for Computational Linguistics, 7:249–266, 2019.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Coqa: A conversational question answering challenge.Transactions of the Association for Computational Linguistics, 7:249–266, 2019

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.846993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.846993Z digest=sha256:dec74e90b6913d0cd659ca9b31565c023521fa38714f29ba5bf51d1d9bd58de1

Observation 93d8bf7c-4959-4bc8-8d67-58a6147bd157 · outbound

This paper cites Hanson-wright inequality and sub-gaussian concentra- tion.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Hanson-wright inequality and sub-gaussian concentra- tion

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:25.738176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:21.943014Z digest=sha256:1d20dab6edab7340b5f3264015abb28f5c91dd7901d48c5ad510df8e849804ba

Observation 8a99de4a-d23d-4050-850c-d43e52fedb8c · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:25.616849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:22.054272Z digest=sha256:1875dc5cd9b791af70e096bf92441f8399e2e00514c299f7a8f7ee7976d5bc43

Observation 9e15ebc8-9549-4bb6-a683-c712d4cb7235 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:22.125724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:22.125724Z digest=sha256:f783dd00137193a15bd17444de6577d3a17f19fe665ea320e17a4d470f3b1deb

Observation b0a527a4-b230-468f-bd55-1992104ff61b · outbound

This paper cites QuIP#: Even better LLM quantization with hadamard incoherence and lattice codebooks.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache QuIP#: Even better LLM quantization with hadamard incoherence and lattice codebooks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:25.408508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:22.217465Z digest=sha256:7e00e9409cd44bc10182ae37ba141137bf66ce618e4421d6e1a7b2f5d734367d

Observation 07b96d21-7971-43bc-ac17-0773479f1a63 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:22.302830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:22.302830Z digest=sha256:150a57ae59c9b328004daafda35a467201290aeb3a183c55835fbd316cb09ece

Observation 50e1376b-fe23-4b48-b5f8-6a7a9f9b03e5 · outbound

This paper cites BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:22.378319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:22.378319Z digest=sha256:1fac85c8a1362e2d28f205d3a2c0d8ac998934125159dd0256997169ec4d2778

Observation 487e206a-3751-47e3-9ec8-a14135340265 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Chain-of-thought prompting elicits reasoning in large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:22.452938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:22.452938Z digest=sha256:7faaeaff8a6b6befb16adb2f9913a7f65cc331a37774c656c7f5594ed19ec50f

Observation d96754ce-5019-4912-aa00-8454d5ad2387 · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:22.512307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:22.512307Z digest=sha256:c416b74a0a003d1b5ab82842792bc680f8fa111ec30f0703a9aaa7d4f6b28c6f

Observation 87d1acc8-019d-4a99-bedb-7d94516415f0 · outbound

This paper cites Benchmarking the Reliability of Post-training Quantization: a Particular Focus on Worst-case Performance.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Benchmarking the Reliability of Post-training Quantization: a Particular Focus on Worst-case Performance

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:42:23.446611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:22.560543Z digest=sha256:27212d622557d3528415ef684b0e7a1b758e8337121928cc410c3fd29189a00c

Observation 96f64d5c-4832-4640-93ca-2772e53b30f5 · outbound

This paper cites Kv cache is 1 bit per channel: Efficient large language model inference with coupled quantization.Advances in Neural Information Processing Systems, 37:3304–3331, 2024.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Kv cache is 1 bit per channel: Efficient large language model inference with coupled quantization.Advances in Neural Information Processing Systems, 37:3304–3331, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:25.245202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:22.657980Z digest=sha256:a0dc31a87a5ebfce47b1cddba3a9fe7ce27ddd1ca56b97067794d998632bda97

Observation b3f8a2f2-0842-4cc6-93e4-361c804c15e6 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:25.047251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:22.696914Z digest=sha256:518842802c4f9b78e92d9a00010515d59949b701c812ccade398933e9d0d8bb7

Observation 4c769c61-a318-4598-9404-46aaa563c585 · outbound

This paper cites Decompose Σ = Cov(X) =D+A , D= diag(Σ),A ii = 0,∥A∥ F ≤Γ.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Decompose Σ = Cov(X) =D+A , D= diag(Σ),A ii = 0,∥A∥ F ≤Γ

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:24.860495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:22.744073Z digest=sha256:01acad1d3000e6907bee2ce18bf3a4aa33513954b519a4a4f7ecf543da4a1ce0

Observation 005a9e7c-b858-409e-88b6-4d08833ba59f · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:24.627060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:22.835376Z digest=sha256:12d2d2ee1a789ae53a8219acdd59ad493fff1797c3d95fc362a6734591156789

Observation 9ae2d4ba-a8d4-4bcd-a9df-c8af1065cd99 · outbound

This paper cites Then hTDh= 1−¯ε, f(h) :=h TAh= 1 d sTAs, and Var(Yi) = (1−¯ε) +f(h).

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Then hTDh= 1−¯ε, f(h) :=h TAh= 1 d sTAs, and Var(Yi) = (1−¯ε) +f(h)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:24.481106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:22.891916Z digest=sha256:4c4e2f429619c46bd9d76d0dafc7b71be29b46a8bb7897d3b62424f44b145efb

Observation 3ff43d2c-dac1-43d0-9560-f3cb61434e5a · outbound

This paper cites Applying the Hanson-Wright inequality [34], for anyu >0 Pr |sTAs|> u ≤2 exp −c u2/Γ2 wherecis a universal constant.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Applying the Hanson-Wright inequality [34], for anyu >0 Pr |sTAs|> u ≤2 exp −c u2/Γ2 wherecis a universal constant

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:24.295084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:22.978513Z digest=sha256:36f65448a17abc651f31b6ed53dfcc3bf83d4738833f5ef0545c7765baa5fb52

Observation 94c89684-f3fd-4a5a-9186-136648659cb3 · outbound

This paper cites Exponent becomes −ln(2/α) ; hence Pr(|f(h)|> t)≤α.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Exponent becomes −ln(2/α) ; hence Pr(|f(h)|> t)≤α

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:24.086439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:23.036088Z digest=sha256:2651fdb754a7821686f755567469dea86ec100b7cbbced1f1f4e841d912fac8b

Observation 9145d042-9aea-4cae-9d61-0018c71de493 · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 50

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:42:23.342602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:42:23.114204Z digest=sha256:c46948dbdf1ee0a2e1e11679f0f2864bd9f0fa0c13c777fe7c14f50676dd2904

Pith citing papers

No inbound Pith citation observations are available.