Pith. sign in

REVIEW 2 cited by

NF4 Isn't Information Theoretically Optimal (and that's Good)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.06965 v2 pith:D3G3THXH submitted 2023-06-12 cs.LG

classification cs.LG
keywords blockimprovedinformationoptimalquantizationsizestheoreticallyabsmax-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This note shares some simple calculations and experiments related to absmax-based blockwise quantization, as used in Dettmers et al., 2023. Their proposed NF4 data type is said to be information theoretically optimal for representing normally distributed weights. I show that this can't quite be the case, as the distribution of the values to be quantized depends on the block-size. I attempt to apply these insights to derive an improved code based on minimizing the expected L1 reconstruction error, rather than the quantile based method. This leads to improved performance for larger quantization block sizes, while both codes perform similarly at smaller block sizes.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pushing the Limits of Large Language Model Quantization via the Linearity Theorem

    cs.LG 2024-11 conditional novelty 7.0 of 10

    A new theorem and method (HIGGS) make per-layer quantization error a reliable predictor of final model perplexity, enabling state-of-the-art data-free and dynamic bit-width LLM compression.

  2. any4: Learned 4-bit Numeric Representation for LLMs

    cs.LG 2025-07 conditional novelty 5.0 of 10

    any4 learns a per-row 16-value codebook for 4-bit LLM weight quantization via activation-weighted k-means, beating int4/fp4/nf4 on perplexity and matching preprocessing methods like AWQ and GPTQ.

Pith tools