Pith. sign in

With Shared Microexponents, A Little Shifting Goes a Long Way

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

This paper introduces Block Data Representations (BDR), a framework for exploring and evaluating a wide spectrum of narrow-precision formats for deep learning. It enables comparison of popular quantization standards, and through BDR, new formats based on shared microexponents (MX) are identified, which outperform other state-of-the-art quantization approaches, including narrow-precision floating-point and block floating-point. MX utilizes multiple levels of quantization scaling with ultra-fine scaling factors based on shared microexponents in the hardware. The effectiveness of MX is demonstrated on real-world models including large-scale generative pretraining and inferencing, and production-scale recommendation systems.

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

any4: Learned 4-bit Numeric Representation for LLMs

cs.LG · 2025-07-07 · conditional · novelty 5.0

any4 learns a per-row 16-value codebook for 4-bit LLM weight quantization via activation-weighted k-means, beating int4/fp4/nf4 on perplexity and matching preprocessing methods like AWQ and GPTQ.

citing papers explorer

Showing 1 of 1 citing paper.

  • any4: Learned 4-bit Numeric Representation for LLMs cs.LG · 2025-07-07 · conditional · none · ref 40 · internal anchor

    any4 learns a per-row 16-value codebook for 4-bit LLM weight quantization via activation-weighted k-means, beating int4/fp4/nf4 on perplexity and matching preprocessing methods like AWQ and GPTQ.