A quantization framework using KLT-enhanced and smooth-fused rotations lets Mamba models run at 8-bit precision with near-full accuracy and at 4-bit weights with moderate loss.
Quip: 2-bit quantization of large language models with guarantees
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
A quantization framework using KLT-enhanced and smooth-fused rotations lets Mamba models run at 8-bit precision with near-full accuracy and at 4-bit weights with moderate loss.