Keeping two or three spike-heavy projections in FP16 or FP8 while quantizing the rest to 8 bits improves LLaMA-family quantization over SmoothQuant in the per-tensor regime.
Advances in Neural Information Processing Systems36, 34278–34294 (2023)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models
Keeping two or three spike-heavy projections in FP16 or FP8 while quantizing the rest to 8 bits improves LLaMA-family quantization over SmoothQuant in the per-tensor regime.