A 25% pruning plus 4-bit quantization configuration retains roughly 20% more benchmark performance than 3-bit quantization alone at matching theoretical compression rates, across two LLMs.
Compressing LLMs: The truth is rarely pure and never simple
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Semantic Retention and Extreme Compression in LLMs: Can We Have Both?
A 25% pruning plus 4-bit quantization configuration retains roughly 20% more benchmark performance than 3-bit quantization alone at matching theoretical compression rates, across two LLMs.