A 25% pruning plus 4-bit quantization configuration retains roughly 20% more benchmark performance than 3-bit quantization alone at matching theoretical compression rates, across two LLMs.
GPTQ: Accurate post-training quantization for generative pre-trained transformers
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Semantic Retention and Extreme Compression in LLMs: Can We Have Both?
A 25% pruning plus 4-bit quantization configuration retains roughly 20% more benchmark performance than 3-bit quantization alone at matching theoretical compression rates, across two LLMs.