A pruning-plus-quantization framework that uses Bayesian optimization to choose per-layer bit widths reports roughly 30 percent memory savings on 7B-13B LLMs at approximately unchanged zero-shot accuracy.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
QPruner: Probabilistic Decision Quantization for Structured Pruning in Large Language Models
A pruning-plus-quantization framework that uses Bayesian optimization to choose per-layer bit widths reports roughly 30 percent memory savings on 7B-13B LLMs at approximately unchanged zero-shot accuracy.