SAAP scores LLM structures with a weighted fusion of two importance measures, prunes the most volatile units, and recovers performance with grouped quantized low-rank fine-tuning.
Pruning Large Language Models via Accuracy Predictor
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Large language models(LLMs) containing tens of billions of parameters (or even more) have demonstrated impressive capabilities in various NLP tasks. However, substantial model size poses challenges to training, inference, and deployment so that it is necessary to compress the model. At present, most model compression for LLMs requires manual design of pruning features, which has problems such as complex optimization pipeline and difficulty in retaining the capabilities of certain parts of the model.Therefore, we propose a novel pruning approach: firstly, a training set of a certain number of architecture-accuracy pairs is established, and then a non-neural model is trained as an accuracy predictor. Using the accuracy predictor to further optimize the search space and search, the optimal model can be automatically selected. Experiments show that our proposed approach is effective and efficient. Compared with the baseline, the perplexity(PPL) on Wikitext2 and PTB dropped by 9.48% and 5,76% respectively, and the average accuracy of MMLU increased by 6.28%.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2024 1verdicts
REJECT 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Adaptive Pruning for Large Language Models with Structural Importance Awareness
SAAP scores LLM structures with a weighted fusion of two importance measures, prunes the most volatile units, and recovers performance with grouped quantized low-rank fine-tuning.