A survey of knowledge distillation, quantization, and pruning for compressing LLMs to run on resource-constrained edge devices, with no new experimental results.
A Survey of Quantization Methods for Efficient Neural Network Inference,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques
A survey of knowledge distillation, quantization, and pruning for compressing LLMs to run on resource-constrained edge devices, with no new experimental results.