A third-party benchmark reports that CompactifAI-compressed Llama 3.1 8B consumes 30-39% less inference energy than the full model with roughly comparable Ragas quality scores, based on 104 self-designed questions.
(Under LLAMA 3.1 community license agreement https://huggingface.co/meta-llama/Llama-3.1-8B/blob/ main/LICENSE) · Hugging Face
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Accuracy and Consumption analysis from a compressed model by CompactifAI from Multiverse Computing
A third-party benchmark reports that CompactifAI-compressed Llama 3.1 8B consumes 30-39% less inference energy than the full model with roughly comparable Ragas quality scores, based on 104 self-designed questions.