A third-party benchmark reports that CompactifAI-compressed Llama 3.1 8B consumes 30-39% less inference energy than the full model with roughly comparable Ragas quality scores, based on 104 self-designed questions.
https:// multiversecomputing.com/compactifai
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Accuracy and Consumption analysis from a compressed model by CompactifAI from Multiverse Computing
A third-party benchmark reports that CompactifAI-compressed Llama 3.1 8B consumes 30-39% less inference energy than the full model with roughly comparable Ragas quality scores, based on 104 self-designed questions.