A third-party benchmark reports that CompactifAI-compressed Llama 3.1 8B consumes 30-39% less inference energy than the full model with roughly comparable Ragas quality scores, based on 104 self-designed questions.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Accuracy and Consumption analysis from a compressed model by CompactifAI from Multiverse Computing
A third-party benchmark reports that CompactifAI-compressed Llama 3.1 8B consumes 30-39% less inference energy than the full model with roughly comparable Ragas quality scores, based on 104 self-designed questions.