FMwork shows that sampling power-of-two batch sizes and capping output length at 128 tokens reproduces a full Llama 3.1 8B benchmark sweep within about 3-5% error at up to 24x lower experimental cost.
Llm-inference-bench: Inference benchmark- ing of large language models on ai accelerators
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
baseline 1
citation-polarity summary
fields
cs.PF 1years
2025 1verdicts
CONDITIONAL 1roles
baseline 1polarities
baseline 1representative citing papers
citing papers explorer
-
Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
FMwork shows that sampling power-of-two batch sizes and capping output length at 128 tokens reproduces a full Llama 3.1 8B benchmark sweep within about 3-5% error at up to 24x lower experimental cost.