Inference backend choice alone can shift LLM benchmark scores by up to 0.055 mean absolute divergence and change which questions are answered correctly, even under greedy decoding.
InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 8086–8098
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
What We Observe as LLM Behavior Can Be a Side-effect of Inference Backend
Inference backend choice alone can shift LLM benchmark scores by up to 0.055 mean absolute divergence and change which questions are answered correctly, even under greedy decoding.