Model rankings in LLM evaluation are budget-dependent: the best-performing model changes with the token generation budget on all three benchmarks tested.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation
Model rankings in LLM evaluation are budget-dependent: the best-performing model changes with the token generation budget on all three benchmarks tested.