Arm-specific adjustment with a pre-chosen early-training log statistic narrows confidence intervals for model performance differences at sufficient run budgets, while broad automatic selection from the log pool adds noise.
Accounting for variance in machine learning benchmarks
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Can Training Logs Make Model Comparisons More Precise?
Arm-specific adjustment with a pre-chosen early-training log statistic narrows confidence intervals for model performance differences at sufficient run budgets, while broad automatic selection from the log pool adds noise.