Pith. sign in

Accuracy on the wrong line: On the pitfalls of noisy data for out-of-distribution generalisation

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

"Accuracy-on-the-line" is a widely observed phenomenon in machine learning, where a model's accuracy on in-distribution (ID) and out-of-distribution (OOD) data is positively correlated across different hyperparameters and data configurations. But when does this useful relationship break down? In this work, we explore its robustness. The key observation is that noisy data and the presence of nuisance features can be sufficient to shatter the Accuracy-on-the-line phenomenon. In these cases, ID and OOD accuracy can become negatively correlated, leading to "Accuracy-on-the-wrong-line". This phenomenon can also occur in the presence of spurious (shortcut) features, which tend to overshadow the more complex signal (core, non-spurious) features, resulting in a large nuisance feature space. Moreover, scaling to larger datasets does not mitigate this undesirable behavior and may even exacerbate it. We formally prove a lower bound on Out-of-distribution (OOD) error in a linear classification model, characterizing the conditions on the noise and nuisance features for a large OOD error. We finally demonstrate this phenomenon across both synthetic and real datasets with noisy data and nuisance features.

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Consensus-Driven Active Model Selection

cs.LG · 2025-07-31 · conditional · novelty 7.0

CODA uses consensus-based priors and Bayesian updating to select the best candidate model with far fewer labels than prior active model selection methods, beating them on 18 of 26 benchmark tasks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Consensus-Driven Active Model Selection cs.LG · 2025-07-31 · conditional · none · ref 57 · internal anchor

    CODA uses consensus-based priors and Bayesian updating to select the best candidate model with far fewer labels than prior active model selection methods, beating them on 18 of 26 benchmark tasks.