{"id":"189577bb-6b53-4e32-ab72-dad99a1a8b62","arxiv_id":"2412.05152","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A unifying taxonomy and formal definition that connects shortcut learning, spurious correlations, Clever Hans behavior, and confounders across detection, mitigation, and datasets.","lead":"This paper proposes a common vocabulary and an organizing taxonomy for the many names given to machine learning models that cheat by using unintended clues, such as \"shortcuts\", \"spurious correlations\", \"Clever Hans behavior\", and \"confounders\". It maps existing detection and mitigation methods into one framework and lists datasets built to study the problem.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The unifying formal definition depends on an unspecified relevance oracle, so shortcut status is not determinate without external task specification; the taxonomy survives, but the formal-unification claim is overstated.","rationale":"The reader's weakest_assumption correctly identifies that the relevant/spurious boundary is the load-bearing point of the formal definition. My stress-test sharpens this into a concrete determinacy problem: Frelevant is not derived from the task, data, or model, so the definition is a schema rather than a closed formalization. This matters because the paper's strongest claim is that a single formal definition can unify shortcut learning. The concern is partially mitigated by the authors' own acknowledgment in Sec. 5.5 and Sec. 8, but the abstract and taxonomy introduction do not carry that caveat prominently enough. The survey's independent value remains: it assembles a broad method map, connects related fields, and compiles datasets, and these contributions do not depend on the formal definition being fully operational. Therefore I would not reject the paper, but I would condition acceptance on reframing the definition as task-relative and explicitly stating that the taxonomy does not resolve contested relevance judgments. This is a modification to the central claim rather than to the survey's practical usefulness.","tokens_in":36118,"tokens_out":8990,"duration_ms":102243,"concrete_test":"Take a public toxicity or hate-speech dataset and train one fixed classifier. Specify two alternative Frelevant sets that are both consistent with the task label: (A) only semantic content features are relevant, and (B) semantic content plus explicit profanity tokens are relevant. For the same trained model, measure dependence on profanity tokens via feature ablation or input perturbation. Apply the Sec. 2 definition under each specification: under (A) the model exhibits a shortcut, under (B) it does not. If the classification flips while the model and data are unchanged, the formal definition is not determinate without a relevance oracle, confirming that the unification claim requires an external value judgment.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"In Sec. 2, the paper defines spurious correlations relative to a set Frelevant of features that are 'considered relevant to solve the task (in the intended way)'. No formal criterion, procedure, or oracle for obtaining Frelevant is provided; it is an external input to the definition. As a result, the definition cannot by itself classify any actual model behavior as a shortcut: two reasonable task specifications can yield opposite verdicts for the same model. For example, in hate speech detection, if profanity is included in Frelevant then reliance on profanity is legitimate, whereas if Frelevant is restricted to semantic intent, the same reliance is a shortcut. The authors acknowledge the difficulty of determining relevant versus spurious features in Sec. 5.5 and Sec. 8, but the abstract and Sec. 4 still present the definition as the unifying foundation of the taxonomy. Because the taxonomy categories in Sec. 5 and Sec. 6 are organized by methodology rather than derived from this definition, the survey's practical value is not destroyed; however, the central claim of providing a single formal definition that unifies the field is weaker than stated. The definition is better described as a definition schema parameterized by an externally supplied notion of relevance.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey and taxonomy paper on shortcut learning, spurious correlations, and confounders in machine learning. The authors propose a formal definition of shortcuts in terms of a ground-truth distribution, a distorted sampling distribution, a feature set, a correlation function, and a notion of 'relevant' features. They distinguish world-induced from sampling-induced spurious correlations, relate shortcuts to Clever Hans behavior, confounders in causality, distribution shift, bias, and adversarial backdoors, and then organize existing detection and mitigation methods into a two-level taxonomy. They also compile a table of datasets containing explicit spurious correlations and close with open challenges. The paper's central claim is that this is the first unified, general taxonomy of shortcut learning, supported by the formal definition and by the breadth of literature covered.","tokens_in":36303,"tokens_out":4649,"duration_ms":51066,"significance":"If the organizational claims hold, the paper is a valuable contribution: it provides a shared vocabulary for a fragmented field, a structured map of detection and mitigation methods across vision, language, medical imaging, and other domains, a useful comparison of prior surveys in Table 1, and a compendium of datasets in Table 4. The connections drawn to causality, fairness, and security are informative and largely accurate, and the discussion of open challenges (e.g., multiple shortcuts, generative models, pretraining-finetuning) is a useful agenda. The paper does not present new experimental results, so its contribution is primarily conceptual and organizational. The formal definition is best understood as a definition schema: it is parameterized by an externally supplied notion of relevant features and by an unspecified correlation function. The taxonomy itself does not depend on the formal definition being fully operational, because the categories in Sections 5 and 6 are organized by methodology, but the abstract and Section 4 present the definition as the unifying foundation, which overstates what the manuscript establishes.","major_comments":[{"comment":"The definition of a spurious correlation is parameterized by the externally supplied set F_relevant, characterized only as features 'considered relevant to solve the task (in the intended way)'. No criterion, procedure, or oracle for obtaining F_relevant is provided, and Sec. 2.1 itself notes that classifying by habitat rather than bird characteristics would change which features are relevant. Consequently, the definition alone cannot determine whether a given model's behavior is a shortcut; two reasonable task specifications can yield opposite verdicts for the same model. The paper acknowledges this difficulty in Sec. 5.5 and Sec. 8, but the abstract and Sec. 4 still present the definition as a formal foundation that 'unifies' the field. I recommend explicitly presenting the definition as a schema parameterized by a task-specific relevance specification, and moving this qualification into the abstract and Section 4.","section":"Sec. 2, 'Spurious Correlations and Shortcuts'"},{"comment":"The correlation function c: F x F -> [0,1] is left as an unspecified primitive, and the paper does not define what it means for a model to 'use' a correlation 'as the basis for its decision-making'. As written, the formal definition cannot be instantiated or tested on a concrete model without additional choices, such as a specific correlation measure on raw pixels or a behavioral criterion for 'reliance'. Please either specify at least one intended instantiation and a formal criterion for reliance, or explicitly state that the definition is conceptual rather than operational; the current text mixes both readings.","section":"Sec. 2, 'Features and Correlations'"},{"comment":"The 'perfect' shortcut category is defined inconsistently. Section 7 defines a perfect shortcut as one that 'occurs in only a single class and in all such samples', while Section 8 says perfect shortcuts are 'present in all samples'. These are different failure regimes: class-conditional presence versus global presence across all classes. The inconsistency affects the dataset classifications in Table 4 and the discussion of method limitations in Section 8, and it should be resolved by aligning the two definitions and stating which datasets in Table 4 fall into which regime.","section":"Sec. 7 vs. Sec. 8"}],"minor_comments":[{"comment":"The first line of the paper contains a spacing error: 'Na vigating Shortcuts' should be 'Navigating Shortcuts'.","section":"Title/first line"},{"comment":"The sentence 'when B is intervened. On.' is broken by a line break; it should read 'when B is intervened on.'","section":"Sec. 3.3"},{"comment":"The phrase 'more robust to shorcuts' is a typo and should read 'more robust to shortcuts'.","section":"Sec. 6.2.1"},{"comment":"The modality header 'Hyperspectal Vision' should be 'Hyperspectral Vision', and the P2S dataset size '2,3k' should use the same decimal convention as the rest of the table (e.g., '2.3k' or '2,300').","section":"Table 4"},{"comment":"Reference [140] (Qiu et al.) is missing a publication year and appears as '[n. d.]'; please complete the bibliographic entry.","section":"References"},{"comment":"The small text in the taxonomy figures is difficult to read in the preprint version; please ensure the final figures are legible at print resolution.","section":"Figures 5 and 6"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it is a solid, well-organized survey that actually delivers a usable map of shortcut detection and mitigation, plus a handy dataset compendium. Second, its central formal claim is weaker than the abstract suggests: the definition of a shortcut depends on an externally supplied set of relevant features, so it is a definition schema rather than a determinate criterion. The taxonomy itself does not depend on that formalism, and it survives the weakness.\n\nWhat is genuinely new: the authors bring together four related but fragmented research threads—shortcuts, spurious correlations, Clever Hans behavior, and confounders—under one taxonomy. They connect shortcut learning to distribution shift, bias, causality, and adversarial security, which is rarely done in one place. Their comparison table of prior surveys is fair and useful. The dataset table, with decoy vs. specialized distinctions and shortcut-strength categories, is a practical reference for anyone running experiments. The prose is clear, and the structure is easy to navigate.\n\nWhere the soft spots are: the formal definition in Sec. 2 relies on F_relevant, a set of features deemed relevant to the task, with no formal criterion or procedure for obtaining it. The authors know this—they acknowledge in Sec. 5.5 and Sec. 8 that deciding relevance is hard—but the abstract and Sec. 4 still present the definition as the unifying foundation. That is an overstatement. For any given model, two reasonable task specifications can classify the same behavior differently. The definition is a schema parameterized by a human- or task-supplied relevance judgment. The silver lining is that the taxonomy categories in Secs. 5–6 are organized by methodology, not derived from the formal definition, so the practical value holds. This is a moderate weakness in framing, not a load-bearing flaw.\n\nMinor quibble: the claim to be the first general taxonomy is a bit strong given Geirhos et al.'s earlier general survey, though this paper's scope and structure do go well beyond it. The citation pattern is fine—self-citations appear as method examples, not as evidence for the definition.\n\nBottom line: this paper deserves a serious referee. It is the kind of reference I would point a new student to, and I would cite it in work on shortcut mitigation. My recommendation: accept after minor revision, with a request to soften the formal-unification claim and explicitly state that the definition depends on an external relevance specification.","headline":"A genuinely useful survey and taxonomy of shortcut learning whose formal definition is best read as a schema—the taxonomy stands on its own merits.","tokens_in":747,"tokens_out":981,"would_cite":true,"duration_ms":25571,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that shortcut learning can be unified under one formal definition and a single taxonomy spanning detection, mitigation, and datasets.","keywords":["shortcut learning","spurious correlations","Clever Hans behavior","confounders","shortcut detection","shortcut mitigation","taxonomy","benchmark datasets"],"falsifier":"An annotation study on a standard benchmark: if human experts cannot reach agreement on a stable set of relevant features for a task such as ImageNet or a chest X-ray dataset, then the definition cannot decide whether a model's behavior is a shortcut, undermining the taxonomy's primary organizing criterion.","tokens_in":35923,"feed_emoji":"🧭","tokens_out":6211,"duration_ms":60892,"temperature":0.7,"pith_summary":"The paper sets out to unify the scattered literature on shortcuts, spurious correlations, Clever Hans behavior, and confounders under a single formal definition: a shortcut occurs when a model bases its decision on a spurious correlation rather than on relevant features. It claims that this definition, together with a distinction between world-induced and sampling-induced spurious correlations, is enough to organize the entire field into a taxonomy of detection and mitigation methods plus a curated set of benchmark datasets. A sympathetic reader would care because, if the taxonomy holds, methods developed under different names in different application areas become comparable and transferable, and open gaps such as handling multiple or perfect shortcuts and non-classification tasks become visible in one map. The paper is a survey and position statement rather than a new algorithm.","feed_headline":"One definition unifies shortcut-learning research","feed_subtitle":"A new taxonomy maps detection and mitigation methods and shows where the field has gaps.","key_machinery":"The machinery is the formal definition itself. Concretely, the paper defines a feature set $F$, a task $T: F_{\\mathrm{input}} \\to F_{\\mathrm{target}}$, and a correlation function $c$; a correlation between a non-relevant feature $f_i \\notin F_{\\mathrm{relevant}}$ and a target feature is spurious, and a shortcut is model behavior that relies on such spurious correlations. This definition does the work of unifying the terminology: Clever Hans behavior is recast as shortcut reliance, the causal confounder is identified as a common cause that generates a spurious correlation, and adversarial backdoor triggers are treated as induced spurious features. The taxonomy then hangs off this definition, with detection and mitigation categories distinguished by where they intervene and by the assumptions they make.","core_discovery":"On its own terms, this paper's central claim is that the terms shortcut, spurious correlation, Clever Hans behavior, and confounder describe one phenomenon, and that a formal definition can capture it. Given a task mapping input features to target features, a correlation between a non-relevant input feature and a target feature is spurious, and a shortcut appears when a model relies on such a correlation for its decisions. The paper separates two origins: world-induced spurious correlations, which exist in the ground-truth distribution, such as waterbirds tending to appear on water, and sampling-induced ones, which arise from a distorted data collection process, such as photographer tags appearing only on waterbird images. Building on that definition, it proposes a taxonomy that splits the field into shortcut detection, via model utility, perturbation, explainability, and causality, and shortcut mitigation, at the dataset, model, and inference levels, and it compiles datasets with explicit spurious correlations, classifying shortcut strength as perfect, semi-perfect, or soft. The intended payoff is a shared vocabulary and a structured map that lets results from one research thread be transferred to another.","pith_inferences":["If the relevant/spurious boundary is genuinely task-relative, then shortcut mitigation is inseparable from task specification; this suggests that future benchmarks may need to ship with explicit, possibly formal, task specifications rather than just labels.","The taxonomy implies a concrete transfer experiment the paper does not run: take a mitigation method proven in vision, such as explanation-based regularization, and evaluate it on backdoor-defense benchmarks, and vice versa, to test whether the unification holds empirically.","The perfect/semi-perfect/soft dataset categorization suggests a testable scaling hypothesis: mitigation method success should correlate with shortcut strength category across the compendium; this could be checked by running a standardized suite of methods over the listed datasets.","The world-induced versus sampling-induced distinction points to different mitigation strategies, data curation and provenance fixes for sampling-induced shortcuts versus reweighting or representation learning for world-induced ones, a division the paper describes but does not formally evaluate."],"forward_implications":["Researchers using the terms shortcut, spurious correlation, Clever Hans, and confounder can map individual methods onto one taxonomy, making cross-domain method transfer explicit.","The taxonomy exposes each method's hidden assumptions, such as shortcut features being easier to learn or the existence of minority groups, so comparisons between methods can be made on assumption match rather than only on benchmark accuracy.","The dataset compendium, with shortcut strength rated perfect, semi-perfect, or soft, gives benchmark selectors a principled basis for matching datasets to method capabilities, for example by showing that group-robustness methods cannot recover from perfect shortcuts.","Connections to causality and security imply that confounder-adjustment tools and backdoor defenses can be imported as shortcut detection and mitigation techniques, expanding the available toolbox without new method development.","The map makes visible underexplored territory: multiple co-occurring shortcuts, shortcuts in generative models, and detection and mitigation beyond image classification."],"supporting_citations":[{"why":"Supplies the informal shortcut concept, an unintended solution that performs well on training data, which the formal definition extends.","marker":"[60]"},{"why":"Provides a prior formalization of spurious correlations as correlations between non-predictive input features and targets, which the paper generalizes.","marker":"[210]"},{"why":"Provides the waterbird running example and the group-robustness framing used throughout the taxonomy.","marker":"[150]"},{"why":"Establishes Clever Hans behavior in machine learning and heatmap-clustering detection, a key taxonomy category.","marker":"[94]"},{"why":"Supplies the bias taxonomy of historical, sampling, and measurement bias used to explain shortcut origins.","marker":"[176]"},{"why":"Gives the selection-bias account behind sampling-induced spurious correlations.","marker":"[196]"},{"why":"Introduces invariant risk minimization, the basis for environment-splitting mitigation methods.","marker":"[12]"},{"why":"Provides a prior typology for shortcut mitigation that this survey positions itself against and extends.","marker":"[57]"}],"fun_headline_variants":["Unifying taxonomy defines shortcuts, spurious correlations, confounders","One framework ties together shortcut, spurious correlation, confounder","New taxonomy maps shortcut detection and mitigation","Shortcut learning: from scattered terms to a single definition"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The definition presumes that for each task one can specify which input features are the relevant ones and tell them apart from spurious ones; if the intended solution is unknown or disputed, the classification of model behavior as a shortcut loses its footing.","fun_headline_variants_meta":{"raw":{"variants":["Unifying taxonomy defines shortcuts, spurious correlations, confounders","One framework ties together shortcut, spurious correlation, confounder","New taxonomy maps shortcut detection and mitigation","Shortcut learning: from scattered terms to a single definition"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00042,"raw_usage":{"total_tokens":2156,"prompt_tokens":933,"completion_tokens":1223,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":1157}},"tokens_in":549,"tokens_out":1223,"duration_ms":9756,"temperature":1.0,"reasoning_tokens":1157,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:48:47.077161+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An annotation study on a standard benchmark: if human experts cannot reach agreement on a stable set of relevant features for a task such as ImageNet or a chest X-ray dataset, then the definition cannot decide whether a model's behavior is a shortcut, undermining the taxonomy's primary organizing criterion.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the selection-bias account behind sampling-induced spurious correlations."}],"review_version":1}