Pith. sign in

REVIEW 3 major objections 6 minor 3 cited by

Navigating Shortcuts, Spurious Correlations, and Confounders: From Origins via Detection to Mitigation

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper argues that shortcut learning can be unified under one formal definition and a single taxonomy spanning detection, mitigation, and datasets.

desk verdict A genuinely useful survey and taxonomy of shortcut learning whose formal definition is best read as a schema—the taxonomy stands on its own merits. read the letter →

arxiv 2412.05152 v1 pith:5RGN6ATY submitted 2024-12-06 cs.LG cs.AI

classification cs.LGcs.AI
keywords shortcutlearningspuriouscorrelationsCleverHansbehaviorconfoundersdetectionmitigationtaxonomybenchmarkdatasets
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to unify the scattered literature on shortcuts, spurious correlations, Clever Hans behavior, and confounders under a single formal definition: a shortcut occurs when a model bases its decision on a spurious correlation rather than on relevant features. It claims that this definition, together with a distinction between world-induced and sampling-induced spurious correlations, is enough to organize the entire field into a taxonomy of detection and mitigation methods plus a curated set of benchmark datasets. A sympathetic reader would care because, if the taxonomy holds, methods developed under different names in different application areas become comparable and transferable, and open gaps such as handling multiple or perfect shortcuts and non-classification tasks become visible in one map. The paper is a survey and position statement rather than a new algorithm.

What carries the argument

The machinery is the formal definition itself. Concretely, the paper defines a feature set $F$, a task $T: F_{\mathrm{input}} \to F_{\mathrm{target}}$, and a correlation function $c$; a correlation between a non-relevant feature $f_i \notin F_{\mathrm{relevant}}$ and a target feature is spurious, and a shortcut is model behavior that relies on such spurious correlations. This definition does the work of unifying the terminology: Clever Hans behavior is recast as shortcut reliance, the causal confounder is identified as a common cause that generates a spurious correlation, and adversarial backdoor triggers are treated as induced spurious features. The taxonomy then hangs off this definition, with detection and mitigation categories distinguished by where they intervene and by the assumptions they make.

What would settle it

An annotation study on a standard benchmark: if human experts cannot reach agreement on a stable set of relevant features for a task such as ImageNet or a chest X-ray dataset, then the definition cannot decide whether a model's behavior is a shortcut, undermining the taxonomy's primary organizing criterion.

Watch

Extended reading notes

Core claim

On its own terms, this paper's central claim is that the terms shortcut, spurious correlation, Clever Hans behavior, and confounder describe one phenomenon, and that a formal definition can capture it. Given a task mapping input features to target features, a correlation between a non-relevant input feature and a target feature is spurious, and a shortcut appears when a model relies on such a correlation for its decisions. The paper separates two origins: world-induced spurious correlations, which exist in the ground-truth distribution, such as waterbirds tending to appear on water, and sampling-induced ones, which arise from a distorted data collection process, such as photographer tags appearing only on waterbird images. Building on that definition, it proposes a taxonomy that splits the field into shortcut detection, via model utility, perturbation, explainability, and causality, and shortcut mitigation, at the dataset, model, and inference levels, and it compiles datasets with explicit spurious correlations, classifying shortcut strength as perfect, semi-perfect, or soft. The intended payoff is a shared vocabulary and a structured map that lets results from one research thread be transferred to another.

Load-bearing premise

The definition presumes that for each task one can specify which input features are the relevant ones and tell them apart from spurious ones; if the intended solution is unknown or disputed, the classification of model behavior as a shortcut loses its footing.

Editorial extensions

If this is right

  • Researchers using the terms shortcut, spurious correlation, Clever Hans, and confounder can map individual methods onto one taxonomy, making cross-domain method transfer explicit.
  • The taxonomy exposes each method's hidden assumptions, such as shortcut features being easier to learn or the existence of minority groups, so comparisons between methods can be made on assumption match rather than only on benchmark accuracy.
  • The dataset compendium, with shortcut strength rated perfect, semi-perfect, or soft, gives benchmark selectors a principled basis for matching datasets to method capabilities, for example by showing that group-robustness methods cannot recover from perfect shortcuts.
  • Connections to causality and security imply that confounder-adjustment tools and backdoor defenses can be imported as shortcut detection and mitigation techniques, expanding the available toolbox without new method development.
  • The map makes visible underexplored territory: multiple co-occurring shortcuts, shortcuts in generative models, and detection and mitigation beyond image classification.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the relevant/spurious boundary is genuinely task-relative, then shortcut mitigation is inseparable from task specification; this suggests that future benchmarks may need to ship with explicit, possibly formal, task specifications rather than just labels.
  • The taxonomy implies a concrete transfer experiment the paper does not run: take a mitigation method proven in vision, such as explanation-based regularization, and evaluate it on backdoor-defense benchmarks, and vice versa, to test whether the unification holds empirically.
  • The perfect/semi-perfect/soft dataset categorization suggests a testable scaling hypothesis: mitigation method success should correlate with shortcut strength category across the compendium; this could be checked by running a standardized suite of methods over the listed datasets.
  • The world-induced versus sampling-induced distinction points to different mitigation strategies, data curation and provenance fixes for sampling-induced shortcuts versus reweighting or representation learning for world-induced ones, a division the paper describes but does not formally evaluate.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This manuscript is a survey and taxonomy paper on shortcut learning, spurious correlations, and confounders in machine learning. The authors propose a formal definition of shortcuts in terms of a ground-truth distribution, a distorted sampling distribution, a feature set, a correlation function, and a notion of 'relevant' features. They distinguish world-induced from sampling-induced spurious correlations, relate shortcuts to Clever Hans behavior, confounders in causality, distribution shift, bias, and adversarial backdoors, and then organize existing detection and mitigation methods into a two-level taxonomy. They also compile a table of datasets containing explicit spurious correlations and close with open challenges. The paper's central claim is that this is the first unified, general taxonomy of shortcut learning, supported by the formal definition and by the breadth of literature covered.

Significance. If the organizational claims hold, the paper is a valuable contribution: it provides a shared vocabulary for a fragmented field, a structured map of detection and mitigation methods across vision, language, medical imaging, and other domains, a useful comparison of prior surveys in Table 1, and a compendium of datasets in Table 4. The connections drawn to causality, fairness, and security are informative and largely accurate, and the discussion of open challenges (e.g., multiple shortcuts, generative models, pretraining-finetuning) is a useful agenda. The paper does not present new experimental results, so its contribution is primarily conceptual and organizational. The formal definition is best understood as a definition schema: it is parameterized by an externally supplied notion of relevant features and by an unspecified correlation function. The taxonomy itself does not depend on the formal definition being fully operational, because the categories in Sections 5 and 6 are organized by methodology, but the abstract and Section 4 present the definition as the unifying foundation, which overstates what the manuscript establishes.

major comments (3)
  1. [Sec. 2, 'Spurious Correlations and Shortcuts'] The definition of a spurious correlation is parameterized by the externally supplied set F_relevant, characterized only as features 'considered relevant to solve the task (in the intended way)'. No criterion, procedure, or oracle for obtaining F_relevant is provided, and Sec. 2.1 itself notes that classifying by habitat rather than bird characteristics would change which features are relevant. Consequently, the definition alone cannot determine whether a given model's behavior is a shortcut; two reasonable task specifications can yield opposite verdicts for the same model. The paper acknowledges this difficulty in Sec. 5.5 and Sec. 8, but the abstract and Sec. 4 still present the definition as a formal foundation that 'unifies' the field. I recommend explicitly presenting the definition as a schema parameterized by a task-specific relevance specification, and moving this qualification into the abstract and Section 4.
  2. [Sec. 2, 'Features and Correlations'] The correlation function c: F x F -> [0,1] is left as an unspecified primitive, and the paper does not define what it means for a model to 'use' a correlation 'as the basis for its decision-making'. As written, the formal definition cannot be instantiated or tested on a concrete model without additional choices, such as a specific correlation measure on raw pixels or a behavioral criterion for 'reliance'. Please either specify at least one intended instantiation and a formal criterion for reliance, or explicitly state that the definition is conceptual rather than operational; the current text mixes both readings.
  3. [Sec. 7 vs. Sec. 8] The 'perfect' shortcut category is defined inconsistently. Section 7 defines a perfect shortcut as one that 'occurs in only a single class and in all such samples', while Section 8 says perfect shortcuts are 'present in all samples'. These are different failure regimes: class-conditional presence versus global presence across all classes. The inconsistency affects the dataset classifications in Table 4 and the discussion of method limitations in Section 8, and it should be resolved by aligning the two definitions and stating which datasets in Table 4 fall into which regime.
minor comments (6)
  1. [Title/first line] The first line of the paper contains a spacing error: 'Na vigating Shortcuts' should be 'Navigating Shortcuts'.
  2. [Sec. 3.3] The sentence 'when B is intervened. On.' is broken by a line break; it should read 'when B is intervened on.'
  3. [Sec. 6.2.1] The phrase 'more robust to shorcuts' is a typo and should read 'more robust to shortcuts'.
  4. [Table 4] The modality header 'Hyperspectal Vision' should be 'Hyperspectral Vision', and the P2S dataset size '2,3k' should use the same decimal convention as the rest of the table (e.g., '2.3k' or '2,300').
  5. [References] Reference [140] (Qiu et al.) is missing a publication year and appears as '[n. d.]'; please complete the bibliographic entry.
  6. [Figures 5 and 6] The small text in the taxonomy figures is difficult to read in the preprint version; please ensure the final figures are legible at print resolution.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a survey whose formal definition is a definitional schema parameterized by an external relevance notion, and whose taxonomy is compiled from external methods rather than derived from the definition.

full rationale

The paper is a survey and taxonomy, not a derivation with fitted parameters. Its central statement, 'A shortcut appears when a model uses a spurious correlation as the basis for its decision-making, i.e., relies on spurious instead of relevant features' (Sec. 2), defines spurious correlations relative to an externally supplied set Frelevant of features 'considered relevant to solve the task (in the intended way)'. This is a definitional dependency, not a circular reduction: Frelevant is not derived from shortcut behavior, and shortcut status is not fitted to any data. The same definition with a different Frelevant yields a different classification, a limitation the paper itself acknowledges ('determining which are relevant and which are spurious can be challenging', Sec. 2.1; 'it is often challenging to decide what features are spurious', Sec. 5). The taxonomy categories in Secs. 4-6 are organized by methodology (model utility, perturbation, XAI, causality; dataset, model, inference time) rather than derived from the definition, so the survey's organizing content is independent of whether the definition is accepted. Self-citations such as Friedrich et al. [57], Teso and Kersting [179], Stammer et al. [169], and Steinmann et al. [172] appear as entries in the method and dataset tables and as pointers to more detailed prior work; they are not used to justify the definition, to exclude alternative definitions, or to supply a uniqueness theorem. There is no fitted input renamed as a prediction, no imported uniqueness claim, and no known result relabeled as new: Table 1 explicitly positions the contribution against external surveys, and the definition is attributed to Geirhos et al. [60] and contrasted with Ye et al. [210]. The difficulty of specifying Frelevant is a real limitation of the definition's scope, but it is a correctness or applicability concern, not circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The formalization rests on conceptual assumptions rather than fitted quantities: the existence of a ground-truth distribution, the specifiability of relevant features, and the adequacy of a symmetric correlation function. No free parameters or invented entities are introduced; the taxonomy is a synthesis of existing literature.

assumptions (3)
  • domain assumption There exists a ground-truth distribution Pgt(x) distinct from the observed distribution P(x).
    Sec. 2 introduces Pgt as an ideal distribution; the distinction between world-induced and sampling-induced shortcuts depends on this separation.
  • domain assumption For each task T, a set of relevant input features Frelevant is well defined.
    Sec. 2 defines spurious correlations relative to Frelevant; the whole definition of a shortcut depends on this being knowable, though the paper acknowledges it is challenging.
  • domain assumption A symmetric correlation function c: F x F to [0,1] adequately represents the dependencies relevant for shortcut behavior.
    Sec. 2 formalizes spurious correlations through c; the taxonomy inherits this abstraction and its limitations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Navigating Shortcuts, Spurious Correlations, and Confounders: From Origins via Detection to Mitigation." pith.science (2026). https://pith.science/paper/5RGN6ATY

@misc{pith2026241205152,
  author       = {Pith},
  title        = {Pith review of: Navigating Shortcuts, Spurious Correlations, and Confounders: From Origins via Detection to Mitigation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5RGN6ATY}},
  note         = {Machine review of arXiv:2412.05152}
}
read the original abstract

Shortcuts, also described as Clever Hans behavior, spurious correlations, or confounders, present a significant challenge in machine learning and AI, critically affecting model generalization and robustness. Research in this area, however, remains fragmented across various terminologies, hindering the progress of the field as a whole. Consequently, we introduce a unifying taxonomy of shortcut learning by providing a formal definition of shortcuts and bridging the diverse terms used in the literature. In doing so, we further establish important connections between shortcuts and related fields, including bias, causality, and security, where parallels exist but are rarely discussed. Our taxonomy organizes existing approaches for shortcut detection and mitigation, providing a comprehensive overview of the current state of the field and revealing underexplored areas and open challenges. Moreover, we compile and classify datasets tailored to study shortcut learning. Altogether, this work provides a holistic perspective to deepen understanding and drive the development of more effective strategies for addressing shortcuts in machine learning.

Figures

Figures reproduced from arXiv: 2412.05152 by the authors.

Figure 1
Figure 1. Models across different settings are susceptible to shortcuts. Models trained on data containing spurious correlations may rely on unintended features for decision-making. These shortcuts can manifest across various domains and tasks, significantly affecting model performance and generalization. Running Example: Classifying Birds into Landbirds and Waterbirds For this illustrative example, let us assume that we have… view at source ↗
Figure 2
Figure 2. Overview where spurious correlations can appear. Given our established example of classifying birds into landbirds and waterbirds (based on their characteristics), the environment is a spurious feature naturally occurring in the world (i). The distorted (⇝) sampling process can then induce spurious correlations through, for example, photographer tags (ii). (i) We provide formal definitions of shortcuts, unifying and… view at source ↗
Figure 3
Figure 3. Overview of the waterbirds example in the context of causality. In the data, we have access to the bird’s characteristics and its environment and we want to predict whether the bird is a landbird or waterbird. If we assume that the bird’s characteristics (i.e., its appearance and abilities) cause both its environment and whether it is a waterbird, environment and the target label are correlated in the data (while no… view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Overview of our taxonomy on shortcut learning. We categorize approaches into the two areas of shortcut detection and shortcut mitigation. Detailed information, including all subcategories, is provided in [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Taxonomy of shortcut detection approaches. A comprehen￾sive breakdown of shortcut detection methods, organized into method￾ological subcategories. The underlying assumption of the following works is that shortcuts are easier to learn than the relevant features. Additio…
Figure 6
Figure 6. Figure 6: Detailed taxonomy of shortcut mitigation ap￾proaches. An in-depth representation of shortcut mitigation strategies, categorizing methods along their main applica￾tion level and their mitigation strategy. The first set of strategies focuses on mitigating shortcuts direc…
Figure 7
Figure 7. Figure 7: Dataset Selection Tradeoffs In the previous sections, we provided a comprehensive overview of shortcut detection and mitigation methods. To validate the effectiveness of these methods in practice, it is necessary to test them on different datasets. Unlike standard mach…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Artificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Task-conditioned foveated masks used as auxiliary attention loss during fine-tuning substantially raise OOD success of robotic foundation models by aligning policy attention to action-critical regions.

  2. VISION: Robust and Interpretable Code Vulnerability Detection Leveraging Counterfactual Augmentation

    cs.AI 2025-08 conditional novelty 6.0 of 10

    LLM-generated counterfactual code pairs with flipped vulnerability labels, used to train a GNN, sharply improve CWE-20 detection and attribution on the released CWE-20-CFA benchmark.

  3. Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Neural Concept Verifier trains image classifiers so predictions must rely on small, verifiable subsets of extracted concepts rather than raw pixel masks.

Reference graph

Works this paper leans on

230 extracted references · 44 canonical work pages · cited by 3 Pith papers

  1. [1]

    Hervé Abdi. 2007. The Kendall rank correlation coe fficient. Encyclopedia of measurement and statistics 2 (2007), 508–510

  2. [2]

    Julius Adebayo, Michael Muelly, Harold Abelson, and Been Kim. 2022. Post hoc explanations may be ineffective for detecting unknown spurious correlation. In International conference on learning representations

  3. [3]

    Ehsan Adeli, Qingyu Zhao, Adolf Pfefferbaum, Edith V Sullivan, Li Fei-Fei, Juan Carlos Niebles, and Kilian M Pohl. 2021. Representation learning with statistical independence to mitigate bias. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2513–2523

  4. [4]

    Yossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas, and Joseph Keshet. 2018. Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by Backdooring. In USENIX Security Symposium. 1615–1631

  5. [5]

    Mohammed Adnan, Yani Ioannou, Kenyon Tsai, Angus Galloway, Hamid Tizhoosh, Rahul G Krishnan, and Graham W. Taylor. 2024. Detecting Shortcuts using Mutual Information. 20 A preprint - December 9, 2024

  6. [6]

    Vedika Agarwal, Rakshith Shetty, and Mario Fritz. 2020. Towards causal vqa: Revealing and reducing spurious correlations by invariant and covariant semantic editing. In Proceedings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition. 9690–9698

  7. [7]

    Goldgof, Rahul Paul, Dmitry Goldgof, and Lawrence O

    Kaoutar Ben Ahmed, Gregory M. Goldgof, Rahul Paul, Dmitry Goldgof, and Lawrence O. Hall. 2021. Discovery of a Generalization Gap of Convolutional Neural Networks on COVID-19 X-Rays Classification. IEEE Access 9 (2021), 72970 – 72979

  8. [8]

    Kaoutar Ben Ahmed, Lawrence O Hall, Dmitry B Goldgof, and Ryan Fogarty. 2022. Achieving multisite generalization for cnn-based disease diagnosis models by mitigating shortcut learning. IEEE Access 10 (2022), 78726–78738

Show all 230 references
  1. [9]

    Saleema Amershi, Maya Cakmak, William Bradley Knox, and Todd Kulesza. 2014. Power to the people: The role of humans in interactive machine learning. AI magazine 35, 4 (2014), 105–120

  2. [10]

    Christopher J Anders, Leander Weber, David Neumann, Wojciech Samek, Klaus-Robert Müller, and Sebastian Lapuschkin. 2022. Finding and removing clever hans: Using explanation methods to debug and improve deep models. Information Fusion 77 (2022), 261–295

  3. [11]

    Md Rifat Arefin, Yan Zhang, Aristide Baratin, Francesco Locatello, Irina Rish, Dianbo Liu, and Kenji Kawaguchi

  4. [12]

    Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz. 2019. Invariant risk minimization. arXiv preprint arXiv:1907.02893 (2019)

  5. [13]

    Vijay Arya, Rachel KE Bellamy, Pin-Yu Chen, Amit Dhurandhar, Michael Hind, Samuel C Hoffman, Stephanie Houde, Q Vera Liao, Ronny Luss, Aleksandra Mojsilovi´c, et al. 2022. Ai explainability 360: Impact and design. In Proceedings of the AAAI Conference on Artificial Intelligenc...

  6. [14]

    Saeid Asgari, Aliasghar Khani, Fereshte Khani, Ali Gholami, Linh Tran, Ali Mahdavi Amiri, and Ghassan Hamarneh. 2022. Masktune: Mitigating spurious correlations by forcing to explore. Advances in Neural Information Processing Systems 35 (2022), 23284–23296

  7. [15]

    Gregor Bachmann and Vaishnavh Nagarajan. 2024. The pitfalls of next-token prediction. arXiv preprint arXiv:2403.06963 (2024)

  8. [16]

    Shortcuts

    Imon Banerjee, Kamanasish Bhattacharjee, John L Burns, Hari Trivedi, Saptarshi Purkayastha, Laleh Seyyed- Kalantari, Bhavik N Patel, Rakesh Shiradkar, and Judy Gichoya. 2023. “Shortcuts” causing bias in radiology artificial intelligence: causes, evaluation and mitigation. Jour...

  9. [17]

    Joseph, and J

    Marco Barreno, Blaine Nelson, Russell Sears, Anthony D. Joseph, and J. D. Tygar. 2006. Can machine learning be secure?. In Symposium on Information, Computer and Communications Security (ASIACCS). 16–25

  10. [18]

    Pedro RAS Bassi, Sergio SJ Dertkigil, and Andrea Cavalli. 2024. Improving deep neural network generalization and robustness to background bias via layer-wise relevance propagation optimization. Nature Communications 15, 1 (2024), 291

  11. [19]

    Sara Beery, Grant Van Horn, and Pietro Perona. 2018. Recognition in terra incognita. In Proceedings of the European conference on computer vision (ECCV). 456–473

  12. [20]

    Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, and Aylin Caliskan. 2023. Easily accessible text-to-image generation amplifies demographic stereotypes at large scale. In Proceedings of the...

  13. [21]

    Abeba Birhane, Sepehr Dehdashtian, Vinay Prabhu, and Vishnu Boddeti. 2024. The Dark Side of Dataset Scaling: Evaluating Racial Classification in Multimodal Models. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency. 1229–1244

  14. [22]

    Abeba Birhane, Sanghyun Han, Vishnu Boddeti, Sasha Luccioni, et al. 2024. Into the laion’s den: Investigating hate in multimodal datasets. Advances in Neural Information Processing Systems 36 (2024)

  15. [23]

    Franziska Boenisch. 2021. A Systematic Review on Model Watermarking for Neural Networks. Frontiers Big Data 4 (2021)

  16. [24]

    Andrea Bontempelli, Stefano Teso, Katya Tentori, Fausto Giunchiglia, and Andrea Passerini. 2023. Concept-level Debugging of Part-Prototype Networks. arXiv:2205.15769 [cs.LG]

  17. [25]

    Samuele Bortolotti, Emanuele Marconato, Tommaso Carraro, Paolo Morettin, Emile van Krieken, Antonio Vergari, Stefano Teso, and Andrea Passerini. 2024. A Neuro-Symbolic Benchmark Suite for Concept Quality and Reasoning Shortcuts. arXiv:2406.10368 [cs.LG] 21 A preprint - Decembe...

  18. [26]

    Simona Bottani, Ninon Burgos, Aurélien Maire, Dario Saracino, Sebastian Ströer, Didier Dormont, Olivier Colliot, Alzheimer’s Disease Neuroimaging Initiative, APPRIMAGE Study Group, et al. 2023. Evaluation of MRI-based machine learning approaches for computer-aided diagnosis of...

  19. [27]

    Alexander Brown, Nenad Tomasev, Jan Freyberg, Yuan Liu, Alan Karthikesalingam, and Jessica Schrouff. 2023. Detecting shortcut learning for fair medical AI using shortcut testing. Nature Communications 14, 1 (2023), 4314

  20. [28]

    Florian Peter Busch, Roshni Kamath, Rupert Mitchell, Wolfgang Stammer, Kristian Kersting, and Martin Mundt

  21. [29]

    Kirill Bykov, Laura Kopf, and Marina M-C Höhne. 2023. Finding Spurious Correlations with Function-Semantic Contrast Analysis. In World Conference on Explainable Artificial Intelligence. Springer, 549–572

  22. [30]

    arXiv:2402.06434 [cs.LG]

    Where is the Truth? The Risk of Getting Confounded in a Continual World. arXiv:2402.06434 [cs.LG]

  23. [31]

    Rwiddhi Chakraborty, Adrian Sletten, and Michael C Kampffmeyer. 2024. ExMap: Leveraging Explainability Heatmaps for Unsupervised Group Robustness to Spurious Correlations. In Proceedings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition. 12017–12026

  24. [32]

    Brandon Carter, Siddhartha Jain, Jonas W Mueller, and David Gifford. 2021. Overinterpretation reveals image classification model pathologies. Advances in Neural Information Processing Systems 34 (2021), 15395–15407

  25. [33]

    Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song. 2017. Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning. arXiv preprint arXiv:1712.05526 (2017)

  26. [34]

    Kushal Chauhan, Rishabh Tiwari, Jan Freyberg, Pradeep Shenoy, and Krishnamurthy Dvijotham. 2023. Inter- active concept bottleneck models. In Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 37. 5948–5955

  27. [35]

    Jun Cheng, Wei Huang, Shuangliang Cao, Ru Yang, Wei Yang, Zhaoqiang Yun, Zhijian Wang, and Qianjin Feng

  28. [36]

    Yimeng Chen, Ruibin Xiong, Zhi-Ming Ma, and Yanyan Lan. 2022. When does group invariant learning survive spurious correlations? Advances in Neural Information Processing Systems 35 (2022), 7038–7051

  29. [37]

    Pattarawat Chormai, Jan Herrmann, Klaus-Robert Müller, and Grégoire Montavon. 2024. Disentangled explana- tions of neural network predictions by finding relevant subspaces. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)

  30. [38]

    Noel Codella, Veronica Rotemberg, Philipp Tschandl, M Emre Celebi, Stephen Dusza, David Gutman, Brian Helba, Aadi Kalloo, Konstantinos Liopyris, Michael Marchetti, et al . 2019. Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin i...

  31. [39]

    Bhusan Chettri. 2023. The clever hans e ffect in voice spoofing detection. In 2022 IEEE Spoken Language Technology Workshop (SLT). IEEE, 577–584

  32. [40]

    Nikolay Dagaev, Brett D Roads, Xiaoliang Luo, Daniel N Barry, Kaustubh R Patil, and Bradley C Love. 2023. A too-good-to-be-true prior to reduce shortcut reliance. Pattern recognition letters166 (2023), 164–171

  33. [41]

    Corentin Dancette, Remi Cadene, Damien Teney, and Matthieu Cord. 2021. Beyond question-based biases: As- sessing multimodal shortcut learning in visual question answering. InProceedings of the IEEE/CVF International Conference on Computer Vision. 1574–1583

  34. [42]

    Elliot Creager, Jörn-Henrik Jacobsen, and Richard Zemel. 2021. Environment inference for invariant learning. In International Conference on Machine Learning. PMLR, 2189–2200

  35. [43]

    Quentin Delfosse, Jannis Blüml, Bjarne Gregori, and Kristian Kersting. 2024. HackAtari: Atari Learning Environments for Robust and Continual Reinforcement Learning. arXiv preprint arXiv:2406.03997 (2024)

  36. [44]

    Quentin Delfosse, Sebastian Sztwiertnia, Mark Rothermel, Wolfgang Stammer, and Kristian Kersting. 2024. Interpretable Concept Bottlenecks to Align Reinforcement Learning Agents. In The Thirty-eighth Annual Conference on Neural Information Processing Systems

  37. [45]

    DeGrave, Joseph D

    Alex J. DeGrave, Joseph D. Janizek, and Su-In Lee. 2021. AI for radiographic COVID-19 detection selects shortcuts over signal. Nat. Mach. Intell. 3, 7 (2021), 610–619

  38. [46]

    Yihe Deng, Yu Yang, Baharan Mirzasoleiman, and Quanquan Gu. 2024. Robust learning with progressive data expansion against spurious correlation. Advances in neural information processing systems 36 (2024). 22 A preprint - December 9, 2024

  39. [47]

    Wenhao Ding, Laixi Shi, Yuejie Chi, and Ding Zhao. 2024. Seeing is not believing: Robust reinforcement learning against spurious correlation. Advances in Neural Information Processing Systems 36 (2024)

  40. [48]

    Li Deng. 2012. The MNIST Database of Handwritten Digit Images for Machine Learning Research [Best of the Web]. IEEE Signal Processing Magazine 29, 6 (2012), 141–142

  41. [49]

    Maximilian Dreyer, Frederik Pahde, Christopher J Anders, Wojciech Samek, and Sebastian Lapuschkin. 2024. From hope to safety: Unlearning biases of deep models via gradient penalization in latent space. In Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 38....

  42. [50]

    Mengnan Du, Fengxiang He, Na Zou, Dacheng Tao, and Xia Hu. 2023. Shortcut Learning of Large Language Models in Natural Language Understanding. Commun. ACM (2023)

  43. [51]

    Varun Dogra, Sahil Verma, Marcin Wo´ zniak, Jana Shafi, Muhammad Fazal Ijaz, et al. 2024. Shortcut Learning Explanations for Deep Natural Language Processing: A Survey on Dataset Biases. IEEE Access (2024)

  44. [52]

    Louisa Fay, Erick Cobos, Bin Yang, Sergios Gatidis, and Thomas Küstner. 2023. Avoiding shortcut-learning by mutual information minimization in deep learning-based image processing. IEEE Access 11 (2023), 64070– 64086

  45. [53]

    Tao Feng, Lizhen Qu, and Gholamreza Haffari. 2023. Less is more: Mitigate spurious correlations for open- domain dialogue response generation models by causal discovery. Transactions of the Association for Computa- tional Linguistics 11 (2023), 511–530

  46. [54]

    Shaohua Fan, Xiao Wang, Chuan Shi, Peng Cui, and Bai Wang. 2023. Generalizing graph neural networks on out-of-distribution graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence(2023)

  47. [55]

    Felix Friedrich, Manuel Brack, Lukas Struppek, Dominik Hintersdorf, Patrick Schramowski, Sasha Luccioni, and Kristian Kersting. 2023. Fair di ffusion: Instructing text-to-image generation models on fairness. arXiv preprint arXiv:2302.10893 (2023)

  48. [56]

    Felix Friedrich, Wolfgang Stammer, Patrick Schramowski, and Kristian Kersting. 2023. Revision Transformers: Instructing Language Models to Change Their Values.. In ECAI. 756–763

  49. [57]

    Pierre Fernandez, Guillaume Couairon, Hervé Jégou, Matthijs Douze, and Teddy Furon. 2023. The Stable Signature: Rooting Watermarks in Latent Diffusion Models. In International Conference on Computer Vision (ICCV). 22409–22420

  50. [58]

    Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed. 2024. Bias and fairness in large language models: A survey.Computational Linguistics (2024), 1–79

  51. [59]

    Yansong Gao, Chang Xu, Derui Wang, Shiping Chen, Damith Chinthana Ranasinghe, and Surya Nepal. 2019. STRIP: a defence against trojan attacks on deep neural networks. In Proceedings of the 35th Annual Computer Security Applications Conference, ACSAC 2019, San Juan, PR, USA, Dec...

  52. [60]

    Felix Friedrich, Wolfgang Stammer, Patrick Schramowski, and Kristian Kersting. 2023. A typology for exploring the mitigation of shortcut behaviour. Nature Machine Intelligence 5, 3 (2023), 319–330

  53. [61]

    Wichmann, and Wieland Brendel

    Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel. 2022. ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. arXiv:1811.12231 [cs.CV]

  54. [62]

    Soumya Suvra Ghosal and Yixuan Li. 2024. Are vision transformers robust to spurious correlations?International Journal of Computer Vision 132, 3 (2024), 689–709

  55. [63]

    Zemel, Wieland Brendel, Matthias Bethge, and Felix A

    Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard S. Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. 2020. Shortcut learning in deep neural networks. Nature Machine Intelligence (2020)

  56. [64]

    Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg. 2017. BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain. arXiv preprint arXiv:1708.06733 (2017)

  57. [65]

    Junfeng Guo, Yiming Li, Xun Chen, Hanqing Guo, Lichao Sun, and Cong Liu. 2023. SCALE-UP: An Efficient Black-box Input-level Backdoor Detection via Analyzing Scaled Prediction Consistency. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, ...

  58. [66]

    Sindhu CM Gowda, Shalmali Joshi, Haoran Zhang, and Marzyeh Ghassemi. 2021. Pulling up by the causal bootstraps: Causal data augmentation for pre-training debiasing. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management. 606–616

  59. [67]

    Curran, and Brian Mac Namee

    Misgina Tsighe Hagos, Kathleen M. Curran, and Brian Mac Namee. 2022. Impact of Feedback Type on Explanatory Interactive Learning. Springer International Publishing, 127–137

  60. [68]

    Andreas Philipp Hassler, Ernestina Menasalvas, Francisco José García-García, Leocadio Rodríguez-Mañas, and Andreas Holzinger. 2019. Importance of medical data preprocessing in predictive modeling and risk factor discovery for the frailty syndrome. BMC medical informatics and d...

  61. [69]

    Avani Gupta and PJ Narayanan. 2024. A survey on Concept-based Approaches For Model Improvement. arXiv preprint arXiv:2403.14566 (2024). 23 A preprint - December 9, 2024

  62. [70]

    Lisa Anne Hendricks, Kaylee Burns, Kate Saenko, Trevor Darrell, and Anna Rohrbach. 2018. Women also snowboard: Overcoming bias in captioning models. In Proceedings of the European conference on computer vision (ECCV). 771–787

  63. [71]

    Xanh Ho, Johannes Mario Meissner, Saku Sugawara, and Akiko Aizawa. 2022. A survey on measuring and mitigating reasoning shortcuts in machine reading comprehension. arXiv preprint arXiv:2209.01824 (2022)

  64. [72]

    Yue He, Zheyan Shen, and Peng Cui. 2019. Towards Non-I.I.D. Image Classification: A Dataset and Baselines. arXiv:1906.02899 [cs.CV]

  65. [73]

    Brian Hu, Paul Tunison, Brandon RichardWebster, and Anthony Hoogs. 2023. Xaitk-saliency: An open source explainable ai toolkit for saliency. In Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 37. 15760–15766

  66. [74]

    Maximilian Idahl, Lijun Lyu, Ujwal Gadiraju, and Avishek Anand. 2021. Towards Benchmarking the Utility of Explanations for Model Debugging. In Proceedings of the First Workshop on Trustworthy Natural Language Processing, Yada Pruksachatkun, Anil Ramakrishna, Kai-Wei Chang, Sat...

  67. [75]

    Floris Holstege, Bram Wouters, Noud Van Giersbergen, and Cees Diks. 2023. Removing Spurious Concepts from Neural Network Representations via Joint Subspace Estimation. arXiv preprint arXiv:2310.11991 (2023)

  68. [76]

    Arthur Jacot, Franck Gabriel, and Clément Hongler. 2018. Neural tangent kernel: Convergence and generalization in neural networks. Advances in neural information processing systems 31 (2018)

  69. [77]

    Zhihua Jin, Xingbo Wang, Furui Cheng, Chunhui Sun, Qun Liu, and Huamin Qu. 2023. ShortcutLens: A visual analytics approach for exploring shortcuts in natural language understanding dataset. IEEE Transactions on Visualization and Computer Graphics (2023)

  70. [78]

    Pavel Izmailov, Polina Kirichenko, Nate Gruver, and Andrew G Wilson. 2022. On feature learning in the presence of spurious correlations. Advances in Neural Information Processing Systems 35 (2022), 38516–38532

  71. [79]

    Rickard Karlsson and Jesse Krijthe. 2024. Detecting hidden confounding in observational data using multiple environments. Advances in Neural Information Processing Systems 36 (2024)

  72. [80]

    Jacob Kauffmann, Lukas Ruff, Grégoire Montavon, and Klaus-Robert Müller. 2020. The clever Hans effect in anomaly detection. arXiv preprint arXiv:2006.10609 (2020)

  73. [81]

    Lawrence Zitnick, and Ross Girshick

    Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2016. CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning. arXiv:1612.06890 [cs.CV]

  74. [82]

    Nayeong Kim, Juwon Kang, Sungsoo Ahn, Jungseul Ok, and Suha Kwak. 2024. Improving Robustness to Multiple Spurious Correlations by Multi-Objective Optimization. ICML (2024)

  75. [83]

    John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. 2023. A Watermark for Large Language Models. arXiv preprint arXiv:2301.10226 (2023)

  76. [84]

    Corentin Kervadec, Grigory Antipov, Moez Baccouche, and Christian Wolf. 2021. Roses are red, violets are blue... but should vqa expect them to?. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2776–2785

  77. [85]

    Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang

  78. [86]

    Earnshaw, Imran S

    Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton A. Earnshaw, Imran S. Haque, Sara Beery, Jure Leskovec, A...

  79. [87]

    Polina Kirichenko, Pavel Izmailov, and Andrew Gordon Wilson. 2022. Last layer re-training is sufficient for robustness to spurious correlations. arXiv preprint arXiv:2204.02937 (2022)

  80. [88]

    Maurice Kraus, David Steinmann, Antonia Wüst, Andre Kokozinski, and Kristian Kersting. 2024. Right on Time: Revising Time Series Models by Constraining their Explanations. arXiv preprint arXiv:2402.12921 (2024)

  81. [89]

    Meelis Kull and Peter Flach. 2014. Patterns of dataset shift. In First international workshop on learning over multiple contexts (LMCE) at ECML-PKDD, V ol. 5

  82. [90]

    Abhinav Kumar, Amit Deshpande, and Amit Sharma. 2024. Causal effect regularization: automated detection and removal of spurious correlations. Advances in Neural Information Processing Systems 36 (2024)

  83. [91]

    Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsallakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, et al . 2020. Captum: A unified and generic model interpretability library for pytorch. arXiv preprint arXiv:2009.0...

  84. [92]

    Tyler LaBonte, Vidya Muthukumar, and Abhishek Kumar. 2024. Towards last-layer retraining for group robustness with fewer annotations. Advances in Neural Information Processing Systems 36 (2024)

  85. [93]

    Lauro Langosco Di Langosco, Jack Koch, Lee D Sharkey, Jacob Pfau, and David Krueger. 2022. Goal Misgeneralization in Deep Reinforcement Learning. In Proceedings of the 39th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 162), Kamal...

  86. [94]

    Sebastian Lapuschkin, Stephan Wäldchen, Alexander Binder, Grégoire Montavon, Wojciech Samek, and Klaus- Robert Müller. 2019. Unmasking Clever Hans predictors and assessing what machines really learn. Nature communications 10, 1 (2019), 1096

  87. [95]

    JuneHyoung Kwon, Eunju Lee, Yunsung Cho, and YoungBin Kim. 2024. Learning to Detour: Shortcut Mitigating Augmentation for Weakly Supervised Semantic Segmentation. In Proceedings of the IEEE /CVF Winter Conference on Applications of Computer Vision. 819–828

  88. [96]

    Yoonho Lee, Michelle S Lam, Helena Vasconcelos, Michael S Bernstein, and Chelsea Finn. 2024. Clarify: Improving Model Robustness With Natural Language Corrections. arXiv preprint arXiv:2402.03715 (2024)

  89. [97]

    Haoxin Li, Yuan Liu, Hanwang Zhang, and Boyang Li. 2023. Mitigating and Evaluating Static Bias of Action Representations in the Background and the Foreground. arXiv:2211.12883 [cs.CV]

  90. [98]

    Yi Li and Nuno Vasconcelos. 2019. Repair: Removing representation bias by dataset resampling. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9572–9581

  91. [99]

    Jungsoo Lee, Eungyeup Kim, Juyoung Lee, Jihyeon Lee, and Jaegul Choo. 2021. Learning debiased represen- tation via disentangled feature augmentation. Advances in Neural Information Processing Systems 34 (2021), 25123–25133

  92. [100]

    Zhiheng Li, Anthony Hoogs, and Chenliang Xu. 2022. Discover and mitigate unknown biases with debiasing alternate networks. In European Conference on Computer Vision. Springer, 270–288

  93. [101]

    Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2021. Towards understanding and mitigating social biases in language models. In International Conference on Machine Learning. PMLR, 6565–6576

  94. [102]

    Weixin Liang and James Zou. 2022. Metashift: A dataset of datasets for evaluating contextual distribution shifts and training conflicts. arXiv preprint arXiv:2202.06523 (2022)

  95. [103]

    Zhiheng Li, Ivan Evtimov, Albert Gordo, Caner Hazirbas, Tal Hassner, Cristian Canton Ferrer, Chenliang Xu, and Mark Ibrahim. 2023. A whac-a-mole dilemma: Shortcuts come in multiples where mitigating one amplifies others. In Proceedings of the IEEE/CVF Conference on Computer Vi...

  96. [104]

    Phillip Lippe, Sara Magliacane, Sindy Löwe, Yuki M Asano, Taco Cohen, and Efstratios Gavves. 2023. BISCUIT: Causal Representation Learning from Binary Interactions. In Proceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intelligence (Proceedings of Machine ...

  97. [105]

    Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg. 2018. Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks. In Research in Attacks, Intrusions, and Defenses (RAID) , V ol. 11050. 273–294

  98. [106]

    Ruyang Liu, Hao Liu, Ge Li, Haodi Hou, TingHao Yu, and Tao Yang. 2022. Contextual debiasing for visual recognition with causal mechanisms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12755–12765

  99. [107]

    Lorenz Linhardt, Klaus-Robert Müller, and Grégoire Montavon. 2024. Preemptively pruning Clever-Hans strategies in deep neural networks. Information Fusion 103 (2024), 102094

  100. [108]

    Zhili Liu, Kai Chen, Yifan Zhang, Jianhua Han, Lanqing Hong, Hang Xu, Zhenguo Li, Dit-Yan Yeung, and James Kwok. 2024. Implicit Concept Removal of Diffusion Models. arXiv:2310.05873 [cs.CV]

  101. [109]

    Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 2015. Deep Learning Face Attributes in the Wild. In Proceedings of International Conference on Computer Vision (ICCV)

  102. [110]

    Adian Liusie, Vatsal Raina, Vyas Raina, and Mark Gales. 2022. Analyzing Biases to Spurious Correlations in Text Classification Tasks. InProceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joi...

  103. [111]

    Yuntao Liu, Yang Xie, and Ankur Srivastava. 2017. Neural Trojans. In2017 IEEE International Conference on Computer Design, ICCD 2017, Boston, MA, USA, November 5-8, 2017. IEEE Computer Society, 45–48. 25 A preprint - December 9, 2024

  104. [112]

    Xu Luo, Longhui Wei, Liangjian Wen, Jinrong Yang, Lingxi Xie, Zenglin Xu, and Qi Tian. 2021. Rectifying the shortcut learning of background for few-shot learning. Advances in Neural Information Processing Systems 34 (2021), 13073–13085

  105. [113]

    Aengus Lynch, Gbètondji JS Dovonon, Jean Kaddour, and Ricardo Silva. 2023. Spawrious: A benchmark for fine control of spurious correlation biases. arXiv preprint arXiv:2303.05470 (2023)

  106. [114]

    Chong Ma, Lin Zhao, Yuzhong Chen, Lei Guo, Tuo Zhang, Xintao Hu, Dinggang Shen, Xi Jiang, and Tianming Liu. 2023. Rectify vit shortcut learning by visual saliency. IEEE Transactions on Neural Networks and Learning Systems (2023)

  107. [115]

    Luyang Luo, Dunyuan Xu, Hao Chen, Tien-Tsin Wong, and Pheng-Ann Heng. 2022. Pseudo bias-balanced learning for debiased chest x-ray classification. In International conference on medical image computing and computer-assisted intervention. Springer, 621–631

  108. [116]

    Jie Ma, Pinghui Wang, Dechen Kong, Zewei Wang, Jun Liu, Hongbin Pei, and Junzhou Zhao. 2024. Robust Visual Question Answering: Datasets, Methods, and Future Challenges. IEEE Transactions on Pattern Analysis and Machine Intelligence 46, 8 (2024), 5575–5594

  109. [117]

    Maggie Makar, Ben Packer, Dan Moldovan, Davis Blalock, Yoni Halpern, and Alexander D’Amour. 2022. Causally motivated shortcut removal using auxiliary labels. In International Conference on Artificial Intelligence and Statistics. PMLR, 739–766

  110. [118]

    Chengzhi Mao, Augustine Cha, Amogh Gupta, Hao Wang, Junfeng Yang, and Carl V ondrick. 2021. Generative interventions for causal learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3947–3956

  111. [119]

    Chong Ma, Lin Zhao, Yuzhong Chen, Sheng Wang, Lei Guo, Tuo Zhang, Dinggang Shen, Xi Jiang, and Tianming Liu. 2023. Eye-gaze-guided vision transformer for rectifying shortcut learning. IEEE Transactions on Medical Imaging 42, 11 (2023), 3384–3394

  112. [120]

    Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A survey on bias and fairness in machine learning. ACM computing surveys (CSUR) 54, 6 (2021), 1–35

  113. [121]

    Olli Miettinen. 1974. Confounding and effect-modification. American Journal of Epidemiology 100, 5 (1974), 350–353

  114. [122]

    Matthias Minderer, Olivier Bachem, Neil Houlsby, and Michael Tschannen. 2020. Automatic shortcut removal for self-supervised representation learning. In International Conference on Machine Learning. PMLR, 6927–6937

  115. [123]

    Emanuele Marconato, Stefano Teso, Antonio Vergari, and Andrea Passerini. 2024. Not all neuro-symbolic concepts are created equal: Analysis and mitigation of reasoning shortcuts. Advances in Neural Information Processing Systems 36 (2024)

  116. [124]

    Alfredo Morabia. 2011. History of the modern epidemiological concept of confounding.Journal of Epidemiology & Community Health 65, 4 (2011), 297–300

  117. [125]

    Nicolas M Müller, Simon Roschmann, Shahbaz Khan, Philip Sperl, and Konstantin Böttinger. 2023. Shortcut Detection with Variational Autoencoders. arXiv preprint arXiv:2302.04246 (2023)

  118. [126]

    Nicolas M Müller, Simon Roschmann, Shahbaz Khan, Philip Sperl, and Konstantin Böttinger. 2024. Shortcut detection with variational autoencoders. In 2024 International Joint Conference on Neural Networks (IJCNN). IEEE, 1–7

  119. [127]

    Mazda Moayeri, Wenxiao Wang, Sahil Singla, and Soheil Feizi. 2023. Spuriosity rankings: sorting data to measure and mitigate biases. Advances in Neural Information Processing Systems 36 (2023), 41572–41600

  120. [128]

    Leann Myers and Maria J Sirois. 2004. Spearman correlation coe fficients, differences between. Encyclopedia of statistical sciences 12 (2004)

  121. [129]

    Vaishnavh Nagarajan, Anders Andreassen, and Behnam Neyshabur. 2020. Understanding the failure modes of out-of-distribution generalization. arXiv preprint arXiv:2010.15775 (2020)

  122. [130]

    Junhyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee, and Jinwoo Shin. 2020. Learning from failure: De-biasing classifier from biased classifier. Advances in Neural Information Processing Systems 33 (2020), 20673–20684

  123. [131]

    Yasmine Mustafa and Tie Luo. 2024. Unmasking Dementia Detection by Masking Input Gradients: A JSM Approach to Model Interpretability and Precision. In Pacific-Asia Conference on Knowledge Discovery and Data Mining. Springer, 75–90. 26 A preprint - December 9, 2024

  124. [132]

    Meike Nauta, Ricky Walsh, Adam Dubowski, and Christin Seifert. 2021. Uncovering and correcting shortcut learning in machine learning models for skin cancer diagnosis. Diagnostics 12, 1 (2021), 40

  125. [133]

    Fahimeh Hosseini Noohdani, Parsa Hosseini, Aryan Yazdan Parast, Hamidreza Yaghoubi Araghi, and Mahdieh Soleymani Baghshah. 2024. Decompose-and-Compose: A Compositional Approach to Mitigating Spurious Correlation. In Proceedings of the IEEE/CVF Conference on Computer Vision and...

  126. [134]

    Harsha Nori, Samuel Jenkins, Paul Koch, and Rich Caruana. 2019. Interpretml: A unified framework for machine learning interpretability. arXiv preprint arXiv:1909.09223 (2019)

  127. [135]

    Junhyun Nam, Jaehyung Kim, Jaeho Lee, and Jinwoo Shin. 2022. Spread spurious attribute: Improving worst-group accuracy with spurious attribute estimation. arXiv preprint arXiv:2204.02070 (2022)

  128. [136]

    Judea Pearl et al. 2000. Models, reasoning and inference. Cambridge, UK: CambridgeUniversityPress 19, 2 (2000), 3

  129. [137]

    Gregory Plumb, Marco Tulio Ribeiro, and Ameet Talwalkar. 2021. Finding and fixing spurious patterns with explanations. arXiv preprint arXiv:2106.02112 (2021)

  130. [138]

    Jorge A Portal-Diaz, Orlando Lovelle-Enríquez, Marlen Perez-Diaz, José D Lopez-Cabrera, Osmany Reyes- Cardoso, and Ruben Orozco-Morales. 2022. New patch-based strategy for COVID-19 automatic identification using chest x-ray images. Health and Technology 12, 6 (2022), 1117–1132

  131. [139]

    OpenAI. 2024. GPT-4 Technical Report

  132. [140]

    GuanWen Qiu, Da Kuang, and Surbhi Goel. [n. d.]. Complexity Matters: Feature Learning in the Presence of Spurious Correlations. In Forty-first International Conference on Machine Learning

  133. [141]

    Joaquin Quiñonero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D Lawrence. 2022. Dataset shift in machine learning. Mit Press

  134. [142]

    Ruggero Ragonesi, Pietro Morerio, and Vittorio Murino. 2023. Learning unbiased classifiers from biased data with meta-learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1–9

  135. [143]

    Ximing Qiao, Yukun Yang, and Hai Li. 2019. Defending Neural Backdoors via Generative Distribution Modeling. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, B...

  136. [144]

    Christian Reimers, Jakob Runge, and Joachim Denzler. 2020. Determining the Relevance of Features for Deep Neural Networks. In European Conference on Computer Vision

  137. [145]

    Qibing Ren, Yiting Chen, Yichuan Mo, Qitian Wu, and Junchi Yan. 2022. Dice: Domain-attack invariant causal learning for improved data privacy protection and adversarial robustness. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 1483–1492

  138. [146]

    Caleb Robinson, Anusua Trivedi, Marian Blazes, Anthony Ortiz, Jocelyn Desbiens, Sunil Gupta, Rahul Dodhia, Pavan K Bhatraju, W Conrad Liles, Aaron Lee, et al. 2021. Deep learning models for COVID-19 chest x-ray classification: Preventing shortcut learning using feature disenta...

  139. [147]

    Abhilasha Ravichander, Joe Stacey, and Marek Rei. 2023. When and Why Does Bias Mitigation Work?. In Findings of the Association for Computational Linguistics: EMNLP 2023. 9233–9247

  140. [148]

    Andrew Slavin Ross, Michael C Hughes, and Finale Doshi-Velez. 2017. Right for the right reasons: Training differentiable models by constraining their explanations. arXiv preprint arXiv:1703.03717 (2017). 27 A preprint - December 9, 2024

  141. [149]

    Berg, and Li Fei-Fei

    Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. 2015. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Visio...

  142. [150]

    Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang. 2019. Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization. arXiv preprint arXiv:1911.08731 (2019)

  143. [151]

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10684–10695

  144. [152]

    Shiori Sagawa, Aditi Raghunathan, Pang Wei Koh, and Percy Liang. 2020. An investigation of why overpa- rameterization exacerbates spurious correlations. In International Conference on Machine Learning. PMLR, 8346–8356

  145. [153]

    clever Hans phenomenon

    Laasya Samhita and Hans J Gross. 2013. The “clever Hans phenomenon” revisited.Communicative & integrative biology 6, 6 (2013), e27122

  146. [154]

    Esha Sarkar, Yousif Alkindi, and Michail Maniatakos. 2020. Backdoor Suppression in Neural Networks using Input Fuzzing and Majority V oting.IEEE Design & Test 37, 2 (2020), 103–110

  147. [155]

    Shiori Sagawa, Pang Wei Koh, Tony Lee, Irena Gao, Sang Michael Xie, Kendrick Shen, Ananya Kumar, Weihua Hu, Michihiro Yasunaga, Henrik Marklund, Sara Beery, Etienne David, Ian Stavness, Wei Guo, Jure Leskovec, Kate Saenko, Tatsunori Hashimoto, Sergey Levine, Chelsea Finn, and ...

  148. [156]

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. 2022. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural I...

  149. [157]

    Philip Sedgwick. 2012. Pearson’s correlation coe fficient. Bmj 345 (2012)

  150. [158]

    Seonguk Seo, Joon-Young Lee, and Bohyung Han. 2022. Information-theoretic bias reduction via causal view of spurious correlation. In Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 36. 2180–2188

  151. [159]

    Patrick Schramowski, Wolfgang Stammer, Stefano Teso, Anna Brugger, Franziska Herbert, Xiaoting Shao, Hans-Georg Luigs, Anne-Katrin Mahlein, and Kristian Kersting. 2020. Making deep neural networks right for the right scientific reasons by interacting with their explanations. N...

  152. [160]

    Amrith Setlur, Saurabh Garg, Virginia Smith, and Sergey Levine. 2024. Prompting is a Double-Edged Sword: Improving Worst-Group Robustness of Foundation Models. InForty-first International Conference on Machine Learning

  153. [161]

    Xiaoting Shao, Arseny Skryagin, Wolfgang Stammer, Patrick Schramowski, and Kristian Kersting. 2021. Right for better reasons: Training differentiable models by constraining their influence functions. In Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 35. 9533–9540

  154. [162]

    Ruoqi Shen, Sebastien Bubeck, and Suriya Gunasekar. 2022. Data Augmentation as Feature Manipulation. In Proceedings of the 39th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 162), Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csa...

  155. [163]

    Preethi Seshadri, Sameer Singh, and Yanai Elazar. 2024. The Bias Amplification Paradox in Text-to-Image Generation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)

  156. [164]

    Qingyi Si, Fandong Meng, Mingyu Zheng, Zheng Lin, Yuanxin Liu, Peng Fu, Yanan Cao, Weiping Wang, and Jie Zhou. 2022. Language Prior Is Not the Only Shortcut: A Benchmark for Shortcut Learning in VQA. arXiv:2210.04692 [cs.CV]

  157. [165]

    Karan Sikka, Indranil Sur, Anirban Roy, Ajay Divakaran, and Susmit Jha. 2023. Detecting trojaned dnns using counterfactual attributions. In 2023 IEEE International Conference on Assured Autonomy (ICAA). IEEE, 76–85

  158. [166]

    David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al. 2016. Mastering the game of Go with deep neural networks and tree search. nature 529, 7587 (2016...

  159. [167]

    Robik Shrestha, Kushal Kafle, and Christopher Kanan. 2022. An investigation of critical issues in bias mitigation techniques. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 1943–1954

  160. [168]

    Peter Spirtes and Kun Zhang. 2016. Causal discovery and inference: concepts and recent methodological advances. In Applied informatics, V ol. 3. Springer, 1–28

  161. [169]

    Wolfgang Stammer, Patrick Schramowski, and Kristian Kersting. 2021. Right for the right concept: Revising neuro-symbolic concepts by interacting with their explanations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 3619–3629

  162. [170]

    Wolfgang Stammer, Antonia Wüst, David Steinmann, and Kristian Kersting. 2024. Neural Concept Binder. Advances in Neural Information Processing Systems (2024)

  163. [171]

    David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al . 2018. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science ...

  164. [172]

    David Steinmann, Wolfgang Stammer, Felix Friedrich, and Kristian Kersting. 2023. Learning to intervene on concept bottlenecks. arXiv preprint arXiv:2308.13453 (2023)

  165. [173]

    Hentschel, Clifton Poth, Dominik Hintersdorf, and Kristian Kersting

    Lukas Struppek, Martin B. Hentschel, Clifton Poth, Dominik Hintersdorf, and Kristian Kersting. 2023. Leverag- ing Diffusion-Based Image Variations for Robust Training on Poisoned Data. Neural Information Processing Systems (NeurIPS) - Workshop on Backdoors in Deep Learning: Th...

  166. [174]

    Bob L Sturm. 2014. A simple method to determine if a music information retrieval system is a "horse". IEEE Transactions on Multimedia 16, 6 (2014), 1636–1644

  167. [175]

    Ryan Steed, Swetasudha Panda, Ari Kobren, and Michael Wick. 2022. Upstream mitigation is not all you need: Testing the bias transfer hypothesis in pre-trained language models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1:...

  168. [176]

    Harini Suresh and John Guttag. 2021. A framework for understanding sources of harm throughout the machine learning life cycle. In Proceedings of the 1st ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization. 1–9

  169. [177]

    Kamil Szyc, Tomasz Walkowiak, and Henryk Maciejewski. 2021. Checking robustness of representations learned by deep neural networks. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer, 399–414

  170. [178]

    Damien Teney, Ehsan Abbasnedjad, and Anton van den Hengel. 2020. Learning what makes a difference from counterfactual examples and gradient supervision. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part X 16. Springer, 580–599

  171. [179]

    Zechen Sun, Yisheng Xiao, Juntao Li, Yixin Ji, Wenliang Chen, and Min Zhang. 2024. Exploring and Mitigating Shortcut Learning for Generative Large Language Models. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Ev...

  172. [180]

    Kshitiz Tiwari, Shuhan Yuan, and Lu Zhang. 2022. Robust Hate Speech Detection via Mitigating Spurious Correlations. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Compu- tational Linguistics and the 12th International Joint Conference o...

  173. [181]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023)

  174. [182]

    Philipp Tschandl, Cliff Rosendahl, and Harald Kittler. 2018. The HAM10000 Dataset: A Large Collection of Multi-Source Dermatoscopic Images of Common Pigmented Skin Lesions. Scientific Data 5 (08 2018)

  175. [183]

    Stefano Teso and Kristian Kersting. 2019. Explanatory interactive machine learning. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society. 239–245

  176. [184]

    Rahul Venkataramani, Parag Dutta, Vikram Melapudi, and Ambedkar Dukkipati. 2024. Causal Feature Alignment: Learning to Ignore Spurious Background Features. In Proceedings of the IEEE /CVF Winter Conference on Applications of Computer Vision. 4666–4674

  177. [185]

    David Wallis and Irène Buvat. 2022. Clever Hans e ffect found in a widely used brain tumour MRI dataset. Medical image analysis 77 (2022), 102368

  178. [186]

    Jinqiang Wang, Rui Hu, Chaoquan Jiang, Rui Hu, and Jitao Sang. 2022. Counterexample Contrastive Learning for Spurious Correlation Elimination. In Proceedings of the 30th ACM International Conference on Multimedia. 4930–4938. 29 A preprint - December 9, 2024

  179. [187]

    Victor Veitch, Alexander D’Amour, Steve Yadlowsky, and Jacob Eisenstein. 2021. Counterfactual invariance to spurious correlations in text classification. Advances in neural information processing systems 34 (2021), 16196–16208

  180. [188]

    Shunxin Wang, Christoph Brune, Raymond Veldhuis, and Nicola Strisciuglio. 2023. DFM-X: Augmentation by leveraging prior knowledge of shortcut learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 129–138

  181. [189]

    Shunxin Wang, Raymond Veldhuis, Christoph Brune, and Nicola Strisciuglio. 2023. What do neural networks learn in image classification? a frequency shortcut perspective. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 1433–1442

  182. [190]

    Xinyi Wang, Wenhu Chen, Michael Saxon, and William Yang Wang. 2021. Counterfactual maximum likelihood estimation for training deep networks. Advances in Neural Information Processing Systems 34 (2021), 25072– 25085

  183. [191]

    Jiaxuan Wang, Sarah Jabbour, Maggie Makar, Michael Sjoding, and Jenna Wiens. 2022. Learning concept credible models for mitigating shortcuts. Advances in neural information processing systems 35 (2022), 33343– 33356

  184. [192]

    Zhao Wang and Aron Culotta. 2021. Robustness to spurious correlations in text classification via automatically generated counterfactuals. In Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 35. 14024– 14031

  185. [193]

    Sheng-Lun Wei, Cheng-Kuang Wu, Hen-Hsen Huang, and Hsin-Hsi Chen. 2024. Unveiling Selection Biases: Ex- ploring Order and Token Sensitivity in Large Language Models. InFindings of the Association for Computational Linguistics: ACL 2024

  186. [194]

    Olivia Wiles, Sven Gowal, Florian Stimberg, Sylvestre-Alvise Rebuffi, Ira Ktena, Krishnamurthy Dvijotham, and Ali Taylan Cemgil. 2022. A Fine-Grained Analysis on Distribution Shift. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, Ap...

  187. [195]

    Yining Wang, Junjie Sun, Chenyue Wang, Mi Zhang, and Min Yang. 2024. Navigate Beyond Shortcuts: Debiased Learning through the Lens of Neural Collapse. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12322–12331

  188. [196]

    Christopher Winship and Robert D Mare. 1992. Models for sample selection bias. Annual review of sociology 18, 1 (1992), 327–350

  189. [197]

    Dongxian Wu and Yisen Wang. 2021. Adversarial Neuron Pruning Purifies Backdoored Deep Models. In Conference on Neural Information Processing Systems (NeurIPS). 16913–16925

  190. [198]

    Jiaying Wu and Bryan Hooi. 2022. Probing spurious correlations in popular event-based rumor detection benchmarks. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 274–290

  191. [199]

    Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technolo...

  192. [200]

    Antonia Wüst, Wolfgang Stammer, Quentin Delfosse, Devendra Singh Dhami, and Kristian Kersting. 2024. Pix2Code: Learning to Compose Neural Visual Concepts as Programs. In The 40th Conference on Uncertainty in Artificial Intelligence

  193. [201]

    2017.Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms

    Han Xiao, Kashif Rasul, and Roland V ollgraf. 2017.Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms. arXiv:cs.LG/1708.07747 [cs.LG]

  194. [202]

    Gunter, and Bo Li

    Xiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov, Carl A. Gunter, and Bo Li. 2021. Detecting AI Trojans Using Meta Neural Analysis. In 42nd IEEE Symposium on Security and Privacy, SP 2021, San Francisco, CA, USA, 24-27 May 2021. IEEE, 103–120

  195. [203]

    Shirley Wu, Mert Yuksekgonul, Linjun Zhang, and James Zou. 2023. Discover and cure: Concept-aware mitigation of spurious correlation. In International Conference on Machine Learning. PMLR, 37765–37786

  196. [204]

    Chandra, Monika Janda, Peter Soyer, and Zongyuan Ge

    Siyuan Yan, Zhen Yu, Xuelin Zhang, Dwarikanath Mahapatra, Shekhar S. Chandra, Monika Janda, Peter Soyer, and Zongyuan Ge. 2023. Towards Trustable Skin Cancer Diagnosis via Rewriting Model’s Decision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...

  197. [205]

    Wanqian Yang, Polina Kirichenko, Micah Goldblum, and Andrew G Wilson. 2022. Chroma-vae: Mitigating shortcut learning with generative classifiers. Advances in Neural Information Processing Systems 35 (2022), 20351–20365. 30 A preprint - December 9, 2024

  198. [206]

    Yu Yang, Eric Gan, Gintare Karolina Dziugaite, and Baharan Mirzasoleiman. 2024. Identifying spurious biases early in training through the lens of simplicity bias. In International Conference on Artificial Intelligence and Statistics. PMLR, 2953–2961

  199. [207]

    Mingfu Xue, Yinghao Wu, Zhiyu Wu, Yushu Zhang, Jian Wang, and Weiqiang Liu. 2023. Detecting backdoor in deep neural networks via intentional adversarial perturbations. Inf. Sci. 634 (2023), 564–577

  200. [208]

    Yao-Yuan Yang, Chi-Ning Chou, and Kamalika Chaudhuri. 2022. Understanding rare spurious correlations in neural networks. arXiv preprint arXiv:2202.05189 (2022)

  201. [209]

    Haotian Ye, James Zou, and Linjun Zhang. 2023. Freeze then train: Towards provable representation learning under spurious correlations and feature noise. In International Conference on Artificial Intelligence and Statistics. PMLR, 8968–8990

  202. [210]

    Wenqian Ye, Guangtao Zheng, Xu Cao, Yunsheng Ma, Xia Hu, and Aidong Zhang. 2024. Spurious Correlations in Machine Learning: A Survey. arXiv preprint arXiv:2402.12715 (2024)

  203. [211]

    Yu Yang, Besmira Nushi, Hamid Palangi, and Baharan Mirzasoleiman. 2023. Mitigating spurious correlations in multi-modal models during fine-tuning. In International Conference on Machine Learning. PMLR, 39365– 39379

  204. [212]

    Yu Yuan, Lili Zhao, Kai Zhang, Guangting Zheng, and Qi Liu. 2024. Do llms overcome shortcut learning? an evaluation of shortcut challenges in large language models. arXiv preprint arXiv:2410.13343 (2024)

  205. [213]

    Yu Yuan, Lili Zhao, Kai Zhang, Guangting Zheng, and Qi Liu. 2024. Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing

  206. [214]

    Samira Zare and Hien Van Nguyen. 2022. Removal of confounders via invariant risk minimization for medical diagnosis. In International Conference on Medical Image Computing and Computer-Assisted Intervention . Springer, 578–587

  207. [215]

    Sriram Yenamandra, Pratik Ramesh, Viraj Prabhu, and Judy Hoffman. 2023. Facts: First amplify correlations and then slice to discover bias. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 4794–4804

  208. [216]

    Kun Zhang, Shaoan Xie, Ignavier Ng, and Yujia Zheng. 2024. Causal Representation Learning from Multiple Distributions: A General Setting. arXiv:2402.05052 [cs.LG]

  209. [217]

    Lily H Zhang and Rajesh Ranganath. 2023. Robustness to spurious correlations improves semantic out-of- distribution detection. In Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 37. 15305–15312

  210. [218]

    Min Zhang, Haoxuan Li, Fei Wu, and Kun Kuang. 2024. Metacoco: A new few-shot classification benchmark with spurious correlation. arXiv preprint arXiv:2404.19644 (2024)

  211. [219]

    Zech, Marcus A

    John R. Zech, Marcus A. Badgeley, Manway Liu, Anthony B. Costa, Joseph J. Titano, and Eric Karl Oermann

  212. [220]

    Xiaoyu Zhang, Rohit Gupta, Ajmal Mian, Nazanin Rahnavard, and Mubarak Shah. 2021. Cassandra: Detecting Trojaned Networks From Adversarial Perturbations. IEEE Access 9 (2021), 135856–135867

  213. [221]

    Xi Zhang, Feifei Zhang, and Changsheng Xu. 2024. NExT-OOD: Overcoming Dual Multiple-Choice VQA Biases. IEEE Transactions on Pattern Analysis and Machine Intelligence46, 4 (2024), 1913–1931

  214. [222]

    Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2017. Men also like shopping: Reducing gender bias amplification using corpus-level constraints. arXiv preprint arXiv:1707.09457 (2017)

  215. [223]

    Qingyu Zhao, Ehsan Adeli, and Kilian M Pohl. 2020. Training confounder-free deep learning models for medical applications. Nature communications 11, 1 (2020), 6010

  216. [224]

    Xingxuan Zhang, Peng Cui, Renzhe Xu, Linjun Zhou, Yue He, and Zheyan Shen. 2021. Deep stable learning for out-of-distribution generalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5372–5382

  217. [225]

    Chunting Zhou, Xuezhe Ma, Paul Michel, and Graham Neubig. 2021. Examining and combating spurious features under distribution shift. In International Conference on Machine Learning. PMLR, 12857–12867. 31

  218. [229]

    Jiayun Zheng and Maggie Makar. 2022. Causally motivated multi-shortcut identification and removal. Advances in Neural Information Processing Systems 35 (2022), 12800–12812

  219. [2015]

    PLOS ONE 10 (10 2015)

    Enhanced Performance of Brain Tumor Classification via Tumor Region Augmentation and Partition. PLOS ONE 10 (10 2015)

  220. [2018]

    PLOS Medicine 15, 11 (2018), e1002683

    Variable Generalization Performance of a Deep Learning Model to Detect Pneumonia in Chest Radiographs: A Cross-Sectional Study. PLOS Medicine 15, 11 (2018), e1002683

  221. [2020]

    In International conference on machine learning

    Concept bottleneck models. In International conference on machine learning. PMLR, 5338–5348

  222. [2024]

    arXiv preprint arXiv:2402.13368 (2024)

    Unsupervised Concept Discovery Mitigates Spurious Correlations. arXiv preprint arXiv:2402.13368 (2024)

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.