Pith. sign in

REVIEW 3 major objections 6 minor 226 references

Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper claims that distribution shift and AI safety align through one-to-one mathematical correspondences that let methods from either field transfer to the other.

desk verdict A useful survey with an overstated claim: the 'formal reducibility' between distribution shift and safety issues doesn't survive contact with dirty-label backdoor attacks. read the letter →

arxiv 2505.22829 v2 pith:DHM4HBH4 submitted 2025-05-28 cs.LG cs.AI

classification cs.LGcs.AI
keywords distributionshiftAIsafetyspuriouscorrelationselectionbiaslabelinvariantlearningalgorithmicfairnessbackdoorattacks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that distribution shift and AI safety are not loosely connected topics but align through specific mathematical relationships. It categorizes distribution shifts by cause (selection bias, spurious correlation, and label shift, with finer subtypes) and AI safety issues (security, fairness, trustworthiness, and democracy), then pairs each shift cause with a concrete safety issue either because methods for one achieve the other or because the two definitions formally reduce to one another. The payoff would be a transfer of methodology: invariant learning and out-of-distribution tools become fairness and alignment tools, while backdoor defenses and uncertainty scores become spurious-correlation and distribution-shift tools. A sympathetic reader would take the paper as an organizing map that makes cross-field method transfer systematic rather than anecdotal.

What carries the argument

The load-bearing machine is a set of definitions that encode each shift cause as a conditional independence statement about features, and each safety issue as the same kind of statement with roles renamed. Definition 5 (environmental change) requires $Y \perp\!\!\perp E \mid C(X)$ while $Y$ is not independent of $E$ given the spurious feature $Z(X)$; Definition 15 (generalized label shift) requires $P(\Phi(X) \mid Y, E)$ to be invariant across environments; Definition 16 (open-set label shift) formalizes classes that appear only in the target domain. The argument proceeds by renaming variables: the environment $E$ becomes a protected attribute $R$, the content feature $C(X)$ becomes a predictive score $S(X)$, a backdoor trigger becomes a spurious feature $Z(X)$, and so on. When the renamed definitions coincide, the paper concludes that methods developed for one side can be mutually adapted to the other.

What would settle it

One concrete check is to construct a two-environment classification problem satisfying Definition 5 in which a representation $\Phi(X)$ satisfies the environment invariance constraint but the Bayes classifier built on it does not satisfy test fairness because $\Phi(X)$ is not a minimal sufficient statistic for $Y$; if such a case exists, the claimed equivalence between environmental change and test fairness breaks and the Section 4.1.1 reduction would need qualification.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes two types of connection between specific causes of distribution shift and fine-grained AI safety issues. First, methods that address a particular shift cause can achieve a corresponding safety goal, for example invariant learning under environmental change yields test fairness, and data pruning is the dual of handling individual selection bias. Second, certain shift definitions and safety definitions reduce to each other: environmental change becomes test fairness or inner alignment when the environment variable is read as a protected attribute or as a trigger, generalized label shift becomes equalized odds, and open-set label shift becomes uncertainty quantification. The paper presents these as definitional equivalences, not merely analogies, and reviews the literature on both sides to show that the method families already resemble each other.

Load-bearing premise

The whole alignment rests on assuming that every distribution shift of interest admits a clean split into invariant content features $C(X)$ and environment-sensitive spurious features $Z(X)$ with $Y$ conditionally independent of the environment given $C(X)$; if real data lack such a decomposition, or it cannot be identified, the formal reductions in Section 4 do not go through.

Editorial extensions

If this is right

  • Invariant learning methods such as invariant risk minimization and invariant causal prediction could be applied directly to enforce test fairness and well-calibration, because the conditional-invariance objective coincides with sufficiency-based fairness when the environment is the protected attribute.
  • Data pruning and importance reweighting become two views of the same problem, so pruning scores could serve as reweighting weights and reweighting weights could guide which points to prune.
  • Backdoor defenses that identify trigger-related patterns could be repurposed to detect and remove spurious features under environmental change, while spurious-correlation robust training could become a backdoor defense.
  • Methods that enforce generalized label shift, including class-conditional distribution matching and contrastive learning, can be used to satisfy equalized odds with respect to the environment as a protected attribute.
  • Open-set label shift and uncertainty quantification are mutually informing: threshold-based unseen-class detection is an uncertainty score, and uncertainty scores can flag inputs whose labels were absent during training.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the reductions in Section 4.1.1 hold as stated, then in binary settings test fairness and invariant learning are the same condition, which would mean any improved invariant learning algorithm immediately yields a fairness algorithm; this is a transfer neither community currently exploits.
  • The clean split into content features $C(X)$ and spurious features $Z(X)$ in Definition 5 is an assumption, and the paper itself notes in Section 4.2 that the distinction is not binary, so before the mappings are used in practice one would need to show when such features exist and are identifiable from data.
  • A testable extension is to benchmark backdoor defenses on standard spurious-correlation datasets and spurious-correlation methods on poisoned datasets, measuring whether the transferred methods preserve both clean accuracy and worst-group or attack-success metrics.
  • The framework suggests that newer shift types, such as performative shift, will also have safety counterparts, but the paper does not map them; finding those counterparts would test whether the one-to-one alignment extends beyond the six pairings treated here.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes a systematic framework for relating specific causes of distribution shift to concrete AI safety issues. It defines six fine-grained shift types under selection bias, spurious correlation, and label shift, and maps them onto safety topics including fairness, security, trustworthiness, and democracy. For each pairing, it claims either that shift-mitigation methods can help achieve a safety goal (type 1) or that the two problems can be formally reduced to one another (type 2), and it reviews relevant literature. The abstract and Section 6 present formal reducibility as a headline contribution.

Significance. If the type-2 reductions were fully established, the paper would provide a valuable unifying bridge between two active research areas, enabling methodological transfer and a shared vocabulary. The paper's strengths are its breadth, its explicit definitions (e.g., Definitions 5, 13, 15), and its careful cataloging of related work, including useful elementary observations such as the link between invariant conditional distributions and test fairness and between generalized label shift and equalized odds. However, the central formal claims are not supported by proofs: the most important reduction, from backdoor poisoning attacks to environmental change, fails for dirty-label attacks, and the key decomposition in Definition 5 is assumed rather than derived. The paper is therefore best read as a conceptual and methodological survey, not as a set of proven reductions.

major comments (3)
  1. [Section 4.1.3] The claimed formal reduction of backdoor poisoning attacks to environmental change does not hold for dirty-label attacks. Definition 5 requires both Y not perpendicular E given Z(X) for a spurious feature and Y perpendicular E given C(X) for an invariant content feature; the text verifies only the first condition for the trigger Z. In the dirty-label case of Definition 13, poisoned samples are drawn as P(X) times an indicator of Y=y_t, so for any candidate C, P(Y|C(X)=c,E=e_p) is a mixture of the clean label distribution and a point mass at y_t, while P(Y|C(X)=c,E=e_r) is the clean distribution alone. These differ whenever p is in (0,1) and the clean distribution is not degenerate, so no C(X) satisfies the required invariance. The footnote noting that some attacks may not fully satisfy Definition 13 does not repair the argument, because the reduction claim is made for the general case that includes dirty-label attacks.
  2. [Definition 5 and Section 4.2] The type-2 reductions rest on the assumed existence of an invariant content/spurious decomposition of X, but no existence, identifiability, or uniqueness argument is supplied. Definition 5 quantifies over 'there exist C(X) and Z(X)', and Section 4.2 later states that the distinction between C(X) and Z(X) is not binary. Without conditions under which such a decomposition exists, the claimed formal reductions to test fairness, inner alignment, and backdoor attacks are conditional on an unverified premise. The paper should either prove the existence for each setting it discusses or explicitly state the additional assumptions needed, and it should temper the abstract's claim of formal reducibility accordingly.
  3. [Section 3.1.1] The claimed duality between data pruning and individual selection bias is asserted through an interpretive choice rather than derived. The text says 'if we interpret P_original as the target distribution and P_selected as the source distribution,' but in individual selection bias both source and target distributions are given, whereas in data pruning the selected distribution is an algorithm's choice and the objective is performance on the original distribution, not on a pre-specified target. This is a useful analogy, but it does not establish the formal 'dual problem' relationship claimed, and the conclusion that methods from the two areas are interchangeable should be correspondingly qualified.
minor comments (6)
  1. [Definition 10] The definition title contains a typo: 'Eqalized Odds' should be 'Equalized Odds'.
  2. [Section 4.1.3] The text contains a typo: 'posioned environment' should be 'poisoned environment'.
  3. [Section 5.1] The sentence 'Noted that we adopt Definition 15' should read 'Note that we adopt Definition 15'.
  4. [Section 4.1.2] In the discussion of deceptive alignment, 'mesa-objectives are awared of' should be 'mesa-objectives are aware of'.
  5. [Abstract] The abstract contains a grammatical error: 'two types connections' should be 'two types of connections'.
  6. [Figure 2] The small text in Figure 2 may be difficult to read in print; a higher-resolution version or larger font would improve clarity.

Circularity Check

3 steps flagged · score 5.0 of 10

Type-2 'formal reductions' are largely definitional identities obtained by renaming variables; the survey and type-1 links remain independent, so circularity is partial, not severe.

  1. self definitional [Section 4.1.1 (Test Fairness vs Environmental Change)]
    "In this setting, if we consider the environment variable 𝐸 as representing the protected attribute 𝑅, then the representation Φ(𝑋) in invariant learning is formally consistent with the predictive score 𝑆(𝑋) in test fairness, and can thus be regarded as an extension of sufficiency at the feature level."

    Definition 5 defines environmental change via the existence of a content feature C(X) with P(Y|C(X),E=e_i)=P(Y|C(X),E=e_j) for all environments, while Definition 7 defines test fairness as P(Y=1|S=s,R=r1)=P(Y=1|S=s,R=r2). Substituting R=E and S=C(X) makes the two conditions syntactically identical. The claimed 'formal reduction' is therefore an equality by variable renaming in the paper's own definitions, not the outcome of an independent derivation or of evidence beyond those definitions.

  2. self definitional [Section 5.1.1 (Generalized Label Shift vs Equalized Odds)]
    "When using the same classifier 𝜔, GLS also ensures that 𝑃(𝑓(𝑋)|𝑌,𝐸=𝑒) remains invariant across different environments 𝑒, where 𝑓=𝜔◦Φ. As illustrated in the Figure 7, this is equivalent to the predicted output ˆ𝑌 satisfying equalized odds (Definition 10) with respect to the environment 𝐸."

    Definition 15 postulates a feature Φ(X) satisfying P(Φ(X)|Y,E=e_i)=P(Φ(X)|Y,E=e_j), and Definition 10 defines equalized odds as P(Yhat=1|Y=c,R=r1)=P(Yhat=1|Y=c,R=r2). Setting Yhat=f(X), R=E, and writing the invariance condition at the output level makes the two statements the same equation. The claimed equivalence is thus a notational relabeling of the generalized label shift condition, so the safety property is built into the distribution-shift definition rather than derived from it.

1 more flagged steps
  1. renaming known result [Section 3.1.1 (Data Pruning as a dual of Individual Selection Bias)]
    "From this definition, there is a natural correspondence between addressing individual selection bias and data pruning, if we interpret 𝑃original(𝑋,𝑌) as the target distribution and 𝑃selected(𝑋,𝑌) as the source distribution."

    The reported 'duality' is manufactured by labeling the pruned data distribution as the source and the original distribution as the target. After that labeling, data pruning's objective (a model trained on the selected subset performs well on the original distribution) is exactly the individual-selection-bias objective (a model trained on the source performs well on the target). The connection is an interpretive renaming of the same objective, not an independently established equivalence between two distinct problems.

full rationale

The type-1 connections (methods for a specific shift type help achieve a safety goal) are literature reviews and are not circular: they report existing methods and are not derived from the paper's own definitions. No parameters are fitted, and the paper does not present numerical predictions obtained from its own outputs. The central type-2 'formal reductions' are, however, partly self-definitional: the paper defines environmental change (Definition 5), invariant learning, and generalized label shift (Definition 15) in conditional-independence language, and the corresponding fairness definitions (test fairness, equalized odds) express the same conditional-independence conditions with different variable names. The paper itself acknowledges the relabeling (e.g., 'if we consider E as R'), so the reductions are transparent rather than hidden, but they are nonetheless equalities by construction. The data-pruning duality is likewise created by labeling P_original as the target and P_selected as the source. None of these steps involves fitted inputs, and no load-bearing self-citation chain is present: the authors' self-references (e.g., [105], [106], [139], [140], [163], [191], [220]) are ordinary methodological citations and do not justify the formal equivalences. The backdoor-to-environmental-change claim in Section 4.1.3 is not counted as circularity because it is a validity gap rather than a reduction-to-inputs: the paper verifies only the spurious-feature half of Definition 5 and does not establish the required invariant content feature for dirty-label attacks; this is a correctness risk but not a circular derivation. Overall, because the central type-2 connections are definitional identities while the type-1 survey content is independent and the relabelings are disclosed, a moderate circularity score is appropriate.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no fitted parameters. Its central claims rest on definitional constructs: the content/spurious feature decomposition in Definition 5, the source/target interpretation in the data pruning duality, and the class-conditional invariance interpretation of GLS. These are useful modeling conventions, but they are not independently established facts, so the question of whether the mappings are discovered or manufactured is open.

assumptions (4)
  • ad hoc to paper Environmental change admits an invariant content/spurious feature decomposition (Definition 5).
    Section 4.1 asserts that every environmental change has C(X), Z(X) with Y independent of E given C(X) and Y not independent given Z(X); no existence or identifiability proof is given, yet Sections 4.1.1, 4.1.2, and 4.1.3 build on it.
  • ad hoc to paper Data pruning and individual selection bias can be related by interpreting the original distribution as target and the selected distribution as source.
    Section 3.1.1 calls data pruning a dual problem to individual selection bias via this interpretive choice; the correspondence is true by construction rather than by a demonstrated reduction.
  • domain assumption Generalized label shift together with a shared classifier yields equalized odds with respect to the environment.
    Section 5.1.1 states that GLS, with a fixed classifier, makes the model output satisfy equalized odds; this is a definitional consequence, not an empirical finding, and it depends on the feature Phi(X) actually carrying all class-conditional invariance.
  • ad hoc to paper The Bayes-error-preserving feature extractor class F_e is non-empty and the content feature lies in the intersection over all environments.
    Definition 5 posits a content feature in the intersection of F_e across environments; no argument guarantees such a feature exists in realistic datasets, and the paper does not discuss identification conditions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies." pith.science (2026). https://pith.science/paper/DHM4HBH4

@misc{pith2026250522829,
  author       = {Pith},
  title        = {Pith review of: Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DHM4HBH4}},
  note         = {Machine review of arXiv:2505.22829}
}
read the original abstract

This paper bridges distribution shift and AI safety through a comprehensive analysis of their conceptual and methodological synergies. While prior discussions often focus on narrow cases or informal analogies, we establish two types connections between specific causes of distribution shift and fine-grained AI safety issues: (1) methods addressing a specific shift type can help achieve corresponding safety goals, or (2) certain shifts and safety issues can be formally reduced to each other, enabling mutual adaptation of their methods. Our findings provide a unified perspective that encourages deeper integration between distribution shift and AI safety research.

Figures

Figures reproduced from arXiv: 2505.22829 by the authors.

Figure 1
Figure 1. Connections between distribution shift causes and AI safety issues. Dashed arrows: Methods addressing a specific cause of [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Summary of safety issues related to distribution shift covered in this paper. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Illustration of the inclusion relationships among different causes of distribution shift. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: An illustration of individual selection bias and group selection bias. The blue curves represent the input distribution of the [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Data pruning is a dual problem of addressing indi [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Illustration of the negative impact of environmental change on a model trained via empirical risk minimization (ERM), using [PITH_FULL_IMAGE:figures/full_fig_p012_6.png]
Figure 7
Figure 7. Figure 7: Summary of group fairness and their relationships with different causes of distribution shift. For risk-based fairness notions, [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: The relationship between inner alignment and outer alignment. [PITH_FULL_IMAGE:figures/full_fig_p018_8.png]
Figure 9
Figure 9. Figure 9: The connection between backdoor poisoning attacks and environmental change. By injecting a trigger, both dirty-label attacks [PITH_FULL_IMAGE:figures/full_fig_p020_9.png]
Figure 10
Figure 10. Figure 10: Difference between environmental change (Definition 5) and content change (Definition 14) in the Waterbirds example. [PITH_FULL_IMAGE:figures/full_fig_p022_10.png]
Figure 11
Figure 11. Figure 11: An illustration of label shift and open-set label shift (OSLS). Label shift is the general setting where the proportions of labels [PITH_FULL_IMAGE:figures/full_fig_p026_11.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

226 extracted references · 32 canonical work pages

  1. [1]

    Faruk Ahmed, Yoshua Bengio, Harm Van Seijen, and Aaron Courville. 2020. Systematic generalisation with group invariant predictions. In International Conference on Learning Representations

  2. [2]

    Kartik Ahuja, Ethan Caballero, Dinghuai Zhang, Jean-Christophe Gagnon-Audet, Yoshua Bengio, Ioannis Mitliagkas, and Irina Rish. 2021. Invariance principle meets information bottleneck for out-of-distribution generalization.Advances in Neural Information Processing Systems34 (2021), 3438– 3450. Manuscript submitted to ACM 28 Liu et al

  3. [3]

    Kartik Ahuja, Karthikeyan Shanmugam, Kush Varshney, and Amit Dhurandhar. 2020. Invariant risk minimization games. InInternational Conference on Machine Learning. PMLR, 145–155

  4. [4]

    Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. 2016. Concrete problems in AI safety.arXiv preprint arXiv:1606.06565(2016)

  5. [5]

    Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz. 2019. Invariant risk minimization.arXiv preprint arXiv:1907.02893(2019)

  6. [6]

    Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. 2021. A general language assistant as a laboratory for alignment.arXiv preprint arXiv:2112.00861(2021)

  7. [7]

    Yang Bai, Gaojie Xing, Hongyan Wu, Zhihong Rao, Chuan Ma, Shiping Wang, Xiaolei Liu, Yimin Zhou, Jiajia Tang, Kaijun Huang, et al. 2024. Backdoor Attack and Defense on Deep Learning: A Survey.IEEE Transactions on Computational Social Systems(2024)

  8. [8]

    Mahsa Baktashmotlagh, Mehrtash T Harandi, Brian C Lovell, and Mathieu Salzmann. 2013. Unsupervised domain adaptation by domain invariant projection. InProceedings of the IEEE international conference on computer vision. 769–776

Show all 226 references
  1. [9]

    Stephen Bates, Anastasios Angelopoulos, Lihua Lei, Jitendra Malik, and Michael Jordan. 2021. Distribution-free, risk-controlling prediction sets. Journal of the ACM (JACM)68, 6 (2021), 1–34

  2. [10]

    Emanuel Ben-Baruch, Adam Botach, Igor Kviatkovsky, Manoj Aggarwal, and Gérard Medioni. 2024. Distilling the Knowledge in Data Pruning. arXiv preprint arXiv:2403.07854(2024)

  3. [11]

    Aharon Ben-Tal, Dick Den Hertog, Anja De Waegenaere, Bertrand Melenberg, and Gijs Rennen. 2013. Robust solutions of optimization problems affected by uncertain probabilities.Management Science59, 2 (2013), 341–357

  4. [12]

    Leonard Bereska and Efstratios Gavves. 2024. Mechanistic Interpretability for AI Safety–A Review.arXiv preprint arXiv:2404.14082(2024)

  5. [13]

    Battista Biggio and Fabio Roli. 2018. Wild patterns: Ten years after the rise of adversarial machine learning. InProceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security. 2154–2156

  6. [14]

    Karsten M Borgwardt, Arthur Gretton, Malte J Rasch, Hans-Peter Kriegel, Bernhard Schölkopf, and Alex J Smola. 2006. Integrating structured biological data by kernel maximum mean discrepancy.Bioinformatics22, 14 (2006), e49–e57

  7. [15]

    Ruichu Cai, Zijian Li, Pengfei Wei, Jie Qiao, Kun Zhang, and Zhifeng Hao. 2019. Learning disentangled semantic representation for domain adaptation. InIJCAI: proceedings of the conference, Vol. 2019. NIH Public Access, 2060

  8. [16]

    Andres Carranza, Dhruv Pai, Rylan Schaeffer, Arnuv Tandon, and Sanmi Koyejo. 2023. Deceptive alignment monitoring.arXiv preprint arXiv:2307.10569(2023)

  9. [17]

    Shiyu Chang, Yang Zhang, Mo Yu, and Tommi Jaakkola. 2020. Invariant rationalization. InInternational Conference on Machine Learning. PMLR, 1448–1458

  10. [18]

    Annie S Chen, Yoonho Lee, Amrith Setlur, Sergey Levine, and Chelsea Finn. 2023. Project and probe: Sample-efficient domain adaptation by interpolating orthogonal features.arXiv preprint arXiv:2302.05441(2023)

  11. [19]

    Bryant Chen, Wilka Carvalho, Nathalie Baracaldo, Heiko Ludwig, Benjamin Edwards, Taesung Lee, Ian Molloy, and Biplav Srivastava. 2018. Detecting backdoor attacks on deep neural networks by activation clustering.arXiv preprint arXiv:1811.03728(2018)

  12. [20]

    Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song. 2017. Targeted backdoor attacks on deep learning systems using data poisoning. arXiv preprint arXiv:1712.05526(2017)

  13. [21]

    Yimeng Chen, Ruibin Xiong, Zhi-Ming Ma, and Yanyan Lan. 2022. When does group invariant learning survive spurious correlations?Advances in Neural Information Processing Systems35 (2022), 7038–7051

  14. [22]

    Mathieu Chevalley, Charlotte Bunne, Andreas Krause, and Stefan Bauer. 2022. Invariant causal mechanisms through distribution matching.arXiv preprint arXiv:2206.11646(2022)

  15. [23]

    Seun-An Choe, Ah-Hyung Shin, Keon-Hee Park, Jinwoo Choi, and Gyeong-Moon Park. 2024. Open-set domain adaptation for semantic segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 23943–23953

  16. [24]

    Alexandra Chouldechova. 2017. Fair prediction with disparate impact: A study of bias in recidivism prediction instruments.Big data5, 2 (2017), 153–163

  17. [25]

    Cody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman, Peter Bailis, Percy Liang, Jure Leskovec, and Matei Zaharia. 2020. Selection via Proxy: Efficient Data Selection for Deep Learning. InInternational Conference on Learning Representations. https://openrevie...

  18. [26]

    Elliot Creager, Jörn-Henrik Jacobsen, and Richard Zemel. 2021. Environment inference for invariant learning. InInternational Conference on Machine Learning. PMLR, 2189–2200

  19. [27]

    Edwige Cyffers, Muni Sreenivas Pydi, Jamal Atif, and Olivier Cappé. 2024. Optimal classification under performative distribution shift.arXiv preprint arXiv:2411.02023(2024)

  20. [28]

    Yihe Deng, Yu Yang, Baharan Mirzasoleiman, and Quanquan Gu. 2023. Robust learning with progressive data expansion against spurious correlation. Advances in neural information processing systems36 (2023), 1390–1402

  21. [29]

    Lauro Langosco Di Langosco, Jack Koch, Lee D Sharkey, Jacob Pfau, and David Krueger. 2022. Goal misgeneralization in deep reinforcement learning. InInternational Conference on Machine Learning. PMLR, 12004–12019

  22. [30]

    Wenhao Ding, Laixi Shi, Yuejie Chi, and Ding Zhao. 2024. Seeing is not believing: Robust reinforcement learning against spurious correlation. Advances in Neural Information Processing Systems36 (2024). Manuscript submitted to ACM Bridging Distribution Shift and AI Safety: Conc...

  23. [31]

    Andrea Dittadi, Frederik Träuble, Francesco Locatello, Manuel Wüthrich, Vaibhav Agrawal, Ole Winther, Stefan Bauer, and Bernhard Schölkopf

  24. [32]

    2020.Machine learning in finance

    Matthew F Dixon, Igor Halperin, and Paul Bilokon. 2020.Machine learning in finance. Vol. 1170. Springer

  25. [33]

    Bao Gia Doan, Ehsan Abbasnejad, and Damith C Ranasinghe. 2020. Februus: Input purification defense against trojan attacks on deep neural network systems. InProceedings of the 36th Annual Computer Security Applications Conference. 897–912

  26. [34]

    Finale Doshi-Velez and Been Kim. 2017. Towards a rigorous science of interpretable machine learning.arXiv preprint arXiv:1702.08608(2017)

  27. [35]

    Cian Eastwood, Shashank Singh, Andrei L Nicolicioiu, Marin Vlastelica Pogančić, Julius von Kügelgen, and Bernhard Schölkopf. 2024. Spuriosity didn’t kill the classifier: Using invariant predictions to harness spurious features.Advances in Neural Information Processing Systems36 (2024)

  28. [36]

    Tom Everitt and Marcus Hutter. 2018. The alignment problem for Bayesian history-based reinforcement learners.Under submission(2018)

  29. [37]

    Tongtong Fang, Nan Lu, Gang Niu, and Masashi Sugiyama. 2020. Rethinking importance weighting for deep learning under distribution shift. Advances in neural information processing systems33 (2020), 11996–12007

  30. [38]

    Bo Fu, Zhangjie Cao, Mingsheng Long, and Jianmin Wang. 2020. Learning to detect open classes for universal domain adaptation. InComputer vision–ECCV 2020: 16th European conference, glasgow, UK, August 23–28, 2020, proceedings, part XV 16. Springer, 567–583

  31. [39]

    2013.Introduction to statistical pattern recognition

    Keinosuke Fukunaga. 2013.Introduction to statistical pattern recognition. Elsevier

  32. [40]

    Marco Fumero, Florian Wenzel, Luca Zancato, Alessandro Achille, Emanuele Rodolà, Stefano Soatto, Bernhard Schölkopf, and Francesco Locatello

  33. [41]

    Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario March, and Victor Lempitsky. 2016. Domain-adversarial training of neural networks.Journal of machine learning research17, 59 (2016), 1–35

  34. [42]

    Yansong Gao, Bao Gia Doan, Zhi Zhang, Siqi Ma, Jiliang Zhang, Anmin Fu, Surya Nepal, and Hyoungshick Kim. 2020. Backdoor attacks and countermeasures on deep learning: A comprehensive review.arXiv preprint arXiv:2007.10760(2020)

  35. [43]

    Yansong Gao, Change Xu, Derui Wang, Shiping Chen, Damith C Ranasinghe, and Surya Nepal. 2019. Strip: A defence against trojan attacks on deep neural networks. InProceedings of the 35th annual computer security applications conference. 113–125

  36. [44]

    Saurabh Garg, Sivaraman Balakrishnan, and Zachary Lipton. 2022. Domain adaptation under open set label shift.Advances in Neural Information Processing Systems35 (2022), 22531–22546

  37. [45]

    ZongYuan Ge, Sergey Demyanov, Zetao Chen, and Rahil Garnavi. 2017. Generative openmax for multi-class open set classification.arXiv preprint arXiv:1707.07418(2017)

  38. [46]

    Wichmann, and Wieland Brendel

    Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel. 2019. ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness.. InInternational Conference on Learning Representations...

  39. [47]

    Muhammad Ghifary, David Balduzzi, W Bastiaan Kleijn, and Mengjie Zhang. 2016. Scatter component analysis: A unified framework for domain adaptation and domain generalization.IEEE transactions on pattern analysis and machine intelligence39, 7 (2016), 1414–1430

  40. [48]

    Behnam Gholami, Pritish Sahu, Ognjen Rudovic, Konstantinos Bousmalis, and Vladimir Pavlovic. 2020. Unsupervised multi-target domain adaptation: An information theoretic approach.IEEE Transactions on Image Processing29 (2020), 3993–4002

  41. [49]

    Alice Gomstyn and Alexandra Jonker. 2024. Democratizing AI: What does it mean and how does it work? https://www.ibm.com/think/insights/ democratizing-ai Online; accessed 2025-02-22

  42. [50]

    Alice Gomstyn and Alexandra Jonker. 2024. Democratizing AI: What does it mean and how does it work? https://www.ibm.com/think/insights/ democratizing-ai Online; accessed 2025-03-12

  43. [51]

    Mingming Gong, Kun Zhang, Tongliang Liu, Dacheng Tao, Clark Glymour, and Bernhard Schölkopf. 2016. Domain adaptation with conditional transferable components. InInternational conference on machine learning. PMLR, 2839–2848

  44. [52]

    Sorin Grigorescu, Bogdan Trasnea, Tiberiu Cocias, and Gigel Macesanu. 2020. A survey of deep learning techniques for autonomous driving. Journal of field robotics37, 3 (2020), 362–386

  45. [53]

    Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg. 2017. Badnets: Identifying vulnerabilities in the machine learning model supply chain. arXiv preprint arXiv:1708.06733(2017)

  46. [54]

    Ruocheng Guo, Pengchuan Zhang, Hao Liu, and Emre Kiciman. 2021. Out-of-distribution prediction with invariant risk minimization: The limitation and an effective fix.arXiv preprint arXiv:2101.07732(2021)

  47. [55]

    Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell. 2017. The off-switch game. InWorkshops at the Thirty-First AAAI Conference on Artificial Intelligence

  48. [56]

    Moritz Hardt, Eric Price, and Nati Srebro. 2016. Equality of opportunity in supervised learning.Advances in neural information processing systems 29 (2016)

  49. [57]

    Tatsunori Hashimoto, Megha Srivastava, Hongseok Namkoong, and Percy Liang. 2018. Fairness without demographics in repeated loss minimization. InInternational Conference on Machine Learning. PMLR, 1929–1938

  50. [58]

    Muyang He, Shuo Yang, Tiejun Huang, and Bo Zhao. 2024. Large-scale dataset pruning with dynamic uncertainty. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7713–7722

  51. [59]

    Wenchong He and Zhe Jiang. 2023. A survey on uncertainty quantification methods for deep neural networks: An uncertainty source perspective. arXiv preprint arXiv:2302.13425(2023). Manuscript submitted to ACM 30 Liu et al

  52. [60]

    Xuanli He, Qiongkai Xu, Jun Wang, Benjamin Rubinstein, and Trevor Cohn. 2023. Mitigating backdoor poisoning attacks through the lens of spurious correlation.arXiv preprint arXiv:2305.11596(2023)

  53. [61]

    Christina Heinze-Deml, Jonas Peters, and Nicolai Meinshausen. 2018. Invariant causal prediction for nonlinear models.Journal of Causal Inference 6, 2 (2018), 20170016

  54. [62]

    Jose Hernandez-Orallo, Fernando Martinez-Plumed, Santiago Escobar, and Pablo A. M. Casares. 2022. Distributional shift — The TAILOR Handbook of Trustworthy AI. https://prafra.github.io/jupyter-book-TAILOR-D3.2/T3.2/distributional_shift.html Online; accessed 2025-02-22

  55. [63]

    Mengting Hu, Zhen Zhang, Shiwan Zhao, Minlie Huang, and Bingzhe Wu. 2023. Uncertainty in natural language processing: Sources, quantification, and applications.arXiv preprint arXiv:2306.04459(2023)

  56. [64]

    Yupeng Hu, Wenxin Kuang, Zheng Qin, Kenli Li, Jiliang Zhang, Yansong Gao, Wenjia Li, and Keqin Li. 2021. Artificial intelligence security: Threats and countermeasures.ACM Computing Surveys (CSUR)55, 1 (2021), 1–36

  57. [65]

    Andong Hua, Jindong Gu, Zhiyu Xue, Nicholas Carlini, Eric Wong, and Yao Qin. 2024. Initialization matters for adversarial transfer learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 24831–24840

  58. [66]

    Kunzhe Huang, Yiming Li, Baoyuan Wu, Zhan Qin, and Kui Ren. 2022. Backdoor defense via decoupling the training process.arXiv preprint arXiv:2202.03423(2022)

  59. [67]

    Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M Ziegler, Tim Maxwell, Newton Cheng, et al. 2024. Sleeper agents: Training deceptive llms that persist through safety training.arXiv preprint arXiv:2401.05566(2024)

  60. [68]

    Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. 2019. Risks from learned optimization in advanced machine learning systems.arXiv preprint arXiv:1906.01820(2019)

  61. [69]

    Eyke Hüllermeier and Willem Waegeman. 2021. Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods. Machine learning110, 3 (2021), 457–506

  62. [70]

    Badr Youbi Idrissi, Martin Arjovsky, Mohammad Pezeshki, and David Lopez-Paz. 2022. Simple data balancing achieves competitive worst-group- accuracy. InConference on Causal Learning and Reasoning. PMLR, 336–351

  63. [71]

    2020.Information technology — Artificial intelligence — Overview of trustworthiness in artificial intelligence

    International Organization for Standardization and International Electrotechnical Commission. 2020.Information technology — Artificial intelligence — Overview of trustworthiness in artificial intelligence. Standard ISO/IEC TR 24028:2020. International Organization for Standard...

  64. [72]

    Pavel Izmailov, Polina Kirichenko, Nate Gruver, and Andrew G Wilson. 2022. On feature learning in the presence of spurious correlations.Advances in Neural Information Processing Systems35 (2022), 38516–38532

  65. [73]

    Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, et al. 2023. Ai alignment: A comprehensive survey.arXiv preprint arXiv:2310.19852(2023)

  66. [74]

    Wengong Jin, Regina Barzilay, and Tommi Jaakkola. 2020. Domain extrapolation via regret minimization.arXiv preprint arXiv:2006.039083 (2020)

  67. [75]

    Guoliang Kang, Lu Jiang, Yi Yang, and Alexander G Hauptmann. 2019. Contrastive adaptation network for unsupervised domain adaptation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4893–4902

  68. [76]

    Kimmo Karkkainen and Jungseock Joo. 2021. Fairface: Face attribute dataset for balanced race, gender, and age for bias measurement and mitigation. InProceedings of the IEEE/CVF winter conference on applications of computer vision. 1548–1558

  69. [77]

    Ori Katz, Ronen Talmon, and Uri Shaham. 2024. Supervised Domain Adaptation Based on Marginal and Conditional Distributions Alignment. Transactions on Machine Learning Research(2024)

  70. [78]

    Timo Kaufmann, Paul Weng, Viktor Bengs, and Eyke Hüllermeier. 2023. A survey of reinforcement learning from human feedback.arXiv preprint arXiv:2312.14925(2023)

  71. [79]

    Davinder Kaur, Suleyman Uslu, Kaley J Rittichier, and Arjan Durresi. 2022. Trustworthy artificial intelligence: a review.ACM computing surveys (CSUR)55, 2 (2022), 1–38

  72. [80]

    Polina Kirichenko, Pavel Izmailov, and Andrew Gordon Wilson. 2022. Last layer re-training is sufficient for robustness to spurious correlations. arXiv preprint arXiv:2204.02937(2022)

  73. [81]

    Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, et al. 2021. Wilds: A benchmark of in-the-wild distribution shifts. InInternational conference on machine learni...

  74. [82]

    Masanori Koyama and Shoichiro Yamaguchi. 2020. When is invariance useful in an out-of-distribution generalization problem?arXiv preprint arXiv:2008.01883(2020)

  75. [83]

    David Krueger, Ethan Caballero, Joern-Henrik Jacobsen, Amy Zhang, Jonathan Binas, Dinghuai Zhang, Remi Le Priol, and Aaron Courville. 2021. Out-of-distribution generalization via risk extrapolation (rex). InInternational conference on machine learning. PMLR, 5815–5826

  76. [84]

    Kun Kuang, Ruoxuan Xiong, Peng Cui, Susan Athey, and Bo Li. 2020. Stable prediction with model misspecification and agnostic distribution shift. Proceedings of the AAAI Conference on Artificial Intelligence34, 04 (Apr. 2020), 4485–4492

  77. [85]

    Sébastien Lachapelle, Tristan Deleu, Divyat Mahajan, Ioannis Mitliagkas, Yoshua Bengio, Simon Lacoste-Julien, and Quentin Bertrand. 2023. Synergies between disentanglement and sparsity: Generalization and identifiability in multi-task learning. InInternational Conference on Ma...

  78. [86]

    Phuc H Le-Khac, Graham Healy, and Alan F Smeaton. 2020. Contrastive representation learning: A framework and review.Ieee Access8 (2020), 193907–193934. Manuscript submitted to ACM Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies 31

  79. [87]

    Jungsoo Lee, Eungyeup Kim, Juyoung Lee, Jihyeon Lee, and Jaegul Choo. 2021. Learning debiased representation via disentangled feature augmentation.Advances in Neural Information Processing Systems34 (2021), 25123–25133

  80. [88]

    Yoonho Lee, Huaxiu Yao, and Chelsea Finn. 2023. Diversify and disambiguate: Out-of-distribution robustness via disagreement. InThe Eleventh International Conference on Learning Representations

  81. [89]

    Daniel Levy, Yair Carmon, John C Duchi, and Aaron Sidford. 2020. Large-scale methods for distributionally robust optimization.Advances in Neural Information Processing Systems33 (2020), 8847–8860

  82. [90]

    Bo Li, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, and Bowen Zhou. 2023. Trustworthy AI: From principles to practices.Comput. Surveys55, 9 (2023), 1–46

  83. [91]

    Bo Li, Yifei Shen, Yezhen Wang, Wenzhen Zhu, Colorado Reed, Dongsheng Li, Kurt Keutzer, and Han Zhao. 2022. Invariant Information Bottleneck for Domain Generalization.Proceedings of the AAAI Conference on Artificial Intelligence36, 7 (2022), 7399–7407

  84. [92]

    Shuang Li, Shiji Song, Gao Huang, Zhengming Ding, and Cheng Wu. 2018. Domain invariant and class discriminative feature learning for visual domain adaptation.IEEE transactions on image processing27, 9 (2018), 4260–4273

  85. [93]

    Xuhong Li, Haoyi Xiong, Xingjian Li, Xuanyu Wu, Xiao Zhang, Ji Liu, Jiang Bian, and Dejing Dou. 2022. Interpretable deep learning: Interpretation, interpretability, trustworthiness, and beyond.Knowledge and Information Systems64, 12 (2022), 3197–3234

  86. [94]

    Ya Li, Mingming Gong, Xinmei Tian, Tongliang Liu, and Dacheng Tao. 2018. Domain generalization via conditional invariant representations. In Proceedings of the AAAI conference on artificial intelligence, Vol. 32

  87. [95]

    Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma. 2021. Anti-backdoor learning: Training clean models on poisoned data. Advances in Neural Information Processing Systems34 (2021), 14900–14912

  88. [96]

    Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma. 2021. Neural attention distillation: Erasing backdoor triggers from deep neural networks.arXiv preprint arXiv:2101.05930(2021)

  89. [97]

    Ya Li, Xinmei Tian, Mingming Gong, Yajing Liu, Tongliang Liu, Kun Zhang, and Dacheng Tao. 2018. Deep domain generalization via conditional invariant adversarial networks. InProceedings of the European conference on computer vision (ECCV). 624–639

  90. [98]

    Yiming Li, Tongqing Zhai, Baoyuan Wu, Yong Jiang, Zhifeng Li, and Shutao Xia. 2020. Rethinking the trigger of backdoor attack.arXiv preprint arXiv:2004.04692(2020)

  91. [99]

    Yudong Li, Shigeng Zhang, Weiping Wang, and Hong Song. 2023. Backdoor attacks to deep learning models and countermeasures: A survey.IEEE Open Journal of the Computer Society4 (2023), 134–146

  92. [100]

    Zachary Lipton, Yu-Xiang Wang, and Alexander Smola. 2018. Detecting and correcting for label shift with black box predictors. InInternational conference on machine learning. PMLR, 3122–3130

  93. [101]

    Evan Z Liu, Behzad Haghgoo, Annie S Chen, Aditi Raghunathan, Pang Wei Koh, Shiori Sagawa, Percy Liang, and Chelsea Finn. 2021. Just train twice: Improving group robustness without training group information. InInternational Conference on Machine Learning. PMLR, 6781–6792

  94. [102]

    Haochen Liu, Yiqi Wang, Wenqi Fan, Xiaorui Liu, Yaxin Li, Shaili Jain, Yunhao Liu, Anil Jain, and Jiliang Tang. 2022. Trustworthy ai: A computational perspective.ACM Transactions on Intelligent Systems and Technology14, 1 (2022), 1–59

  95. [103]

    Jiashuo Liu, Zheyuan Hu, Peng Cui, Bo Li, and Zheyan Shen. 2021. Heterogeneous risk minimization. InInternational Conference on Machine Learning. PMLR, 6804–6814

  96. [104]

    Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg. 2018. Fine-pruning: Defending against backdooring attacks on deep neural networks. In International symposium on research in attacks, intrusions, and defenses. Springer, 273–294

  97. [105]

    Litian Liu and Yao Qin. 2024. Fast Decision Boundary based Out-of-Distribution Detector.ICML(2024)

  98. [106]

    Litian Liu and Yao Qin. 2025. Detecting Out-of-Distribution through the Lens of Neural Collapse.IEEE Conference on Computer Vision and Pattern Recognition (CVPR)(2025)

  99. [107]

    Lydia T Liu, Max Simchowitz, and Moritz Hardt. 2019. The implicit fairness criterion of unconstrained learning. InInternational Conference on Machine Learning. PMLR, 4051–4060

  100. [108]

    Yingqi Liu, Wen-Chuan Lee, Guanhong Tao, Shiqing Ma, Yousra Aafer, and Xiangyu Zhang. 2019. Abs: Scanning neural networks for back-doors by artificial brain stimulation. InProceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security. 1265–1282

  101. [109]

    Yiran Liu, Xiaoang Xu, Zhiyi Hou, and Yang Yu. 2024. Causality Based Front-door Defense Against Backdoor Attack on Language Models. In Forty-first International Conference on Machine Learning

  102. [110]

    Mingsheng Long, Yue Cao, Jianmin Wang, and Michael Jordan. 2015. Learning transferable features with deep adaptation networks. InInternational conference on machine learning. PMLR, 97–105

  103. [111]

    Mingsheng Long, Jianmin Wang, Guiguang Ding, Jiaguang Sun, and Philip S Yu. 2014. Transfer joint matching for unsupervised domain adaptation. InProceedings of the IEEE conference on computer vision and pattern recognition. 1410–1417

  104. [112]

    Clare Lyle, Amy Zhang, Minqi Jiang, Joelle Pineau, and Yarin Gal. 2021. Resolving causal confusion in reinforcement learning via robust exploration. InSelf-Supervision for Reinforcement Learning Workshop-ICLR, Vol. 2021

  105. [113]

    Wenao Ma, Cheng Chen, Shuang Zheng, Jing Qin, Huimao Zhang, and Qi Dou. 2022. Test-time adaptation with calibration of medical image classification nets for label distribution shift. InInternational Conference on Medical Image Computing and Computer-Assisted Intervention. Spri...

  106. [114]

    Divyat Mahajan, Shruti Tople, and Amit Sharma. 2021. Domain generalization using causal matching. InInternational conference on machine learning. PMLR, 7313–7324. Manuscript submitted to ACM 32 Liu et al

  107. [115]

    Adyasha Maharana, Prateek Yadav, and Mohit Bansal. 2024. $\mathbb{D}^2$ Pruning: Message Passing for Balancing Diversity & Difficulty in Data Pruning. InThe Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=thbtoAkCe9

  108. [116]

    Maggie Makar, Ben Packer, Dan Moldovan, Davis Blalock, Yoni Halpern, and Alexander D’Amour. 2022. Causally motivated shortcut removal using auxiliary labels. InInternational Conference on Artificial Intelligence and Statistics. PMLR, 739–766

  109. [117]

    Schulze Buschoff, Robert Geirhos, and Felix A

    Kristof Meding, Luca M. Schulze Buschoff, Robert Geirhos, and Felix A. Wichmann. 2022. Trivial or Impossible — dichotomous data difficulty masks model differences (on ImageNet and beyond). InInternational Conference on Learning Representations. https://openreview.net/forum?id=...

  110. [118]

    Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A survey on bias and fairness in machine learning. ACM computing surveys (CSUR)54, 6 (2021), 1–35

  111. [119]

    Milton Llera Montero, Casimir JH Ludwig, Rui Ponte Costa, Gaurav Malhotra, and Jeffrey Bowers. 2020. The role of disentanglement in generalisation. InInternational Conference on Learning Representations

  112. [120]

    Saeid Motiian, Marco Piccirilli, Donald A Adjeroh, and Gianfranco Doretto. 2017. Unified deep supervised domain adaptation and generalization. InProceedings of the IEEE international conference on computer vision. 5715–5725

  113. [121]

    Krikamol Muandet, David Balduzzi, and Bernhard Schölkopf. 2013. Domain generalization via invariant feature representation. InInternational conference on machine learning. PMLR, 10–18

  114. [122]

    Jens Müller, Robert Schmier, Lynton Ardizzone, Carsten Rother, and Ullrich Köthe. 2021. Learning robust models using the principle of independent causal mechanisms. InDAGM German Conference on Pattern Recognition. Springer, 79–110

  115. [123]

    Junhyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee, and Jinwoo Shin. 2020. Learning from Failure: Training Debiased Classifier from Biased Classifier. InAdvances in Neural Information Processing Systems

  116. [124]

    A Tuan Nguyen, Toan Tran, Yarin Gal, and Atilim Gunes Baydin. 2021. Domain invariant representation learning with domain density transforma- tions.Advances in Neural Information Processing Systems34 (2021), 5264–5275

  117. [125]

    Jianmo Ni, Jiacheng Li, and Julian McAuley. 2019. Justifying Recommendations using Distantly-Labeled Reviews and Fine-Grained Aspects. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natura...

  118. [126]

    Changdae Oh, Heeji Won, Junhyuk So, Taero Kim, Yewon Kim, Hosik Choi, and Kyungwoo Song. 2022. Learning fair representation via distributional contrastive disentanglement. InProceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 1295–1305

  119. [127]

    Matteo Pagliardini, Martin Jaggi, François Fleuret, and Sai Praneeth Karimireddy. 2022. Agree to disagree: Diversity through disagreement for better transferability.arXiv preprint arXiv:2202.04414(2022)

  120. [128]

    Sinno Jialin Pan, James T Kwok, Qiang Yang, et al. 2008. Transfer learning via dimensionality reduction.. InAAAI, Vol. 8. 677–682

  121. [129]

    Sinno Jialin Pan, Ivor W Tsang, James T Kwok, and Qiang Yang. 2010. Domain adaptation via transfer component analysis.IEEE transactions on neural networks22, 2 (2010), 199–210

  122. [130]

    Peter S Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks. 2024. AI deception: A survey of examples, risks, and potential solutions.Patterns5, 5 (2024)

  123. [131]

    Cato Pauling, Michael Gimson, Muhammed Qaid, Ahmad Kida, and Basel Halak. 2022. A tutorial on adversarial learning attacks and countermeasures. arXiv preprint arXiv:2202.10377(2022)

  124. [132]

    Xingchao Peng, Zijun Huang, Ximeng Sun, and Kate Saenko. 2019. Domain agnostic learning with disentangled representations. InInternational conference on machine learning. PMLR, 5102–5112

  125. [133]

    Xingchao Peng, Yichen Li, and Kate Saenko. 2020. Domain2vec: Domain embedding for unsupervised domain adaptation. InComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16. Springer, 756–774

  126. [134]

    Juan Perdomo, Tijana Zrnic, Celestine Mendler-Dünner, and Moritz Hardt. 2020. Performative prediction. InInternational Conference on Machine Learning. PMLR, 7599–7609

  127. [135]

    Jonas Peters, Peter Bühlmann, and Nicolai Meinshausen. 2016. Causal inference by using invariant prediction: identification and confidence intervals.Journal of the Royal Statistical Society Series B: Statistical Methodology78, 5 (2016), 947–1012

  128. [136]

    Hoang Phan, Andrew Gordon Wilson, and Qi Lei. 2024. Controllable Prompt Tuning For Balancing Group Distributional Robustness.Proceedings of Machine Learning Research235 (2024), 40667–40687

  129. [137]

    Aahlad Puli, Nitish Joshi, Yoav Wald, He He, and Rajesh Ranganath. 2022. Nuisances via negativa: Adjusting for spurious correlations via data augmentation.arXiv preprint arXiv:2210.01302(2022)

  130. [138]

    Aahlad Puli, Lily H Zhang, Eric K Oermann, and Rajesh Ranganath. 2021. Out-of-distribution generalization in the presence of nuisance-induced spurious correlations.arXiv preprint arXiv:2107.00520(2021)

  131. [139]

    Yao Qin, Xuezhi Wang, Balaji Lakshminarayanan, Ed H Chi, and Alex Beutel. 2023. What are effective labels for augmented data? improving calibration and robustness with autolabel. In2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML). IEEE, 365–376

  132. [140]

    Yao Qin, Chiyuan Zhang, Ting Chen, Balaji Lakshminarayanan, Alex Beutel, and Xuezhi Wang. 2022. Understanding and improving robustness of vision transformers through patch-based negative augmentation.Advances in Neural Information Processing Systems35 (2022), 16276–16289

  133. [141]

    Shikai Qiu, Andres Potapczynski, Pavel Izmailov, and Andrew Gordon Wilson. 2023. Simple and fast group robustness by automatic feature reweighting. InInternational Conference on Machine Learning. PMLR, 28448–28467. Manuscript submitted to ACM Bridging Distribution Shift and AI...

  134. [142]

    Tim Räz. 2021. Group fairness: Independence revisited. InProceedings of the 2021 ACM conference on fairness, accountability, and transparency. 129–137

  135. [143]

    Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. 2019. Do imagenet classifiers generalize to imagenet?. InInternational conference on machine learning. PMLR, 5389–5400

  136. [144]

    Mateo Rojas-Carulla, Bernhard Schölkopf, Richard Turner, and Jonas Peters. 2018. Invariant models for causal transfer learning.Journal of Machine Learning Research19, 36 (2018), 1–34

  137. [145]

    Elan Rosenfeld, Pradeep Ravikumar, and Andrej Risteski. 2022. Domain-adjusted regression or: Erm may already learn features sufficient for out-of-distribution generalization.arXiv preprint arXiv:2202.06856(2022)

  138. [146]

    Sorawit Saengkyongam, Nikolaj Thams, Jonas Peters, and Niklas Pfister. 2023. Invariant policy learning: A causal perspective.IEEE transactions on pattern analysis and machine intelligence45, 7 (2023), 8606–8620

  139. [147]

    Hashimoto, and Percy Liang

    Shiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, and Percy Liang. 2020. Distributionally Robust Neural Networks. InInternational Conference on Learning Representations. https://openreview.net/forum?id=ryxGuJrFvS

  140. [148]

    Aniruddha Saha, Akshayvarun Subramanya, and Hamed Pirsiavash. 2020. Hidden trigger backdoor attacks. InProceedings of the AAAI conference on artificial intelligence, Vol. 34. 11957–11965

  141. [149]

    Berkman Sahiner, Weijie Chen, Ravi K Samala, and Nicholas Petrick. 2023. Data drift in medical machine learning: implications and potential remedies.The British Journal of Radiology96, 1150 (2023), 20220878

  142. [150]

    Kuniaki Saito, Donghyun Kim, Stan Sclaroff, and Kate Saenko. 2020. Universal domain adaptation through self supervision.Advances in neural information processing systems33 (2020), 16282–16292

  143. [151]

    Shibani Santurkar, Dimitris Tsipras, and Aleksander Madry. 2020. Breeds: Benchmarks for subpopulation shift.arXiv preprint arXiv:2008.04859 (2020)

  144. [152]

    Iqbal H Sarker, Md Hasan Furhad, and Raza Nowrozy. 2021. Ai-driven cybersecurity: an overview, security intelligence modeling and research directions.SN Computer Science2, 3 (2021), 173

  145. [153]

    Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017. Grad-cam: Visual explanations from deep networks via gradient-based localization. InProceedings of the IEEE international conference on computer vision. 618–626

  146. [154]

    Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton. 2022. Goal misgeneralization: Why correct specifications aren’t enough for correct goals.arXiv preprint arXiv:2210.01790(2022)

  147. [155]

    K Shailaja, Banoth Seetharamulu, and MA Jabbar. 2018. Machine learning in healthcare: A review. In2018 Second international conference on electronics, communication and aerospace technology (ICECA). IEEE, 910–914

  148. [156]

    Minglai Shao, Dong Li, Chen Zhao, Xintao Wu, Yujie Lin, and Qin Tian. 2024. Supervised algorithmic fairness in distribution shifts: A survey. arXiv preprint arXiv:2402.01327(2024)

  149. [157]

    Guangyu Shen, Yingqi Liu, Guanhong Tao, Shengwei An, Qiuling Xu, Siyuan Cheng, Shiqing Ma, and Xiangyu Zhang. 2021. Backdoor scanning for deep neural networks through k-arm optimization. InInternational Conference on Machine Learning. PMLR, 9525–9536

  150. [158]

    Hua Shen, Tiffany Knearem, Reshmi Ghosh, Kenan Alkiek, Kundan Krishna, Yachuan Liu, Ziqiao Ma, Savvas Petridis, Yi-Hao Peng, Li Qiwei, et al

  151. [159]

    Jian Shen, Yanru Qu, Weinan Zhang, and Yong Yu. 2018. Wasserstein distance guided representation learning for domain adaptation. InProceedings of the AAAI conference on artificial intelligence, Vol. 32

  152. [160]

    Tianhao Shen, Renren Jin, Yufei Huang, Chuang Liu, Weilong Dong, Zishan Guo, Xinwei Wu, Yan Liu, and Deyi Xiong. 2023. Large language model alignment: A survey.arXiv preprint arXiv:2309.15025(2023)

  153. [161]

    Zheyan Shen, Peng Cui, Tong Zhang, and Kun Kunag. 2020. Stable learning via sample reweighting. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 5692–5699

  154. [162]

    Yucheng Shi, Mengnan Du, Xuansheng Wu, Zihan Guan, Jin Sun, and Ninghao Liu. 2023. Black-box backdoor defense via zero-shot image purification.Advances in Neural Information Processing Systems36 (2023), 57336–57366

  155. [163]

    Zhouxing Shi, Nicholas Carlini, Ananth Balashankar, Ludwig Schmidt, Cho-Jui Hsieh, Alex Beutel, and Yao Qin. 2023. Effective robustness against natural distribution shifts for models with different training data.Advances in Neural Information Processing Systems36 (2023), 73543–73558

  156. [164]

    Hidetoshi Shimodaira. 2000. Improving predictive inference under covariate shift by weighting the log-likelihood function.Journal of statistical planning and inference90, 2 (2000), 227–244

  157. [165]

    Changjian Shui, Zijian Li, Jiaqi Li, Christian Gagné, Charles X Ling, and Boyu Wang. 2021. Aggregating from multiple target-shifted sources. In International Conference on Machine Learning. PMLR, 9638–9648

  158. [166]

    Herbert A Simon. 1954. Spurious correlation: A causal interpretation.Journal of the American statistical Association49, 267 (1954), 467–479

  159. [167]

    Anoopkumar Sonar, Vincent Pacelli, and Anirudha Majumdar. 2021. Invariant policy optimization: Towards stronger generalization in reinforcement learning. InLearning for Dynamics and Control. PMLR, 21–33

  160. [168]

    Hossein Souri, Liam Fowl, Rama Chellappa, Micah Goldblum, and Tom Goldstein. 2022. Sleeper agent: Scalable hidden trigger backdoors for neural networks trained from scratch.Advances in Neural Information Processing Systems35 (2022), 19165–19178

  161. [169]

    Benjamin Stoler, Ingrid Navarro, Meghdeep Jana, Soonmin Hwang, Jonathan Francis, and Jean Oh. 2024. SafeShift: Safety-informed distribution shifts for robust trajectory prediction in autonomous driving. In2024 IEEE Intelligent Vehicles Symposium (IV). IEEE, 1179–1186. Manuscri...

  162. [170]

    Xinwei Sun, Botong Wu, Xiangyu Zheng, Chang Liu, Wei Chen, Tao Qin, and Tie-Yan Liu. 2021. Recovering latent causal factor for generalization to distributional shifts.Advances in Neural Information Processing Systems34 (2021), 16846–16859

  163. [171]

    Harry Surden. 2021. Machine learning and law: An overview.Research handbook on big data law(2021), 171–184

  164. [172]

    Remi Tachet des Combes, Han Zhao, Yu-Xiang Wang, and Geoffrey J Gordon. 2020. Domain adaptation with conditional distribution matching and generalized label shift.Advances in Neural Information Processing Systems33 (2020), 19276–19289

  165. [173]

    Haoru Tan, Sitong Wu, Fei Du, Yukang Chen, Zhibin Wang, Fan Wang, and Xiaojuan Qi. 2024. Data pruning via moving-one-sample-out.Advances in Neural Information Processing Systems36 (2024)

  166. [174]

    Damien Teney, Ehsan Abbasnejad, Simon Lucey, and Anton Van den Hengel. 2022. Evading the simplicity bias: Training a diverse set of models discovers solutions with superior ood generalization. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition....

  167. [175]

    Damien Teney, Ehsan Abbasnejad, and Anton van den Hengel. 2021. Unshuffling data for improved generalization in visual question answering. In Proceedings of the IEEE/CVF international conference on computer vision. 1417–1427

  168. [176]

    Brandon Tran, Jerry Li, and Aleksander Madry. 2018. Spectral signatures in backdoor attacks.Advances in neural information processing systems31 (2018)

  169. [177]

    Loc Truong, Chace Jones, Brian Hutchinson, Andrew August, Brenda Praggastis, Robert Jasper, Nicole Nichols, and Aaron Tuor. 2020. Systematic evaluation of backdoor data poisoning attacks on image classifiers. InProceedings of the IEEE/CVF conference on computer vision and patt...

  170. [178]

    Alexander Turner, Dimitris Tsipras, and Aleksander Madry. 2019. Label-consistent backdoor attacks.arXiv preprint arXiv:1912.02771(2019)

  171. [179]

    Victor Veitch, Alexander D’Amour, Steve Yadlowsky, and Jacob Eisenstein. 2021. Counterfactual invariance to spurious correlations in text classification.Advances in neural information processing systems34 (2021), 16196–16208

  172. [180]

    Sahil Verma and Julia Rubin. 2018. Fairness definitions explained. InProceedings of the international workshop on software fairness. 1–7

  173. [181]

    Yoav Wald, Amir Feder, Daniel Greenfeld, and Uri Shalit. 2021. On calibration and out-of-domain generalization.Advances in neural information processing systems34 (2021), 2215–2227

  174. [182]

    Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y Zhao. 2019. Neural cleanse: Identifying and mitigating backdoor attacks in neural networks. In2019 IEEE symposium on security and privacy (SP). IEEE, 707–723

  175. [183]

    Jindong Wang, Xixu Hu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Haojun Huang, Wei Ye, Xiubo Geng, et al. 2023. On the robustness of chatgpt: An adversarial and out-of-distribution perspective.arXiv preprint arXiv:2302.12095(2023)

  176. [184]

    Lun Wang, Zaynah Javed, Xian Wu, Wenbo Guo, Xinyu Xing, and Dawn Song. 2021. Backdoorl: Backdoor attack against competitive reinforcement learning.arXiv preprint arXiv:2105.00579(2021)

  177. [185]

    Shanshan Wang, Yiyang Chen, Zhenwei He, Xun Yang, Mengzhu Wang, Quanzeng You, and Xingyi Zhang. 2023. Disentangled representation learning with causality for unsupervised domain adaptation. InProceedings of the 31st ACM International Conference on Multimedia. 2918–2926

  178. [186]

    Tianlu Wang, Rohit Sridhar, Diyi Yang, and Xuezhi Wang. 2021. Identifying and mitigating spurious correlations for improving robustness in nlp models.arXiv preprint arXiv:2110.07736(2021)

  179. [187]

    Xin Wang, Hong Chen, Zihao Wu, Wenwu Zhu, et al. 2024. Disentangled representation learning.IEEE Transactions on Pattern Analysis and Machine Intelligence(2024)

  180. [188]

    Yixu Wang, Yan Teng, Kexin Huang, Chengqi Lyu, Songyang Zhang, Wenwei Zhang, Xingjun Ma, Yu-Gang Jiang, Yu Qiao, and Yingchun Wang

  181. [189]

    Zhenting Wang, Hailun Ding, Juan Zhai, and Shiqing Ma. 2022. Training with more confidence: Mitigating injected and natural backdoors during training.Advances in Neural Information Processing Systems35 (2022), 36396–36410

  182. [190]

    Zhonghao Wang, Yunchao Wei, Rogerio Feris, Jinjun Xiong, Wen-Mei Hwu, Thomas S Huang, and Honghui Shi. 2020. Alleviating semantic-level shift: A semi-supervised domain adaptation method for semantic segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and ...

  183. [191]

    Tao Wen, Zihan Wang, Quan Zhang, and Qi Lei. 2025. Elastic Representation: Mitigating Spurious Correlations for Group Robustness.arXiv preprint arXiv:2502.09850(2025)

  184. [192]

    Fake Alignment: Are LLMs Really Aligned Well?arXiv preprint arXiv:2311.05915(2023)

  185. [193]

    Keru Wu, Yuansi Chen, Wooseok Ha, and Bin Yu. 2023. Prominent Roles of Conditionally Invariant Components in Domain Adaptation: Theory and Algorithms.arXiv preprint arXiv:2309.10301(2023)

  186. [194]

    Minghao Xu, Jian Zhang, Bingbing Ni, Teng Li, Chengjie Wang, Qi Tian, and Wenjun Zhang. 2020. Adversarial domain adaptation with domain mixup. InProceedings of the AAAI conference on artificial intelligence, Vol. 34. 6502–6509

  187. [195]

    Shuo Yang, Zeke Xie, Hanyu Peng, Min Xu, Mingming Sun, and Ping Li. 2023. Dataset Pruning: Reducing Training Data by Examining Generalization Influence. InThe Eleventh International Conference on Learning Representations. https://openreview.net/forum?id=4wZiAXD29TQ

  188. [196]

    Robert Williamson and Aditya Menon. 2019. Fairness risk measures. InInternational conference on machine learning. PMLR, 6786–6797

  189. [197]

    Wenkai Yang, Yankai Lin, Peng Li, Jie Zhou, and Xu Sun. 2021. Rethinking stealthiness of backdoor attack against nlp models. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Languag...

  190. [198]

    Yu Yang, Eric Gan, Gintare Karolina Dziugaite, and Baharan Mirzasoleiman. 2024. Identifying spurious biases early in training through the lens of simplicity bias. InInternational Conference on Artificial Intelligence and Statistics. PMLR, 2953–2961

  191. [199]

    Huaxiu Yao, Yu Wang, Sai Li, Linjun Zhang, Weixin Liang, James Zou, and Chelsea Finn. 2022. Improving out-of-distribution robustness via selective augmentation. InInternational Conference on Machine Learning. PMLR, 25407–25437

  192. [200]

    Wanqian Yang, Polina Kirichenko, Micah Goldblum, and Andrew G Wilson. 2022. Chroma-vae: Mitigating shortcut learning with generative classifiers.Advances in Neural Information Processing Systems35 (2022), 20351–20365

  193. [201]

    Kaichao You, Mingsheng Long, Zhangjie Cao, Jianmin Wang, and Michael I Jordan. 2019. Universal domain adaptation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2720–2729

  194. [202]

    Werner Zellinger, Thomas Grubinger, Edwin Lughofer, Thomas Natschläger, and Susanne Saminger-Platz. 2017. Central moment discrepancy (cmd) for domain-invariant representation learning.arXiv preprint arXiv:1702.08811(2017)

  195. [203]

    Werner Zellinger, Bernhard A Moser, Thomas Grubinger, Edwin Lughofer, Thomas Natschläger, and Susanne Saminger-Platz. 2019. Robust unsupervised domain adaptation for neural networks via moment alignment.Information Sciences483 (2019), 174–191

  196. [204]

    Wenqian Ye, Guangtao Zheng, Xu Cao, Yunsheng Ma, Xia Hu, and Aidong Zhang. 2024. Spurious correlations in machine learning: A survey. arXiv preprint arXiv:2402.12715(2024)

  197. [205]

    Runtian Zhai, Chen Dan, Zico Kolter, and Pradeep Ravikumar. 2021. Doro: Distributional and outlier robust optimization. InInternational Conference on Machine Learning. PMLR, 12345–12355

  198. [206]

    Shengfang Zhai, Qingni Shen, Xiaoyi Chen, Weilong Wang, Cong Li, Yuejian Fang, and Zhonghai Wu. 2023. Ncl: Textual backdoor defense using noise-augmented contrastive learning. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)....

  199. [207]

    Hanlin Zhang, Yi-Fan Zhang, Weiyang Liu, Adrian Weller, Bernhard Schölkopf, and Eric P Xing. 2022. Towards principled disentanglement for domain generalization. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 8024–8034

  200. [208]

    Yi Zeng, Minzhou Pan, Hoang Anh Just, Lingjuan Lyu, Meikang Qiu, and Ruoxi Jia. 2023. Narcissus: A practical clean-label backdoor attack with limited information. InProceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security. 771–785

  201. [209]

    Michael Zhang and Christopher Ré. 2022. Contrastive adapters for foundation model group robustness.Advances in Neural Information Processing Systems35 (2022), 21682–21697

  202. [210]

    Michael Zhang, Nimit S Sohoni, Hongyang R Zhang, Chelsea Finn, and Christopher Ré. 2022. Correct-n-contrast: A contrastive approach for improving robustness to spurious correlations.arXiv preprint arXiv:2203.01517(2022)

  203. [211]

    Xingxuan Zhang, Peng Cui, Renzhe Xu, Linjun Zhou, Yue He, and Zheyan Shen. 2021. Deep stable learning for out-of-distribution generalization. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5372–5382

  204. [212]

    Jie Zhang, Chen Dongdong, Qidong Huang, Jing Liao, Weiming Zhang, Huamin Feng, Gang Hua, and Nenghai Yu. 2022. Poison ink: Robust and invisible backdoor attack.IEEE Transactions on Image Processing31 (2022), 5691–5705

  205. [213]

    Yi-Fan Zhang, Zhang Zhang, Da Li, Zhen Jia, Liang Wang, and Tieniu Tan. 2022. Learning domain invariant representations for generalizable person re-identification.IEEE Transactions on Image Processing32 (2022), 509–523

  206. [214]

    Zaixi Zhang, Qi Liu, Zhicai Wang, Zepu Lu, and Qingyong Hu. 2023. Backdoor defense via deconfounded representation learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12228–12238

  207. [215]

    Shanshan Zhao, Mingming Gong, Tongliang Liu, Huan Fu, and Dacheng Tao. 2020. Domain generalization via entropy regularization.Advances in neural information processing systems33 (2020), 16096–16107

  208. [216]

    Xingxuan Zhang, Zekai Xu, Renzhe Xu, Jiashuo Liu, Peng Cui, Weitao Wan, Chong Sun, and Chen Li. 2022. Towards domain generalization in object detection.arXiv preprint arXiv:2203.14387(2022)

  209. [217]

    Jiayun Zheng and Maggie Makar. 2022. Causally motivated multi-shortcut identification and removal.Advances in Neural Information Processing Systems35 (2022), 12800–12812

  210. [218]

    Kecheng Zheng, Jiawei Liu, Wei Wu, Liang Li, and Zheng-jun Zha. 2021. Calibrated feature decomposition for generalizable person re-identification. arXiv preprint arXiv:2111.13945(2021)

  211. [219]

    Rui Zheng, Wei Shen, Yuan Hua, Wenbin Lai, Shihan Dou, Yuhao Zhou, Zhiheng Xi, Xiao Wang, Haoran Huang, Tao Gui, et al. 2023. Improving generalization of alignment with human preferences through group invariant learning.arXiv preprint arXiv:2310.11971(2023)

  212. [220]

    Shihao Zhao, Xingjun Ma, Xiang Zheng, James Bailey, Jingjing Chen, and Yu-Gang Jiang. 2020. Clean-label backdoor attacks on video recognition models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 14443–14452

  213. [221]

    Jianlong Zhou, Fang Chen, and Andreas Holzinger. 2020. Towards explainability for AI fairness. InInternational workshop on extending explainable AI beyond deep models and classifiers. Springer, 375–386

  214. [222]

    Mingli Zhu, Shaokui Wei, Li Shen, Yanbo Fan, and Baoyuan Wu. 2023. Enhancing fine-tuning based backdoor defense with sharpness-aware minimization. InProceedings of the IEEE/CVF International Conference on Computer Vision. 4466–4477. Manuscript submitted to ACM

  215. [224]

    Ziliang Samuel Zhong, Xiang Pan, and Qi Lei. 2025. Bridging Domains with Approximately Shared Features. InThe 28th International Conference on Artificial Intelligence and Statistics

  216. [2020]

    On the transfer of disentangled representations in realistic settings.arXiv preprint arXiv:2010.14407(2020)

  217. [2023]

    Leveraging sparse and shared feature activations for disentangled representation learning.Advances in Neural Information Processing Systems 36 (2023), 27682–27698

  218. [2024]

    Towards Bidirectional Human-AI Alignment: A Systematic Review for Clarifications, Framework, and Future Directions.arXiv preprint arXiv:2406.09264(2024)

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.