Pith. sign in

REVIEW 3 major objections 2 minor 54 references

A multimodal framework uses gradient reversal to unlearn demographic attributes and detect mild cognitive impairment with higher accuracy and smaller gaps across sex and language groups.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-26 21:59 UTC pith:WWWGLUE7

load-bearing objection The paper applies cross-modal fusion plus gradient reversal to shrink demographic gaps in multilingual MCI detection on TAUKADIAL and PREPARE, but the abstract leaves the key assumption untested. the 3 major comments →

arxiv 2606.18571 v1 pith:WWWGLUE7 submitted 2026-06-17 cs.LG cs.CLcs.SDeess.AS

Fair Cognitive Impairment Detection Through Unlearning

classification cs.LG cs.CLcs.SDeess.AS
keywords mild cognitive impairmentunlearninggradient reversalmultimodal fusionfairnessspeech analysisdemographic bias
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper presents a method that fuses speech, text, and image inputs while applying gradient reversal to stop the shared embedding from encoding patient sex and language. The aim is to stop models from using demographic correlations that produce unequal accuracy across patient subgroups. If the approach works, screening tools could maintain performance regardless of a patient's background. Experiments on the TAUKADIAL and PREPARE benchmarks show gains over prior multimodal and multilingual baselines together with reduced subgroup disparities. The same unlearning step also yields representations that transfer more effectively between datasets.

Core claim

The central claim is that cross-model fusion combined with gradient-reversal unlearning produces a shared embedding supporting superior MCI classification accuracy on TAUKADIAL and PREPARE while substantially shrinking performance differences between male and female patients and across languages, and that the resulting representations transfer more robustly across datasets.

What carries the argument

Gradient reversal applied during training to discourage the shared multimodal embedding from encoding task-irrelevant demographic attributes such as sex and language.

Load-bearing premise

Demographic attributes like sex and language contain no information genuinely useful for MCI detection and can be removed without lowering overall classification performance.

What would settle it

A controlled test in which the same models trained with demographic unlearning show lower MCI classification accuracy than the unregularized baseline on a dataset where sex and language are uncorrelated with impairment labels.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Higher MCI classification accuracy than existing multilingual and multimodal baselines on the evaluated benchmarks.
  • Substantially smaller performance gaps across patient subgroups defined by sex and language.
  • Improved transfer performance when representations learned on one dataset are applied to another.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same unlearning step could be tested on other speech-based medical classification tasks that currently exhibit demographic performance gaps.
  • Deployed systems using this method might require less per-group recalibration when moved to new populations.
  • One could measure whether the unlearned embeddings preserve all clinically relevant acoustic or linguistic cues by checking accuracy on auxiliary tasks that do depend on demographics.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The paper proposes a multimodal framework for detecting Mild Cognitive Impairment (MCI) from spontaneous speech that fuses speech, text, and image modalities via cross-model fusion and applies gradient reversal unlearning to discourage the shared embedding from encoding demographic attributes (sex and language). Evaluated on the TAUKADIAL and PREPARE benchmarks, it claims to outperform state-of-the-art multilingual and multimodal baselines in MCI classification accuracy while substantially reducing performance gaps across subgroups, and to improve robustness under cross-dataset transfer.

Significance. If the empirical claims hold under rigorous validation, the work could advance fairness-aware multimodal learning for medical screening applications by demonstrating that gradient-reversal unlearning can simultaneously boost task performance and reduce demographic disparities. The emphasis on multilingual spontaneous-speech data and cross-dataset transfer addresses practically relevant challenges in scalable cognitive assessment.

major comments (3)
  1. [§3.2] §3.2 (Unlearning via Gradient Reversal): The central claim that demographic unlearning removes only task-irrelevant information rests on the untested assumption that sex- and language-linked speech features carry no genuine MCI signal. No post-unlearning probe classifier accuracies on demographic attributes are reported to confirm that reversal succeeded without residual leakage or excessive suppression of useful features.
  2. [§4.2] §4.2 and §4.3 (Experimental Results): The reported outperformance and gap reduction lack accompanying statistical significance tests, confidence intervals, or error analysis across the TAUKADIAL/PREPARE splits. Without these, it is impossible to determine whether gains exceed baseline variance or arise from post-hoc hyperparameter choices.
  3. [§4.4] §4.4 (Ablation Studies): No ablation is presented on the gradient reversal strength hyperparameter or on the contribution of each modality to the fairness-performance trade-off. This leaves open whether the claimed improvements are robust or sensitive to specific implementation choices.
minor comments (2)
  1. [Abstract] The abstract states quantitative superiority without any numbers or statistical details; this should be corrected to include at least the key accuracy and fairness metrics.
  2. [§3] Notation for the shared embedding and the reversal loss term is introduced without an explicit equation reference in the method section, making it harder to follow the implementation.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for the constructive feedback. The comments identify valuable opportunities to strengthen the empirical support for our unlearning claims and the robustness of the reported results. We address each major comment below.

read point-by-point responses
  1. Referee: §3.2 (Unlearning via Gradient Reversal): The central claim that demographic unlearning removes only task-irrelevant information rests on the untested assumption that sex- and language-linked speech features carry no genuine MCI signal. No post-unlearning probe classifier accuracies on demographic attributes are reported to confirm that reversal succeeded without residual leakage or excessive suppression of useful features.

    Authors: We agree that direct verification of unlearning success via probes would provide stronger evidence. While the observed gains in both MCI accuracy and subgroup fairness already indicate that task-relevant signal is retained, we will add post-unlearning probe experiments in the revision. These will report sex and language classification accuracies on the shared embeddings before and after gradient reversal to quantify residual demographic encoding. revision: yes

  2. Referee: §4.2 and §4.3 (Experimental Results): The reported outperformance and gap reduction lack accompanying statistical significance tests, confidence intervals, or error analysis across the TAUKADIAL/PREPARE splits. Without these, it is impossible to determine whether gains exceed baseline variance or arise from post-hoc hyperparameter choices.

    Authors: We acknowledge this limitation in the current presentation. In the revised manuscript we will report statistical significance tests (paired t-tests or McNemar’s test across multiple random seeds), 95% confidence intervals, and standard deviations computed over repeated runs on the TAUKADIAL and PREPARE splits to establish that the improvements exceed baseline variance. revision: yes

  3. Referee: §4.4 (Ablation Studies): No ablation is presented on the gradient reversal strength hyperparameter or on the contribution of each modality to the fairness-performance trade-off. This leaves open whether the claimed improvements are robust or sensitive to specific implementation choices.

    Authors: We will incorporate the requested ablations. The revision will include (i) a sweep over the gradient-reversal strength hyperparameter showing its effect on the accuracy–fairness trade-off and (ii) modality-ablation results that isolate the contribution of speech, text, and image inputs to both MCI classification and demographic gap reduction. revision: yes

Circularity Check

0 steps flagged

No circularity: empirical method with standard unlearning, results from evaluation not by construction

full rationale

The paper proposes a multimodal MCI detection framework using cross-model fusion and gradient reversal for demographic unlearning. Claims of outperforming baselines and reducing subgroup gaps rest on experimental results on TAUKADIAL and PREPARE, not on any derivation, fitted parameter renamed as prediction, or self-citation chain. No equations appear in the provided text, and the unlearning step follows established adversarial training patterns without reducing the fairness or performance gains to tautological inputs. This is a standard empirical ML contribution with independent content from benchmarks.

Axiom & Free-Parameter Ledger

0 free parameters · 2 axioms · 0 invented entities

The central claim rests on standard supervised learning assumptions plus the domain assumption that demographic variables can be treated as removable nuisance factors without collateral damage to the MCI signal. No free parameters or invented entities are introduced in the abstract.

axioms (2)
  • domain assumption Demographic attributes are task-irrelevant for MCI detection and can be adversarially removed without harming predictive performance on the target task.
    Invoked by the choice of gradient reversal to discourage encoding of sex and language in the shared embedding.
  • domain assumption The TAUKADIAL and PREPARE benchmarks provide representative samples for evaluating both accuracy and subgroup fairness.
    The evaluation and transfer claims depend on these datasets being appropriate proxies for real-world use.

pith-pipeline@v0.9.1-grok · 5675 in / 1309 out tokens · 15499 ms · 2026-06-26T21:59:34.479733+00:00 · methodology

0 comments
read the original abstract

Mild Cognitive Impairment (MCI) is a medical condition characterized by a noticeable decline in memory, language, or thinking abilities. MCI detection from spontaneous speech is promising for scalable screening. However, learned models often exploit demographic cues correlated with labels, resulting in a large performance gap across subgroups. We present a multimodal framework that combines (i) cross-model fusion between modalities (speech, text, and image), and (ii) unlearning using gradient reversal that discourages the shared embedding from encoding task-irrelevant demographic attributes. Evaluated on the multilingual benchmarks TAUKADIAL and PREPARE, our method outperforms the state-of-the-art multilingual and multimodal baseline in MCI classification while substantially reducing the performance gap across patient subgroups (sex and language). We further analyze transfer across datasets, showing that demographic unlearning helps learn more robust representations for MCI detection.

Figures

Figures reproduced from arXiv: 2606.18571 by Hadi Amiri, Jiali Cheng, William Nguyen.

Figure 1
Figure 1. Figure 1: Architecture of FMD. 1) The cross-modal (CM) fusion module followed by a feed-forward network (FFN) supports richer and finer-grained modality interaction and fusion compared to standard concatenation. 2) The unlearning (UL) component removes task-irrelevant biases from the model, yielding fairer and more robust performance. This is achieved using an auxiliary demographic classifier fDemo that identifies s… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

54 extracted references · 8 canonical work pages · 1 internal anchor

  1. [1]

    Introduction Speech-based assessment is a promising approach for screening cognitive impairment because spontaneous speech reflects cog- nitive and linguistic changes in lexical choice and diversity, syn- tactic complexity, disfluencies, and prosody [1, 2]. However, real-world clinical speech datasets are typically small, heteroge- neous, and demographica...

  2. [2]

    Fair Cognitive Impairment Detection Through Unlearning

    Related Work MCI detection: Existing work has developed models to detect MCI and Alzheimer’s disease from speech signals using seman- tic features [18], linguistic features [19, 20], training data aug- mentation [21, 22], and prompt learning [23]. [24] provides a review of LLM applications in dementia care, analyzes survey results from individuals with de...

  3. [3]

    unlearning

    Method FMD consists of two tightly coupled components: 1) a multi- modal MCI classifier with cross-modal fusion to produce accu- rate MCI diagnosis, and 2) an unlearning module that removes demographic information from the learned representations to mitigate bias. 3.1. Multimodal MCI Detection via Cross-Modal Fusion Our architecture consists of unimodal e...

  4. [4]

    Datasets We use the following two datasets for our experiments

    Experiments 4.1. Datasets We use the following two datasets for our experiments. • TAUKADIAL [37]: a dataset of 387 samples with two la- bels: normal subjects (NC) and patients with Mild Cognitive Impairment (MCI). • PREPARE [17]: a dataset of 1644 samples with three labels: NC, MCI, and ADRD (Alzheimer’s Disease and Related De- mentias). The data statist...

  5. [5]

    - CM” and “-UL

    Results 5.1. Main Results FMD improves overall performance.Across both datasets, FMD consistently achieves higher overall F1 scores than all baselines. On TAUKADIAL, FMDLang (where language is used Table 1:Dataset statistics. Label TAUKADIAL (n= 387) PREPARE (n= 1,644) Sex Language Sex Language F M En Non-En F M En Non-En NC 102 63 63 102 539 371 759 151 ...

  6. [6]

    Conclusion We study demographic bias in Mild Cognitive Impairment (MCI) detection, where models rely on spurious demographic attributes and show substantial performance disparities across different patient subgroups. We propose FMD, a framework that combines cross-modal representation fusion and an un- learning designed to discourage demographic informati...

  7. [7]

    All technical content, analyses, and conclusions were produced by the authors and verified for accuracy

    Generative AI Use Disclosure We used LLMs to assist with editing and polishing our writing for clarity and concision. All technical content, analyses, and conclusions were produced by the authors and verified for accuracy

  8. [8]

    Artifi- cial intelligence, speech, and language processing approaches to monitoring alzheimer’s disease: a systematic review,

    S. De la Fuente Garcia, C. W. Ritchie, and S. Luz, “Artifi- cial intelligence, speech, and language processing approaches to monitoring alzheimer’s disease: a systematic review,”Journal of Alzheimer’s Disease, vol. 78, no. 4, pp. 1547–1574, 2020

  9. [9]

    Detecting Cognitive Decline Using Speech Only: The ADReSSo Challenge,

    S. Luz, F. Haider, S. de la Fuente, D. Fromm, and B. MacWhin- ney, “Detecting Cognitive Decline Using Speech Only: The ADReSSo Challenge,” inInterspeech 2021, 2021, pp. 3780–3784

  10. [10]

    CogniV oice: Multi- modal and Multilingual Fusion Networks for Mild Cognitive Im- pairment Assessment from Spontaneous Speech,

    J. Cheng, M. Elgaar, N. Vakil, and H. Amiri, “CogniV oice: Multi- modal and Multilingual Fusion Networks for Mild Cognitive Im- pairment Assessment from Spontaneous Speech,” inInterspeech 2024, 2024, pp. 4308–4312

  11. [11]

    Speechcare: dynamic multimodal modeling for cognitive screening in diverse linguistic and speech task contexts,

    H. Azadmaleki, Y . Haghbin, S. Rashidi, M. J. Momeni Nezhad, A. Zolnour, and M. Zolnoori, “Speechcare: dynamic multimodal modeling for cognitive screening in diverse linguistic and speech task contexts,”npj Digital Medicine, vol. 8, no. 1, p. 677, 2025

  12. [12]

    Bias and fairness in self-supervised acoustic representations for cognitive impairment detection,

    K. Gulzar, K. Riedhammer, E. N ¨oth, A. K. Maier, and P. A. P ´erez-Toro, “Bias and fairness in self-supervised acoustic representations for cognitive impairment detection,” 2026. [Online]. Available: https://arxiv.org/abs/2603.02937

  13. [13]

    Don‘t take the easy way out: Ensemble based methods for avoiding known dataset biases,

    C. Clark, M. Yatskar, and L. Zettlemoyer, “Don‘t take the easy way out: Ensemble based methods for avoiding known dataset biases,” inProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), K. Inui, J. Jiang, V . Ng, and X. Wan, Eds. Hon...

  14. [14]

    End- to-end bias mitigation by modelling biases in corpora,

    R. Karimi Mahabadi, Y . Belinkov, and J. Henderson, “End- to-end bias mitigation by modelling biases in corpora,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, Eds. Online: Association for Computational Linguistics, Jul. 2020, pp. 8706–8716. [Online]. Availa...

  15. [15]

    Guide the learner: Controlling product of experts debiasing method based on token attribution similarities,

    A. Modarressi, H. Amirkhani, and M. T. Pilehvar, “Guide the learner: Controlling product of experts debiasing method based on token attribution similarities,” inProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, A. Vlachos and I. Augenstein, Eds. Dubrovnik, Croatia: Association for Computational Li...

  16. [16]

    Towards debiasing NLU models from unknown biases,

    P. A. Utama, N. S. Moosavi, and I. Gurevych, “Towards debiasing NLU models from unknown biases,” inProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), B. Webber, T. Cohn, Y . He, and Y . Liu, Eds. Online: Association for Computational Linguistics, Nov. 2020, pp. 7597–7610. [Online]. Available: https://aclantholo...

  17. [17]

    Learning from others’ mistakes: Avoiding dataset biases without modeling them,

    V . Sanh, T. Wolf, Y . Belinkov, and A. M. Rush, “Learning from others’ mistakes: Avoiding dataset biases without modeling them,” inInternational Conference on Learning Representations,

  18. [18]

    Available: https://openreview.net/forum?id= Hf3qXoiNkR

    [Online]. Available: https://openreview.net/forum?id= Hf3qXoiNkR

  19. [19]

    Kernel-whitening: Overcome dataset bias with isotropic sentence embedding,

    S. Gao, S. Dou, Q. Zhang, and X. Huang, “Kernel-whitening: Overcome dataset bias with isotropic sentence embedding,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y . Goldberg, Z. Kozareva, and Y . Zhang, Eds. Abu Dhabi, United Arab Emirates: Association for Computational Linguistics, Dec. 2022, pp. 4112–4122. ...

  20. [20]

    Towards stable natural language understanding via information entropy guided debiasing,

    L. Du, X. Ding, Z. Sun, T. Liu, B. Qin, and J. Liu, “Towards stable natural language understanding via information entropy guided debiasing,” inProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki, Eds. Toronto, Canada: Association for Computational Lingu...

  21. [21]

    Robust natural language understanding with residual attention debiasing,

    F. Wang, J. Y . Huang, T. Yan, W. Zhou, and M. Chen, “Robust natural language understanding with residual attention debiasing,” inFindings of the Association for Computational Linguistics: ACL 2023, A. Rogers, J. Boyd-Graber, and N. Okazaki, Eds. Toronto, Canada: Association for Computational Linguistics, Jul. 2023, pp. 504–519. [Online]. Available: https...

  22. [22]

    Debiasing masks: A new framework for shortcut mitigation in NLU,

    J. M. Meissner, S. Sugawara, and A. Aizawa, “Debiasing masks: A new framework for shortcut mitigation in NLU,” inProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y . Goldberg, Z. Kozareva, and Y . Zhang, Eds. Abu Dhabi, United Arab Emirates: Association for Computational Linguistics, Dec. 2022, pp. 7607–7613. [Online...

  23. [23]

    FairFlow: Mitigating dataset biases through undecided learning for natural language understanding,

    J. Cheng and H. Amiri, “FairFlow: Mitigating dataset biases through undecided learning for natural language understanding,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, Eds. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 21 960–21 975. ...

  24. [24]

    Connected speech-based cognitive assessment in chinese and english,

    S. D. L. F. Garcia, F. Haider, D. Fromm, B. MacWhinney, A. Lanzi, Y .-N. Chang, C.-J. Chou, Y .-C. Liuet al., “Connected speech-based cognitive assessment in chinese and english,”arXiv preprint arXiv:2406.10272, 2024

  25. [25]

    Prepare: Pioneering research for early prediction of alzheimer’s and related dementias eureka challenge,

    “Prepare: Pioneering research for early prediction of alzheimer’s and related dementias eureka challenge,” https://www.drivendata. org/competitions/group/nih-nia-alzheimers-adrd-competition/, 2023

  26. [26]

    Linguistic features extracted by GPT-4 improve Alzheimer‘s disease detection based on spontaneous speech,

    J. Heitz, G. Schneider, and N. Langer, “Linguistic features extracted by GPT-4 improve Alzheimer‘s disease detection based on spontaneous speech,” inProceedings of the 31st Interna- tional Conference on Computational Linguistics, O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, and S. Schockaert, Eds. Abu Dhabi, UAE: Association for Comp...

  27. [27]

    A digital language coherence marker for monitoring dementia,

    D. Gkoumas, A. Tsakalidis, and M. Liakata, “A digital language coherence marker for monitoring dementia,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association for Computational Linguistics, Dec. 2023, pp. 16 021–16 034. [Online]. Available: https://aclantho...

  28. [28]

    Reformulating NLP tasks to capture longitudinal manifestation of language disorders in people with dementia

    D. Gkoumas, M. Purver, and M. Liakata, “Reformulating NLP tasks to capture longitudinal manifestation of language disorders in people with dementia.” inProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association for Computational Linguistics, Dec. 2023, pp. 15 904–15...

  29. [29]

    Two directions for clinical data generation with large language models: Data-to-label and label-to-data,

    R. Li, X. Wang, and H. Yu, “Two directions for clinical data generation with large language models: Data-to-label and label-to-data,” inFindings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association for Computational Linguistics, Dec. 2023, pp. 7129–7143. [Online]. Available: https://ac...

  30. [30]

    CDA: A contrastive data augmentation method for Alzheimer‘s disease detection,

    J. Duan, F. Wei, J. Liu, H. Li, T. Liu, and J. Wang, “CDA: A contrastive data augmentation method for Alzheimer‘s disease detection,” inFindings of the Association for Computational Linguistics: ACL 2023, A. Rogers, J. Boyd-Graber, and N. Okazaki, Eds. Toronto, Canada: Association for Computa- tional Linguistics, Jul. 2023, pp. 1819–1826. [Online]. Availa...

  31. [31]

    Domain adaptation via prompt learning for Alzheimer‘s detection,

    S. Farzana and N. Parde, “Domain adaptation via prompt learning for Alzheimer‘s detection,” inFindings of the Association for Computational Linguistics: EMNLP 2024, Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, Eds. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 15 963–15 976. [Online]. Available: https://aclanthology.org/ 20...

  32. [32]

    Introduction to large language models (llms) for dementia care and research,

    M. S. Treder, S. Lee, and K. A. Tsvetanov, “Introduction to large language models (llms) for dementia care and research,”Frontiers in Dementia, vol. 3, p. 1385303, 2024

  33. [33]

    On the social bias of speech self-supervised models,

    Y .-C. Lin, T.-Q. Lin, H.-C. Lin, A. T. Liu, and H.-y. Lee, “On the social bias of speech self-supervised models,”Proc. Interspeech 2024, 4638-4642, 2024

  34. [34]

    Emo- bias: A large scale evaluation of social bias on speech emotion recognition,

    Y .-C. Lin, H. Wu, H.-C. Chou, C.-C. Lee, and H.-y. Lee, “Emo- bias: A large scale evaluation of social bias on speech emotion recognition,”Proc. Interspeech 2024, 2024

  35. [35]

    Automatic classification of news subjects in broadcast news: Ap- plication to a gender bias representation analysis,

    V . Pelloin, L. Dodson, ´E. Chapuis, N. Herv ´e, and D. Doukhan, “Automatic classification of news subjects in broadcast news: Ap- plication to a gender bias representation analysis,” inProc. Inter- speech 2024, 2024, pp. 3055–3059

  36. [36]

    Exploring sources of racial bias in automatic speech recognition through the lens of rhythmic varia- tion,

    L.-F. Lai and N. Holliday, “Exploring sources of racial bias in automatic speech recognition through the lens of rhythmic varia- tion,” inProc. Interspeech 2023, 2023

  37. [37]

    A con- trastive learning approach to mitigate bias in speech models,

    A. Koudounas, F. Giobergia, E. Pastor, E. Baraliset al., “A con- trastive learning approach to mitigate bias in speech models,” in Interspeech 2024. ISCA, 2024, pp. 827–831

  38. [38]

    Mitigating bias against non-native accents,

    Y . Zhang, Y . Zhang, B. M. Halpern, T. Patel, and O. Scharen- borg, “Mitigating bias against non-native accents,” inProceedings of the Annual Conference of the International Speech Communi- cation Association, INTERSPEECH, vol. 2022, 2022, pp. 3168– 3172

  39. [39]

    Au- tomatic children speech sound disorder detection with age and speaker bias mitigation,

    G. Kim, Y . Eom, S. S. Sung, S. Ha, T.-J. Yoon, and J. So, “Au- tomatic children speech sound disorder detection with age and speaker bias mitigation,” inProc. Interspeech 2024, 2024, pp. 1420–1424

  40. [40]

    Revealing con- founding biases: A novel benchmarking approach for aggregate- level performance metrics in health assessments,

    S. Goria, R. Polle, S. Fara, and N. Cummins, “Revealing con- founding biases: A novel benchmarking approach for aggregate- level performance metrics in health assessments,” inProc. Inter- speech 2024, 2024, pp. 1440–1444

  41. [41]

    Un- veil multi-picture descriptions for multilingual mild cognitive impairment detection via contrastive learning,

    K. Qi, J. Cheng, Y . Zhu, H. Amiri, and X. Liang, “Un- veil multi-picture descriptions for multilingual mild cognitive impairment detection via contrastive learning,”arXiv preprint arXiv:2505.17067, 2025

  42. [42]

    Speech Unlearning,

    J. Cheng and H. Amiri, “Speech Unlearning,” inInterspeech 2025, 2025, pp. 3209–3213

  43. [43]

    Unsupervised domain adaptation by backpropagation,

    Y . Ganin and V . Lempitsky, “Unsupervised domain adaptation by backpropagation,” inInternational conference on machine learn- ing. PMLR, 2015, pp. 1180–1189

  44. [44]

    Bengio, J

    Y . Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” inProceedings of the 26th Annual International Conference on Machine Learning, ser. ICML ’09. New York, NY , USA: Association for Computing Machinery, 2009, p. 41–48. [Online]. Available: https://doi.org/10.1145/1553374.1553380

  45. [45]

    The interspeech 2024 taukadial challenge: Multilingual mild cognitive impairment detection with multimodal approach,

    B. Barrera-Altuna, D. Lee, Z. Zarnaz, J. Han, and S. Kim, “The interspeech 2024 taukadial challenge: Multilingual mild cognitive impairment detection with multimodal approach,” inProc. Inter- speech 2024, 2024, pp. 967–971

  46. [46]

    Robust speech recognition via large-scale weak su- pervision,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak su- pervision,” inICML. PMLR, 2023

  47. [47]

    https://huggingface.co/google-bert/bert-base-multilingual-cased, 2018

  48. [48]

    Sigmoid loss for language image pre-training,

    X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” inProceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 11 975–11 986

  49. [49]

    Last layer re-training is sufficient for robustness to spuri- ous correlations,

    P. Kirichenko, P. Izmailov, and A. G. Wilson, “Last layer re-training is sufficient for robustness to spuri- ous correlations,” inThe Eleventh International Confer- ence on Learning Representations, 2023. [Online]. Available: https://openreview.net/forum?id=Zb6c8A-Fghk

  50. [50]

    Change is hard: A closer look at subpopulation shift,

    Y . Yang, H. Zhang, D. Katabi, and M. Ghassemi, “Change is hard: A closer look at subpopulation shift,”arXiv preprint arXiv:2302.12254, 2023

  51. [51]

    Do students debias like teachers? on the distillability of bias mitigation methods,

    J. Cheng, C. Agarwal, and H. Amiri, “Do students debias like teachers? on the distillability of bias mitigation methods,”arXiv preprint arXiv:2510.26038, 2025

  52. [52]

    AST: Audio Spectrogram Transformer,

    Y . Gong, Y .-A. Chung, and J. Glass, “AST: Audio Spectrogram Transformer,” inInterspeech, 2021

  53. [53]

    Unsupervised cross-lingual representation learning for speech recognition,

    A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli, “Unsupervised cross-lingual representation learning for speech recognition,”Interspeech, 2021

  54. [54]

    Last layer re- training is sufficient for robustness to spurious correlations,

    P. Kirichenko, P. Izmailov, and A. G. Wilson, “Last layer re- training is sufficient for robustness to spurious correlations,”arXiv preprint arXiv:2204.02937, 2022