Pith. sign in

REVIEW 3 major objections 4 minor 32 references

Linguistic relatedness does not improve cross-lingual transfer in large multilingual ASR once one hour of target-language data is available.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 08:31 UTC pith:CJE2DV5A

load-bearing objection Broad, transparent negative result on relatedness in ASR transfer, but the 'in every setting' claim hinges on a hand-picked Δ=5 WER bound and would flip at Δ=2. the 3 major comments →

arxiv 2607.04814 v2 pith:CJE2DV5A submitted 2026-07-06 cs.CL cs.AI

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition

classification cs.CL cs.AI
keywords cross-lingual transferautomatic speech recognitionlinguistic relatednesslow-resource African languagesmultilingual ASRsequential fine-tuningequivalence testingword error rate
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks whether pre-training an ASR system on a language related to a low-resource target reduces how much target-language data is needed. To isolate this, it sequentially fine-tunes four large multilingual ASR models on a related or unrelated auxiliary language and then on the target, across six factor combinations, two Africa-centric corpora, and ten languages in three families. The consistent result is that any transfer advantage from relatedness disappears after as little as one hour of target-language fine-tuning; models with related, unrelated, or no pre-adaptation become practically equivalent. This matters because it suggests that genealogical or typological similarity alone is not a reliable guide for choosing auxiliary languages when adapting ASR to low-resource African languages.

Core claim

Pre-adaptation on a related auxiliary language produces no practically meaningful improvement in target-language word error rate over unrelated pre-adaptation or no pre-adaptation once one hour of target-language data is used, and this holds in every experimental factor. Before target fine-tuning, pre-adapted models can beat the unmodified baseline, but related and unrelated auxiliaries are practically equivalent, showing the gain is generic pre-adaptation, not relatedness. The finding was replicated across full-model fine-tuning and parameter-efficient low-rank adaptation, across two Africa-centric corpora, and across four ASR models of different scale and architecture, with all comparisons

What carries the argument

The central mechanism is two-stage sequential fine-tuning: a base multilingual ASR model is fine-tuned on an auxiliary language and then on the target language, compared against target-only fine-tuning. Relatedness is operationalized through genealogical family membership (Nilotic, Bantu, Cushitic), validated by corpus-level transcript-similarity measures (Fréchet distance and character n-gram TF-IDF) that show within-family clustering. The conclusions rest on an equivalence-testing procedure: differences in corpus-level WER are bootstrapped over test utterances, and any difference whose 90% confidence interval lies within ±5 absolute WER points is classified as practically equivalent.

Load-bearing premise

The conclusion rests on the chosen equivalence bound of Δ = 5 WER percentage points; if differences of 2 to 4 points are practically meaningful, some reported gaps at one hour of target data would count as real relatedness gains and the central claim would not stand.

What would settle it

Run the same six-factor protocol with an equivalence bound of Δ = 2 WER points and repeated fine-tuning seeds per condition; if the 90% bootstrap confidence intervals for related-vs-baseline at one hour fall wholly outside ±2 in several factors, the paper's central claim would be false. The specific comparisons to check are Factor 3 (−3.43 [−3.72, −3.13]) and Factor 6 OmniASR (+4.04 [+3.79, +4.29]), whose intervals already cross tighter bounds.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Relatedness-guided sequential fine-tuning is unlikely to be an effective strategy for extending large multilingual ASR to low-resource languages, since even one hour of target data erases any relatedness advantage.
  • Without any target data, pre-adaptation can help, but the benefit is not specific to related languages: unrelated auxiliaries give the same transfer gain.
  • The results hold across two fine-tuning methods (full and LoRA), two corpora, and four model sizes and architectures, suggesting the pattern is not an artifact of one model or dataset.
  • The practical implication for ASR development is to collect even a small amount of target-language speech rather than spend annotation effort on auxiliary related languages, since one hour is enough to match related-language pre-adaptation.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the equivalence threshold were set lower (e.g., 2 to 3 WER points), some reported one-hour differences—such as the −3.43-point related-vs-baseline gap in Factor 3—would become meaningful, so the paper's negative conclusion is sensitive to the choice of bound.
  • The near-universal convergence after one hour suggests that large multilingual ASR models already encode enough shared structure that any target fine-tuning dominates the prior auxiliary signal; relatedness may matter more in truly zero-shot settings or with much smaller models.
  • A direct extension would test whether relatedness helps below one hour of target data, at, say, 5 to 30 minutes, or in languages from families absent from model pre-training, where the auxiliary signal might persist longer.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper asks whether fine-tuning a multilingual ASR model first on an auxiliary language that is linguistically related to a low-resource target, and then on the target, yields practically meaningful improvements over fine-tuning on the target alone or on an unrelated auxiliary. The authors embed the two-stage sequential fine-tuning recipe of Pillai et al. in a large grid: six factors, two Africa-centric corpora, ten languages from three families, and four ASR models (Whisper Small, Whisper Large v3, XLS-R, OmniASR). They compare related versus unrelated auxiliary pre-adaptation and related pre-adaptation versus a target-only baseline, using corpus-level WER differences with 90% bootstrap confidence intervals and a TOST-style equivalence bound of Δ = ±5 absolute WER points. The central finding is that while pre-adaptation can help before any target data is used, once one hour of target speech is available all conditions become practically equivalent, and therefore linguistic relatedness alone does not reliably predict transfer gains.

Significance. The question is timely and practically important: if true, the result would caution against using genealogical relatedness as the main criterion for selecting auxiliary languages when extending large multilingual ASR systems to low-resource languages. The experimental design is a genuine strength: it is broad, controlled, and transparently reported, with full per-condition WER tables, confidence intervals, code release under GPL-3.0, and an explicit limitations section. The bootstrap/TOST machinery is appropriate for the stated comparison, and the authors are careful to separate "pre-adaptation in general" from "pre-adaptation due to relatedness." The main weakness is that the headline claim is stated in universal terms ("in every setting") but depends on a hand-chosen equivalence bound and on single training runs per condition; both choices are acknowledged but not stress-tested. The paper is a solid empirical contribution, but the strength of the conclusion currently exceeds what the reported evidence can support without additional sensitivity analysis.

major comments (3)
  1. [Section 4, Tables 3–4] The central claim is operationalized through Δ = ±5 WER points, stated in Section 4 without independent justification or sensitivity analysis. The claim is not robust to plausible alternative bounds. For example, in Table 4 at 1 hour, Factor 3 (Dholuo target, related auxiliary Kalenjin) reports Δ = −3.43 [−3.72, −3.13]; under a Δ = 2 bound this interval lies wholly below −2 and would be classified as a practically meaningful improvement, directly contradicting "in every setting." Similarly, in Table 3 at 1 hour, Factor 6 Whisper Large shows Δ = −2.55 [−2.74, −2.35], which would also be meaningful under Δ = 2. Since the authors provide no external anchor for the 5-point threshold and no sensitivity analysis over smaller bounds, the universal form of the conclusion is not supported by the data as presented. This is load-bearing because the conclusion is defined by the bound, not discovered
  2. [Section 4, bootstrap procedure] All confidence intervals in Tables 3 and 4 are obtained by bootstrapping over test utterances from a single trained model per condition. The intervals therefore capture test-set sampling variability only, not fine-tuning stochasticity. With one seed per condition, the observed point differences—and hence the practical-equivalence classifications—could shift under different initializations. Given that the paper's conclusion is about the effects of a training procedure (pre-adaptation), the uncertainty in the training process itself is directly relevant. The authors should either run multiple seeds per condition (at least for a representative subset) and report the across-seed variance, or explicitly caveat the conclusion as conditional on single runs. This issue is acknowledged in the Limitations but is not resolved by the current analysis.
  3. [Section 3.3 and Factor 6, Figure 5] For OmniASR, the paper states that the model has pre-exposure to all AfriVoices KE languages except Maasai, which includes the target Kalenjin and the related auxiliary Dholuo. This makes the Factor 6 OmniASR condition a weak test of cross-lingual transfer: the model has already seen the target language during pretraining, so relatedness-driven transfer is confounded with target-language memorization. The paper notes this difference in prior exposure but still includes OmniASR in the "across all models" generalization claim. The authors should either analyze OmniASR separately, or restrict the universal claim to models without target pretraining coverage.
minor comments (4)
  1. [Abstract / Conclusion] The phrase "in every setting" is stronger than what the analysis supports, especially given the sensitivity to Δ and the single-run design. Consider softening to "in the settings we tested, and under the stated Δ = 5 bound."
  2. [Table 2] The "None" column includes WER values above 100, which is explained in the caption but may still confuse readers. A footnote or explicit marker would help.
  3. [Section 4, TOST description] The paper says the equivalence criterion is "inherited from a two one-sided tests (TOST) procedure," but the reported classification based on whether the CI lies wholly inside ±5 is more directly a "confidence-interval inclusion" rule. The TOST framing adds little and may imply a formal hypothesis test with error control that the interval-based procedure does not fully provide.
  4. [Section 5, Factor 4] The dual-auxiliary Factor 4 comparison in Table 3 uses the Factor 1 unrelated-auxiliary grouping as the reference, but the caption does not state whether those reference models were retrained in the same batch or are simply carried over from Factor 1. Please clarify.

Circularity Check

0 steps flagged

No significant circularity: the central claim is an empirical comparison, not a fitted prediction or self-citation chain.

full rationale

The paper's central conclusion—that related-language pre-adaptation yields no practically meaningful WER improvement once one hour of target data is used—is an empirical finding, not a derivation from its own inputs. WER values come from held-out target test sets under different fine-tuning conditions, and the comparison is operationalized with a stated equivalence bound of Δ = ±5 WER percentage points (Section 4). This bound is not fitted to the data or to the outcome; it is a declared modeling choice, explicitly acknowledged in the Limitations: 'Our conclusions assume an equivalence bound of ∆ = 5 WER points as the threshold for practical meaningfulness.' The bootstrap confidence intervals could have fallen outside the bounds; the fact that they did not is an empirical result. No fitted parameter is later renamed as a prediction, and no equation defines the result in terms of itself. The paper's self-citations (e.g., Wanzare et al., 2026; Mbogho et al., 2025) concern dataset construction and corpus statistics, not the theoretical load-bearing claim, and do not assume the conclusion. The concerns about sensitivity to Δ, single training runs, and the exclusion of fine-tuning stochasticity from confidence intervals are robustness or uncertainty issues, not circularity. The paper itself transparently records these limitations. Because the central claim is conditional on an explicit, non-fitted criterion and is contradicted by no internal construction, no circular step is present.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

The central claim rests on the chosen equivalence bound, the family-based definition of relatedness, and the assumption that transcript similarity approximates phonological/morphological similarity. Free parameters are experimental design choices (data volumes, bound) plus LoRA hyperparameters; no invented entities are introduced.

free parameters (3)
  • Practical equivalence bound Δ = 5 WER percentage points
    Hand-chosen TOST bound in Section 4; all ‘practically meaningful’ conclusions are statements about this threshold.
  • Target-language data volume grid = 1, 10, 70 hours (25 for WaxalNLP)
    Selected data volumes define the ‘once one hour is available’ result; they are experiment-design choices, not derived.
  • LoRA rank and learning rate per data volume = (4,1e-4), (8,1e-4), (8,7e-5)
    Ad hoc settings for Factor 2, scaled by data volume and tuned on validation data; not the central transfer comparison.
axioms (4)
  • domain assumption Genealogical family membership (Nilotic/Bantu/Cushitic) is a valid proxy for multidimensional linguistic relatedness.
    Used throughout to define related vs unrelated auxiliary-target pairs (Sections 3.1–3.2); supported only by broad typological characterization and corpus-level transcript similarity, which the authors note is limited by scarce linguistic documentation.
  • domain assumption Text-based transcript similarity captures phonological and morphological similarity because the languages use largely phonemic orthographies.
    Section 3.2; underpins the corpus-level validation of language groupings, though tonal/pitch features are not transcribed.
  • standard math Non-parametric bootstrap over test utterances gives valid confidence intervals for group WER differences.
    Section 4, 5,000 resamples; standard, but covers only test-sampling error, not training-run variance.
  • domain assumption Two-stage sequential fine-tuning is an appropriate vehicle for measuring cross-lingual transfer.
    Section 4; adopted from Pillai et al. (2025). The conclusion is about this specific strategy, not all possible transfer mechanisms.

pith-pipeline@v1.3.0-alltime-deepseek · 14061 in / 10448 out tokens · 111342 ms · 2026-08-02T08:31:19.905016+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition." pith.science (2026). https://pith.science/paper/CJE2DV5A

@misc{pith2026260704814,
  author       = {Pith},
  title        = {Pith review of: Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CJE2DV5A}},
  note         = {Machine review of arXiv:2607.04814}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection at scale. A promising direction is to leverage the linguistic relatedness between a low-resource target language and languages previously seen by a model to reduce the volume of target-language data needed for effective adaptation. Although this approach has proven reliable for text-based models, its effectiveness in the speech domain remains contested. We employ a systematic controlled experimental design spanning six factors, two Africa-centric corpora, and four large ASR models, sequentially adapting on a related auxiliary language followed by the target to isolate whether linguistic relatedness reliably predicts cross-lingual transfer gains across these conditions. In every setting, pre-adaptation on related auxiliary languages yields no practically meaningful improvements once as little as one hour of target-language data is available, suggesting that relatedness alone may not reliably predict transfer gains in large multilingual ASR, or constitute an effective strategy for extending such models to low-resource languages.

Figures

Figures reproduced from arXiv: 2607.04814 by Andrei Florian, Biatus Maina Kamau, Boniface Mwau, Christiane Fellbaum, Cynthia Jayne Amol, Happy Buzaaba, Hope Kerubo Ombaba, Lilian Diana Awuor Wanzare, Xiaoyu Cui.

Figure 1
Figure 1. Figure 1: Experimental design for the first factor us [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Corpus-level pairwise similarity results across three metrics for all languages in the AfriVoices KE (top) [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: WER on the Kalenjin test set across all first-factor models. [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: WER on the target language test set as a function of target-language fine-tuning hours across factors two [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: WER on the Kalenjin test set as a function [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

32 extracted references · 3 canonical work pages

  1. [1]

    Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages =

    On the Cross-lingual Transferability of Monolingual Representations , author =. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages =. 2020 , publisher =

  2. [2]

    2111.09296 , archivePrefix =

    Babu, Arun and Wang, Changhan and Tjandra, Andros and Lakhotia, Kushal and Xu, Qiantong and Goyal, Naman and Singh, Kritika and von Platen, Patrick and Saraf, Yatharth and Pino, Juan and Baevski, Alexei and Conneau, Alexis and Auli, Michael , year =. 2111.09296 , archivePrefix =

  3. [3]

    Transactions of the Association for Computational Linguistics , volume =

    Enriching Word Vectors with Subword Information , author =. Transactions of the Association for Computational Linguistics , volume =. 2017 , url =

  4. [4]

    Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages =

    Unsupervised Cross-lingual Representation Learning at Scale , author =. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages =. 2020 , publisher =

  5. [5]

    2602.02734 , archivePrefix =

    Diack, Abdoulaye and Nelson, Perry and Agbesi, Kwaku and Nakalembe, Angela and MohamedKhair, MohamedElfatih and Dube, Vusumuzi and Siyavora, Tavonga and Venugopalan, Subhashini and Hickey, Jason and Okonkwo, Uche and Bapna, Abhishek and Wiafe, Isaac and Helegah, Raynard Dodzi and Atsakpo, Elikem Doe and Nutrokpor, Charles and Winful, Fiifi Baffoe Payin an...

  6. [6]

    Information Processing & Management , volume =

    Zero-shot Cross-lingual Transfer Language Selection Using Linguistic Similarity , author =. Information Processing & Management , volume =. 2023 , url =

  7. [7]

    2017 , url =

    Heusel, Martin and Ramsauer, Hubert and Unterthiner, Thomas and Nessler, Bernhard and Hochreiter, Sepp , booktitle =. 2017 , url =

  8. [8]

    and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , booktitle =

    Hu, Edward J. and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , booktitle =. 2022 , url =

  9. [9]

    The State and Fate of Linguistic Diversity and Inclusion in the

    Joshi, Pratik and Santy, Sebastin and Budhiraja, Amar and Bali, Kalika and Choudhury, Monojit , booktitle =. The State and Fate of Linguistic Diversity and Inclusion in the. 2020 , publisher =

  10. [10]

    Omnilingual

    Keren, Gil and Kozhevnikov, Artyom and Meng, Yen and Ropers, Christophe and Setzler, Matthew and Wang, Skyler and Adebara, Ife and Auli, Michael and Balioglu, Can and Chan, Kevin and Cheng, Chierh and Chuang, Joe and Droof, Caley and Duppenthaler, Mark and Duquenne, Paul-Ambroise and Erben, Alexander and Gao, Cynthia and Gonzalez, Gabriel Mejia and Lyu, K...

  11. [11]

    Low Resource

    Khare, Shreya and Mittal, Ashish and Diwan, Anuj and Sarawagi, Sunita and Jyothi, Preethi and Bharadwaj, Samarth , booktitle =. Low Resource. 2021 , publisher =

  12. [12]

    Social Psychological and Personality Science , volume =

    Equivalence Tests: A Practical Primer for t Tests, Correlations, and Meta-Analyses , author =. Social Psychological and Personality Science , volume =. 2017 , url =

  13. [13]

    Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing , pages =

    From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual Transformers , author =. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing , pages =. 2020 , publisher =

  14. [14]

    Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages =

    Choosing Transfer Languages for Cross-Lingual Learning , author =. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages =. 2019 , publisher =

  15. [15]

    Exploration of

    Liu, Yunfei and Yang, Xiaojia and Qu, Dan , journal =. Exploration of. 2024 , url =

  16. [16]

    Building Low-resource

    Mbogho, Audrey and Awuor, Quin and Kipkebut, Andrew and Wanzare, Lilian and Oloo, Vivian , journal =. Building Low-resource. 2025 , url =

  17. [17]

    Nekoto, Wilhelmina and Marivate, Vukosi and Matsila, Tshinondiwa and Fasubaa, Timi and Fagbohungbe, Taiwo and Akinola, Solomon Oluwole and Muhammad, Shamsuddeen and Kabongo Kabenamualu, Salomon and Osei, Salomey and Sackey, Freshia and Niyongabo, Rubungo Andre and Macharm, Ricky and Ogayo, Perez and Ahia, Orevaoghene and Berhe, Musie Meressa and Adeyemi, ...

  18. [18]

    Information Processing & Management , volume =

    Adapting Multilingual Speech Representation Model for a New, Under-resourced Language through Multilingual Fine-tuning and Continued Pretraining , author =. Information Processing & Management , volume =. 2023 , url =

  19. [19]

    Masakhane -- Machine Translation For

    Orife, Iroro and Kreutzer, Julia and Sibanda, Blessing and Whitenack, Daniel and Siminyu, Kathleen and Martinus, Laura and Ali, Jamiil Toure and Abbott, Jade and Marivate, Vukosi and Kabongo, Salomon and Meressa, Musie and Murhabazi, Espoir and Ahia, Orevaoghene and van Biljon, Elan and Ramkilowan, Arshath and Akinfaderin, Adewale and. Masakhane -- Machin...

  20. [20]

    2021 , url =

    Pillutla, Krishna and Swayamdipta, Swabha and Zellers, Rowan and Thickstun, John and Welleck, Sean and Choi, Yejin and Harchaoui, Zaid , booktitle =. 2021 , url =

  21. [21]

    2022 , eprint =

    Robust Speech Recognition via Large-Scale Weak Supervision , author =. 2022 , eprint =

  22. [22]

    Sentence-

    Reimers, Nils and Gurevych, Iryna , booktitle =. Sentence-. 2019 , publisher =

  23. [23]

    2024 , publisher =

    Song, Zheshu and Zhuo, Jianheng and Yang, Yifan and Ma, Ziyang and Zhang, Shixiong and Chen, Xie , booktitle =. 2024 , publisher =. doi:10.21437/Interspeech.2024-892 , issn =

  24. [24]

    2604.08448 , archivePrefix =

    Wanzare, Lilian and Amol, Cynthia and Maina, Ezekiel and Odhiambo, Nelson and Kerubo, Hope and Misula, Leila and Oloo, Vivian and Mboya, Rennish and Onkoba, Edwin and Ombui, Edward and Muguro, Joseph and wa Maina, Ciira and Kipkebut, Andrew and Otom, Alfred Omondi and Kang'ethe, Ian Ndung'u and Kanyi, Angela Wambui and Omwenga, Brian Gichana , year =. 260...

  25. [25]

    2019 , publisher =

    Wu, Shijie and Dredze, Mark , booktitle =. 2019 , publisher =

  26. [26]

    Are All Languages Created Equal in Multilingual

    Wu, Shijie and Dredze, Mark , booktitle =. Are All Languages Created Equal in Multilingual. 2020 , publisher =

  27. [27]

    2023 , url =

    Zhang, Qingru and Chen, Minshuo and Bukharin, Alexander and He, Pengcheng and Cheng, Yu and Chen, Weizhu and Zhao, Tuo , booktitle =. 2023 , url =

  28. [28]

    Interspeech 2021 , year =

    Unsupervised Cross-Lingual Representation Learning for Speech Recognition , author =. Interspeech 2021 , year =. doi:10.21437/Interspeech.2021-329 , url =

  29. [29]

    Proceedings of the 6th Workshop on Research in Computational Linguistic Typology and Multilingual NLP , month = mar, year =

    Predicting Positive Transfer for Improved Low-Resource Speech Recognition Using Acoustic Pseudo-Tokens , author =. Proceedings of the 6th Workshop on Research in Computational Linguistic Typology and Multilingual NLP , month = mar, year =. doi:10.18653/v1/2024.sigtyp-1.13 , url =

  30. [30]

    Interspeech 2023 , year =

    Comparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages , author =. Interspeech 2023 , year =. doi:10.21437/Interspeech.2023-1061 , url =

  31. [31]

    Elizabeth and Chambers, Summer and Kelley, Matthew and Walther, G

    Liebl, J. Elizabeth and Chambers, Summer and Kelley, Matthew and Walther, G. Aspects of Selecting the Right. Proceedings of the Ninth Workshop on the Use of Computational Methods in the Study of Endangered Languages (. 2026 , address =. doi:10.18653/v1/2026.computel-1.16 , url =

  32. [32]

    ACM Transactions on Asian and Low-Resource Language Information Processing , publisher =

    Multistage Fine-tuning Strategies for Automatic Speech Recognition in Low-resource Languages , author =. ACM Transactions on Asian and Low-Resource Language Information Processing , publisher =. 2025 , note =. doi:10.1145/3813800 , url =