REVIEW 3 major objections 5 minor 114 references
Addressing Pitfalls in Auditing Practices of Automatic Speech Recognition Technologies: A Case Study of People with Aphasia
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Standard speech-to-text audits systematically understate how badly aphasia speech is transcribed.
desk verdict Robust WER audit of six ASR systems on aphasia speech shows a real disparity, but the 'only Whisper hallucinates' claim rests on spot checks and needs to be hedged or systematically verified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a three-part audit design. First, text standardization is varied across five levels, from the original transcript through removal of fillers, fragments, repeated words, and repeated phrases; the paper shows WER and even service rankings change with that choice, and it uses the aphasia community’s stated preference for the most cleaned version as the primary reference. Second, performance is disaggregated by aphasia fluency, Boston-classification type, gender, race, and acoustic covariates—nonvocal duration share (a dysfluency proxy from voice activity detection) and background noise energy—analysed with regressions clustered on speakers. Third, evaluation uses a metric suite beyond WER$, = (S+I+D)/N$: CER, BLEU, ROUGE-1/2/L, METEOR, WIL, RIL, insertion rate, and a manually verified binary hallucination indicator. Each part is designed to catch a class of error that the other parts miss.
What would settle it
Re-run the audit keeping the segments the pipeline dropped (those containing unintelligible words), score them against a reference that preserves or clinically adjudicates those words, and check whether the 6–10 percentage point WER gap and the 53-of-56 hallucination concentration persist. If the gap shrinks substantially, the reported disparity is an artifact of ground-truth cleaning rather than a property of the ASR services.
Extended reading notes
Core claim
Across six commercial ASR services, transcriptions of speech from people with aphasia consistently have WERs 6–10 percentage points higher than control speakers (for example, 0.17 versus 0.09 for the worst-performing service and 0.12 versus 0.06 for the best), with all differences significant at $p<0.001$. Among 56 confirmed hallucinations in the open-weight Whisper model, 53 occurred for aphasia speakers, and an audio-manipulation experiment produced more Whisper hallucinations for aphasia speech than for control speech. The methodological claim is that standard audit practices hide this harm: a single text-standardization choice can reverse which service ranks best, treating “aphasia” as one group hides that non-fluent and Global aphasia are far worse (average WER 0.21 and 0.305, versus 0.07 for controls), and WER cannot distinguish fabricated hallucinated content from ordinary insertion errors. The paper concludes that audits should vary standardization in line with community preferences, disaggregate subgroups and acoustic covariates, and report a metric suite that includes hallucination rate.
Load-bearing premise
The cleaned human transcriptions used as the reference for both groups are correct and unbiased, even though every audio segment containing a word marked unintelligible was thrown out, and that removal may exclude the most severe aphasia speech and distort the measured disparity.
Editorial extensions
If this is right
- If the central claim holds, single-metric, single-standardization audits can produce unstable service rankings: the paper shows Whisper significantly outperforming Amazon under minimal cleaning but Amazon significantly outperforming Whisper under the community-preferred cleaning.
- Disaggregation by aphasia type changes the conclusion: non-fluent and Global aphasia show much higher WERs, so any audit that reports only “aphasia versus control” conceals the speakers who need the most accurate transcription.
- Hallucination rate should be a standard reporting metric, because WER and insertion rate cannot distinguish fabricated content from stutter-like repetitions, and in this study hallucinations were nearly exclusive to aphasia speakers.
- Acoustic covariates (nonvocal pause share and background noise) are measurable confounders that raise both WER and hallucination likelihood, but the aphasia indicator remains the dominant factor after adjustment.
- Community-preferred cleaning is not in tension with competitive WER: for half of the services, removing repeated words made no significant WER difference, so audits need not sacrifice comparability to honor user preferences.
Reading between the lines
- The three pitfalls likely generalize beyond aphasia: any speech community with disfluencies, non-standard dialects, or clinician-specific transcription norms should be audited with multiple standardizations, disaggregated subgroups, and hallucination metrics, but that extension is ours, not the paper’s.
- Because the paper’s ground-truth cleaning removes every segment containing an unintelligible word, the reported WER gap is probably a lower bound; an audit that retains or clinically adjudicates those segments could find an even larger disparity.
- A testable extension of the hallucination result is to apply the paper’s audio-manipulation treatments (leading silence, white noise, early cutoff) to other large ASR models; the paper shows these manipulations raise Whisper hallucination rates and affect aphasia speech more, but whether other models behave the same is unknown.
- The community-preference finding suggests ASR standardization should become a user-facing option rather than a hidden default, since preferences varied with recovery stage and purpose; this design implication goes beyond what the paper demonstrates.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a community-driven auditing framework for ASR systems and demonstrates it in a case study comparing six commercial ASR services on AphasiaBank speech from people with aphasia and a control group. It identifies three pitfalls in standard ASR audits—fixed text standardization, aggregate-only demographic comparisons, and reliance on a single metric (WER)—and presents evidence for each: WER disparities between aphasia and control speakers across all six services, heterogeneity of performance across aphasia subtypes and acoustic covariates, and Whisper-specific hallucinations that are largely concentrated in aphasia speakers. The paper also reports a small community survey indicating that aphasia speakers prefer heavily cleaned transcriptions, and it makes reproducibility claims backed by a public GitHub repository.
Significance. If the findings hold, the paper makes a valuable contribution to fair and accessible ASR auditing. The central WER disparity is supported by multiple robustness checks: it holds across six services, on matched and unmatched samples, under weighted and unweighted aggregation, and in regression models with clustered standard errors. The disaggregation by aphasia type and acoustic features is a useful corrective to monolithic group comparisons, and the community-engagement component is a strength, even if small. The paper also ships code and gives detailed preprocessing descriptions, which supports reproducibility. The main caveat is that one load-bearing claim—that hallucinations occur only in Whisper—rests on a weaker evidentiary basis than the rest of the empirical analysis.
major comments (3)
- [§6 and Appendix A.8.2] The claim that 'no instances of hallucinations in the other ASR services' were found is not supported by the described methodology. The manual hallucination review was applied only to 1,198 Whisper candidate files selected via percentile thresholds on WER, BLEU, ROUGE, METEOR, WIL, RIL, CER, and Insertion Rate; the other five services were only spot-checked 'throughout the cleaning process.' A hallucination in Amazon, Google, Microsoft, AssemblyAI, or Rev AI that does not push a file past those thresholds would not enter the reviewed set, and the spot checks are not described as service-blind or systematic. Because the abstract and Pitfall 3 treat hallucinations as a distinct class of generative-AI errors, the exclusivity result is load-bearing. I recommend either conducting the same systematic manual review on the other services' candidate files or explicitly weakening the claim to 'Whisper hallucinated in our data; we did not systematically verify absence in other services.'
- [Appendix A.1.1] The ground-truth cleaning pipeline removes any word marked unintelligible ('xxx') and then drops every audio segment containing such a token. This is likely to exclude the most severely affected aphasia speech, since unintelligible tokens are a hallmark of severe aphasia. The paper does not report how many segments were excluded per group or provide a sensitivity analysis that retains or imputes these segments. This selection could attenuate or otherwise distort the measured WER disparity and the hallucination concentration. I ask the authors to quantify the exclusion counts by group and, if feasible, show that the main results are robust to an alternative treatment of unintelligible segments (e.g., retaining them as error tokens or analyzing the excluded subset separately).
- [§2.2 and §3] The 'standard audit' results in §3 are not based on default configurations for two of the six services: Google uses the Chirp model and Microsoft uses the Azure Continuous model, both chosen after initial testing revealed poor default performance. This is a reasonable engineering decision, but it means the headline comparison is not strictly a comparison of out-of-the-box systems. The paper should state more prominently that the results are model-specific rather than service-default-specific, and should report the default-model WERs (or at least the observed failure modes) so readers can assess how much the exceptions affect the rank ordering and the disparity estimates.
minor comments (5)
- [Table 1] The checkmark notation in Table 1 is hard to parse (e.g., '✓(✓)✓' appears as a single cell entry); a legend or separate columns for 'default' versus 'available but not default' would improve clarity.
- [§4.2 and Figure 2] The text says 'Wilcoxon signed rank tests' in one place and 'Wilcoxon rank-sum tests' in another; these are different tests, and the figure caption should specify which was used for the pairwise comparisons.
- [Appendix A.2.2] The description of the whisper_normalizer-based cleaning is detailed, but the statement that 'we remove additional filler words not removed by the whisper_normalizer' lists filler tokens in a comma-separated inline list; presenting them as a table or code block would make the exact token set easier to reproduce.
- [References] References [49] and [50] appear to be the same survey paper duplicated with different formatting; please merge or distinguish them.
- [Appendix A.8.2] The sentence describing hallucination traits lists 'repetitions not present in the audio file' as a hallucination indicator; since WER insertions from stutters are explicitly distinguished from hallucinations earlier in §6, please clarify how this indicator was applied to avoid overlap with ordinary disfluency insertions.
Circularity Check
No significant circularity: the central WER disparity is measured against external AphasiaBank ground truth and is robust across standardization variants; self-citations are methodological benchmarks, not load-bearing inputs.
full rationale
The paper's central claims are empirical and self-contained. The WER disparity between aphasia and control speakers is computed by comparing six commercial ASR outputs against AphasiaBank's externally curated CHAT transcriptions, using standard string-matching metrics (WER, CER, BLEU, etc.); no fitted parameter is renamed as a prediction. The community-preferred RFFRR standardization is used as the main preprocessing, but the authors explicitly report that the aphasia-control disparity persists across all standardization variants (Figure 2 and Appendix Tables 7-8), so the headline result does not reduce to that input choice. The propensity-score matching and regression analyses are descriptive controls, not circular derivations: the aphasia coefficient remains significant after conditioning on acoustic and demographic covariates. Hallucination findings are new manual audits of 1,198 metric-selected files, with the rate explicitly labeled as a lower bound and confirmed hallucinations enumerated; the assertion that other services showed no hallucinations is supported only by spot checks, which is an evidentiary limitation rather than a circularity. Self-citations (e.g., Koenecke et al. 2024 for hallucination taxonomy, Choi and Mei 2025 for the absence of automated detectors) supply definitions and prior benchmarks but do not constitute the load-bearing evidence for the current results, which rest on the paper's own data collection and manual review. No equation or claim is shown to be equivalent to its own inputs by construction.
Assumptions & free parameters
free parameters (4)
- Propensity matching caliper =
0.13
- Background noise RMS threshold =
0.01
- Hallucination review percentile cutoffs =
90th percentile for lower-is-better metrics, 10th percentile for higher-is-better metrics
- Audio inclusion cutoffs =
max 4 minutes, min 2 seconds, min 4 ground-truth words
assumptions (6)
- domain assumption AphasiaBank CHAT transcriptions are accurate and programmatically cleanable without systematic bias.
- domain assumption Removing unintelligible tokens and segments containing them does not bias the aphasia-control comparison.
- domain assumption Silero VAD nonvocal duration share is a valid proxy for speech dysfluency.
- domain assumption Six of seven surveyed C.H.A.T. participants' preference for maximal text cleaning generalizes to the aphasia community.
- domain assumption Manual hallucination review, guided by the taxonomy of Koenecke et al. (2024), correctly distinguishes hallucinations from mistranscriptions.
- standard math Standard clustered-regression assumptions hold for the WER and hallucination models.
Cite this review
Pith. "Pith review of Addressing Pitfalls in Auditing Practices of Automatic Speech Recognition Technologies: A Case Study of People with Aphasia." pith.science (2026). https://pith.science/paper/5DVHWL32
@misc{pith2026250608846,
author = {Pith},
title = {Pith review of: Addressing Pitfalls in Auditing Practices of Automatic Speech Recognition Technologies: A Case Study of People with Aphasia},
year = {2026},
howpublished = {\url{https://pith.science/paper/5DVHWL32}},
note = {Machine review of arXiv:2506.08846}
}
read the original abstract
Automatic Speech Recognition (ASR) systems' growing use warrants robust auditing approaches to ensure equitable transcription quality, especially for people with speech disorders like aphasia who disproportionately depend on ASR. While academic and industry audits have revealed performance disparities across user populations, standard auditing practices often overlook nuances that risk masking harm to marginalized groups. We identify three common pitfalls in standard ASR audits: (1) adhering to one method of text standardization, which can mask variance in ASR performance and ignore the standardization preferences of marginalized communities; (2) displaying high-level demographic findings without considering performance disparities by nuanced intersectional subgroups, or conditioning on relevant acoustic properties; and (3) reporting only one gold-standard metric (Word Error Rate), which inadequately quantifies common generative AI errors like hallucinations. We propose a holistic auditing framework addressing these pitfalls, and in a case study of six popular ASR systems, find consistently worse ASR performance for speakers with aphasia relative to a control group. We call on practitioners to implement these robust, community-driven ASR auditing practices better suited for the rapidly changing ASR landscape.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Tanel Alumäe and Allison Koenecke. 2025. Striving for open-source and equitable speech-to-speech translation
2025
-
[2]
Amberscript. 2022. Transcription Guidelines. https://www.amberscript.com/en/academy/transcription-guidelines/ Accessed: 2024-11- 12
2022
-
[3]
Giuseppe Attanasio, Beatrice Savoldi, Dennis Fucci, Dirk Hovy, et al. 2024. Twists, humps, and pebbles: multilingual speech recognition models exhibit gender performance gaps. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics
2024
-
[4]
Max Bain, Jaesung Huh, Tengda Han, and Andrew Zisserman. 2023. WhisperX: Time-Accurate Speech Transcription of Long-Form Audio. InProceedings of Interspeech 2023. 4489–4493
2023
-
[5]
Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. InProceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization. 65–72
2005
-
[6]
Solon Barocas, Anhong Guo, Ece Kamar, Jacquelyn Krones, Meredith Ringel Morris, Jennifer Wortman Vaughan, W Duncan Wadsworth, and Hanna Wallach. 2021. Designing disaggregated evaluations of ai systems: Choices, considerations, and tradeoffs. InProceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society. 368–378
2021
-
[7]
Loïc Barrault, Yu-An Chung, Mariano Cora Meglioli, David Dale, Ning Dong, Paul-Ambroise Duquenne, Hady Elsahar, Hongyu Gong, Kevin Heffernan, John Hoffman, et al. 2023. SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.arXiv preprint arXiv:2308.11596(2023)
arXiv 2023
-
[8]
1996.Aphasia: A clinical perspective
David Frank Benson and Alfredo Ardila. 1996.Aphasia: A clinical perspective. Oxford University Press, New York, NY
1996
Show all 114 references
-
[9]
Tuba Bircan and Duha Ceylan. 2024. Machine Discriminating: Automated Speech Recognition Biases in Refugee Interviews.Journal of Immigrant & Refugee Studies(2024), 1–16
2024
-
[10]
Hervé Bredin, Ruiqing Yin, Juan Manuel Coria, Gregory Gelly, Pavel Korshunov, Marvin Lavechin, Diego Fustes, Hadrien Titeux, Wassim Bouaziz, and Marie-Philippe Gill. 2020. Pyannote. audio: neural building blocks for speaker diarization. InICASSP 2020-2020 IEEE International co...
2020
-
[11]
2023.Feminist AI: Critical Perspectives on Algorithms, Data, and Intelligent Machines
Jude Browne, Stephen Cave, Eleanor Drage, and Kerry McInerney. 2023.Feminist AI: Critical Perspectives on Algorithms, Data, and Intelligent Machines. Oxford University Press, United Kingdom
2023
-
[12]
Garance Burke and Hilke Schellmann. 2024. Researchers say an AI-powered transcription tool used in hospitals invents things no one ever said.The Associated Press(2024). https://apnews.com/article/ai-artificial-intelligence-health-business- 90020cdf5fa16c79ca2e5b6c4c9bbb14
2024
-
[13]
Kelly Caine. 2016. Local standards for sample size at CHI. InProceedings of the 2016 CHI conference on human factors in computing systems. 981–992
2016
-
[14]
Joan A Casey, Rachel Morello-Frosch, Daniel J Mennitt, Kurt Fristrup, Elizabeth L Ogburn, and Peter James. 2017. Race/ethnicity, socioeconomic status, residential segregation, and spatial variation in noise exposure in the contiguous United States.Environmental health perspect...
2017
-
[15]
1998.Nothing about us without us: Disability oppression and empowerment
James I Charlton. 1998.Nothing about us without us: Disability oppression and empowerment. University of California Press, Berkeley, CA
1998
-
[16]
Anna Seo Gyeong Choi and Katelyn Xiaoying Mei. 2025. What are AI hallucinations? Why AIs sometimes make things up.The Conversation(21 March 2025). https://theconversation.com/what-are-ai-hallucinations-why-ais-sometimes-make-things-up-242896 Accessed: June 6, 2025
2025
-
[17]
Wu Chou, C-H Lee, B-H Juang, and Frank K. Soong. 1994. A minimum error rate pattern recognition approach to speech recognition. International Journal of Pattern Recognition and Artificial Intelligence8, 01 (1994), 5–31
1994
-
[18]
Chris Code and Brian Petheram. 2011. Delivering for aphasia.International Journal of Speech-Language Pathology13, 1 (Feb. 2011), 3–10
2011
-
[19]
European Commission. 2021. Proposal for a Regulation laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) and amending certain Union legislative acts. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:52021PC0206 Accessed: 2024-10-01
2021
-
[20]
Equal Employment Opportunity Commission. 1978. Uniform Guidelines on Employee Selection Procedures (1978). https://www.eeoc. gov/laws/guidance/four-fifths-rule Accessed: 2024-10-01. Addressing Auditing Pitfalls in ASR Technologies FAccT ’26, June 25–28, 2026, Montreal, QC, Canada
1978
-
[21]
New York City Council. 2021. NYC Local Law 144 of 2021: Automated Employment Decision Tools. https://www.nyc.gov/assets/dca/ downloads/pdf/workers/Local-Law-144-AEDT-Summary.pdf Accessed: 2024-10-01
2021
-
[22]
Laura M Dale, Sophie Goudreau, Stephane Perron, Martina S Ragettli, Marianne Hatzopoulou, and Audrey Smargiassi. 2015. Socioeco- nomic status and environmental noise exposure in Montreal, Canada.BMC public health15 (2015), 1–8
2015
-
[23]
Deepgram. 2024. Why switch from OpenAI Whisper? https://deepgram.com/compare/openai-vs-deepgram-alternative
2024
-
[24]
Fernando Delgado, Stephen Yang, Michael Madaio, and Qian Yang. 2023. The participatory turn in ai design: Theoretical foundations and the current state of practice. InProceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization. 1–23
2023
-
[25]
Wesley Hanwen Deng, Boyuan Guo, Alicia Devrio, Hong Shen, Motahhare Eslami, and Kenneth Holstein. 2023. Understanding Practices, Challenges, and Opportunities for User-Engaged Algorithm Auditing in Industry Practice. InProceedings of the 2023 CHI Conference on Human Factors in...
2023
-
[26]
Wesley Hanwen Deng, Manish Nagireddy, Michelle Seng Ah Lee, Jatinder Singh, Zhiwei Steven Wu, Kenneth Holstein, and Haiyi Zhu
-
[27]
Mackenzie E Fama, Erin Lemonds, and Galya Levinson. 2022. The subjective experience of word-finding difficulties in people with aphasia: A thematic analysis of interview data.American Journal of Speech-Language Pathology31, 1 (2022), 3–11
2022
-
[28]
Siyuan Feng, Bence Mark Halpern, Olya Kudina, and Odette Scharenborg. 2024. Towards inclusive automatic speech recognition. Computer Speech & Language84 (2024), 101567
2024
-
[29]
Siyuan Feng, Olya Kudina, Bence Mark Halpern, and Odette Scharenborg. 2021. Quantifying Bias in Automatic Speech Recognition. arXiv:2103.15122 [eess.AS] https://arxiv.org/abs/2103.15122
2021 arXiv
-
[30]
Kathleen C Fraser, Jed A Meltzer, Naida L Graham, Carol Leonard, Graeme Hirst, Sandra E Black, and Elizabeth Rochon. 2014. Automated classification of primary progressive aphasia subtypes from narrative speech transcripts.cortex55 (2014), 43–60
2014
-
[31]
Kathleen C Fraser, Frank Rudzicz, Naida Graham, and Elizabeth Rochon. 2013. Automatic speech recognition in the diagnosis of primary progressive aphasia. InProceedings of the fourth workshop on speech and language processing for assistive technologies. 47–54
2013
-
[32]
Rita Frieske and Bertram E Shi. 2024. Hallucinations in Neural Automatic Speech Recognition: Identifying Errors and Hallucinatory Models.arXiv preprint arXiv:2401.015720 (2024)
2024 arXiv
-
[33]
Sanchit Gandhi, Patrick von Platen, and Alexander M Rush. 2023. Distil-whisper: Robust knowledge distillation via large-scale pseudo labelling.arXiv preprint arXiv:2311.00430(2023)
2023 arXiv
-
[34]
Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, and Jonathan Herzig. 2024. Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 7765–7784
2024
-
[35]
Mengzhe Geng, Xurong Xie, Zi Ye, Tianzi Wang, Guinan Li, Shujie Hu, Xunying Liu, and Helen Meng. 2022. Speaker adaptation using spectro-temporal deep features for dysarthric and elderly speech recognition.IEEE/ACM Transactions on Audio, Speech, and Language Processing30 (2022)...
2022
-
[36]
Laurence Gillick and Stephen J Cox. 1989. Some statistical issues in the comparison of speech recognition algorithms. InInternational Conference on Acoustics, Speech, and Signal Processing,. IEEE, 532–535
1989
-
[37]
Sharon Goldwater, Dan Jurafsky, and Christopher D Manning. 2010. Which words are hard to recognize? Prosodic, lexical, and disfluency factors that increase speech recognition error rates.Speech Communication52, 3 (2010), 181–200
2010
-
[38]
2001.BDAE: The Boston Diagnostic Aphasia Examination(3rd ed.)
Harold Goodglass, Edith Kaplan, and Sandra Weintraub. 2001.BDAE: The Boston Diagnostic Aphasia Examination(3rd ed.). Lippincott Williams & Wilkins, Philadelphia, PA
2001
-
[39]
Goodman and Julia Tréhu
Ellen P. Goodman and Julia Tréhu. 2022. AI Audit-Washing and Accountability. Available at: https://www.gmfus.org/news/ai-audit- washing-and-accountability. [Accessed 15-11-2024]
2022
-
[40]
J Ross Graham, Shelialah Pereira, and Robert Teasell. 2011. Aphasia and return to work in younger stroke survivors.Aphasiology25, 8 (2011), 952–960
2011
-
[41]
Ziwei Gu, Jing Nathan Yan, and Jeffrey M Rzeszotarski. 2021. Understanding user sensemaking in machine learning fairness assessment systems. InProceedings of the Web Conference 2021. 658–668
2021
-
[42]
Camille Harris, Chijioke Mgbahurike, Neha Kumar, and Diyi Yang. 2024. Modeling gender and dialect bias in automatic speech recognition. InFindings of the Association for Computational Linguistics: EMNLP 2024. 15166–15184
2024
-
[43]
Xiaodong He, Li Deng, and Alex Acero. 2011. Why word error rate is not a good metric for speech recognizer training for the speech translation task?. In2011 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 5632–5635
2011
-
[44]
Annika Heuser, Tyler Kendall, Miguel del Rio, Quinten McNamara, Nishchal Bhandari, Corey Miller, and Migüel Jetté. 2024. Quantifica- tion of stylistic differences in human-and ASR-produced transcripts of African American English. InProceedings of Interspeech 2024. 4538–4542
2024
-
[45]
Takuya Higuchi, Nobutaka Ito, Takuya Yoshioka, and Tomohiro Nakatani. 2016. Robust MVDR beamforming using time-frequency masks for online/offline ASR in noise. In2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, FAccT ’26, June 25–28...
2016
-
[46]
HuggingFace. 2024. Open ASR Leaderboard. https://huggingface.co/spaces/hf-audio/open_asr_leaderboard
2024
-
[47]
Will Hughes, John; Williams. 2023. Introducing Ursa from Speechmatics. https://www.speechmatics.com/ursa
2023
-
[48]
Maliha Jahan, Priyam Mazumdar, Thomas Thebaud, Mark Hasegawa-Johnson, Jesús Villalba, Najim Dehak, and Laureano Moro- Velazquez. 2025. Unveiling performance bias in asr systems: A study on gender, age, accent, and more. InICASSP 2025-2025 IEEE International Conference on Acous...
2025
-
[49]
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation.Comput. Surveys55, 12 (2023), 1–38
2023
-
[50]
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of Hallucination in Natural Language Generation.Comput. Surveys55, 12 (March 2023), 1–38. doi:10.1145/3571730
2023 doi
-
[51]
Alexander Johnson, Kevin Everson, Vijay Ravi, Anissa Gladney, Mari Ostendorf, and Abeer Alwan. 2022. Automatic Dialect Density Estimation for African American english. InProceedings of Interspeech 2022. 1283–1287
2022
-
[52]
Aura Kagan. 1998. Supported conversation for adults with aphasia: Methods and resources for training conversation partners. Aphasiology12, 9 (1998), 816–830
1998
-
[53]
Mayank Kaushik, Matthew Trinkle, and Ahmad Hashemi-Sakhtsari. 2010. Automatic detection and removal of disfluencies from spontaneous speech. InProceedings of the Australasian International Conference on Speech Science and Technology (SST), Vol. 70. 109
2010
-
[54]
Finn Kensing and Jeanette Blomberg. 1998. Participatory design: Issues and concerns.Computer supported cooperative work (CSCW)7 (1998), 167–185
1998
-
[55]
David S Knopman. 2011. Regional cerebral dysfunction: higher mental functions. 2270–2274 pages
2011
-
[56]
Allison Koenecke, Anna Seo Gyeong Choi, Katelyn X Mei, Hilke Schellmann, and Mona Sloane. 2024. Careless whisper: Speech-to-text hallucination harms. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency. 1672–1681
2024
-
[57]
Allison Koenecke, Andrew Nam, Emily Lake, Joe Nudell, Minnie Quartey, Zion Mengesha, Connor Toups, John R Rickford, Dan Jurafsky, and Sharad Goel. 2020. Racial disparities in automated speech recognition.Proceedings of the National Academy of Sciences 117, 14 (2020), 7684–7689
2020
-
[58]
Allison Koenecke, John-Jose Nunez, Anaïs Rameau, and Irene Y. Chen. 2025. Perspective: Listening to Users when Auditing Medical AI Scribes.Machine Learning for Health(2025)
2025
-
[59]
2025.ACM TechBrief: Automated Speech Recognition
Allison Koenecke, Niranjan Sivakumar, Jingjin Li, and Shaomei Wu. 2025.ACM TechBrief: Automated Speech Recognition. ACM. doi:10.1145/3779316
2025 doi
-
[60]
Korbinian Kuhn, Verena Kersken, Benedikt Reuter, Niklas Egger, and Gottfried Zimmermann. 2024. Measuring the accuracy of automatic speech recognition solutions.ACM Transactions on Accessible Computing16, 4 (2024), 1–23
2024
-
[61]
Petri Laukka, Clas Linnman, Fredrik Åhs, Anna Pissiota, Örjan Frans, Vanda Faria, Åsa Michelgård, Lieuwe Appel, Mats Fredrikson, and Tomas Furmark. 2008. In a nervous voice: Acoustic analysis and perception of anxiety in social phobics’ speech.Journal of Nonverbal Behavior32 (...
2008
-
[62]
Duc Le, Keli Licata, and Emily Mower Provost. 2018. Automatic quantitative analysis of spontaneous aphasic speech.Speech Communication100 (2018), 1–12
2018
-
[63]
2024.Aphasia
H Le and MY Lui. 2024.Aphasia. StatPearls Publishing, Treasure Island (FL). https://www.ncbi.nlm.nih.gov/books/NBK559315/ [Updated 2023 Mar 27]
2024
-
[64]
Colin Lea, Zifang Huang, Jaya Narain, Lauren Tooley, Dianna Yee, Dung Tien Tran, Panayiotis Georgiou, Jeffrey P Bigham, and Leah Findlater. 2023. From user perceptions to technical improvement: Enabling people who stutter to better use speech recognition. In Proceedings of the...
2023
-
[65]
Jaime B Lee, Laura E Kinsey, and Leora R Cherney. 2024. Typing versus handwriting: A preliminary investigation of modality effects in the writing output of people with aphasia.American Journal of Speech-Language Pathology33, 6S (2024), 3422–3430
2024
-
[66]
Min Kyung Lee, Nina Grgić-Hlača, Michael Carl Tschantz, Reuben Binns, Adrian Weller, Michelle Carney, and Kori Inkpen. 2020. Human-centered approaches to fair and responsible AI. InExtended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems. 1–8
2020
-
[67]
Patrick Loeber. 2024. Universal-2 vs OpenAI’s Whisper: Comparing Speech-to-Text models in real-world use cases. https://www. assemblyai.com/blog/comparing-universal-2-and-openai-whisper/ AssemblyAI
2024
-
[68]
Hidalgo Lopez, Shelly Sandeep, MaKayla Wright, Grace M
Julio C. Hidalgo Lopez, Shelly Sandeep, MaKayla Wright, Grace M. Wandell, and Anthony B. Law. 2023. Quantifying and Improving the Performance of Speech Recognition Systems on Dysphonic Speech.Otolaryngology–Head and Neck Surgery168, 5 (Jan. 2023), 1130–1138
2023
-
[69]
2021.Tools for Analyzing Talk Part 1: The CHAT Transcription Format
Brian MacWhinney. 2021.Tools for Analyzing Talk Part 1: The CHAT Transcription Format. Technical Report. Carnegie Mellon University, Pittsburgh, PA. doi:10.21415/3mhn-0z89
2021 doi
-
[70]
MacWhinney, D
B. MacWhinney, D. Fromm, M. Forbes, and A. Holland. 2011. AphasiaBank: Methods for studying discourse.Aphasiology25 (2011), 1286–1307. Addressing Auditing Pitfalls in ASR Technologies FAccT ’26, June 25–28, 2026, Montreal, QC, Canada
2011
-
[71]
Brian MacWhinney, Davida Fromm, Margaret Forbes, and Audrey Holland. 2011. AphasiaBank: Methods for studying discourse. Aphasiology25, 11 (2011), 1286–1307
2011
-
[72]
Joshua L Martin and Kelly Elizabeth Wright. 2023. Bias in automatic speech recognition: The case of African American language. Applied Linguistics44, 4 (2023), 613–630
2023
-
[73]
Brian McFee, Colin Raffel, Dawen Liang, Daniel PW Ellis, Matt McVicar, Eric Battenberg, Oriol Nieto, et al. 2024. librosa: Audio and Music Signal Analysis in Python. https://librosa.org/ Version 0.10.2
2024
-
[74]
I don’t think these devices are very culturally sensitive
Zion Mengesha, Courtney Heldreth, Michal Lahav, Juliana Sublewski, and Elyse Tuennerman. 2021. “I don’t think these devices are very culturally sensitive. ”—Impact of automated speech recognition errors on African Americans.Frontiers in Artificial Intelligence4 (2021), 725911
2021
-
[75]
Jacob Metcalf, Emanuel Moss, Elizabeth Anne Watkins, Ranjit Singh, and Madeleine Clare Elish. 2021. Algorithmic impact assessments and accountability: The co-construction of impacts. InProceedings of the 2021 ACM conference on fairness, accountability, and transparency. 735–746
2021
-
[76]
George Miller. 1955. Note on the bias of information estimates. InProceedings of a conference on the estimation of information flow. 95–100
1955
-
[77]
Masato Mimura, Shinsuke Sakai, and Tatsuya Kawahara. 2021. An end-to-end model from speech to clean transcript for parliamentary meetings. In2021 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC). IEEE, 465–470
2021
-
[78]
Maryam Sadat Mirzaei, Kourosh Meshgi, and Tatsuya Kawahara. 2018. Exploiting automatic speech recognition errors to enhance partial and synchronized caption for facilitating second language listening.Computer Speech & Language49 (2018), 17–36
2018
-
[79]
Meredith Moore, Hemanth Venkateswara, and Sethuraman Panchanathan. 2018. Whistle-blowing ASRs: Evaluating the Need for More Inclusive Speech Recognition Systems. InProceedings of Interspeech 2018. 466–470
2018
-
[80]
Andrew Cameron Morris, Viktoria Maier, and Phil D Green. 2004. From WER and RIL to MER and WIL: improved evaluation measures for connected speech recognition.. InProceedings of Interspeech 2004. 2765–2768
2004
-
[81]
Dena Mujtaba, Nihar Mahapatra, Megan Arney, J Yaruss, Hope Gerlach-Houck, Caryn Herring, and Jia Bin. 2024. Lost in transcription: Identifying and quantifying the accuracy biases of automatic speech recognition systems against disfluent speech. InProceedings of the 2024 Confer...
2024
-
[82]
Arun Narayanan and DeLiang Wang. 2013. The role of binary mask patterns in automatic speech recognition in background noise. The Journal of the Acoustical Society of America133, 5 (2013), 3083–3093
2013
-
[83]
Brandon Nguy, Yina M Quique, Robert Cavanaugh, and William S Evans. 2022. Representation in aphasia research: An examination of US treatment studies published between 2009 and 2019.American Journal of Speech-Language Pathology31, 3 (2022), 1424–1430
2022
-
[84]
Naizeth Núñez Macías, Martina Hielscher-Fastabend, and Hendrik Buschmeier. 2023. Use and acceptance of voice assistants among people with aphasia in Germany.Frontiers in Communication8 (2023), 1176475
2023
-
[85]
Orestis Papakyriakopoulos, Anna Seo Gyeong Choi, William Thong, Dora Zhao, Jerone Andrews, Rebecca Bourke, Alice Xiang, and Allison Koenecke. 2023. Augmented datasheets for speech datasets and ethical decision-making. InProceedings of the 2023 ACM Conference on Fairness, Accou...
2023
-
[86]
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. InProceedings of the 40th International Conference on Machine Learning. PMLR, 28492–28518
2023
-
[87]
Inioluwa Deborah Raji, I Elizabeth Kumar, Aaron Horowitz, and Andrew Selbst. 2022. The fallacy of AI functionality. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency. 959–972
2022
-
[88]
Inioluwa Deborah Raji, Andrew Smart, Rebecca N White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes. 2020. Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing. InProceeding...
2020
-
[89]
matrix design
Jeremy A Rassen, Daniel H Solomon, Robert J Glynn, and Sebastian Schneeweiss. 2011. Simultaneously assessing intended and unintended treatment effects of multiple treatment options: a pragmatic “matrix design”.Pharmacoepidemiology and drug safety20, 7 (2011), 675–683
2011
-
[90]
Vikas Raunak, Arul Menezes, and Marcin Junczys-Dowmunt. 2021. The Curious Case of Hallucinations in Neural Machine Translation. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 1172–1183
2021
-
[91]
Christophe Ris and Stephane Dupont. 2001. Assessing local noise level estimation methods: Application to noise robust ASR.Speech communication34, 1-2 (2001), 141–158
2001
-
[92]
James Robert and Marc Webbie. 2018. Pydub. Available at: http://pydub.com/
2018
-
[93]
Ana Rodrigues, Rita Santos, Jorge Abreu, Pedro Beça, Pedro Almeida, and Sílvia Fernandes. 2019. Analyzing the performance of ASR systems: The effects of noise, distance to the device, age and gender. InProceedings of the XX International Conference on Human Computer Interactio...
2019
-
[94]
Amrit Romana, Minxue Niu, Matthew Perez, and Emily Mower Provost. 2024. FluencyBank Timestamped: An Updated Data Set for Disfluency Detection and Automatic Intended Speech Recognition.Journal of Speech, Language, and Hearing Research67, 11 (2024), 1–13
2024
-
[95]
Shinimol Salim, Syed Shahnawazuddin, and Waquar Ahmad. 2023. Automatic speaker verification system for dysarthric speakers using prosodic features and out-of-domain data augmentation.Applied Acoustics210 (2023), 109412
2023
-
[96]
Shakeel A Sheikh, Md Sahidullah, Fabrice Hirsch, and Slim Ouni. 2022. Machine learning for stuttering identification: Review, challenges and future directions.Neurocomputing514 (2022), 385–402
2022
-
[97]
Matthew A Siegler and Richard M Stern. 1995. On the effects of speech rate in large vocabulary speech recognition systems. In1995 international conference on acoustics, speech, and signal processing, Vol. 1. IEEE, 612–615
1995
-
[98]
Mona Sloane and Emanuel Moss. 2023. Assessing the Assessment: Comparing Algorithmic Impact Assessments and AI Audits.SSRN Electronic Journal(2023). doi:10.2139/ssrn.4486259
2023 doi
-
[99]
Mona Sloane, Hilke Schellmann, Katelyn Xiaoying Mei, Anna Seo Gyeong Choi, and Allison Koenecke. 2026. The case for stakeholder- driven AI auditing in automatic speech recognition.Nature Machine Intelligence(2026), 1–2
2026
-
[100]
Nathalie A Smuha. 2021. Beyond the individual: governing AI’s societal harm.Internet Policy Review10, 3 (2021)
2021
-
[101]
Joseph Paul Stemberger and Barbara May Bernhardt. 2020. Phonetic transcription for speech-language pathology in the 21st century. Folia Phoniatrica et Logopaedica72, 2 (2020), 75–83
2020
-
[102]
Silero Team. 2021. Silero vad: pre-trained enterprise-grade voice activity detector (vad), number detector and language classifier. Retrieved March31 (2021), 2023
2021
-
[103]
Suramya Tomar. 2006. Converting video formats with FFmpeg.Linux journal2006, 146 (2006), 10
2006
-
[104]
Ravichander Vipperla, Steve Renals, and Joe Frankel. 2010. Ageing voices: The effect of changes in voice parameters on ASR performance. EURASIP Journal on Audio, Speech, and Music Processing2010 (2010), 1–10
2010
-
[105]
Laurin Wagner, Bernhard Thallinger, and Mario Zusag. 2024. CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions. InProceedings of Interspeech 2024. 1265–1269
2024
-
[106]
Bejamin Wallace-Wells. 2024. John Fetterman’s War. Available at https://www.newyorker.com/magazine/2024/07/01/john-fettermans- war (2024/09/23)
2024
-
[107]
Angelina Wang, Aaron Hertzmann, and Olga Russakovsky. 2024. Benchmark suites instead of leaderboards for evaluating AI fairness. Patterns5, 11 (2024)
2024
-
[108]
Ye-Yi Wang, Alex Acero, and Ciprian Chelba. 2003. Is word error rate a good indicator for spoken language understanding accuracy. In 2003 IEEE workshop on automatic speech recognition and understanding (IEEE Cat. No. 03EX721). IEEE, 577–582
2003
-
[109]
Alicia Beckford Wassink, Cady Gansen, and Isabel Bartholomew. 2022. Uneven success: automatic speech recognition and ethnicity- related dialects.Speech Communication140 (2022), 50–70
2022
-
[110]
Lauren Werner, Gaojian Huang, and Brandon J Pitts. 2019. Automated speech recognition systems and older adults: a literature review and synthesis. InProceedings of the Human Factors and Ergonomics Society Annual Meeting, Vol. 63. SAGE Publications Sage CA: Los Angeles, CA, 42–46
2019
-
[111]
John Wiseman. 2024. py-webrtcvad: Python interface to the WebRTC Voice Activity Detector. Available at: https://github.com/ wiseman/py-webrtcvad
2024
-
[112]
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. BERTScore: Evaluating Text Generation with BERT. InInternational Conference on Learning Representations
2019
-
[113]
Speaker 0: xxx. Speaker 1:xxx
Robin Zhao, Anna S.G. Choi, Allison Koenecke, and Anaïs Rameau. 2024. Quantification of Automatic Speech Recognition System Performance on d/Deaf and Hard of Hearing Speech.The Laryngoscope135, 1 (Aug. 2024), 191–197. A Appendix We describe our data filtering, standardization,...
2024
-
[2022]
InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency
Exploring how machine learning practitioners (try to) use fairness toolkits. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency. 473–484
2022
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.