REVIEW 3 major objections 4 minor 32 references
Gender Representation in French Broadcast Corpora and Its Impact on ASR Performance
T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A French broadcast-trained ASR system produces significantly higher word error rates for women than for men, with the gap driven by occasional speakers and spontaneous speech.
desk verdict Useful role-based descriptive study of gender imbalance in French broadcast corpora, but the causal claim about ASR performance is not supported by the observational analysis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key analytical device is the speaker-role categorization: a speaker is an Anchor when he or she exceeds 1% of a show's speech turns and 1% of its total speech time, and Punctual when below both thresholds. Crossing this role with gender isolates who carries the performance gap. The paper also uses Wilcoxon rank-sum tests on episode-level WER values, and explains the pattern through speaker adaptation in the decoding pipeline: fMLLR-adapted features give speakers with more data a stronger adaptation, so Anchors recover from acoustic mismatch while Punctual speakers, who are disproportionately women, do not.
What would settle it
Decode the same evaluation set without the speaker-adaptation step and compare the gender WER gaps: the paper's explanation predicts the Punctual-speaker gap should shrink or disappear when adaptation data is no longer an advantage. Alternatively, retrain the acoustic model on a gender-balanced subset of the same shows; if the women-versus-men gap among Punctual speakers persists at a similar size, gender imbalance in training data is not the main driver.
Extended reading notes
Core claim
The central empirical claim is that an ASR system trained on roughly 100 hours of gender-unbalanced French broadcast speech exhibits a gender bias in word error rate. Averaged over the evaluation set, women show 42.9% WER against 34.3% for men, with medians of 29% and 25% respectively. The paper then shows this overall gap is carried by specific subpopulations: among Punctual speakers, average WER is 49.04% for women and 38.56% for men; among Anchors, the difference is not significant ($p=0.173$). Similarly, on spontaneous speech women average 61.29% WER versus 46.51% for men, while on prepared speech the gap is not significant. The authors conclude that gender imbalance in training data translates into biased ASR performance, but that speaker role and speech spontaneity are entangled factors that determine where the bias appears.
Load-bearing premise
The significance tests treat each episode-level word error rate as an independent observation even though the same speaker can appear in several episodes; if those values are correlated within a speaker, the reported $p$-values could be too small.
Editorial extensions
If this is right
- If the claim holds, ASR systems trained on existing French broadcast corpora carry a measurable performance penalty for women's speech, so evaluations reporting only aggregate WER will hide a systematic error skew.
- The Anchor versus Punctual distinction predicts that the gender gap will appear mainly in data with many non-professional, occasional speakers rather than in news-reader-style speech.
- The results imply that rebalancing training data or ensuring sufficient adaptation data for female speakers are concrete remedies for reducing gender bias in French ASR.
- Because spontaneous speech amplifies the gap, the bias may grow worse as ASR expands toward conversational and entertainment content without rebalanced training data.
Reading between the lines
- As an extension of the paper's logic, the show-level role definition means the same person could be Anchor in one show and Punctual in another; testing whether stable voice characteristics or show-level factors drive the gap would refine the causal story.
- The paper's own proposed confirmation is to decode without speaker adaptation; if the Punctual-speaker gender gap persists under that condition, the explanation would shift from adaptation-data quantity to the acoustic model itself.
- The same role-based methodology could be applied to other underrepresented groups in broadcast data, such as dialectal, non-native, or older speakers, where unequal show presence may create analogous WER gaps.
- A practical extension is to stratify benchmark reporting by speaker role and gender, so future French ASR work would routinely expose rather than average away such performance disparities.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies gender representation in four widely used French broadcast corpora (ESTER1, ESTER2, ETAPE, REPERE) and examines whether the gender imbalance in the training data is associated with the performance of an ASR system. The authors first report that women account for 33.16% of speakers and 22.57% of speech time in the training data, and that this under-representation is stronger among Anchor speakers than among Punctual speakers. They then evaluate a Kaldi HMM-DNN ASR system trained on roughly 100 hours of prepared speech, computing episode-level WER on a 70-hour evaluation set stratified by gender, speaker role, and speech type. The headline results are a pooled female/male WER of 42.9% vs 34.3% (p<0.001), a significant gap among Punctual speakers (49.04% vs 38.56%), a significant gap for spontaneous speech (61.29% vs 46.51%), and no significant gap among Anchors or for prepared speech. The abstract and conclusion interpret these results causally, stating that the disparity in available data causes performance to decrease on women.
Significance. If the empirical findings hold, this is a useful contribution to the growing literature on fairness in speech and language technology. The descriptive sex-ratio results in Tables 3-5 are straightforward and credible, and they corroborate external reports on women's under-representation in French media. The introduction of the Anchor/Punctual role dichotomy is a simple and reproducible analytic choice that adds a genuinely useful stratification for interpreting ASR performance differences. The paper also makes an honest attempt to discuss speaker adaptation as a potential mechanism. However, the inferential and causal claims are currently stronger than the evidence supports: the statistical tests ignore within-speaker correlation, the pooled gender comparison is not adjusted for show composition despite evidence of between-show heterogeneity in Figure 1, and the proposed causal mechanism is left for future work. With clustered or mixed-effects analysis, show-adjusted estimates, and appropriately hedged causal language, the paper would be a solid empirical contribution.
major comments (3)
- [Section 3.3.2 and Section 4.2.1] The Wilcoxon rank-sum tests in Section 4.2 treat each episode-level WER value as independent, but Section 3.3.2 states that a given speaker can contribute several WER values, one per show occurrence. Under the null, WERs from the same speaker are likely correlated, which makes the reported p-values (e.g., p<0.001 for the pooled female/male comparison) anti-conservative. Because the repeated-speaker structure is especially pronounced for Anchor speakers and frequent Punctual speakers, this could change the role-gender conclusions. Please add a cluster-robust test, a permutation test that resamples speakers, or a mixed-effects model with speaker (and show) as random effects, and report how the significance of the key comparisons changes.
- [Section 4.2.1 and Figure 1] The paper itself reports that within-show gender differences are inconsistent across shows, with some shows favoring women. This is exactly the signature of a show-composition effect: women may be over-represented in harder spontaneous shows, so the pooled 42.9%-vs-34.3% gap need not reflect any systematic female-speech disadvantage. The subsequent analyses by role (Section 4.2.3) and speech type (Section 4.2.4) still pool across shows, so they do not remove this confound. Please provide show-adjusted estimates, for example by including show as a fixed or random effect or by computing within-show gender contrasts, and state whether the gender effect survives such adjustment.
- [Section 5] The causal interpretation in the abstract and conclusion ('causes performance to decrease') is not supported by the experimental design. Section 5 proposes that less fMLLR adaptation data for women causes the larger gap for Punctual speakers, but it explicitly defers the confirming experiment ('A way to confirm our hypothesis would be to reproduce our analysis on WER values obtained without using speaker adapted features'), and no gender-balanced training control is reported. As it stands, the study demonstrates associations, not causation. Please either report the decoding-without-adaptation result and/or a gender-balanced training condition, or revise the abstract, Section 5, and conclusion to use association language such as 'is associated with' or 'is correlated with'.
minor comments (4)
- [Abstract and Conclusion] The abstract's phrase 'causes performance to decrease on women' and the conclusion's wording 'produces gender bias performance' overstate what the observational comparison can establish; even if the statistical issues are resolved, the causal claim should be softened unless a controlled training experiment is reported.
- [Section 2.2] The final sentence of Section 2.2 ends with a bare citation '[13]' after 'an ASR system trained on these data'; this citation seems intended to point to the system described in Section 3.3.1, but as written it is not integrated into a grammatical sentence and should be moved or rephrased.
- [Section 4.2.4] For prepared speech, the text reports p=0.005 immediately after saying 'WER scores are similar between men and women.' Since Section 3.3.2 fixes alpha=0.001, this p-value is not significant at the stated threshold, and the sentence should explicitly say that the difference is not significant at alpha=0.001 to avoid confusion.
- [Conclusion] The conclusion states a 'WER increase of 24% for women compared to men,' but the reported pooled values 42.9% and 34.3% correspond to a relative increase of about 25.1% (and 8.6 percentage points absolute); please reconcile the numbers.
Circularity Check
No significant circularity: the reported WER gap is measured, not fitted, and the self-citation points to a concrete ASR system rather than a load-bearing uniqueness claim.
full rationale
The paper's central empirical claim is an observed performance difference between male and female speakers on a fixed ASR system. The WER values are computed directly from decoder output against reference transcriptions, with no parameter fitted to the gender outcome; the role threshold (1% of turns and speech time) is defined a priori from representation statistics and is independent of WER. The ASR system is cited from prior work [13] sharing an author, but this is a pointer to a concrete Kaldi HMM-DNN system whose outputs are inputs to the analysis, not a premise derived to match the gender result. The causal phrasing in the abstract may outrun the observational, show-confounded design, but that is an identifiability/correctness concern, not circularity: no equation or definition reduces the claimed result to its inputs. The paper itself acknowledges that across shows there is no consistent gender trend (Section 4.2.1), which further shows the headline gap was not constructed to match a preselected conclusion. No circular step can be exhibited, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (1)
- Role threshold for Anchor vs Punctual speakers =
1% of speech turns and 1% of total speech time per show
assumptions (3)
- domain assumption Gender labels in corpus metadata are accurate
- domain assumption Wilcoxon rank-sum test assumption of independent observations holds across episode-level WER values
- domain assumption The four ESTER, ETAPE, and REPERE corpora are representative of French broadcast used for ASR training
Cite this review
Pith. "Pith review of Gender Representation in French Broadcast Corpora and Its Impact on ASR Performance." pith.science (2026). https://pith.science/paper/LB5MAGSS
@misc{pith2026190808717,
author = {Pith},
title = {Pith review of: Gender Representation in French Broadcast Corpora and Its Impact on ASR Performance},
year = {2026},
howpublished = {\url{https://pith.science/paper/LB5MAGSS}},
note = {Machine review of arXiv:1908.08717}
}
read the original abstract
This paper analyzes the gender representation in four major corpora of French broadcast. These corpora being widely used within the speech processing community, they are a primary material for training automatic speech recognition (ASR) systems. As gender bias has been highlighted in numerous natural language processing (NLP) applications, we study the impact of the gender imbalance in TV and radio broadcast on the performance of an ASR system. This analysis shows that women are under-represented in our data in terms of speakers and speech turns. We introduce the notion of speaker role to refine our analysis and find that women are even fewer within the Anchor category corresponding to prominent speakers. The disparity of available data for both gender causes performance to decrease on women. However this global trend can be counterbalanced for speaker who are used to speak in the media when sufficient amount of data is available.
Figures
Reference graph
Works this paper leans on
-
[1]
Ahmed Abbasi, Jingjing Li, Gari Clifford, and Herman Taylor. 2018. Make “Fair- ness by Design" Part of Machine Learning. Harvard Business Review (2018). https://hbr.org/2018/08/make-fairness-by-design-part-of-machine-learning
work page 2018
-
[2]
Martine Adda-Decker and Lori Lamel. 2005. Do speech recognizers prefer female speakers?. In Proceedings of the 9th European Conference on Speech Communication and Technology (INTERSPEECH 2005) . 2205–2208
work page 2005
-
[3]
Zou, Venkatesh Saligrama, and Adam T
Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam T. Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Proceedings of the 30 th Conference on Neural Information Processing Systems (NIPS 2016) . 4349–4357
work page 2016
-
[4]
Danah Boyd and Kate Crawford. 2012. Critical Questions for Big Data: Provo- cations for a Cultural, Technological, and Scholarly Phenomenon. Information, communication & society 15, 5 (2012), 662–679
work page 2012
-
[5]
Joy Buolamwini and Timnit Gebru. 2018. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Proceedings of the Conference on Fairness, Accountability and Transparency (ACM FAT 2018) . 77–91
work page 2018
-
[6]
Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science 356, 6334 (2017), 183–186
work page 2017
-
[7]
Sainathand Yonghui Wu, Rohit Prabhavalkar, Patricl Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J
Chung-Cheng Chiu, Tara N. Sainathand Yonghui Wu, Rohit Prabhavalkar, Patricl Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J. Weiss, Kanishka Rao, Ekaterina Gonina, Navdeep Jaitly, Bo Li, Jan Chorowski, and Michiel Bacchiani. 2018. State- of-the-art speech recognition with sequence-to-sequence models. In Proceedings of the IEEE International Conference on Acou...
work page 2018
-
[8]
Christopher Cieri, David Miller, and Kevin Walker. 2004. The Fisher Corpus: a Resource for the Next Generations of Speech-to-Text.. In LREC, Vol. 4. 69–71
work page 2004
Show all 32 references
-
[9]
Kate Crawford. 2017. The Trouble with Bias. Retrieved July 8, 2019 from https://www.youtube.com/watch?v=fMym_BKWQzk NIPS 2017 Keynote
2017
-
[10]
CSA. 2018. La Représentation des Femmes à la Télévision et à la Radio. Rapport d’Exercice 2017. Retrieved July 8, 2019 from https://en.calameo.com/read/ 00453987548c2c813939e?page=1
2018
-
[11]
Martine De Calmès and Guy Pérennou. 1998. BDLEX: a lexicon for spoken and written French. In Proceedings of the 1 st International Conference on Language Resources and Evaluation (LREC 1998) . 1129–1136
1998
-
[12]
David Doukhan, Jean Carrive, Félicien Vallet, Anthony Larcher, and Sylvain Meignier. 2018. An open-source speaker gender detection framework for mon- itoring gender equality. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP...
2018
-
[13]
Zied Elloumi, Laurent Besacier, Olivier Galibert, Juliette Kahn, and Benjamin Lecouteux. 2018. ASR Performance Prediction on Unseen Broadcast Programs using Convolutional Neural Networks. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Proce...
2018
-
[14]
Sylvain Galliano, Edouard Geoffrois, Djamel Mostefa, Khalid Choukri, Jean- François Bonastre, and Guillaume Gravier. 2005. The ESTER phase II evaluation campaign for the rich transcription of French broadcast news. InProceedings of the 9th European Conference on Speech Communi...
2005
-
[15]
Sylvain Galliano, Guillaume Gravier, and Laura Chaubard. [n. d.]. The ESTER 2 evaluation campaign for the rich transcription of French radio broadcasts. In Pro- ceedings of the 10 th Annual Conference of the International Speech Communication Association (INTERSPEECH 2009) . 2583–2586
2009
-
[16]
Garofolo, Lori F
John S. Garofolo, Lori F. Lamel, William M. Fisher, Jonathan G. Fiscus, and David S. Pallett. 1993. DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1. NASA STI/Recon technical report n 93 (1993)
1993
-
[17]
Aude Giraudel, Matthieu Carré, Valérie Mapelli, Juliette Kahn, Olivier Galibert, and Ludovic Quintard. 2012. The REPERE Corpus: a multimodal corpus for person recognition. In Proceedings of the 8th International Conference on Language Resources and Evaluation (LREC 2012) . 1102–1107
2012
-
[18]
John J Godfrey, Edward C Holliman, and Jane McDaniel. 1992. SWITCHBOARD: Telephone Speech Corpus for Research and Development. In Proceedings of the 1992 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP92), Vol. 1. IEEE, 517–520. Gender Represen...
1992
-
[19]
Sharon Goldwater, Dan Jurafsky, and Christopher D. Manning. 2010. Which words are hard to recognize? Prosodic, lexical, and disfluency factors that increase speech recognition error rates. Speech Communication 52, 3 (2010), 181–200
2010
-
[20]
Guillame Gravier, Gilles Adda, Niklas Paulson, Matthieu Carré, Aude Giraudel, and Olivier Galibert. 2012. The ETAPE corpus for the evaluation of speech- based TV content processing in the French language. In Proceedings of the 8 th International Conference on Language Resource...
2012
-
[21]
Christopher J Leggetter and Philip C Woodland. 1995. Maximum likelihood linear regression for speaker adaptation of continuous density hidden Markov models. Computer speech & language 9, 2 (1995), 171–185
1995
-
[22]
Macharia, L
S. Macharia, L. Ndangam, M. Saboor, E. Franke, S. Parr, and E. Opoku. 2015. Who Makes the News. Global Media Monitoring Project 2015. Retrieved July 8, 2019 from http://cdn.agilitycms.com/who-makes-the-news/Imported/reports_2015/ global/gmmp_global_report_en.pdf
2015
-
[23]
Mann and Donald R
Henry B. Mann and Donald R. Whitney. 1947. On a Test of Whether one of Two Random Variables is Stochastically Larger than the Other. The Annals of Mathematical Statistics (1947), 50–60
1947
-
[24]
Smriti Parsheera. 2018. A Gendered Perspective on Artificial Intelligence. In 2018 ITU Kaleidoscope: Machine Learning for a 5G Future (ITU K) . IEEE, 1–7
2018
-
[25]
Daniel Povey, Arnad Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely. 2011. The Kaldi Speech Recogni- tion Toolkit. (2011)
2011
-
[26]
Marcelo O. R. Prates, Pedro H. C. Avelar, and Luís C Lamb. 2018. Assessing Gender Bias in Machine Translation - A Case Study with Google Translate. (2018). arXiv:1809.02208 http://arxiv.org/abs/1809.02208
2018 arXiv
-
[27]
Andreas Stolcke. 2002. SRILM - An Extensible Language Modeling Toolkit. In Proceedings of the 7 th International Conference on Spoken Language Processing (ICSLP2002). 901–904
2002
-
[28]
Belding, Kai-Wei Chang, and William Yang Wang
Tony Sun, Andrew Gaut, Shirlyn Tang, Yuxin Huang, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth M. Belding, Kai-Wei Chang, and William Yang Wang. 2019. Mitigating Gender Bias in Natural Language Processing: Literature Review. (2019). arXiv:1906.08976 http://arxiv.org/abs/1906.08976
2019 arXiv
-
[29]
Rachael Tatman. 2017. Gender and Dialect Bias in YouTube’s Automatic Captions. In Proceedings of the First ACL Workshop on Ethics in Natural Language Processing . 53–59
2017
-
[30]
Rachael Tatman and Conner Kasten. 2017. Effects of Talker Dialect, Gender & Race on Accuracy of Bing Speech and YouTube Automatic Captions. In Proceedings of the 19th Annual Conference of the International Speech Communication Association (INTERSPEECH 2017). 934–938
2017
-
[31]
Eva Vanmassenhove, Christian Hardmeier, and Andy Way. 2018. Getting Gender Right in Neural Machine Translation. InProceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP 2018) . 3003–3008
2018
-
[32]
Fei-Yue Wang. 2011. AI’s Hall of Fame. IEEE Intelligent Systems 26, 4 (2011), 5–15. https://ieeexplore.ieee.org/document/5968105
2011
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.