{"id":"355e26f3-9e26-480f-a164-30e7bb638df2","arxiv_id":"2411.09431","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Five Whisper model sizes transcribe Dutch female speech more accurately than male speech in most tested conditions, while the measured size and direction of the gender gap varies by dataset and show type.","lead":"This paper tests five Whisper speech recognition models on Dutch speech and finds that word error rates differ between male and female speakers, with women's speech usually transcribed more accurately. It matters for automatic subtitling and broadcast media, where unequal accuracy can mean unequal access for different groups.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's 'bias across all model sizes' is contradicted by the paper's own significance tables; only a subset of model-size and condition comparisons are statistically significant.","rationale":"The reader's verdict was CONDITIONAL, and the reader's rationale already noted that 'not all model sizes show significant bias' and that the abstract overstates the results. My stress-test confirms this as the most load-bearing concern: the paper's headline claim is directly undermined by its own significance tables. I do not disagree with the reader's identified weakest assumption about gender-label accuracy, but I find the abstract overstatement more central and more decisively checkable: it is an internal inconsistency that can be settled by recomputing the reported statistics, whereas label noise is an acknowledged, plausible but less directly testable threat. The concrete test is deliberately narrow: recompute all gender comparisons with correction and count significant models per dataset. If the count is lower than five, the claim 'across all model sizes' is false. This does not change the verdict, because the paper can be repaired by revising the headline and reporting exact statistics, and the evidence for bias in at least the larger models appears credible. The reader's CONDITIONAL verdict stands.","tokens_in":1021,"tokens_out":852,"duration_ms":44486,"concrete_test":"Recompute, from the per-speaker WER data or from the paper's reported counts, the exact p-values and effect-size confidence intervals for every gender comparison in Tables 1 and 2, applying a Benjamini-Hochberg correction over all model x condition x dataset comparisons. Then count, per dataset, how many of the five Whisper models show a significant corrected gender difference. If, as the tables currently suggest, fewer than five models in either dataset reach significance, the abstract's 'across all model sizes' claim fails and must be narrowed to the specific model sizes and conditions that are actually significant.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the abstract, 'substantial disparities in WER among gender groups across all model sizes, with bias identified through statistical testing,' is not supported by the paper's own reported results. In Table 1 (Common Voice), only the medium and large models carry significance markers (p < .01); tiny, base, and small show no significant gender difference. In Table 2 (NPO), significance is scattered and sparse: Tiny is significant only for Eloquent and All; Base for Eloquent and All; Small only for Radio; Medium has no significance marker in any column; Large only for Eloquent and Radio. The 'All' column for Small, Medium, and Large is not significant. Thus, under the reported analyses, gender bias is not found 'across all model sizes' in either dataset. The headline appears to conflate directional WER differences with statistically identified bias, or to generalize from the significant subset. This is an internal inconsistency rather than a disagreement with external consensus: the tables as printed do not support the abstract's sweep. In addition, no multiple-testing correction is reported and exact test statistics are absent, so it is unclear how many of the sparse significant p-values would survive correction. The core finding that the largest models show bias may be robust, but the strongest claim as stated is not.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates five Whisper model sizes (tiny, base, small, medium, large) on Dutch speech from the Common Voice dataset and a Dutch public broadcasting (NPO) dataset. It compares word error rate, character error rate, and BERT-based semantic similarity across binary gender groups, using weighted per-speaker metrics and statistical tests with assumption checks. It also proposes a WER Parity fairness metric with a threshold chosen after discussions with the NPO, and interprets results through a moral framework for quality-of-service harms. The authors report that female speech is generally recognized better than male speech, that larger models tend to be more accurate but also show more significant WER differences, and that the medium model offers a practical performance/efficiency trade-off.","tokens_in":11434,"tokens_out":7079,"duration_ms":67640,"significance":"If the findings are reported accurately, the paper provides a useful empirical contribution to the growing literature on demographic bias in ASR systems, and it is one of the few studies targeting Dutch data in a real-world subtitling context. The methodology is generally careful: per-speaker aggregation reduces statistical dependence, the authors check normality and homogeneity assumptions and use non-parametric alternatives, and the fairness discussion is explicitly grounded in a moral framework. The proposed WER Parity metric is transparent and easy to apply. However, the central claim as stated in the abstract overstates the evidence in the paper's own tables, and the absence of multiple-testing correction and exact test statistics weakens the statistical conclusions. These issues are fixable within the scope of a revision.","major_comments":[{"comment":"The abstract's claim of \"substantial disparities in WER among gender groups across all model sizes, with bias identified through statistical testing\" is not supported by the reported significance tests. In Table 1 only the medium and large models carry significance markers on Common Voice; tiny, base, and small do not. In Table 2 the 'All' column is significant only for tiny and base, the medium model has no significant marker in any column, and several per-category cells are not significant. The Section 4 bullet \"Significant biases were observed across gender attributes, with larger models exhibiting more pronounced disparities\" also overgeneralizes. Please revise the abstract, Section 4, and conclusion to state the actual pattern, for example that statistically significant gender differences appear for the larger Common Voice models and for a sparse subset of NPO conditions.","section":"Abstract; Section 3.1, Table 1; Section 3.2, Table 2; Section 4"},{"comment":"The manuscript does not report the test statistics, degrees of freedom, or exact p-values for any of the significance tests, and no multiple-testing correction is applied despite the large number of comparisons (five model sizes times multiple speech categories, roughly 25 tests in total). With this many comparisons, several p<0.05 results are expected by chance. Please provide a table of exact p-values or confidence intervals, report which specific test was used in each comparison (t-test, Welch, or Mann-Whitney), and apply a multiple-testing correction or explicitly justify treating each comparison as a separate family. This is necessary to evaluate whether the claim of statistically identified bias survives.","section":"Section 2.4; Tables 1 and 2"},{"comment":"Gender labels for the NPO dataset are inferred from speaker names and contextual clues, and the manuscript acknowledges this risk in Section 4.1 but does not validate the labels. Since the group-level WER comparison is the foundation of the bias claim, a material share of mislabeled or non-binary speakers could change the conclusions. Please add a sensitivity analysis or at least quantify the expected labeling accuracy, and discuss the limitation that only binary gender is considered.","section":"Section 2.2; Section 4.1"},{"comment":"The WER Parity threshold of 25% is presented as 'deemed reasonable after discussions with the NPO,' but no sensitivity analysis is reported, and the metric's output is interpreted as 'unfair and relevant differences' versus 'fair.' Because the paper alternates between fairness judgments derived from this threshold and statistical bias testing, the arbitrariness of the threshold weakens the fairness conclusions even though the threshold does not affect the significance tests. Please show how the fairness verdicts in Table 2 change for a range of epsilon values, and clearly separate threshold-based fairness claims from statistically identified predictive bias.","section":"Section 2.4, Eq. (4); Tables 2 and 3"}],"minor_comments":[{"comment":"Equation (4) and the surrounding text contain formatting artifacts such as 'W ERmale' and the typo 'a fairness metrics'; please correct these to 'WER_male' and 'a fairness metric.'","section":"Section 2.4, Eq. (4)"},{"comment":"The word 'proprose' should be 'propose.'","section":"Section 2.3"},{"comment":"The phrase 'relative gender bias of 0.9% (11.4%)' is ambiguous; please clarify that 0.9% is the absolute percentage-point difference and 11.4% is the relative difference, and use the same convention consistently in Section 3.2.","section":"Section 3.1"},{"comment":"The BERT-based similarity calculation does not specify which BERT checkpoint or language model was used for Dutch (e.g., multilingual BERT, BERTje, or another model); without this detail the BSS results are not reproducible.","section":"Section 2.4, Eq. (3)"},{"comment":"Table 3 reports CER values without significance markers, but the text states that CER 'follows the same trend' as WER; either add significance tests for CER or explicitly describe Table 3 as descriptive only.","section":"Section 3.2; Table 3"},{"comment":"The statement 'the large model demonstrates fairness across all categories' in Section 3.2 is based on the WER Parity threshold, while the same model has statistically significant WER differences in the Eloquent and Radio categories; please clarify the distinction between fairness (threshold-based) and bias (significance-based) in this sentence.","section":"Section 3.2; Section 4"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is a straightforward, mostly careful measurement study: five Whisper sizes, Dutch Common Voice plus a real NPO broadcast corpus, WER/CER/BSS, per-speaker aggregation, assumption checks, and non-parametric tests when needed. What is actually new is the NPO dataset and the explicit WER Parity threshold, plus the moral-framework wrapper around quality-of-service harms. The authors cite Fuckner et al. (2023), which already evaluated Whisper on Dutch, so this is a legitimate extension of an established bias-auditing program rather than a first.\n\nCredit where due: per-speaker aggregation is the right move to avoid non-independence; checking normality and homogeneity and switching to Mann-Whitney/Welch is more care than most ASR bias papers take; and they are transparent that BSS did not work and that hallucinations matter. The NPO corpus, imperfect as they admit, is genuinely useful for subtitle-relevant evaluation.\n\nThe soft spots are real but not fatal. The abstract says “substantial disparities … across all model sizes,” but Table 1 shows significance only for medium and large on Common Voice, and Table 2 has significance in only a subset of NPO conditions; several “All” columns are not significant. The directional WER differences are consistent, but the strongest claim is not supported by their own tests. There is no multiple-testing correction, and exact statistics and confidence intervals are missing, so we do not know how many of the sparse p-values would survive. The 25% WER Parity threshold is arbitrary, though they admit it and it does not feed back into the significance tests. Gender labels in the NPO data are inferred from names; they flag this risk but do not validate it. Code and NPO data are not released, which hurts reproducibility.\n\nThe core finding is plausible and probably robust: larger Whisper models show a statistically detectable female advantage on Dutch speech, at least under some conditions. The paper is not a breakthrough, but it is a solid extension that would benefit from revision to align the abstract with the tables, add multiple-testing-aware reporting, and open access to data and code.\n\nRecommended: send it to peer review. It deserves referee time, with the expectation of a revision rather than rejection.","headline":"Worth reading as a careful Dutch-language ASR bias audit, but the abstract oversells the significance pattern in its own tables.","tokens_in":12022,"tokens_out":1984,"would_cite":false,"duration_ms":19491,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Whisper, a state-of-the-art speech recognition model family, shows statistically significant gender-based word error rate disparities on Dutch speech, generally recognizing female voices more accurately than male voices across model sizes.","keywords":["automatic speech recognition","gender bias","Dutch speech","Whisper","word error rate","fairness metrics","quality-of-service harm","WER parity"],"falsifier":"Recompute the gender WER comparisons on a subset of the Dutch broadcaster's data where each speaker's gender is verified by the speakers themselves rather than inferred from names; if the statistically significant female-over-male gaps for the tiny, base, and medium models shrink to non-significance, the central bias claim is an artifact of label noise.","tokens_in":11007,"feed_emoji":"🎙️","tokens_out":7550,"duration_ms":66997,"temperature":0.7,"pith_summary":"The paper tries to establish that Whisper, a state-of-the-art automatic speech recognition model family, is predictively biased by gender on Dutch speech, and that this bias matters as a quality-of-service harm for automatic subtitling. Evaluating five model sizes on an open corpus of read Dutch sentences and on roughly 37 hours of Dutch broadcast TV and radio, the authors compare word error rate, character error rate, and a BERT-based semantic similarity across male and female speakers. They find substantial and statistically significant word error rate disparities between gender groups, with female speech generally recognized more accurately than male speech, and larger models showing more pronounced bias on read speech. The study also proposes a WER Parity metric, a ratio bound on group word error rates meant to flag unfair quality-of-service gaps, grounded in a moral framework that ties fairness measurement to concrete harms. If correct, the findings imply that Dutch automatic subtitling systems built on Whisper will systematically transcribe male and female speakers with unequal accuracy.","feed_headline":"Whisper transcribes female Dutch speech better than male","feed_subtitle":"Five model sizes, tested on read and broadcast Dutch, show significant word-error-rate gaps by gender.","key_machinery":"The central machinery is a two-part evaluation protocol. First, per-speaker WER scores are computed with a word-count-weighted mean and aggregated by speaker ID, removing the dependence that would arise from multiple clips by the same speaker; significance is then tested with t-tests or ANOVA, switching to Mann-Whitney U or Welch ANOVA when normality or homogeneity assumptions fail. Second, the proposed WER Parity metric compares the larger group WER to the smaller one as a ratio bound, flagging unfairness when $\\max(\\mathrm{WER}_{\\mathrm{male}}, \\mathrm{WER}_{\\mathrm{female}}) / \\min(\\mathrm{WER}_{\\mathrm{male}}, \\mathrm{WER}_{\\mathrm{female}}) \\le 1.25$ is violated. The Whisper models themselves are the object under test, and the text normalizer from the original Whisper paper is used to standardize transcripts before scoring.","core_discovery":"On its own terms, the central claim is that gender bias in Whisper is real, measurable, and consistent in direction on Dutch data: across all five model sizes, per-speaker word error rates differ significantly between men and women, with female speech recognized more accurately in most settings (for example, a relative WER advantage of about 10–11% for the medium and large models on read speech, and much larger relative gaps of 20–28% for the tiny and base models on broadcast data). The bias is established through two-sample t-tests and one-way ANOVA after checking normality and variance assumptions, and it persists across the read-speech and broadcast domains, though not in every speech category. A separate fairness check, the proposed WER Parity metric, finds some of these gaps unfair at a 25% relative threshold, particularly for broadcast radio and eloquent TV speech. The paper also argues that the morally relevant harm is quality-of-service: unequal subtitle accuracy that can distort public perception and exclude groups.","pith_inferences":["Because the broadcaster's gender labels are inferred from names, the true bias could be partially confounded with name-associated differences in speaking style or program type; a label-verified re-analysis would tell whether the female advantage survives.","The finding that larger models show more bias on read speech but less on broadcast speech suggests that bias does not scale monotonically with capacity; testing intermediate checkpoints on matched domain data could reveal whether fine-tuning or domain-specific training is the effective lever.","The 25% WER Parity threshold is a convention chosen with the broadcaster rather than derived from a cost model; linking the threshold to the actual utility loss of mis-subtitled content for deaf and hard-of-hearing viewers would place fairness judgments on firmer normative ground.","The same per-speaker aggregation and parity-ratio protocol can be applied to age, accent, or non-binary gender groups on datasets that carry those attributes, which would show whether the female-over-male pattern is specific to gender or a proxy for other speech characteristics."],"forward_implications":["Dutch broadcasters using Whisper for automatic subtitling should expect systematically lower word error rates for female speakers than for male speakers on eloquent TV and radio content, with the gap varying by model size.","Selecting a Whisper model for Dutch subtitling involves a three-way trade-off: larger models improve overall accuracy but can increase gender bias on read speech, while the medium model offers the best accuracy-to-speed balance and the large model is the only one deemed fair across all broadcast categories.","Statistical significance and fairness are distinct: a statistically significant WER gap can be classified as fair under the proposed 25% WER Parity bound, so deployment decisions should report both.","The BERT-based semantic similarity metric used here was unreliable for short Common Voice sentences, so it should not be used as a standalone quality-of-service measure for Dutch ASR without further refinement."],"supporting_citations":[{"why":"Supplies the Whisper model family and the text normalizer used to prepare transcripts before scoring.","marker":"[31]"},{"why":"Supplies the open read-speech corpus with demographic metadata that forms the first test bed.","marker":"[3]"},{"why":"Supplies the moral framework that connects quality-of-service harms to the choice of fairness metric.","marker":"[39]"},{"why":"Establishes the requirement that bias findings state what harm occurs and to whom, which the paper adopts.","marker":"[6]"},{"why":"Provides the equivalence-test logic on which the WER Parity metric is modeled.","marker":"[24]"},{"why":"Documents hallucination harms in Whisper, corroborating the paper's qualitative observation for short segments.","marker":"[21]"},{"why":"Provides the normality test used to check ANOVA and t-test assumptions before significance testing.","marker":"[34]"}],"fun_headline_variants":["Whisper's Dutch transcriptions favor female speech","Gender bias in Whisper: Dutch female speech transcribed better","Whisper shows consistent WER gap favoring female Dutch speakers","Dutch ASR test: Whisper transcribes women more accurately","Whisper's gender bias: female Dutch speech wins on accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire comparison assumes the male/female labels are correct, but the open-corpus labels are self-reported and the broadcaster labels are inferred from speaker names; if a material share of those labels is wrong, the measured WER gaps and the bias conclusion are not trustworthy.","fun_headline_variants_meta":{"raw":{"variants":["Whisper's Dutch transcriptions favor female speech","Gender bias in Whisper: Dutch female speech transcribed better","Whisper shows consistent WER gap favoring female Dutch speakers","Dutch ASR test: Whisper transcribes women more accurately","Whisper's gender bias: female Dutch speech wins on accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000656,"raw_usage":{"total_tokens":2972,"prompt_tokens":881,"completion_tokens":2091,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":497,"completion_tokens_details":{"reasoning_tokens":2007}},"tokens_in":497,"tokens_out":2091,"duration_ms":13447,"temperature":1.0,"reasoning_tokens":2007,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:38:40.691579+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the gender WER comparisons on a subset of the Dutch broadcaster's data where each speaker's gender is verified by the speakers themselves rather than inferred from names; if the statistically significant female-over-male gaps for the tiny, base, and medium models shrink to non-significance, the central bias claim is an artifact of label noise.","supporting_citations":[{"cited_title":"Social psychological and personality science8(4), 355–362 (2017)","cited_arxiv_id":null,"evidence_quote":"Provides the equivalence-test logic on which the WER Parity metric is modeled."},{"cited_title":"In: The 2024 ACM Conference on Fair- ness, Accountability, and Transparency","cited_arxiv_id":null,"evidence_quote":"Documents hallucination harms in Whisper, corroborating the paper's qualitative observation for short segments."},{"cited_title":"Biometrika 52(3/4), 591–611 (1965)","cited_arxiv_id":null,"evidence_quote":"Provides the normality test used to check ANOVA and t-test assumptions before significance testing."}],"review_version":1}