{"id":"3a6620d1-9163-4dc0-bffb-b5db5a338ee5","arxiv_id":"2506.17525","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A quality audit of Common Voice, FLEURS, and VoxPopuli finds serious micro- and macro-level data defects concentrated in less institutionalized languages, with a case study of Taiwanese Southern Min.","lead":"This paper audits three widely used multilingual speech datasets and shows that several languages contain severe quality problems, such as mismatched scripts, wrong language labels, and very short or silent audio. It argues that dataset builders need sociolinguistic awareness and language planning to avoid hiding failures behind misleading evaluation scores.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'strong positive correlation' between institutionalization and dataset quality is asserted but never measured; it rests on an unvalidated qualitative sample without inter-rater reliability or a defined institutionalization variable.","rationale":"The reader's CONDITIONAL verdict already centers on the same weakness, and the stress-test confirms that it is load-bearing. The paper's strongest contribution is the concrete documentation of quality failures in specific datasets: full-corpus quantitative metrics (duration, VAD, speaker diversity), the Norwegian WER experiment, the Arabic and Cantonese classification results, and the nan_tw case study are credible and independently checkable. Those findings support the existence of serious macro-level issues in particular languages. What they do not support, without additional analysis, is the abstract-level generalization that macro-level issues are 'more prevalent' in less-institutionalized languages and that a 'strong positive correlation' exists. The review sample is small, the language selection is not described, and the macro-issue diagnosis is qualitative with no agreement metric. Because the prevalence claim is the part most likely to be quoted and acted upon by dataset users, it needs either a proper operationalization and statistical test or a downgrade to a hypothesis. Since the reader already marked the paper CONDITIONAL on essentially this ground, no verdict change is needed; the concrete test would determine whether the condition can be lifted or should be made stricter.","tokens_in":21099,"tokens_out":3295,"duration_ms":40163,"concrete_test":"Build a transparent coding matrix from Table 16: for all ~40 reviewed languages, record an institutionalization score (e.g., Ethnologue EGIDS level, coded by two annotators) and a binary macro-issue verdict under the paper's four categories, using the same 100-sentence reviews plus the metadata-level checks on script and dialect declarations. Report Cohen's kappa for macro-issue coding, then run Fisher's exact test (or Kendall rank correlation) between EGIDS and macro-issue presence. If the association is non-significant or inter-rater agreement is low, the prevalence/correlation claim should be reduced to case-study status; if it survives with acceptable agreement, the central claim is materially supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim is that macro-level issues are 'more prevalent in less-institutionalized, often under-resourced languages' and that there is a 'strong positive correlation' between institutionalization and dataset quality. That correlation is asserted, not demonstrated: no institutionalization variable is defined or scored, no prevalence table is given, and no correlation statistic is reported. The qualitative evidence consists of case studies (Norwegian, Arabic, Cantonese, Fula, Kabuverdianu, nan_tw) plus the observation that VoxPopuli, whose 16 European Parliament languages are nearly all highly institutionalized, shows no macro issues. Around 40 languages were reviewed by volunteer native speakers at 100 randomly sampled utterances each (Section 1, Table 16), with no inter-rater agreement, no confidence intervals, and no explanation of how the 40 languages were selected. Macro-level issues such as dialect scope (ff_sn, kea_cv) are difficult to detect reliably from 100 sentences, and the table does not record a uniform macro-issue verdict for every reviewed language, so the denominator for 'prevalence' is unclear. Section 8 acknowledges coverage limits, but the unsupported correlation is used in the abstract and introduction as a main result. The pattern may be real, but as reported it is an interpretation consistent with chosen examples rather than a measured association.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript audits three widely used multilingual speech datasets (Mozilla Common Voice 17.0, FLEURS, and VoxPopuli) for data quality issues, dividing them into micro-level issues (short durations, low speech proportion, topic imbalance, limited speaker diversity) and macro-level issues (unspecified writing systems, register/variety confusion, ambiguous dialect scope). It presents case studies for Norwegian Bokmål/Nynorsk, Arabic, Cantonese/Hong Kong Chinese, Fula, Kabuverdianu, and Taiwanese Southern Min, and it proposes a language-planning checklist for future dataset creation. The paper's headline claim is that macro-level issues are more prevalent in less-institutionalized, often under-resourced languages and that there is a 'strong positive correlation' between a language's institutionalization status and dataset quality.","tokens_in":21380,"tokens_out":4126,"duration_ms":53321,"significance":"The concrete case studies are valuable and, for the most part, well supported: the Norwegian WER experiment, the canto-filter analysis of Cantonese subsets, the nan_tw structure analysis, and the references to public issue reports give the audit a credible empirical core. The proposed guidelines for sociolinguistic assessment and proactive language planning are a useful contribution to dataset-creation practice, and the paper explicitly connects dataset construction to language planning in a way that is likely to stimulate debate. If the headline correlation were supported, the paper would have a strong systematic claim with clear consequences for low-resource ASR evaluation. As written, however, that central claim is asserted rather than demonstrated, which both overstates the evidence and obscures the genuinely solid case-study findings.","major_comments":[{"comment":"The claim of a 'strong positive correlation between a language's institutionalization status and its dataset quality' is never measured or operationalized. No institutionalization variable is defined or scored, no per-language prevalence table is provided, and no correlation statistic is reported. The support offered in §4 is qualitative: selected case studies plus the observation that VoxPopuli, whose 16 European Parliament languages are mostly highly institutionalized, shows no macro-level issues. This is not a correlation. The claim appears in the abstract and introduction as a main result, so it is load-bearing; please either define and measure institutionalization and quality across a defined set of languages, or restrict the abstract and introduction to the observed case-study pattern rather than asserting a statistical relationship.","section":"Abstract and §1"},{"comment":"The prevalence claim that macro-level issues are 'more prevalent' in less-institutionalized languages is not supported by the described qualitative protocol. The paper reports that around 40 languages were reviewed by volunteer native speakers at 100 randomly sampled sentences each, but it does not report inter-rater agreement, confidence intervals, how the roughly 40 languages were selected, or a uniform macro-issue verdict for every reviewed language. Macro-level issues such as dialect scope (ff_sn, kea_cv) are difficult to diagnose reliably from 100 sentences, and the appendix table lists languages without recording outcomes. Section 8's coverage limitation is honest, but it does not fix the unsupported prevalence statement, which is used to justify the paper's main claim. Please either provide the missing protocol details and per-language results, or weaken the prevalence statements to 'the reviewed languages exhibit...'.","section":"§1, Table 16, §8"},{"comment":"The classification results in Tables 2, 4, and 5 rest on marker-based scripts whose precision and recall are not reported. This matters because the Norwegian, Arabic, and Cantonese percentages are presented as quantitative evidence for the macro-level claims. The Norwegian WER experiment is a strong independent confirmation for that particular case, but the Arabic and Cantonese prevalence numbers would be more convincing with a validation set, a precision/recall estimate, or a release of the annotated samples used for manual verification. Please add such validation or explicitly label these numbers as approximate heuristic classifications.","section":"§4.1–4.2"}],"minor_comments":[{"comment":"The paper says VoxPopuli has 16 European languages in §2 and in Table 6, but §3.2 and Figure 1 refer to '14 languages' in VoxPopuli; this inconsistency should be corrected, and the speech-proportion analysis should state which languages were excluded.","section":"§2 and Table 6"},{"comment":"The sentence 'While all 14 languages in V oxPopuli have at least 89% speech' is followed by 'the da_dk training set in FLEURS' with broken spacing; please copyedit the section for formatting errors such as 'theda_dktraining set'.","section":"§3.2"},{"comment":"The sentence 'This is significant because the the Guinean variant of Fula has the most speakers' contains a duplicated article; please fix the typo.","section":"§4.3"},{"comment":"The text 'around 12 hours in each language' is misspelled as 'aroudn 12 hours'; please correct.","section":"Appendix B"},{"comment":"The canto-filter package used in the Cantonese analysis is authored by one of the paper's co-authors; the text cites the package but should also state this authorship and affiliation explicitly for transparency.","section":"§4.2.2"}],"recommendation":"major_revision","confidential_remarks":"I see this paper as a potentially valuable contribution if the central claim is reframed to match the evidence. The case-study material is strong enough to warrant publication, but the current abstract and introduction claim a measured correlation that is not actually measured. I would encourage the editor to request either (a) a full operationalization of institutionalization and dataset quality with a reported statistic, or (b) a careful weakening of the abstract and headline claims to 'the inspected languages show a pattern consistent with...' rather than a correlation. This is a fixable framing and reporting issue, not a rejection-level flaw in the underlying observations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the multilingual speech dataset audit. The concrete findings are the real value: FLEURS yue_hk is essentially Standard Written Chinese, not Cantonese; ar_eg is ~98.6% Fusha; MCV17 nn_no contains 8.1% Bokmal; and the nan_tw subset is a dual-script dictionary dump. These are new, specific, and checkable. The Norwegian WER experiment, where a Bokmal model gets 49% WER on the mixed nn_no subset versus 24% on FLEURS nb_no, gives a sharp sense of how these issues distort evaluation. The paper also grounds its case studies in public issue reports and native-speaker review, and the proposed checklist for dataset creators is sensible.\n\nThe soft spots are real but localized. The \"strong positive correlation between institutionalization and dataset quality\" is asserted, not demonstrated. There is no defined institutionalization variable, no scored measure, no reported correlation statistic. The qualitative sample is roughly 100 sentences per language from volunteer native speakers, with no inter-rater agreement and no explanation of how the ~40 languages were selected; macro issues like dialect scope are hard to catch in 100 sentences. So the abstract's headline claim is an interpretation consistent with chosen examples, not a measured association. Section 8 does limit coverage, but the correlation is used in the abstract and intro as a main result. That should be softened or backed with actual measurement.\n\nMinor point: the canto-filter package is by a co-author, though it is open-source and independently usable; the central Cantonese observation is corroborated by public issues, so I don't see a circularity problem. The classification heuristics are transparent marker lists. The quantitative parts present (duration, speech proportion, speaker counts, WER experiment) look fine. The citation pattern is honest; Kreutzer et al. for methodology is credited.\n\nWho is this for: anyone building or using multilingual speech datasets, especially for low-resource languages. It deserves a serious referee. My own verdict would be conditional: accept after the correlation claim is either measured or downgraded, and after the qualitative sampling is described with clearer denominators.\n\nSend it to peer review.","headline":"A genuinely useful dataset audit with concrete new findings, whose headline correlation claim outruns its evidence.","tokens_in":21871,"tokens_out":1539,"would_cite":true,"duration_ms":16761,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Multilingual speech datasets carry serious quality problems concentrated in less-institutionalized languages, making many low-resource benchmark results untrustworthy without per-language inspection.","keywords":["speech dataset quality","multilingual ASR","low-resource languages","sociolinguistics","diglossia","digraphia","language planning","data audit"],"falsifier":"Re-run the audit on the full validated splits of the flagged languages (for example FLEURS yue_hk, FLEURS ar_eg, MCV17 nn_no, and MCV17 nan_tw) with two independent native-speaker annotation teams and reported agreement per language, then compare macro-level issue rates and the WER change after cleaning. If full-split classification finds the issue rates close to zero or the cleaning has no measurable effect on WER, the central prevalence and institutionalization-correlation claims would collapse.","tokens_in":20936,"feed_emoji":"🗣️","tokens_out":9404,"duration_ms":95023,"temperature":0.7,"pith_summary":"The paper sets out to establish that three widely used multilingual speech datasets — Mozilla Common Voice 17.0, FLEURS, and Vox Populi — contain quality problems serious enough to distort automatic speech recognition (ASR) evaluation, and that the most harmful problems are not spread randomly: they concentrate in less-institutionalized, often under-resourced languages. The authors split the problems into micro-level issues, which automatic metrics can detect, and macro-level issues, which require knowledge of a language's sociolinguistic situation and are largely invisible to filters. Through native-speaker review of roughly 100 sentences per language across about 40 languages, plus targeted case studies of Norwegian, Arabic, Cantonese, and Taiwanese Southern Min, they argue that locale labels and raw benchmark scores cannot be trusted without per-language inspection. If the paper is right, dataset building for less-institutionalized languages should be treated as a language-planning activity, not just a technical data-collection task.","feed_headline":"Speech datasets for low-resource languages hide serious flaws","feed_subtitle":"An audit of Common Voice, FLEURS and VoxPopuli links quality gaps to a language's institutional status.","key_machinery":"The audit protocol is the load-bearing instrument. On the quantitative side, the paper computes signal-to-noise ratio, voice activity detection, median utterance duration, median word count, and average hours per speaker for each language subset, which exposes micro-level issues like short clips, silence-heavy recordings, and single-speaker subsets. On the qualitative side, volunteer native speakers reviewed 100 randomly sampled text-plus-audio sentences per language across about 40 languages for coherence, audio-text alignment, dialect, topic domain, and language identification, which exposes macro-level issues. Targeted classifier scripts — a marker-word algorithm for Norwegian Bokmål versus Nynorsk, a marker-word algorithm for Arabic Fusha versus dialect, and the canto-filter package for Cantonese versus Standard Written Chinese — convert sociolinguistic distinctions into testable proportions. The micro/macro taxonomy does the argument's work: micro-level issues can be fixed programmatically, while macro-level issues require language planning decisions about orthography, register, and dialect scope.","core_discovery":"The paper's central claim is that the quality of multilingual speech datasets systematically tracks a language's institutionalization. Micro-level defects — extremely short utterances, low proportions of actual speech, imbalanced topic domains, and lack of speaker diversity — appear across languages but are detectable programmatically and can be mitigated by filtering; macro-level defects — unspecified writing systems in digraphic languages, ambiguous registers in diglossic languages, and underspecified dialect boundaries — are concentrated in less-institutionalized languages and require human linguistic expertise to diagnose. The evidence includes Norwegian subsets that mix Bokmål and Nynorsk despite their locale labels, a FLEURS Arabic subset labeled Egyptian that is almost entirely Modern Standard Arabic, a FLEURS Cantonese subset that contains no Cantonese at all, and a Common Voice Taiwanese Southern Min subset that is dictionary-like, dual-script, and poorly aligned. The paper further shows the evaluation cost: a Norwegian Bokmål ASR model's substitution error rate rises by roughly 25 percentage points on the mixed Nynorsk-labeled subset. What follows is that published WER numbers for many low-resource languages may reflect dataset artifacts rather than model capability.","pith_inferences":["If the institutionalization-quality correlation extends beyond the audited languages, raw WER comparisons across low-resource languages should be treated as provisional until each subset passes a sociolinguistic audit; this is stricter than the paper's own recommendation to add human evaluation.","The paper's framing implies a feedback loop it only gestures at: dataset creation is itself a language-planning intervention, so tracking how contributors' script and register choices shift across Common Voice versions for a digraphic language would provide a measurable way to watch community orthographic norms form.","A testable extension of the paper's argument is that prescribing orthography and register rules before collection would lower downstream WER for a digraphic or diglossic language even when audio quantity is held constant, which would separate macro-level contamination from data scarcity."],"forward_implications":["Norwegian subsets mix the two written standards: MCV17's nn_no contains 8.1% Bokmål sentences and FLEURS's nb_no contains 8.8% Nynorsk, and a Bokmål-trained ASR model's substitution error rate is about 25 percentage points higher on the mixed nn_no subset, so WER comparisons across these datasets are not clean measures of the same language.","FLEURS's yue_hk subset is almost entirely Standard Written Chinese (89.8% of sampled prompts), not Cantonese, meaning model evaluations on that subset can appear successful while producing outputs in the wrong register.","FLEURS's ar_eg subset is 98.6% Modern Standard Arabic rather than Egyptian Arabic, showing that locale labels in diglossic settings can encode the wrong variety.","MCV17's nan_tw subset has 21 hours of raw audio but only 48.3% actual speech, and its text prompts are mostly dictionary-style single words or phrases written redundantly in two scripts with mismatched audio, making it nearly unusable without substantial restructuring.","Dataset creators should conduct sociolinguistic assessment, prescribe orthography and register choices, enforce multi-level quality checks, and release detailed metadata as a standard part of building speech datasets for less-institutionalized languages."],"supporting_citations":[{"why":"Defines Mozilla Common Voice, the volunteer-driven corpus whose 17.0 release is one of the three datasets audited here.","marker":"Ardila et al. (2020)"},{"why":"Defines FLEURS, the few-shot benchmark dataset audited here.","marker":"Conneau et al. (2023)"},{"why":"Defines VoxPopuli, the European Parliament speech corpus audited here.","marker":"Wang et al. (2021)"},{"why":"Supplies the audit methodology for multilingual datasets that this paper adapts from text to speech.","marker":"Kreutzer et al. (2022)"},{"why":"Provides FLoRes-101, the source of FLEURS text prompts, explaining the formal Wikipedia register the audit finds.","marker":"Goyal et al. (2022)"},{"why":"Establishes the concept of diglossia used to analyze Arabic, Cantonese, and other register-divided languages.","marker":"Ferguson (1959)"},{"why":"Establishes the concept of digraphia used to analyze script choices in Serbian, Norwegian, Mongolian, and Taiwanese Southern Min.","marker":"Dale (1980)"},{"why":"Provides the canto-filter package used to classify Cantonese versus Standard Written Chinese prompts.","marker":"Lau et al. (2024)"},{"why":"Documents how lack of speaker diversity in ASR data produces biased downstream performance, supporting the speaker-diversity concern.","marker":"Koenecke et al. (2020)"},{"why":"Introduces Whisper, the downstream model whose evaluation the paper shows can be distorted by these dataset problems.","marker":"Radford et al. (2023)"}],"fun_headline_variants":["Audit: low-resource speech datasets mask model failures","Speech dataset flaws skew low-resource benchmarks","Quality gaps in Common Voice, FLEURS, VoxPopuli track language status","Macro-level flaws in low-resource speech datasets evade simple filters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The prevalence and correlation claims rest on the assumption that 100 randomly sampled sentences per language, judged by volunteer native speakers with no reported inter-rater agreement, represent the quality of each full language subset, even though the paper itself notes in its limitations that many languages remain uninspected.","fun_headline_variants_meta":{"raw":{"variants":["Audit: low-resource speech datasets mask model failures","Speech dataset flaws skew low-resource benchmarks","Quality gaps in Common Voice, FLEURS, VoxPopuli track language status","Macro-level flaws in low-resource speech datasets evade simple filters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000177,"raw_usage":{"total_tokens":1293,"prompt_tokens":947,"completion_tokens":346,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":274}},"tokens_in":563,"tokens_out":346,"duration_ms":3846,"temperature":1.0,"reasoning_tokens":274,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:30:22.872689+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the audit on the full validated splits of the flagged languages (for example FLEURS yue_hk, FLEURS ar_eg, MCV17 nn_no, and MCV17 nan_tw) with two independent native-speaker annotation teams and reported agreement per language, then compare macro-level issue rates and the WER change after cleaning. If full-split classification finds the issue rates close to zero or the cleaning has no measurable effect on WER, the central prevalence and institutionalization-correlation claims would collapse.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines Mozilla Common Voice, the volunteer-driven corpus whose 17.0 release is one of the three datasets audited here."}],"review_version":1}