{"id":"fb318fa3-e6bf-4f06-9656-60131e422dd1","arxiv_id":"2509.07139","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The ML-SUPERB 2.0 Challenge adds accented and dialectal speech evaluation to multilingual ASR, and all five submitted systems beat the baselines.","lead":"This paper presents the ML-SUPERB 2.0 Challenge, a speech recognition benchmark that evaluates multilingual models on 149 languages and 93 accents or dialects through a hidden test server. A reader might care because it measures how well automatic speech recognition works across language varieties, not just high-resource languages.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dialectal test set shares 9 of 13 source corpora with the provided development set, so reported dialectal gains may reflect dev-set adaptation rather than inclusive robustness.","rationale":"The reader's weakest assumption focused on the accuracy of language, accent, and dialect labels in the source corpora. That is a valid data-quality concern, and the paper's cleaning efforts (Section 2.1) show awareness of such issues. However, the more immediate and testable threat to the central claim is the design overlap between the dialectal development set and the hidden test set. Since participants are allowed to use any data and are explicitly given the development set, the same nine corpora appear on both sides of the evaluation. This makes the 'hidden' dialectal test partially exposed in distribution even if the exact utterances are not. The submitted systems could have adapted to corpus-specific acoustic conditions, speaker pools, or transcription conventions without improving genuine dialectal robustness. The paper does not report any per-corpus analysis, so the observed aggregate gains cannot be attributed to inclusive improvements over held-out varieties. This is a stronger and more specific internal-validity concern than the label-noise issue, because it directly follows from the described benchmark construction. The abstract/Section 4.1 numeric inconsistency adds confusion but is secondary to the missing corpus-level breakdown. The verdict remains conditional: the standard multilingual results may be sound, and the submitted systems likely did outperform the baselines on the aggregate metrics, but the dialectal inclusivity claim requires the proposed per-corpus check before it can be accepted as evidence of generalizable progress.","tokens_in":8121,"tokens_out":4517,"duration_ms":40739,"concrete_test":"Request the organizers to release per-corpus dialectal CER and LID for every submission and for the XEUS baseline, specifically separating the 9 development-overlap corpora (refs [25-33]) from the 4 held-out corpora (refs [34-37]). Recompute the best-over-XEUS deltas on the 4 held-out corpora alone. If the deltas shrink substantially or reverse, the dialectal robustness claim in Section 5 is largely an artifact of dev/test corpus overlap rather than inclusive ASR progress. This check only requires re-scoring stored inference outputs on the hidden test set, not new data collection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that all five submitted systems improved over SSL baselines, especially on dialectal data, rests on evaluation against a hidden test set. Section 2.2 states that the dialectal development set is drawn from 9 corpora [25-33] and the hidden test set is drawn from the same 9 corpora plus 4 additional ones [34-37]. Section 3.5 states participants are given this development set and may use any data. Consequently, the development portion is a near-perfect distributional proxy for most of the test set, enabling participants to tune models, select checkpoints, or adapt to corpus-specific speakers, recording conditions, and label conventions. The reported best-over-XEUS dialectal gains (23.0 LID, 30.2 CER in Section 4.1) may therefore overstate general robustness to unseen varieties. Without a per-corpus breakdown separating the 9 shared corpora from the 4 held-out corpora, the inclusivity conclusion in Section 5 is not justified by the aggregate scores. A separate consistency issue also clouds interpretation: the abstract reports 23% LID and 18% CER improvements, while Section 4.1 gives 12.4 and 19.3 for the same general test set; the discrepancy should be resolved for the reader to know which result is the headline.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports on the Interspeech 2025 ML-SUPERB 2.0 Challenge, an open ASR competition in which participants submit systems through a DynaBench-based server and are evaluated on a multilingual test set spanning 149 languages and a newly collected accented/dialectal test set spanning 93 varieties. The authors describe data cleaning decisions, challenge rules, and evaluation metrics (Standard LID/CER, language robustness StD/Worst-15 CER, and dialectal LID/CER), and compare five submissions from three teams against SSL and supervised baselines. They report that all submissions ranked above the best SSL baselines, with per-metric improvements over XEUS of 12.4 in Standard LID, 19.3 in Standard CER, 5.8 in StD, 4.1 in Worst-15 CER, 23.0 in Dialect LID, and 30.2 in Dialect CER.","tokens_in":8341,"tokens_out":4544,"duration_ms":36889,"significance":"If the reported results are taken at face value, the challenge makes a useful contribution: it is one of the broadest multilingual ASR evaluations to date, it includes an online evaluation server that keeps the test set hidden from participants, and it provides evidence that unconstrained community submissions can improve over standard SSL fine-tuning on both standard and dialectal test sets. Strengths include transparent documentation of data-cleaning decisions, a fixed ranking protocol based on average rank across six metrics, and release of the evaluation infrastructure. The empirical claim about dialectal robustness is weakened by the partial overlap between development and test sources, and the headline numbers in the abstract do not match Section 4.1, so the results as currently presented need revision before the significance can be fully assessed.","major_comments":[{"comment":"The abstract states a 23% LID improvement and an 18% CER reduction on the general multilingual test set, and a 15.7% LID improvement on accented/dialectal data. Section 4.1 reports improvements over XEUS of 12.4 in Standard LID, 19.3 in Standard CER, 23.0 in Dialect LID, and 30.2 in Dialect CER, with no 15.7 value appearing anywhere in the paper. The abstract's 23% matches the Dialect LID number, not the general-set LID number; the 18% matches neither Standard CER (19.3) nor Dialect CER (30.2); and 15.7 is unexplained. This makes the headline result ambiguous and must be corrected so the reader knows which metric each number refers to.","section":"Abstract and Section 4.1"},{"comment":"The dialectal development set is drawn from 9 corpora [25-33] and the hidden test set is drawn from the same 9 corpora plus 4 additional ones [34-37]. Because participants receive the development set (Section 3.5) and may use any data, 9 of the 13 test-source corpora are effectively public during system development, making the test set a near-perfect distributional proxy for the development set for those corpora. The reported dialectal gains (23.0 in LID, 30.2 in CER) may therefore overstate robustness to unseen varieties. Please report a per-corpus breakdown separating the 9 shared corpora from the 4 held-out corpora and discuss whether the aggregate conclusion holds on the 4 truly unseen corpora.","section":"Section 2.2 and Section 3.5"}],"minor_comments":[{"comment":"Section 3.2 says participants are tasked with developing systems for 154 languages, while Section 2.1 and the introduction consistently say 149 languages; please align these numbers.","section":"Section 3.2"},{"comment":"The phrase '200+ languages, accents, and dialects' is imprecise: the paper evaluates 149 languages and 93 accents/dialects, which are not both languages; consider writing 'more than 200 language varieties and accents' or stating the two numbers explicitly.","section":"Abstract and Conclusion"},{"comment":"The claim that 'each team had a system submission that ranked 1st in at least 1 metric' would be easier to verify if the per-metric ranks of the five submissions were shown in a table; Figure 3 is hard to read at the level of individual metrics.","section":"Section 4.1"},{"comment":"The ranking is based on only seven systems (five submissions plus two baselines), and no confidence intervals or significance tests are reported; the authors should note that small rank differences are not necessarily meaningful.","section":"Section 3.7"}],"recommendation":"major_revision","confidential_remarks":"Please verify the abstract metrics against the final tables before publication; the discrepancy between the abstract (23%/18%/15.7%) and Section 4.1 (12.4/19.3/23.0/30.2) is publication-blocking. Also ask the authors to provide the per-corpus breakdown for the dialectal test set so the dev/test overlap concern can be assessed quantitatively."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this for the benchmark, not for a sharp new algorithm. The ML-SUPERB 2.0 Challenge adds dialectal varieties and a hidden DynaBench evaluation server to the existing ML-SUPERB line, and the authors did the careful data-cleaning work that makes such a benchmark usable. The core empirical finding—that all five submitted systems beat the SSL baselines on the general multilingual test set—is supported by the tables and is believable. The robustness metrics (Worst 15, StD) are a good addition, and the decision to let participants use any data is the right call for a community challenge.\n\nThe soft spots are in the reporting. The abstract's headline numbers do not match Section 4.1: the abstract claims a 23% LID improvement and 18% CER reduction on the general set, but Section 4.1 reports 12.4 and 19.3 for those metrics (and the 23.0 is actually the Dialect LID improvement). The dialectal LID number is also inconsistent (15.7 in the abstract vs 23.0 in the body). That needs to be fixed.\n\nMore substantively, the dialectal test set is built from 9 of the same corpora as the development set plus only 4 held-out ones. Participants had the development set, so they could adapt to those 9. The paper does not break out scores on the 4 new corpora, so the claim that the results show robustness to unseen varieties is not supported by the aggregate numbers. The authors should either provide per-corpus results or soften the conclusion. This is a real limitation, though it does not invalidate the overall finding that submissions improved over baselines.\n\nMissing error bars are a minor issue for a challenge report. The evaluation methodology is otherwise sound: the test set is hidden, the ranking metrics were fixed in advance, and the baselines are clearly described.\n\nThe paper deserves a serious referee. It is a useful resource for anyone building multilingual or dialectal ASR benchmarks. I would cite it. The inconsistencies are fixable, and the dialectal overlap should be acknowledged or analyzed.","headline":"A solid challenge write-up whose core finding is believable, but the abstract's numbers don't match the body and the dialectal test set shares most of its corpora with the development set, so the inclusivity claims are weaker than advertised.","tokens_in":8913,"tokens_out":3245,"would_cite":true,"duration_ms":26936,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new ASR challenge spans 149 languages and 93 dialects, and every submitted system beat the strongest self-supervised baselines.","keywords":["multilingual ASR","language identification","dialect robustness","benchmark challenge","self-supervised learning","inclusive speech processing","character error rate","ML-SUPERB 2.0"],"falsifier":"Take a random sample of utterances from the dialectal test set, have human annotators verify the dialect label and transcript, and recompute LID and CER on the cleaned subset; if the submitted systems' advantage over XEUS largely disappears, the claimed gains are an artifact of label noise.","tokens_in":7927,"feed_emoji":"🎙️","tokens_out":5018,"duration_ms":39138,"temperature":0.7,"pith_summary":"This paper presents the ML-SUPERB 2.0 Challenge, an ASR evaluation covering 149 languages and 93 accents and dialects, run on a hidden test set through an online server. It claims that all five submitted systems from three teams outperformed the strongest self-supervised baselines, with the best per-metric system cutting dialectal character error rate by 30.2 points and improving dialectal language identification by 23.0 points over the XEUS baseline. The authors see this as evidence that unconstrained community challenges can drive progress toward inclusive speech technology, while also showing that accented and dialectal speech remains the weak spot of current systems.","feed_headline":"All five systems beat baselines in 149-language ASR challenge","feed_subtitle":"Best entry cut dialectal CER by 30.2 points and raised dialectal language ID by 23 points.","key_machinery":"The mechanisms are the challenge protocol and the metric suite. Participants upload model weights and inference code to an online evaluation server built on DynaBench, which runs inference on a hidden test set and returns only aggregate scores, preventing benchmark overfitting. Rankings are computed by averaging each system's rank on six metrics: Standard LID accuracy, Standard CER, CER standard deviation across languages, worst-15 CER, Dialectal LID accuracy, and Dialectal CER. The data pipeline also does linguistic normalization, such as merging Tagalog and Filipino, removing Norwegian due to conflation of written standards, and reconciling ISO codes, because label consistency is required for fair robustness metrics.","core_discovery":"The paper's central claim is that a challenge with no restrictions on training data, architectures, or pretrained models, evaluated on a fully hidden test set spanning 149 languages and 93 language varieties, produces ASR systems that beat strong self-supervised baselines on every metric. The best submission per metric improved over XEUS by 12.4 in Standard LID accuracy, 19.3 in Standard CER, 5.8 in standard deviation of CER, 4.1 in worst-15-languages CER, 23.0 in Dialectal LID, and 30.2 in Dialectal CER. The paper also reports that supervised models like Whisper and OWSM degrade sharply on languages unseen in their training data, and that even the best challenge systems perform considerably worse on accented and dialectal data than on standard varieties.","pith_inferences":["The dialectal gains may come less from novel architectures than from the freedom to curate external training data and ensemble models; a follow-up ablation separating data from modeling would test that.","Since the dialectal test set draws on a finite set of accent corpora, the challenge measures robustness to those particular varieties; a future round with entirely unseen accent sources would reveal how much of the gain is generic.","The seen-versus-unseen collapse of supervised models suggests the challenge could double as an audit tool for which languages are actually represented in a pretraining corpus.","Rank-based aggregation can hide large absolute-score gaps; reporting raw deltas alongside ranks would give a fuller picture of system differences."],"forward_implications":["If the results hold, unconstrained shared tasks become a dependable mechanism for pushing multilingual and dialectal ASR beyond self-supervised fine-tuning baselines.","The large remaining gap on accented and dialectal speech means future benchmarks and models must treat dialect robustness as a separate objective, not a byproduct of language coverage.","The hidden-test-set server design blocks benchmark overfitting, so the measured gains are more likely to transfer to new speech data than gains from open test sets.","The use of average rank across six metrics makes leaderboard positions less sensitive to the different dynamic ranges of CER and LID accuracy.","Comparing supervised and self-supervised baselines shows that training-data language coverage is the dominant factor in cross-lingual performance, so scaling pretraining data may matter more than architecture choice."],"supporting_citations":[{"why":"Supplies the earlier ML-SUPERB multilingual data that was combined and cleaned to build the general test set.","marker":"[11]"},{"why":"Provides the ML-SUPERB 1.0 Challenge design and hidden-set languages that the 2.0 Challenge extends.","marker":"[13]"},{"why":"Defines the ML-SUPERB 2.0 benchmark whose public training and development sets are used for baselines and participant support.","marker":"[14]"},{"why":"Describes DynaBench, the platform that hosts the hidden evaluation server and enforces the submission API.","marker":"[16]"},{"why":"FLEURS is a source corpus whose label inconsistencies (e.g., Serbian orthography) motivated the data-cleaning steps.","marker":"[22]"},{"why":"Common Voice is a crowd-sourced corpus that introduces orthography and ISO-code issues the authors had to resolve.","marker":"[23]"},{"why":"MMS is one of the self-supervised SSL baselines that submitted systems had to outperform.","marker":"[17]"},{"why":"XEUS is the strongest SSL baseline and the reference point for the reported absolute improvements.","marker":"[18]"}],"fun_headline_variants":["ML-SUPERB 2.0: 5 systems beat baselines on 149 languages","ASR challenge: best entry cuts dialect CER by 30.2%","Inclusive ASR: dialectal CER down 30.2% in ML-SUPERB 2.0","5 ASR systems outdo baselines on hidden 149-language test","ML-SUPERB 2.0: 23% better LID, 30% better dialect CER"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole evaluation stands on the source corpora's language, accent, and dialect labels and transcripts being correct; if those labels are wrong or noisy, the rankings and fairness conclusions do not follow.","fun_headline_variants_meta":{"raw":{"variants":["ML-SUPERB 2.0: 5 systems beat baselines on 149 languages","ASR challenge: best entry cuts dialect CER by 30.2%","Inclusive ASR: dialectal CER down 30.2% in ML-SUPERB 2.0","5 ASR systems outdo baselines on hidden 149-language test","ML-SUPERB 2.0: 23% better LID, 30% better dialect CER"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000838,"raw_usage":{"total_tokens":3631,"prompt_tokens":902,"completion_tokens":2729,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":2607}},"tokens_in":518,"tokens_out":2729,"duration_ms":17740,"temperature":1.0,"reasoning_tokens":2607,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:13:19.110086+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of utterances from the dialectal test set, have human annotators verify the dialect label and transcript, and recompute LID and CER on the cleaned subset; if the submitted systems' advantage over XEUS largely disappears, the claimed gains are an artifact of label noise.","supporting_citations":[{"cited_title":"Since all of these models are self-supervised, we develop ASR systems via fine-tuning on the ML-SUPERB 2.0 public set [14]","cited_arxiv_id":null,"evidence_quote":"Supplies the earlier ML-SUPERB multilingual data that was combined and cleaned to build the general test set."},{"cited_title":"The challenge introduces a novel multilin- gual test suite of accented and dialect speech and uses new metrics to test the robustness of ASR systems","cited_arxiv_id":null,"evidence_quote":"Provides the ML-SUPERB 1.0 Challenge design and hidden-set languages that the 2.0 Challenge extends."},{"cited_title":"Wav2vec 2.0: A framework for self- supervised learning of speech representations,","cited_arxiv_id":null,"evidence_quote":"Defines the ML-SUPERB 2.0 benchmark whose public training and development sets are used for baselines and participant support."},{"cited_title":"Robust speech recognition via large-scale weak supervision,","cited_arxiv_id":null,"evidence_quote":"Describes DynaBench, the platform that hosts the hidden evaluation server and enforces the submission API."},{"cited_title":"SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Genera- tive Capabilities,","cited_arxiv_id":null,"evidence_quote":"FLEURS is a source corpus whose label inconsistencies (e.g., Serbian orthography) motivated the data-cleaning steps."},{"cited_title":"A V-SUPERB: A Multi-Task Evaluation Bench- mark for Audio-Visual Representation Models,","cited_arxiv_id":null,"evidence_quote":"Common Voice is a crowd-sourced corpus that introduces orthography and ISO-code issues the authors had to resolve."}],"review_version":2}