{"id":"136aafdd-a528-464a-ac86-ab5cf1129663","arxiv_id":"2603.09725","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"On real noisy Dutch semi-spontaneous speech, five of eight SOTA ASR models reach WER under 22%, yet five single-channel SE methods fail to improve and often degrade ASR.","lead":"Researchers released DRES, a 1.5-hour Dutch semi-spontaneous speech set from 80 speakers recorded in real noisy public indoor spaces, and tested eight ASR systems plus five speech-enhancement methods on it. Several modern ASR models already work reasonably well, but the tested single-channel enhancers did not help recognition and often hurt it.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The strongest claim is an empirical null result on a new realistic Dutch test set. Table 2 and the paired bootstrap tests supply direct evidence; DNSMOS improvements without ASR gains further illustrate the known quality–recognition mismatch. The reader's weakest_assumption (limited generalizability of the chosen SE/ASR sample and single-channel setup) is accurate as a scope caveat and is already acknowledged by the authors. Because the paper does not claim universal failure of all modern SE, only that these five algorithms did not help these eight models on DRES, the caveat does not collapse the central claim. No stronger internal concern (e.g., transcription protocol error, metric misuse, or unstated preprocessing that could reverse the null) appears in the text. Therefore the reader's CONDITIONAL verdict (pending data release and tighter caveats) remains appropriate; no adjustment is required.","tokens_in":13164,"tokens_out":456,"duration_ms":14323,"concrete_test":"After public release of DRES, re-run the exact Table 2 pipeline (same five SE checkpoints, same eight ASR models, same channel-2 16 kHz normalization, same bootstrap) on the released files; if any SE method yields a statistically significant WER reduction (CI excluding 0) for any of the eight models, the null claim would need revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central negative claim—that none of the five tested single-channel SE algorithms improved WER for the eight SOTA ASR models on DRES, and most degraded it—is directly supported by Table 2, the bootstrap significance tests (§3.3, §4.2), and the DNSMOS results (Fig. 3). The reader's weakest_assumption correctly flags limited representativeness (single-channel only, specific SE/ASR sample, uncontrolled distance, Dutch indoor babble), but this is already scoped in the abstract, §1, §5, and the future-work paragraph on multi-channel SE; it does not create an internal inconsistency or undermine the measured null result on this corpus. No hidden assumption, statistical flaw, or contradiction with the reported numbers is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces DRES, a 1.5-hour Dutch semi-spontaneous speech corpus from 80 speakers recorded with a four-channel linear array in four public indoor environments with real background talkers and noise. Orthographic transcripts follow the Jasmin-CGN protocol. The authors evaluate five single-channel SE algorithms (spectral subtraction, spectral noise gating, MetricGAN-OKD, and two SGMSE+ checkpoints) via DNSMOS and eight off-the-shelf SOTA ASR systems (Google Chirp 3/Telephony, Azure, MMS, Whisper-large-v3/turbo, NeMo-nl, CGN-Conformer) via WER on the unenhanced baseline and after each SE method. On the baseline, five of eight ASR models achieve average WER below 22% (best: Google Chirp 3 at 11.2%). None of the SE methods improves WER for any ASR model; most produce statistically significant degradations (paired speaker-based bootstrap). The authors conclude that modern single-channel SE does not help SOTA ASR on this realistic Dutch data, in contrast to recent results on artificial English mixtures, and motivate multi-channel work.","tokens_in":13416,"tokens_out":1190,"duration_ms":9614,"significance":"Real noisy, semi-spontaneous Dutch speech with multi-person-checked transcripts is scarce; DRES fills a clear evaluation gap for both ASR and SE. The central negative result—that none of five widely used single-channel SE algorithms improves WER of eight contemporary ASR systems, while most degrade it—is carefully scoped, supported by Table 2, DNSMOS (Fig. 3), and speaker-based bootstrap tests with 95% CIs, and usefully contrasts with recent English artificial-mixture findings. Strengths include transparent corpus design, multi-location recording, multi-person transcription checking, and explicit future-work plans for multi-channel SE. If the null SE result holds under broader conditions, it has practical implications for whether single-channel SE should be inserted before modern E2E ASR in real indoor babble.","major_comments":[{"comment":"The central claim that modern single-channel SE fails to help SOTA ASR rests on a specific sample of five SE algorithms and eight ASR models evaluated only on channel 2 with uncontrolled 1–1.5 m distance (§3.1–3.2, Table 2, §4.2). While the measured null result on this corpus is solid, the manuscript should more explicitly bound the generalization claim (already partially acknowledged in §5) and, if space permits, add at least one multi-channel or alternative SE baseline, or a short analysis of residual noise/artifact types that may explain the WER degradation despite DNSMOS gains. Without that, the contrast with [40,41] remains suggestive rather than fully diagnostic of language vs. realism vs. algorithm class.","section":null},{"comment":"Speaker-to-array distance is uncontrolled and only subjectively estimated (§2.2). Because distance affects SNR, reverberation, and the relative benefit of SE, the paper should report at least a coarse distance or level distribution (or a sensitivity check) so that readers can judge how representative the acoustic conditions are of the claimed 1–1.5 m scenario. This is load-bearing for interpreting both the absolute WERs and the SE null result.","section":null}],"minor_comments":[{"comment":"Table 2: report per-location WERs after SE (or at least note that the additional analysis found no location-wise improvements) so the reader can verify the claim that SE never helped at any site.","section":null},{"comment":"Fig. 2 / Fig. 3: DNSMOS distributions are clear, but adding mean ± std or median values in the caption or a small table would aid quantitative comparison.","section":null},{"comment":"§2.4: vocabulary size (2,842) and total speech duration after silence removal are useful; a short note on speaking rate or average utterance length would further characterize the semi-spontaneous material.","section":null},{"comment":"Clarify whether amplitude normalization to 0.707 is applied before or after SE and whether it interacts with any of the SE implementations.","section":null},{"comment":"Minor typographical consistency: “SotA” vs “SOTA”, and ensure all model names (e.g., Whisper-large-V3 vs Whisper-large-v3) match the Hugging Face / API identifiers used.","section":null},{"comment":"Release statement: upon acceptance the corpus will be released; a short note on license and expected metadata (speaker demographics already in Table 1) would strengthen the contribution.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a solid, carefully executed resource-plus-evaluation contribution. The null SE result is the most interesting scientific claim and is well supported on this corpus; the main risk is over-generalization beyond the chosen SE/ASR sample and single-channel setting, which the authors already partially flag. Minor revision to tighten the scope language and add a distance/level note (or multi-channel teaser) should be sufficient. Fit for a speech/audio journal is good; the Dutch focus is a feature, not a bug, given the scarcity of real noisy Dutch data."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful thing here is DRES itself plus the measured null. They collected 1.5 h of semi-spontaneous Dutch from 80 speakers in four real public indoor spaces (babble, exhibition noise, etc.) with a four-channel array and careful orthographic transcripts following the Jasmin-CGN protocol. That fills a real gap: almost all public Dutch data is clean, and most SE/ASR noise work still leans on synthetic mixtures.\n\nWhat they do with it is straightforward and well-executed. Eight off-the-shelf SOTA ASR systems (Chirp 3, Azure, Whisper-large-v3 and turbo, MMS, NeMo-nl, a CGN Conformer, etc.) on the unenhanced channel-2 audio give a clear ranking: five models under 22% WER, Chirp 3 at 11.2%. Then five single-channel SE methods (spectral subtraction, noise gating, MetricGAN-OKD, two SGMSE+ checkpoints) raise DNSMOS but either leave WER unchanged or significantly degrade it under speaker-based bootstrap tests. That directly contradicts recent English synthetic-mixture claims for SGMSE+ and is the result people will cite. Statistics and scoping are appropriate for the claims as written; no circularity, no invented metrics.\n\nSoft spots are mostly scope, not internal cracks. Only single-channel (they flag multi-channel for future work), uncontrolled 1–1.5 m distance, 1.5 h total, and a specific sample of SE/ASR systems. The generalization claim is already caveated in the abstract and discussion, so the reader’s “weakest assumption” is real but not load-bearing. Data release is still “upon acceptance,” which is the only practical caveat for reuse.\n\nThis is for anyone building or evaluating Dutch ASR/SE, or anyone tempted to assume English synthetic SE gains transfer. It deserves a serious referee; once the corpus is public it becomes a standard test set. I would engage with it and cite the negative SE result.","headline":"Solid resource paper: new Dutch realistic test set plus a clean negative result that single-channel SE does not help modern ASR on it.","tokens_in":13979,"tokens_out":493,"would_cite":true,"duration_ms":4535,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Modern single-channel speech enhancement does not improve SOTA ASR on real Dutch noisy speech; five of eight models still stay under 22% WER without it.","keywords":["Dutch speech","realistic noisy speech","semi-spontaneous speech","speech recognition","speech enhancement","word error rate","DNSMOS","microphone array"],"falsifier":"A controlled re-run in which a multi-channel or jointly-trained enhancement front-end, or a larger set of SE algorithms, yields a statistically significant WER reduction on the same DRES utterances for at least one of the eight ASR models.","tokens_in":14064,"feed_emoji":"🎙️","tokens_out":654,"duration_ms":6417,"temperature":0.7,"pith_summary":"Most speech-recognition and enhancement research still relies on clean speech mixed with artificial noise. This paper asks what happens when the same state-of-the-art systems are tested on real people speaking Dutch in busy public indoor spaces with genuine background talkers. The authors release DRES, a 1.5-hour semi-spontaneous corpus from 80 speakers recorded with a four-channel microphone array in four different buildings. On the unenhanced recordings, five of eight off-the-shelf ASR systems already achieve average word-error rates below 22 percent. When five well-known single-channel enhancement algorithms (classical and modern) are applied first, none of them improves recognition accuracy; most make it worse even though they raise objective speech-quality scores. The central message is that results obtained on synthetic English mixtures do not automatically transfer to realistic multilingual conditions, so evaluation on genuine noisy speech remains essential.","feed_headline":"Speech enhancers fail to help ASR on real Dutch noise","feed_subtitle":"Five of eight modern recognizers stay under 22% WER without them; all five enhancers hurt or do nothing.","key_machinery":"DRES itself: a 1.5-hour multi-channel Dutch corpus of elicited (free-speech, picture-card, prompt-card) speech from 80 speakers in four real public buildings, used as a realistic test set that exposes the gap between synthetic noise mixtures and genuine acoustic conditions.","core_discovery":"On real Dutch semi-spontaneous speech recorded in public indoor noise, none of five single-channel speech-enhancement algorithms improves the word-error rate of eight state-of-the-art ASR systems; most significantly degrade it. At the same time, five of those ASR systems already achieve average WER below 22 percent on the raw recordings, showing that modern recognizers can be surprisingly robust without enhancement.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["SE fails to aid ASR on real Dutch public noise","No WER drop from five SE models on DRES Dutch speech","Modern ASR hits under 22% WER without SE on DRES","Single-channel enhancers leave ASR unchanged or worse","Realistic Dutch indoor noise: SE brings no ASR gains"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That the five chosen single-channel enhancers and the eight off-the-shelf ASR models, run only on one microphone of a four-channel array at uncontrolled speaker distance, are representative enough of modern practice for the null enhancement result to generalise beyond this Dutch indoor setting.","fun_headline_variants_meta":{"raw":{"variants":["SE fails to aid ASR on real Dutch public noise","No WER drop from five SE models on DRES Dutch speech","Modern ASR hits under 22% WER without SE on DRES","Single-channel enhancers leave ASR unchanged or worse","Realistic Dutch indoor noise: SE brings no ASR gains"]},"model":"grok-4.5","effort":"low","cost_usd":0.004044,"raw_usage":{"total_tokens":1236,"prompt_tokens":747,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":40440000,"prompt_tokens_details":{"text_tokens":747,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":423,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":747,"tokens_out":66,"duration_ms":4722,"temperature":1.0,"reasoning_tokens":423,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T00:02:48.881195+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A controlled re-run in which a multi-channel or jointly-trained enhancement front-end, or a larger set of SE algorithms, yields a statistically significant WER reduction on the same DRES utterances for at least one of the eight ASR models.","supporting_citations":[],"review_version":1}