{"id":"a9283504-e103-450c-9461-121b6e31588c","arxiv_id":"2505.01338","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Distant-microphone, low-latency speech enhancement works better when synthetic room impulse responses link reverberation time to room volume and when the training target keeps early reflections with a 300 ms decay.","lead":"This paper asks whether a single microphone can clean up speech recorded 5 to 10 meters away in big, echoing rooms like theatres and conference halls, in real time. It shows that a speech enhancement model trained with room impulse responses whose reverberation time follows room volume, plus a dereverberation target that keeps early reflections, improves predicted quality scores.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 3.04-vs-2.77 headline is not yet supported: original DFN3 differs in both training data and (5,500) target, intrusive metrics are withheld, and DNSMOS is the only endpoint, without CIs or listening tests.","rationale":"The reader's conditional verdict is reasonable, but the weakest link is not primarily the T60 curve in Eq. (1). That curve is a heuristic for realistic RIR sampling, and the Table 2 gain (PESQ 2.17 vs 2.08) is modest; even if the curve is imperfect, the feasibility and target-design findings would survive. The claim that would change the paper's value is the SOTA result in Table 3. What would have to be true is that DFN3-d.m.'s superiority over DFN3 is due to the proposed pipeline and that DNSMOS OVRL reflects the intended perceptual improvement. The paper leaves both under-supported: the baseline differs in training data and target simultaneously, intrusive metrics for baselines are withheld, no significance or confidence intervals are reported for the 100 test examples, and the authors themselves defer listening tests to future work. A controlled retrain with full metric reporting, plus a small listening test, would settle this. I therefore keep the reader's conditional verdict; the concern is addressable and does not warrant rejection.","tokens_in":9148,"tokens_out":10021,"duration_ms":100739,"concrete_test":"Retrain DFN3 on the exact distant-mic corpus and volume-based RIR sampling used for DFN3-d.m., but with the original (5,500) target; score that checkpoint, the stock DFN3, and DFN3-d.m. on the same 100-example real-RIR test set using DNSMOS OVRL, PESQ/STOI/SI-SDR (referenced both to the corresponding windowed target and to anechoic), and a paired bootstrap CI. If the (5,500) retrained model matches 3.04 or the intrusive metrics reverse the ordering, the Table 3 claim is not attributable to the proposed pipeline. Optionally run a P.808 listening test on a stratified subset; if human OVRL does not reproduce the DNSMOS ordering, the SOTA claim is metric-dependent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim (Table 3: DFN3-d.m. OVRL 3.04 vs original DFN3 2.77) is presented as evidence that the proposed training pipeline is state-of-the-art, but the comparison does not isolate the pipeline's components. The original DFN3 baseline was trained on the standard DNS corpus with target (5,500) (§3.3/§4.3), while DFN3-d.m. is trained on distant-mic RIRs with volume-based T60 sampling and target (0,300). Any of these changes could drive the 0.27 OVRL gap. Section 4.2 does vary the target in isolation, but only on a synthetic gpuRIR test set and only with VCTK-only training; the full-data (5,500) counterpart of DFN3-d.m. is never evaluated in Table 3. The paper also states it withholds PESQ/STOI/SI-SDR for DFN3/FSN+ because \"these values are lower and give a wrong representation,\" and the conclusion defers listening tests to future work. With a 100-example test set and no confidence intervals or paired significance on the DNSMOS difference, the dependence of the headline on one non-intrusive metric is the least secure link in the argument.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses single-channel low-latency speech enhancement in distant microphone scenarios (talker-to-mic distances of 5–10 m in large rooms such as conference rooms and theatres). It makes three main contributions: (i) a volume-based T60 sampling strategy for synthetic RIR generation grounded in architectural acoustics data; (ii) a systematic ANOVA-based study of dereverberation target design, specifically the offset before natural decay and the maximum residual T60 (T60max); (iii) an application of the resulting training pipeline to two real-time architectures (DFN3 and HSTN), reporting DNSMOS improvements over prior state-of-the-art models. Experiments use the DNS 2022 corpus, FRA-RIR and gpuRIR for synthetic RIRs, and a 100-example test set of real distant-microphone RIRs. The main quantitative findings are that volume-based T60 sampling improves PESQ/STOI/DNSMOS over unconstrained sampling (Table 2), that preserving early reflections with offsets up to 80 ms and a T60max of 300 ms is beneficial in distant scenarios (Section 4.2), and that DFN3-d.m. reaches an OVRL DNSMOS of 3.04 versus 2.77 for the original DFN3 (Table 3).","tokens_in":9467,"tokens_out":3549,"duration_ms":34403,"significance":"If the results are valid, the paper fills a genuine gap in the single-channel speech enhancement literature by systematically studying distant microphone scenarios and large rooms, which are underrepresented in prior benchmarks. The volume-based T60 sampling idea is a simple but useful correction to the common practice of independently sampling room dimensions and T60, and the ANOVA-based ablation of dereverberation targets is a structured and welcome empirical contribution. The paper also ships audio examples and supplementary material, which aids reproducibility. However, the headline state-of-the-art claim is not fully supported because the key comparison in Table 3 is confounded by multiple simultaneous changes, and the authors explicitly withhold intrusive metrics for two baselines. The lack of confidence intervals and listening tests further tempers the strength of the perceptual claims. These issues are fixable but require additional experiments and analysis.","major_comments":[{"comment":"The comparison between DFN3-d.m. and the original DFN3 is not an ablation of the proposed pipeline: the two models differ in training data (DNS corpus versus the new distant-mic RIR sets), in the dereverberation target ((5,500) versus (0,300)), and in the test set. Any of these changes could account for the 0.27 OVRL gap. Moreover, the authors explicitly withhold PESQ, STOI, and SI-SDR for DFN3 and FSN+ because \"these values are lower and give a wrong representation,\" which is selective reporting and prevents an apples-to-apples comparison. To support the state-of-the-art claim, the authors should either report all metrics for all models, or run controlled ablations that isolate each component (e.g., DFN3 trained on the standard corpus but with the new RIRs, or with the new target). At minimum, the text should clearly state that the comparison is a pipeline-level comparison and not a claim about model architecture alone.","section":"Section 4.3, Table 3"},{"comment":"The volume-T60 curve T60 = 0.145 ln(V) - 0.165 is central to the improvement reported in Table 2, but the manuscript provides no information about the curve-fit itself: how many data points were used, from which of references [24,25,26], what the fit residuals are, and what error bars or confidence intervals apply. The ±20% variation around the curve is also introduced without justification. If this curve does not represent the actual acoustics of conference rooms and theatres across the volume range 10^2 to 10^5 m^3, the observed gain in Table 2 may not generalize. Please provide the underlying data, fit statistics, and a rationale for the ±20% range, or validate the distribution against measured RIRs from a larger corpus.","section":"Section 2.1, Equation (1)"},{"comment":"The ANOVA-based conclusions about offset and T60max are systematic and well structured, but they rest on a synthetic gpuRIR test set and use DNSMOS as the sole endpoint. No confidence intervals or paired significance values are reported for the DNSMOS differences themselves (the asterisks in Figure 3 indicate pairwise t-test results but not the magnitude uncertainty). Section 5 states that listening tests are future work, which is a significant limitation for a task whose evaluation is inherently perceptual. To strengthen the practical recommendations (e.g., choosing (0,300) over (30,300)), the authors should report confidence intervals for the main effects in Figure 3, and ideally include a small listening experiment on the real test set used in Section 4.3.","section":"Sections 3.1 and 4.2"}],"minor_comments":[{"comment":"Typo: \"Nework Architectures\" should be \"Network Architectures\".","section":"Section 3.2"},{"comment":"The caption reads \"MOS MOS MOS (SIG) (BAK) (OVL)\", which is confusing because the header repeats MOS. Please revise the header to clearly separate the DNSMOS subscales.","section":"Table 2 caption"},{"comment":"The definition of N1 as \"the direct sound\" is ambiguous; specify whether N1 is the sample index of the direct-sound arrival or the end of the direct-sound region.","section":"Section 2.2, Equation (2)"},{"comment":"The text reports that in the small-room distant-microphone condition, \"(0, 300) and (30, 300) obtain the highest OVRL score of 2.86\", while later it says \"(30, 300) obtains the highest OVRL score of 3.02\" for the large room. These numbers are presented without a direct link to Figure 3 panels, which may confuse readers; please clarify which condition each value refers to and ensure consistency with Figure 3.","section":"Section 4.2"},{"comment":"Reference [24] (Harris handbook) lacks volume and page details; please complete the bibliographic information or use a consistent citation style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be of interest to the speech enhancement community and fits the scope of a signal processing conference/journal. The main concern is the selective reporting of metrics in the state-of-the-art comparison: the authors admit to withholding PESQ/STOI/SI-SDR for two baselines because those values are lower, which is problematic for fair comparison and may set a poor precedent. Additionally, the T60-volume curve fit lacks transparency. The ANOVA-based ablation is solid and could be a strong point if the other issues are addressed. I recommend major revision rather than rejection because the core ideas are defensible and the missing experiments are within reach."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere is my read on 2505.01338. The paper tackles a real gap: single-channel, low-latency speech enhancement for distant microphones (5–10 m, conference rooms/theatres), which the SE literature has mostly left alone. The most useful parts are the systematic study of the early-reflection offset and T60max in the windowed dereverberation target, and the simple volume–T60 constraint for generating synthetic RIRs. The offset/T60max work is careful: ANOVA with Bonferroni-corrected post-hoc tests, and the finding that talker-to-mic distance matters more than room size for choosing how much reverb to leave in the target is actionable and, as far as I can tell, holds. The volume–T60 curve is just an interpolation through three textbook data points, not a fitted model with error bars, but the points themselves are from architecture guidelines and the ±20% band is a reasonable engineering choice. Both pieces are worth having.\n\nThe soft spots are real. The headline claim in Table 3 — DFN3-d.m. at 3.04 OVRL versus original DFN3 at 2.77 — comes from models trained with different data and different targets ('0,300' versus '5,500'). That comparison does not isolate the pipeline's contribution. Worse, the paper openly says it withholds PESQ/STOI/SI-SDR for DFN3 and FullSubNet+ because those values are lower and would misrepresent performance. That is selective reporting, and it cuts against the credibility of the SOTA statement. The test set is only 100 examples from 37 RIRs, there are no confidence intervals or significance tests on the key DNSMOS differences, and no listening tests — the paper itself defers them to future work. The volume-based sampling gain in Table 2 (PESQ 2.08 to 2.17) is small and untested. None of this sinks the practical contribution, but it does mean the paper's strongest claim is not yet supported.\n\nThe citation pattern is fine, and the supplementary material with audio examples is a plus. This paper is for anyone training enhancement models for lecture capture, stage sound, or hearing aids in large rooms; the window-design guidance alone is worth reading.\n\nRecommendation: send it to peer review. It has a real idea and a systematic ablation that deserves referee time, but require a revision that reports all metrics for all baselines, adds confidence intervals or significance tests, and — ideally — retrains the baselines on the same data and target so the pipeline's effect is isolated. A listening test would be ideal, but at minimum the DNSMOS-only conclusions should be softened.\n\nBest","headline":"The paper's practical guidance on how much reverb to leave in the target is solid and useful, but its headline SOTA claim rests on a confounded comparison and selective metrics.","tokens_in":9950,"tokens_out":6646,"would_cite":true,"duration_ms":59457,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper establishes that single-channel, low-latency speech enhancement is feasible when the microphone is 5–10 m from the talker in large rooms, provided the training pipeline simulates room acoustics with a volume-dependent…","keywords":["speech enhancement","dereverberation","low latency","distant microphone","room impulse response","T60","DNSMOS","single channel"],"falsifier":"Take a set of real conference rooms and theatres, measure their volumes and reverberation times, retrain the model with those measured pairs replacing the fitted curve, and compare DNSMOS OVRL on a fixed distant-microphone test set; if the curve-trained model does not beat independent random sampling, the proposed volume–$T_{60}$ coupling is not doing the work.","tokens_in":8971,"feed_emoji":"🎙️","tokens_out":8950,"duration_ms":80877,"temperature":0.7,"pith_summary":"This paper asks whether a single microphone can clean up speech recorded 5 to 10 meters away in large rooms like conference halls and theatres, in real time with low latency. It argues yes: the bottleneck is not the architecture but the training data. Simulating room impulse responses with a volume-dependent reverberation time ($T_{60} = 0.145\\ln(V) - 0.165$, ±20%) produces more realistic training conditions, and choosing a dereverberation target that keeps the first 300 ms of decay rather than removing all reverberation gives the best perceived quality. On a distant-microphone test set the adapted model raises DNSMOS OVRL from 2.77 to 3.04, which would matter for lecture capture, drama, and stage acoustics.","feed_headline":"Distant-mic speech enhancement works at 10 meters","feed_subtitle":"Volume-aware room simulation plus a 300 ms reverb-decay target lifts DNSMOS OVRL from 2.77 to 3.04.","key_machinery":"The argument is carried by two data-generation choices. First, volume-based $T_{60}$ sampling: a curve-fit $T_{60} = 0.145\\ln(V) - 0.165$, drawn from architectural acoustics guidelines and given ±20% random variation, replaces independent random sampling of room dimensions and reverberation time. Second, a windowing rule for the dereverberation target: the target signal is clean speech convolved with the original room impulse response multiplied by a gain window that is unity up to a chosen offset and then decays to −60 dB at a rate set by the target $T_{60}^{\\max}$. The paper scans offsets from 0 to 80 ms and target decays of no-decay, 150, 300, and 500 ms, and selects a 0 ms offset with a 300 ms decay as the training configuration for the final comparisons.","core_discovery":"The central discovery is that realistic room impulse response simulation matters more than model choice for distant-microphone speech enhancement. Randomly sampling $T_{60}$ and room volume independently creates implausible combinations, and networks trained on those learn little new: far-microphone training without volume-based sampling performs about the same as close-microphone training. Coupling $T_{60}$ to volume through the fitted log curve improves PESQ from 2.08 to 2.17 and OVRL DNSMOS from 2.64 to 2.69. The dereverberation target also matters: at distances up to 0.5 m the network should predict the direct sound, but at 5–10 m preserving early reflections (up to 80 ms) and decaying the impulse response to a 300 ms $T_{60}$ target is best, with OVRL scores reaching 2.86–3.02 depending on room. The resulting adapted model scores 3.04 OVRL versus 2.77 for the baseline.","pith_inferences":["A natural extension is to re-fit the volume–$T_{60}$ curve to measured pairs from real conference rooms and theatres; if the fitted curve is replaced by measured data, the paper's comparisons would directly test whether the volume coupling is necessary.","A testable extension is to train one network with two dereverberation targets—direct sound for close microphones and a 300 ms decay for distant microphones—using distance as a conditioning input, since the paper finds distance, not room size, is the controlling variable.","Because the paper relies on DNSMOS and a limited set of measured impulse responses, formal listening tests would clarify whether the preserved early reflections sound natural to human listeners."],"forward_implications":["Training pipelines for large-room enhancement should sample reverberation time conditionally on room volume; independent sampling yields near-zero gain over close-microphone training.","The dereverberation target should be matched to microphone distance: close microphones benefit from predicting the direct sound, while distant microphones benefit from preserving 30–80 ms of early reflections.","Keeping 300 ms of residual reverberation gives the best overall quality at distance; extending the target to 500 ms degrades background suppression, especially in large rooms.","The pipeline works at both 40 ms and 20 ms latency, with the lower-latency model also improving over published baselines on the distant-microphone test set."],"supporting_citations":[{"why":"These three references supply the room-type $T_{60}$-versus-volume data points from which the paper fits the curve $T_{60} = 0.145\\ln(V) - 0.165$.","marker":"[24, 25, 26]"},{"why":"It introduces the windowing approach that shapes room impulse responses to a maximum decay target, the basis of the dereverberation targets tested here.","marker":"[19]"},{"why":"It presents the DeepFilterNet3 model and its (5 ms, 500 ms) dereverberation target, the baseline the adapted distant-microphone model is compared against.","marker":"[20]"},{"why":"It provides the GPU-accelerated room impulse response simulator used for synthetically generated test RIRs with controlled room dimensions and distances.","marker":"[22]"},{"why":"It provides the randomized image-source method used to generate distant-microphone training room impulse responses.","marker":"[30]"},{"why":"It supplies measured room impulse responses used in the distant-microphone test set.","marker":"[31]"},{"why":"It supplies additional measured room impulse responses for the same test set.","marker":"[32]"},{"why":"It defines DNSMOS, the non-intrusive neural metric used to measure overall quality, signal quality, and background suppression.","marker":"[35]"},{"why":"It supplies the clean speech, noise, and room impulse response training data, with the real-time low-latency evaluation constraints adopted by the paper.","marker":"[6]"}],"fun_headline_variants":["Volume-aware simulation boosts distant-mic speech quality","Early reflections key for far-field speech dereverberation","Better room simulation improves distant-mic speech enhancement","Volume-aware room sim lifts DNSMOS to 3.04 at 10 m","Distant-mic speech gains from realistic reverb simulation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's gains rest on a curve linking room volume to reverberation time, fitted to three published guidelines without reported error bars; if real conference and theatre acoustics deviate from that curve, the measured improvements may not transfer.","fun_headline_variants_meta":{"raw":{"variants":["Volume-aware simulation boosts distant-mic speech quality","Early reflections key for far-field speech dereverberation","Better room simulation improves distant-mic speech enhancement","Volume-aware room sim lifts DNSMOS to 3.04 at 10 m","Distant-mic speech gains from realistic reverb simulation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000641,"raw_usage":{"total_tokens":2963,"prompt_tokens":970,"completion_tokens":1993,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":586,"completion_tokens_details":{"reasoning_tokens":1911}},"tokens_in":586,"tokens_out":1993,"duration_ms":12216,"temperature":1.0,"reasoning_tokens":1911,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:20:14.571415+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of real conference rooms and theatres, measure their volumes and reverberation times, retrain the model with those measured pairs replacing the fitted curve, and compare DNSMOS OVRL on a fixed distant-microphone test set; if the curve-trained model does not beat independent random sampling, the proposed volume–$T_{60}$ coupling is not doing the work.","supporting_citations":[{"cited_title":"Speech enhancement with multichannel wiener filter techniques in multimicrophone binaural hearing aids,","cited_arxiv_id":null,"evidence_quote":"It introduces the windowing approach that shapes room impulse responses to a maximum decay target, the basis of the dereverberation targets tested here."},{"cited_title":"An analysis of environment, microphone and data simulation mismatches in robust speech recognition,","cited_arxiv_id":null,"evidence_quote":"It presents the DeepFilterNet3 model and its (5 ms, 500 ms) dereverberation target, the baseline the adapted distant-microphone model is compared against."},{"cited_title":"Fullsubnet+: Channel atten- tion fullsubnet with complex spectrograms for speech enhancement,","cited_arxiv_id":null,"evidence_quote":"It provides the GPU-accelerated room impulse response simulator used for synthetically generated test RIRs with controlled room dimensions and distances."},{"cited_title":"1960, McGraw-Hill New York, 1957","cited_arxiv_id":null,"evidence_quote":"It provides the randomized image-source method used to generate distant-microphone training room impulse responses."},{"cited_title":"Acoustics and architecture in italian catholic churches,","cited_arxiv_id":null,"evidence_quote":"It supplies measured room impulse responses used in the distant-microphone test set."},{"cited_title":"Influence of proportion towards speech intelligibility in mosque’s praying hall,","cited_arxiv_id":null,"evidence_quote":"It supplies additional measured room impulse responses for the same test set."},{"cited_title":"A pitch tracking corpus with evaluation on multipitch tracking sce- nario.,","cited_arxiv_id":null,"evidence_quote":"It defines DNSMOS, the non-intrusive neural metric used to measure overall quality, signal quality, and background suppression."},{"cited_title":"ICASSP 2022 deep noise suppression challenge,","cited_arxiv_id":null,"evidence_quote":"It supplies the clean speech, noise, and room impulse response training data, with the real-time low-latency evaluation constraints adopted by the paper."}],"review_version":1}