REVIEW 2 major objections 1 minor 52 references
High-fidelity room-acoustic simulations reduce word error rates by up to 38 percent compared to geometrical methods.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-07-01 03:33 UTC pith:CWQFCPKP
load-bearing objection Higher-fidelity simulations improve enhancement results in this setup, but the datasets must be shown to differ only in the simulation method itself. the 2 major comments →
Improving multichannel speech enhancement through accurate room-acoustic simulations
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Augmenting training data for SpatialNet with high-fidelity room-acoustic simulations that use advanced acoustic modelling and hybrid wave-based plus geometrical methods produces an up to 38 percent relative reduction in median word error rate on measured evaluation data, relative to training on lower-fidelity geometrical-acoustics data alone.
What carries the argument
The fidelity level of room-acoustic simulation methods (geometrical acoustics versus wave-based and hybrid modelling) used to generate multichannel training data for speech enhancement.
Load-bearing premise
The measured evaluation recordings come from acoustic conditions that are materially better matched by the high-fidelity simulations than by the geometrical simulations, with dataset size, content, and other training variables held constant.
What would settle it
Generate new training sets using the same high-fidelity method but for rooms whose measured impulse responses differ substantially from the evaluation rooms, retrain the models, and check whether the word-error-rate advantage over geometrical simulations disappears.
If this is right
- High-fidelity simulation data produces models that generalize better to real acoustic conditions than geometrical data alone.
- The measured gains appear directly in word error rate when the enhanced signals are passed to a downstream recognizer.
- Hybrid wave-geometrical simulations can be used as a practical route to higher fidelity without full wave-based computation for every source.
Where Pith is reading between the lines
- The same fidelity advantage may appear in related tasks such as source separation or dereverberation that also rely on simulated room data.
- Audio machine-learning pipelines may benefit more from upgrading simulation engines than from further model architecture changes.
- Mismatch between simulated and real room acoustics remains a controllable bottleneck that can be reduced by raising simulation accuracy.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript examines how room-acoustic simulation fidelity affects multichannel speech enhancement by training SpatialNet on datasets augmented via geometrical acoustics, wave-based methods, and a hybrid approach, then evaluating on measured data. It reports that the high-fidelity dataset yields up to 38% relative reduction in median word error rate versus the lower-fidelity alternatives.
Significance. If the result holds, the work would demonstrate that higher-fidelity acoustic modeling in training-data augmentation can produce measurable gains on real measured recordings, with direct implications for data-generation pipelines in deep-learning speech enhancement. The choice to evaluate on held-out measured data (rather than simulated test conditions) is a methodological strength that increases practical relevance.
major comments (2)
- [Section 3] Section 3 (dataset construction): The central claim attributes the WER reduction to simulation fidelity alone. This requires explicit confirmation that the high-fidelity, geometrical, and hybrid datasets are identical in every other respect (room count and geometries, absorption coefficients, source/receiver positions, speech and noise waveforms, total training duration, and augmentation pipelines). The manuscript provides no such statement, so the performance delta cannot yet be isolated to fidelity.
- [Results section] Results section and abstract: The reported 'up to 38 % relative reduction in median word error rate' supplies no statistical significance tests, error bars, number of evaluation rooms or utterances, or controls for room diversity and confounding variables. These omissions prevent assessment of whether the quantitative claim is robust.
minor comments (1)
- [Abstract] Abstract: 'advanced acoustic modelling' is imprecise; replace with the concrete methods (e.g., FDTD, BEM) used for the high-fidelity and hybrid cases.
Simulated Author's Rebuttal
We thank the referee for the constructive comments, which help clarify the isolation of simulation fidelity and strengthen the statistical reporting of our results. We address each major comment below.
read point-by-point responses
-
Referee: [Section 3] Section 3 (dataset construction): The central claim attributes the WER reduction to simulation fidelity alone. This requires explicit confirmation that the high-fidelity, geometrical, and hybrid datasets are identical in every other respect (room count and geometries, absorption coefficients, source/receiver positions, speech and noise waveforms, total training duration, and augmentation pipelines). The manuscript provides no such statement, so the performance delta cannot yet be isolated to fidelity.
Authors: We agree that an explicit statement is required to rigorously isolate the effect of simulation fidelity. All three datasets were generated from the same underlying room geometries, absorption coefficients, source/receiver positions, speech and noise waveforms, total training duration, and post-processing augmentation pipelines, differing solely in the acoustic simulation engine employed. In the revised manuscript we will insert a dedicated paragraph in Section 3 that explicitly enumerates these shared parameters. revision: yes
-
Referee: [Results section] Results section and abstract: The reported 'up to 38 % relative reduction in median word error rate' supplies no statistical significance tests, error bars, number of evaluation rooms or utterances, or controls for room diversity and confounding variables. These omissions prevent assessment of whether the quantitative claim is robust.
Authors: We acknowledge that the current presentation lacks the requested statistical details. The evaluation was performed on a fixed set of held-out measured rooms and utterances; in the revision we will report the exact counts, add bootstrap-derived error bars or confidence intervals around the median WER values, and include a statistical significance test (Wilcoxon signed-rank) comparing the high-fidelity model against the baselines. We will also explicitly state the controls used to ensure room diversity. revision: yes
Circularity Check
No circularity; empirical comparison on held-out measured data
full rationale
The paper reports an empirical comparison: models are trained on datasets generated via different room-acoustic simulation methods (geometrical, wave-based, hybrid) and evaluated on separate measured recordings. The reported WER improvement is a direct experimental outcome, not a derivation that reduces to its inputs by construction. No equations, fitted parameters renamed as predictions, self-citation load-bearing steps, or ansatz smuggling appear in the abstract or described content. The evaluation uses external measured data, satisfying the self-contained benchmark criterion.
Axiom & Free-Parameter Ledger
axioms (1)
- domain assumption Wave-based simulations more accurately reproduce real-room acoustics than geometrical-acoustics simulations.
read the original abstract
Room-acoustic simulations are widely used to augment training data for deep-learning-based speech enhancement. While most pipelines rely on simplified geometrical acoustics, wave-based approaches offer greater physical accuracy. In this work, we examine how simulation fidelity affects multichannel speech enhancement performance. To this end, we train SpatialNet on datasets augmented with different room-acoustic simulation methods and evaluate the resulting models on measured data. We compare lower-fidelity datasets based on geometrical acoustics with a high-fidelity dataset using advanced acoustic modelling and a hybrid combination of wave-based and geometrical acoustics simulations. Training on the high-fidelity dataset results in an up to 38 % relative reduction in median word error rate compared to the lower-fidelity alternatives. These results show that augmentation with high-fidelity room-acoustic simulations directly translates into improved multichannel speech enhancement performance.
Figures
Reference graph
Works this paper leans on
-
[1]
However, this interaction frequently occurs in acous- tically complex environments
Introduction As smart devices become deeply embedded in everyday life, natural language user interfaces (NLUIs) have emerged as a key paradigm for enabling seamless and intuitive human-device in- teraction. However, this interaction frequently occurs in acous- tically complex environments. Across homes, public spaces, and in-vehicle settings, the intellig...
-
[2]
Improving multichannel speech enhancement through accurate room-acoustic simulations
Simulation paradigms for acoustic modelling In the context of data augmentation for speech enhancement, different simulation methods/fidelity may lead to different spa- arXiv:2606.31552v1 [eess.AS] 30 Jun 2026 tial and spectral characteristics in the generated RIRs. Room- acoustic simulation methods are commonly divided into geo- metrical acoustics (GA) a...
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[3]
Experiment setup 3.1. Network configuration We train SpatialNet [5] to investigate the relationship be- tween room-acoustic simulation accuracy and neural network downstream performance. Inspired by the Conformer ar- chitecture [34], SpatialNet integrates convolutional modelling and multi-head self-attention [35] for end-to-end multichan- nel speech enhan...
work page 2023
-
[4]
Results We enhance and transcribe all utterances of LibriCSS-EM6 with each of the trained SpatialNet models and calculate the corre- sponding word error rates (WERs). Figure 2 shows the me- dian WER of all models under the investigated overlap con- ditions, with bootstrapped 95 % confidence intervals. Across all overlap conditions, the model trained with ...
-
[5]
Conclusion This paper investigated different room-acoustic simulation paradigms for augmenting training data in multichannel speech enhancement and analysed the impact of simulation fidelity on downstream performance. We observe a relative reduction in median word error rate of up to 38 % when the training dataset is augmented using high-fidelity room-aco...
-
[6]
Generative AI use disclosure The authors acknowledge the use of ChatGPT 5.2 (accessed in February 2026) to rephrase some sentences for clarity, polish the manuscript’s language, and help to arrange and format ta- bles and figures. All AI-assisted content was reviewed and re- vised by the authors to ensure accuracy and clarity of meaning
work page 2026
-
[7]
A fully convolutional neural network for speech enhancement,
S. R. Park and J. W. Lee, “A fully convolutional neural network for speech enhancement,” inProc. Interspeech, Stockholm, Sweden, 2017, pp. 1993–1997
work page 2017
-
[8]
Speech enhancement with LSTM recurrent neural networks and its application to noise-robust ASR,
F. Weninger, H. Erdogan, S. Watanabe, E. Vincent, J. Le Roux, J. R. Hershey, and B. Schuller, “Speech enhancement with LSTM recurrent neural networks and its application to noise-robust ASR,” inProc. 12th Int. Conf. Latent Var. Anal. Signal Separa- tion, Liberec, Czech Republic, 2015, pp. 91–99
work page 2015
-
[9]
Supervised speech separation based on deep learning: An overview,
D. Wang and J. Chen, “Supervised speech separation based on deep learning: An overview,”IEEE/ACM Trans. Audio Speech Lang. Process., vol. 26, no. 10, pp. 1702–1726, 2018
work page 2018
-
[10]
H. Schröter, A. N. Escalante-B., T. Rosenkranz, and A. Maier, “DeepFilterNet: A low complexity speech enhancement frame- work for full-band audio based on deep filtering,” inProc. IEEE Int. Conf. Acoust. Speech Signal Process. (ICASSP), Singapore, Singapore, 2022, pp. 7407–7411
work page 2022
-
[11]
C. Quan and X. Li, “SpatialNet: Extensively learning spatial in- formation for multichannel joint speech separation, denoising and dereverberation,”IEEE/ACM Trans. Audio Speech Lang. Process., vol. 32, pp. 1310–1323, 2024
work page 2024
-
[12]
R. Haeb-Umbach, T. Nakatani, M. Delcroix, C. Boeddeker, and T. Ochiai, “Microphone array signal processing and deep learn- ing for speech enhancement: Combining model-based and data- driven approaches to parameter estimation and filtering,”IEEE Signal Processing Magazine, vol. 41, no. 6, pp. 12–23, 2024
work page 2024
-
[13]
A time-frequency fusion model for multi-channel speech enhancement,
X. Zeng, S. Xu, and M. Wang, “A time-frequency fusion model for multi-channel speech enhancement,”EURASIP J. Au- dio Speech Music Process., vol. 2024, no. 1, pp. 47:1–12, 2024
work page 2024
-
[14]
ICASSP 2023 Deep Noise Suppression Challenge,
H. Dubey, A. Aazami, V . Gopal, B. Naderi, S. Braun, R. Cutler, A. Ju, M. Zohourian, M. Tang, M. Golestaneh, and R. Aichner, “ICASSP 2023 Deep Noise Suppression Challenge,”IEEE Open J. Signal Process., vol. 5, pp. 725–737, 2024
work page 2023
-
[15]
A study on data augmentation of reverberant speech for robust speech recognition,
T. Ko, V . Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” inProc. IEEE Int. Conf. Acoust. Speech Sig- nal Process. (ICASSP), New Orleans, LA, USA, 2017, pp. 5220– 5224
work page 2017
-
[16]
C. Kim, A. Misra, K. Chin, T. Hughes, A. Narayanan, T. Sainath, and M. Bacchiani, “Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in Google Home,” inProc. Interspeech, Stockholm, Sweden, 2017, pp. 379–383
work page 2017
-
[17]
WHAMR!: Noisy and reverberant single-channel speech sep- aration,
M. Maciejewski, G. Wichern, E. McQuinn, and J. Le Roux, “WHAMR!: Noisy and reverberant single-channel speech sep- aration,” inProc. IEEE Int. Conf. Acoust. Speech Signal Pro- cess. (ICASSP), Online conference, 2020, pp. 696–700
work page 2020
-
[18]
A study on more realistic room simulation for far-field key- word spotting,
E. Bezzam, R. Scheibler, C. Cadoux, and T. Gisselbrecht, “A study on more realistic room simulation for far-field key- word spotting,” inProc. Asia-Pacific Signal Information Pro- cess. Assoc. Annual Summit Conf. (APSIPA ASC), Auckland, New Zealand, 2020, pp. 674–680
work page 2020
-
[19]
R. Arakawa, M. Parvaix, C. Lai, H. Erdogan, and A. Olwal, “Quantifying the effect of simulator-based data augmentation for speech recognition on augmented reality glasses,” inProc. IEEE Int. Conf. Acoust. Speech Signal Process. (ICASSP), Seoul, Korea, 2024, pp. 726–730
work page 2024
-
[20]
How to (virtually) train your speaker localizer,
P. Srivastava, A. Deleforge, A. Politis, and E. Vincent, “How to (virtually) train your speaker localizer,” inProc. Interspeech, Dublin, Ireland, 2023, pp. 1204–1208
work page 2023
-
[21]
E. Gusó, J. Luberadzka, U. Sayin, and X. Serra, “MB-RIRs: a synthetic room impulse response dataset with frequency- dependent absorption coefficients,” inProc. IEEE Workshop Appl. Signal Process. Audio Acoust. (WASPAA), Tahoe City, CA, USA, 2025, pp. 1–5
work page 2025
-
[22]
GW A: A large high-quality acoustic dataset for audio processing,
Z. Tang, R. Aralikatti, A. J. Ratnarajah, and D. Manocha, “GW A: A large high-quality acoustic dataset for audio processing,” in Proc. ACM SIGGRAPH Conf., Vancouver, BC, Canada, article no. 36, 2022
work page 2022
-
[23]
Overview of geometrical room acoustic modeling techniques,
L. Savioja and U. P. Svensson, “Overview of geometrical room acoustic modeling techniques,”J. Acoust. Soc. Am., vol. 138, no. 2, pp. 708–730, 2015
work page 2015
-
[24]
Calculating the acoustical room response by the use of a ray tracing technique,
A. Krokstad, S. Strøm, and S. Sørsdal, “Calculating the acoustical room response by the use of a ray tracing technique,”J. Sound Vib., vol. 8, no. 1, pp. 118–125, 1968
work page 1968
-
[25]
Fifteen years’ experience with computerized ray tracing,
——, “Fifteen years’ experience with computerized ray tracing,” Appl. Acoust., vol. 16, no. 4, pp. 291–312, 1983
work page 1983
-
[26]
Image method for efficiently sim- ulating small-room acoustics,
J. B. Allen and D. A. Berkley, “Image method for efficiently sim- ulating small-room acoustics,”J. Acoust. Soc. Am., vol. 65, no. 4, pp. 943–950, 1979
work page 1979
-
[27]
Extension of the image model to arbitrary polyhedra,
J. Borish, “Extension of the image model to arbitrary polyhedra,” J. Acoust. Soc. Am., vol. 75, no. 6, pp. 1827–1836, 1984
work page 1984
-
[28]
Computation of edge diffraction for more accurate room acoustics auralization,
R. R. Torres, U. P. Svensson, and M. Kleiner, “Computation of edge diffraction for more accurate room acoustics auralization,” J. Acoust. Soc. Am., vol. 109, no. 2, pp. 600–610, 2001
work page 2001
-
[29]
C. Schissler, R. Mehra, and D. Manocha, “High-order diffraction and diffuse reflections for interactive sound propagation in large environments,”ACM Trans. Graph., vol. 33, no. 4, article no. 39, 2014
work page 2014
-
[30]
Computer simulations in room acoustics: Con- cepts and uncertainties,
M. V orländer, “Computer simulations in room acoustics: Con- cepts and uncertainties,”J. Acoust. Soc. Am., vol. 133, no. 3, pp. 1203–1213, 2013
work page 2013
-
[31]
A round robin on room acoustical simulation and auralization,
F. Brinkmann, L. Aspöck, D. Ackermann, S. Lepa, M. V orländer, and S. Weinzierl, “A round robin on room acoustical simulation and auralization,”J. Acoust. Soc. Am., vol. 145, no. 4, pp. 2746– 2760, 2019
work page 2019
-
[32]
Finite difference and finite volume methods for wave-based modelling of room acoustics,
B. Hamilton, “Finite difference and finite volume methods for wave-based modelling of room acoustics,” Ph.D. dissertation, University of Edinburgh, 2016
work page 2016
-
[33]
Time domain room acoustic simulations using the spectral element method,
F. Pind, A. P. Engsig-Karup, C.-H. Jeong, J. S. Hesthaven, M. S. Mejling, and J. Strømann-Andersen, “Time domain room acoustic simulations using the spectral element method,”J. Acoust. Soc. Am., vol. 145, no. 6, pp. 3299–3310, 2019
work page 2019
-
[34]
A review of finite element methods for room acous- tics,
A. G. Prinn, “A review of finite element methods for room acous- tics,”Acoustics, vol. 5, no. 2, pp. 367–395, 2023
work page 2023
-
[35]
Finite-difference time-domain simulation of low-frequency room acoustic problems,
D. Botteldooren, “Finite-difference time-domain simulation of low-frequency room acoustic problems,”J. Acoust. Soc. Am., vol. 98, no. 6, pp. 3302–3308, 1995
work page 1995
-
[36]
Room acoustics modelling in the time-domain with the nodal discontin- uous Galerkin method,
H. Wang, I. Sihar, R. Pagán Muñoz, and M. Hornikx, “Room acoustics modelling in the time-domain with the nodal discontin- uous Galerkin method,”J. Acoust. Soc. Am., vol. 145, no. 4, pp. 2650–2663, 2019
work page 2019
-
[37]
F. Pind, C.-H. Jeong, A. P. Engsig-Karup, J. S. Hesthaven, and J. Strømann-Andersen, “Time-domain room acoustic simulations with extended-reacting porous absorbers using the discontinuous Galerkin method,”J. Acoust. Soc. Am., vol. 148, no. 5, pp. 2851– 2863, 2020
work page 2020
-
[38]
Massively par- allel nodal discontinous Galerkin finite element method simulator for room acoustics,
A. Melander, E. Strøm, F. Pind, A. P. Engsig-Karup, C.-H. Jeong, T. Warburton, N. Chalmers, and J. S. Hesthaven, “Massively par- allel nodal discontinous Galerkin finite element method simulator for room acoustics,”Int. J. High Perform. Comput. Appl., vol. 38, no. 3, pp. 154–174, 2024
work page 2024
-
[39]
S. Siltanen, T. Lokki, and L. Savioja, “Rays or waves? under- standing the strengths and weaknesses of computational room acoustics modeling techniques,” inProc. Int. Symp. Room Acoust. (ISRA), Melbourne, Australia, 2010
work page 2010
-
[40]
Conformer: Convolution-augmented transformer for speech recognition,
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y . Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y . Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in Proc. Interspeech, Shanghai, China, 2020, pp. 5036–5040
work page 2020
-
[41]
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” inProc. 31st Conf. Neural Inf. Process. Syst. (NIPS), Long Beach, CA, USA, 2017
work page 2017
-
[42]
gpuRIR: A python library for room impulse response simulation with GPU accelera- tion,
D. Diaz-Guerra, A. Miguel, and J. R. Beltran, “gpuRIR: A python library for room impulse response simulation with GPU accelera- tion,”Multimed. Tools Appl., vol. 80, no. 4, pp. 5653–5671, 2021
work page 2021
-
[43]
Continuous speech separation: dataset and analysis,
Z. Chen, T. Yoshioka, L. Lu, T. Zhou, Z. Meng, Y . Luo, J. Wu, X. Xiao, and J. Li, “Continuous speech separation: dataset and analysis,” inProc. IEEE Int. Conf. Acoust. Speech Signal Pro- cess. (ICASSP), Online conference, 2020
work page 2020
-
[44]
Lib- rispeech: An ASR corpus based on public domain audio books,
V . Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Lib- rispeech: An ASR corpus based on public domain audio books,” inProc. IEEE Int. Conf. Acoust. Speech Signal Pro- cess. (ICASSP), South Brisbane, QLD, Australia, 2015, pp. 5206– 5210
work page 2015
-
[45]
G. Götz, S. J. Schlecht, and V . Pulkki, “A dataset of higher-order Ambisonic room impulse responses and 3D models measured in a room with varying furniture,” inProc. Int. Conf. Immersive 3D Audio (I3DA), Online conference, 2021
work page 2021
-
[46]
T. McKenzie, L. McCormack, and C. Hold, “Dataset of spa- tial room impulse responses in a variable acoustics room for six degrees-of-freedom rendering and analysis,”arXiv preprint arXiv:2111.11882, 2021
-
[47]
K. Kinoshita, M. Delcroix, T. Yoshioka, T. Nakatani, E. Habets, R. Haeb-Umbach, V . Leutnant, A. Sehr, W. Kellermann, R. Maas, S. Gannot, and B. Raj, “The reverb challenge: A common evalu- ation framework for dereverberation and recognition of reverber- ant speech,” inProc. IEEE Workshop Appl. Signal Process. Audio Acoust. (WASPAA), New Paltz, NY , USA, 2013
work page 2013
-
[48]
Generating nonstation- ary multisensor signals under a spatial coherence constraint,
E. A. P. Habets, I. Cohen, and S. Gannot, “Generating nonstation- ary multisensor signals under a spatial coherence constraint,”J. Acoust. Soc. Am., vol. 124, no. 5, pp. 2911–2917, 2008
work page 2008
-
[49]
The Kaldi speech recog- nition toolkit,
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlí ˇcek, Y . Qian, P. Schwarz, J. Silovský, G. Stemmer, and K. Veselý, “The Kaldi speech recog- nition toolkit,” inProc. IEEE Workshop Autom. Speech Recognit. Underst. (ASRU), Waikoloa Village, HI, USA, 2011
work page 2011
-
[50]
PyKaldi2: Yet an- other speech toolkit based on Kaldi and PyTorch,
L. Lu, X. Xiao, Z. Chen, and Y . Gong, “PyKaldi2: Yet an- other speech toolkit based on Kaldi and PyTorch,”arXiv preprint arXiv:1907.05955, 2019
-
[51]
Speech recognition with deep recurrent neural networks,
A. Graves, A. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” inProc. IEEE Int. Conf. Acoust. Speech Signal Process. (ICASSP), Vancouver, BC, Canada, 2013, pp. 6645–6649
work page 2013
-
[52]
Maxi- mum mutual information estimation of hidden markov model pa- rameters for speech recognition,
L. R. Bahl, P. F. Brown, P. V . de Souza, and R. L. Mercer, “Maxi- mum mutual information estimation of hidden markov model pa- rameters for speech recognition,” inProc. IEEE Int. Conf. Acoust. Speech Signal Process. (ICASSP), Tokyo, Japan, 1986, pp. 49–52
work page 1986
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.