REVIEW 4 major objections 4 minor 203 references
Deep, data-driven modeling of room acoustics: literature review and research perspectives
T0 review · 4 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This review argues that deep, data-driven room acoustics models fall on the same physics-versus-data and geometric-versus-wave axes as classical models, and that boundary-aware networks may close the gap between the two deep-learning camps.
desk verdict A useful two-axis review of deep room-acoustics modeling whose 'gap' between geometry-based and wave-based models is an interesting hypothesis that needs sturdier support than the author's own submitted paper. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a two-axis classification frame plus a theoretical bridge. The frame sorts room acoustics models, before and after deep learning, by whether they are physical or data-driven and by whether they assume geometric or wave-based sound behavior; the paper uses it to show that deep models with physical priors split into geometry-based and wave-based families, with an empty region between them. The bridge is the boundary integral equation, a wave-based formulation of the interior sound field on the room's boundary whose solution asymptotically reduces to geometric acoustics, so it inherently connects the two regimes. The proposed mechanism for crossing the gap is the PIBI-Net, a physics-informed boundary integral network that folds boundary geometry and material properties into both the model structure and the training strategy, so the network learns with the room's surface rather than treating it as an external condition. The framework's work is to convert a scattered literature into a structured map, and the boundary-integral/PIBI-Net pairing is what makes the map's central research direction concrete.
What would settle it
Train a boundary-informed network and an otherwise identical network with all boundary inputs removed on the same corpus of rooms; if the boundary-free network matches it on sound field reconstruction accuracy across frequencies, source positions, and unseen geometries, the claim that boundary information is the key bridge between geometric and wave-based deep models would be falsified.
Extended reading notes
Core claim
The paper's central claim is that deep, data-driven room acoustics models inherit the same conceptual structure as classical physical and data-driven models, and should be classified along the same two axes: physical versus data-driven and geometric versus wave-based. It reviews the field to show that most deep models, borrowed from speech and image processing, lack the space-time structure of acoustic wave propagation, while recent models that include either geometric or wave-based priors have produced their clearest successes in sound field reconstruction. Surveying these, the paper finds an apparent gap: geometry-based deep models and wave-based deep models have developed largely separately, with no deep counterpart to the classical boundary-integral models that already sit between the two regimes. Because the wave-based boundary integral equation asymptotically reduces to geometric acoustics, the review argues that the key to combining geometric and wave-based deep models may be to include boundary information in both the model structure and the training strategy, and it points to physics-informed boundary integral networks as a first promising step in that direction.
Load-bearing premise
The review's load-bearing premise is that the studies it surveys and the two-axis scheme it draws are representative enough to make the gap between ray-based and wave-based deep models real, and that the single boundary-integral-network result generalizes beyond its own experiments.
Editorial extensions
If this is right
- Deep models that ignore wave-propagation structure will keep underperforming on tasks that require room impulse responses or full sound fields, where time delays and space-time relations are the essence.
- A unified class of boundary-informed networks could combine the efficiency of geometric models with the accuracy of wave-based models for sound field reconstruction.
- Sound field reconstruction, not parameter estimation or enhancement, is where physics-informed deep models are currently proving themselves.
- Progress hinges on datasets that reconcile realistic audio scenes, with moving and directional sources and microphones, with accurate labeling of source and receiver positions.
- Understanding why nested nonlinear networks work for a process traditionally modeled as linear and time-invariant remains an open prerequisite for deliberate architecture design.
Reading between the lines
- Editorial inference: the boundary-bridge recipe likely transfers to other wave-physics inverse problems, where ray-based and full-wave models could be reconciled by networks trained with boundary integral losses.
- Editorial inference: the apparent gap may be partly a chronological artifact of a young field; as boundary-informed models mature, they could absorb both geometry-based and wave-based approaches rather than remain a third category.
- Editorial inference: a controlled benchmark with identical rooms and data, comparing PIBI-Nets against pure geometry-conditioned and pure wave-equation-regularized networks, would quantify when boundary information is what actually improves reconstruction.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a literature review of deep, data-driven room acoustics modeling, presented at the 11th Convention of the European Acoustics Association. It introduces a two-dimensional classification scheme along physical/data-driven and geometric/wave-based axes, applies this scheme to both traditional room acoustics models and deep learning (DL) models, and reviews two broad categories of DL models: purely data-driven models and models with geometric or wave-based physical priors. The paper closes with three research perspectives: the need for larger and better labeled datasets, the need to understand why nesting and nonlinearity in deep networks benefit room acoustics modeling, and the claim that there is an apparent gap between geometry-based and wave-based deep room acoustics models. The author suggests that including boundary information in both model structure and training, as exemplified by physics-informed boundary integral networks (PIBI-Nets), may be the key to combining geometric and wave-based information.
Significance. If the proposed classification and the claimed geometry/wave gap are accepted, the paper provides a useful conceptual framework for organizing a rapidly growing literature, and its research perspectives point to concrete directions for future work. The review is broad, with a substantial reference list, and it makes an explicit effort to connect traditional acoustics modeling concepts to the deep learning literature. Strengths include the structured overview in Figure 1, the clear separation of purely data-driven models from models with physical priors, and the honest use of hedged language such as 'apparent gap' and 'may lie.' However, the central research-perspective claim depends on a non-systematically selected literature and on an unpublished, author-authored model description, so the significance of the paper currently rests on a conjecture rather than on demonstrated evidence.
major comments (4)
- [Section 3, Figure 1] The claimed gap between geometry-based and wave-based deep learning models is asserted on the basis of a literature review whose selection is not described: the paper provides no search protocol, inclusion/exclusion criteria, or explicit category-assignment rules. Since the emptiness of the geometry+wave cell in Figure 1 is the central observation from which the key research perspective is derived, this absence is load-bearing. The authors should either provide a systematic methodology (e.g., search databases, search terms, screening criteria, and a table of the categorized papers) or present the gap explicitly as a tentative observation from a curated selection, with the corresponding limitation stated.
- [Section 3, reference [199]] The PIBI-Net model is cited as the 'first and promising result' in the direction of combining geometric and wave-based information, but reference [199] is a submitted, not-yet-published paper by the author, and the text provides only a single-sentence description with no architectural detail, experimental setup, or quantitative results. Because this model is the only occupant of the bridging cell in Figure 1, the reader cannot evaluate whether it actually combines geometric and wave-based behavior. The manuscript should describe the model and its reported results in enough detail to support the claim, cite a publicly available version, or explicitly label the bridge as a conjecture that awaits evidence.
- [Section 2.3] The classification of existing models as geometry-based or wave-based is too coarse to support the gap claim. For instance, INRAS is described as including boundary geometry and NACF as including material properties, and many PINN formulations are trained with boundary-condition losses; these could plausibly be viewed as already combining geometric and wave-based information. The paper does not justify why these models are excluded from the geometry+wave cell. Without clear criteria for what counts as 'geometry-based,' 'wave-based,' and 'boundary-based,' the asserted gap may be an artifact of the author's categorization rather than a property of the literature.
- [Section 3] The argument that the asymptotic geometric interpretation of the boundary integral equation [37] implies that boundary information in a DL model will combine geometric and wave-based priors is an analogy at the level of the continuous operator. It is not evidence that a deep network trained with a boundary-integral loss will exhibit such hybrid behavior, since the optimization dynamics and approximation properties of the network are not governed by the stationary-phase approximation. The paper should separate the mathematical property of the BIE from the empirical hypothesis about PIBI-Nets, and should note that the latter remains untested in publicly available literature.
minor comments (4)
- [Section 2.2] The manuscript contains a formatting artifact consisting of a repeated paragraph and the stray header 'van Waterschoot Part B2 DIORAMA 3' in the middle of Section 2.2; this should be removed.
- [Abstract and Section 2] The abstract states that 'the majority' of deep data-driven room acoustics models lack intrinsic space-time structure, but the paper does not quantify this. Please provide a rough count from the reviewed papers or soften the claim to 'many' or 'most reviewed models.'
- [References and in-text citations] The in-text citation numbering appears inconsistent with the reference list in places (for example, the duplicate passage cites geometry inference with different numbers, and 'Transformer-based models observed to perform below expectations' is cited to [84], which in the reference list is a different work). Please reconcile all citation numbers.
- [Figure 1] Figure 1 would be easier to interpret if the caption clearly stated that the placement of deep learning models in the cells reflects the author's own categorization, and if the meaning of the color coding (blue/green/red/yellow) were explained in the caption itself rather than only in the acknowledgments.
Circularity Check
No circular derivation found: the review's taxonomy and research perspectives are interpretive, not derived from their own inputs.
full rationale
This paper is a literature review; it does not fit parameters to data, derive predictions from first principles, or construct a mathematical derivation chain. The central 'apparent gap' between geometry-based and wave-based deep models (Sec. 3) is an interpretive observation based on the author's classification of cited works (Fig. 1, Sec. 2), not a quantity computed from the inputs; no equation in the paper reduces to another by construction. The forward-looking suggestion that boundary-information-rich models (PIBI-Nets) may bridge the gap is explicitly framed as a perspective ('may lie', 'first and promising result') and cites [37] for the independent mathematical fact that the boundary integral equation has an asymptotic geometric interpretation. Although several references ([36], [37], [198], [199]) come from the author's group and [199] is a submitted paper, this self-citation is not load-bearing in the sense of a theorem whose proof is replaced by the citation: the review does not claim to prove PIBI-Nets work, and the taxonomy would stand even if the cited works were removed. Potential concerns about selection bias or the strength of the evidence for the gap are correctness and completeness issues, not circularity. No circular step meeting the quote-and-reduction standard is present.
Assumptions & free parameters
assumptions (3)
- domain assumption Room acoustics can be usefully classified along the axes of physics versus data-driven and geometric versus wave-based.
- domain assumption The cited literature is representative and sufficiently complete to support the claimed gap between geometry-based and wave-based deep models.
- domain assumption The linear, time-invariant assumption of traditional room acoustics modeling is the correct reference frame for comparing deep nonlinear models.
Cite this review
Pith. "Pith review of Deep, data-driven modeling of room acoustics: literature review and research perspectives." pith.science (2026). https://pith.science/paper/DBVFUTZG
@misc{pith2026250416289,
author = {Pith},
title = {Pith review of: Deep, data-driven modeling of room acoustics: literature review and research perspectives},
year = {2026},
howpublished = {\url{https://pith.science/paper/DBVFUTZG}},
note = {Machine review of arXiv:2504.16289}
}
read the original abstract
Our everyday auditory experience is shaped by the acoustics of the indoor environments in which we live. Room acoustics modeling is aimed at establishing mathematical representations of acoustic wave propagation in such environments. These representations are relevant to a variety of problems ranging from echo-aided auditory indoor navigation to restoring speech understanding in cocktail party scenarios. Many disciplines in science and engineering have recently witnessed a paradigm shift powered by deep learning (DL), and room acoustics research is no exception. The majority of deep, data-driven room acoustics models are inspired by DL-based speech and image processing, and hence lack the intrinsic space-time structure of acoustic wave propagation. More recently, DL-based models for room acoustics that include either geometric or wave-based information have delivered promising results, primarily for the problem of sound field reconstruction. In this review paper, we will provide an extensive and structured literature review on deep, data-driven modeling in room acoustics. Moreover, we position these models in a framework that allows for a conceptual comparison with traditional physical and data-driven models. Finally, we identify strengths and shortcomings of deep, data-driven room acoustics models and outline the main challenges for further research.
Reference graph
Works this paper leans on
-
[199]
F. Miotello et al., “HOMULA-RIR: A room impulse response dataset for tele- conferencing and spatial audio applications acquired through higher-order mi- crophones and uniform linear microphone arrays,” in Proc. 2024 IEEE Int. Conf. Acoust., Speech, Signal Process. Workshops (ICASSPW ’24) , (Seoul, Korea), pp. 795–799, 2024
work page 2024
-
[37]
The use of equivalent source method in computational acoustics,
S. Lee, “The use of equivalent source method in computational acoustics,” J. Comput. Acoust., vol. 25, no. 1, 2017. Article No. 1630001
2017
-
[1]
INTRODUCTION People spend about 90 % of their time indoors [1], hence our auditory system has been trained to perceive and process sound only after it has been “shaped” by the acoustics of our inside living environments. Whereas room acoustics is potentially ben- eficial for perceptual tasks and experiences including human in- door navigation by means of ...
2025
-
[2]
Deep, data-driven modeling of room acoustics: literature review and research perspectives
LITERA TURE REVIEW 2.1 Concise overview of traditional room acoustics models Traditional room acoustics models (i.e., those developed before the deep learning era) can be categorized along two dimen- arXiv:2504.16289v1 [eess.AS] 22 Apr 2025 11th Convention of the European Acoustics Association M´alaga, Spain • 23rd – 26th June 2025 • sions: the first dime...
work page Pith review arXiv 2025
-
[3]
We end this review paper by formulating three prominent perspectives for fu- ture research
RESEARCH PERSPECTIVES From the above literature study, it is clear that deep learning holds great potential for room acoustics modeling. We end this review paper by formulating three prominent perspectives for fu- ture research. Firstly, data availability remains the first and foremost re- quirement in the development of deep, data-driven models. Over the...
2025
-
[4]
The scientific responsibility is assumed by its authors
ACKNOWLEDGEMENTS This research work was carried out at the ESAT Laboratory of KU Leuven, in the frame of KU Leuven internal funds C14/21/075 and C3/23/056, FWO projects G0A0424N and S005525N, and the AI Research Program of the Flemish Govern- ment. The scientific responsibility is assumed by its authors
-
[5]
The National Human Activity Pattern Survey (NHAPS): a resource for assessing exposure to environmental pollutants,
N. E. Klepeis et al., “The National Human Activity Pattern Survey (NHAPS): a resource for assessing exposure to environmental pollutants,” J. Expo. Sci. Environ. Epidemiol., vol. 11, pp. 231–252, 2001
2001
-
[6]
A summary of research investigating echolocation abilities of blind and sighted humans,
A. J. Kolarik et al., “A summary of research investigating echolocation abilities of blind and sighted humans,” Hearing Res., vol. 310, pp. 60–68, 2014
2014
Show all 203 references
-
[7]
Localization of a virtual wall by means of ac- tive echolocation by untrained sighted persons,
D. Pelegr ´ın-Garc´ıa et al. , “Localization of a virtual wall by means of ac- tive echolocation by untrained sighted persons,” Applied Acoustics, vol. 139, pp. 82–92, 2018
2018
-
[8]
Engaging concert hall acoustics is made up of temporal enve- lope preserving reflections,
T. Lokki et al., “Engaging concert hall acoustics is made up of temporal enve- lope preserving reflections,” J. Acoust. Soc. Am., vol. 129, no. 6, pp. EL223– EL228, 2011
2011
-
[9]
The modulation transfer function in room acoustics as a predictor of speech intelligibility,
T. Houtgast and H. J. M. Steeneken, “The modulation transfer function in room acoustics as a predictor of speech intelligibility,” Acta Acustica united with Acustica, vol. 28, no. 1, pp. 66–73, 1973
1973
-
[10]
The cocktail party phenomenon: a review of research on speech intelligibility in multiple-talker conditions,
A. W. Bronkhorst, “The cocktail party phenomenon: a review of research on speech intelligibility in multiple-talker conditions,” Acta Acustica united with Acustica, vol. 86, no. 1, pp. 117–128, 2000
2000
-
[11]
Functionality of hearing aids: state-of-the-art and future model-based solutions,
B. Kollmeier and J. Kiessling, “Functionality of hearing aids: state-of-the-art and future model-based solutions,”Int. J. Audiology, vol. 57, p. S3–S28, 2016
2016
-
[12]
Finite volume time domain room acoustics simulation un- der general impedance boundary conditions,
S. Bilbao et al., “Finite volume time domain room acoustics simulation un- der general impedance boundary conditions,”IEEE/ACM Trans. Audio Speech Language Process., vol. 24, no. 1, pp. 161–173, 2016
2016
-
[13]
FDTD methods for 3-D room acoustics simulation with high-order accuracy in space and time,
B. Hamilton and S. Bilbao, “FDTD methods for 3-D room acoustics simulation with high-order accuracy in space and time,” IEEE/ACM Trans. Audio Speech Language Process., vol. 25, no. 11, pp. 2112–2124, 2017
2017
-
[14]
W. C. Sabine, Collected Papers on Acoustics. Cambridge, MA, USA: Harvard University, 1922
1922
-
[15]
Review of objective room acoustics measures and future needs,
J. S. Bradley, “Review of objective room acoustics measures and future needs,” Applied Acoustics, vol. 72, no. 10, pp. 713–720, 2011
2011
-
[16]
Monaural room acoustic parameters from music and speech,
P. Kendrick et al. , “Monaural room acoustic parameters from music and speech,” J. Acoust. Soc. Am., vol. 124, no. 1, pp. 278–287, 2008
2008
-
[17]
Estimation of room acoustic parameters: The ACE challenge,
J. Eaton et al., “Estimation of room acoustic parameters: The ACE challenge,” IEEE/ACM Trans. Audio Speech Language Process., vol. 24, no. 10, pp. 1681– 1693, 2016
2016
-
[18]
Pole and zero modeling of room trans- fer functions,
J. Mourjopoulos and M. A. Paraskevas, “Pole and zero modeling of room trans- fer functions,” J. Sound Vib., vol. 146, no. 2, pp. 281–302, 1991
1991
-
[19]
Fifty years of artificial reverberation,
V . V ¨alim¨aki et al., “Fifty years of artificial reverberation,” IEEE Trans. Audio Speech Language Process., vol. 20, no. 5, pp. 1421–1448, 2012
2012
-
[20]
State-space estimation of spatially dynamic room im- pulse responses using a room acoustic model-based prior,
K. MacWilliam et al., “State-space estimation of spatially dynamic room im- pulse responses using a room acoustic model-based prior,” Frontiers in Signal Process., vol. 4, 2024
2024
-
[21]
Introduction to the special issue on machine learn- ing in acoustics,
Z. H. Michalopoulou et al., “Introduction to the special issue on machine learn- ing in acoustics,” J. Acoust. Soc. Am., vol. 150, no. 4, pp. 3204–3210, 2021
2021
-
[22]
An overview of machine learning and other data-based meth- ods for spatial audio capture, processing, and reproduction,
M. Cobos et al., “An overview of machine learning and other data-based meth- ods for spatial audio capture, processing, and reproduction,” EURASIP J. Au- dio Speech Music Process., vol. 2022, 2022. Article No. 10
2022
-
[23]
Yu, The estimation of acoustic parameters and representations based on room impulse responses
W. Yu, The estimation of acoustic parameters and representations based on room impulse responses . PhD thesis, Delft University of Technology, The Netherlands, 2024
2024
-
[24]
G ¨otz, Data-driven room-acoustic modelling
G. G ¨otz, Data-driven room-acoustic modelling. PhD thesis, Aalto University, Finland, 2024
2024
-
[25]
Karakonstantis, Data-driven methods for large-scale sound field acquisition and analysis
X. Karakonstantis, Data-driven methods for large-scale sound field acquisition and analysis. PhD thesis, Technical University of Denmark, Denmark, 2024
2024
-
[26]
Kuttruff, Room Acoustics
H. Kuttruff, Room Acoustics. Spon Press, 5 ed., 2009
2009
-
[27]
Spatial impulse response rendering I: Analysis and synthesis,
J. Merimaa and V . Pulkki, “Spatial impulse response rendering I: Analysis and synthesis,” J. Audio Eng. Soc., vol. 53, no. 12, pp. 1115–1127, 2005
2005
-
[28]
Common acoustical pole and zero modeling of room transfer functions,
Y . Haneda, S. Makino, and Y . Kaneda, “Common acoustical pole and zero modeling of room transfer functions,” IEEE Trans. Speech Audio Process. , vol. 2, no. 2, pp. 320–328, 1994
1994
-
[29]
A scalable algorithm for physically motivated and sparse approximation of room impulse responses with orthonormal basis functions,
G. Vairetti et al., “A scalable algorithm for physically motivated and sparse approximation of room impulse responses with orthonormal basis functions,” IEEE/ACM Trans. Audio Speech Language Process., vol. 25, no. 7, pp. 1547– 1561, 2017
2017
-
[30]
High-precision parallel graphic equal- izer,
J. R ¨am¨o, V . V¨alim¨aki, , and B. Bank, “High-precision parallel graphic equal- izer,” IEEE/ACM Trans. Audio Speech Language Process. , vol. 22, no. 12, pp. 1894–1904, 2014
1904
-
[31]
Optimally regularized adaptive filtering algorithms for room acoustic signal enhancement,
T. van Waterschoot, G. Rombouts, and M. Moonen, “Optimally regularized adaptive filtering algorithms for room acoustic signal enhancement,” Signal Processing, vol. 88, no. 3, pp. 594–611, 2008
2008
-
[32]
Spatial decomposition method for room impulse responses,
S. Tervo et al., “Spatial decomposition method for room impulse responses,” J. Audio Eng. Soc., vol. 61, no. 1/2, pp. 17–28, 2013
2013
-
[33]
Array technology for acous- tic wave field analysis in enclosures,
A. J. Berkhout, D. de Vries, and J. J. Sonke, “Array technology for acous- tic wave field analysis in enclosures,” J. Acoust. Soc. Am. , vol. 102, no. 5, pp. 2757–2770, 1997
1997
-
[34]
Space-time-frequency processing of acoustic wave fields: Theory, algorithms, and applications,
F. Pinto and M. Vetterli, “Space-time-frequency processing of acoustic wave fields: Theory, algorithms, and applications,” IEEE Trans. Signal Process. , vol. 58, no. 9, pp. 4608–4620, 2010
2010
-
[35]
D. P. Jarrett, E. A. P. Habets, and P. A. Naylor, Theory and Applications of Spherical Microphone Array Processing. Springer, 2017
2017
-
[36]
A unified approach to finite and boundary element discretization in linear time-harmonic acoustics,
S. Marburg, “A unified approach to finite and boundary element discretization in linear time-harmonic acoustics,” inComputational Acoustics of Noise Prop- agation in Fluids - Finite and Boundary Element Methods (S. Marburg and B. Nolte, eds.), Springer, 2008
2008
-
[38]
Room impulse response interpolation using a sparse spatio-temporal representation of the sound field,
N. Antonello et al. , “Room impulse response interpolation using a sparse spatio-temporal representation of the sound field,” IEEE/ACM Trans. Audio Speech Language Process., vol. 25, no. 10, pp. 1929–1941, 2017
1929
-
[39]
Joint acoustic localization and dereverberation through plane wave decomposition and sparse regularization,
N. Antonello et al., “Joint acoustic localization and dereverberation through plane wave decomposition and sparse regularization,”IEEE/ACM Trans. Audio Speech Language Process., vol. 27, no. 12, pp. 1893–1905, 2019
1905
-
[40]
A state-space framework for the boundary integral equation,
R. Ali et al., “A state-space framework for the boundary integral equation,” tech. rep., KU Leuven, Leuven, Belgium, 2025
2025
-
[41]
Relating wave-based and geometric acoustics using a stationary phase approximation,
R. Ali et al., “Relating wave-based and geometric acoustics using a stationary phase approximation,” in Proc. Forum Acusticum 2023, (Turin, Italy), 2023
2023
-
[42]
Overview of geometrical room acoustic modeling techniques,
L. Savioja and P. Svensson, “Overview of geometrical room acoustic modeling techniques,” J. Acoust. Soc. Am., vol. 138, no. 2, pp. 708–730, 2015
2015
-
[43]
Image method for efficiently simulating small- room acoustics,
J. B. Allen and D. A. Berkley, “Image method for efficiently simulating small- room acoustics,” J. Acoust. Soc. Am., vol. 65, no. 4, pp. 943–950, 1979
1979
-
[44]
Calculating the acoustical room re- sponse by the use of a ray tracing technique,
A. Krokstad, S. Strøm, and S. Sørsdal, “Calculating the acoustical room re- sponse by the use of a ray tracing technique,” J. Sound Vib. , vol. 8, no. 1, pp. 118–125, 1968
1968
-
[45]
A beam tracing method for interactive architectural acoustics,
T. Funkhouser et al. , “A beam tracing method for interactive architectural acoustics,” J. Acoust. Soc. Am., vol. 115, no. 2, pp. 739–756, 2004
2004
-
[46]
Digital delay networks for designing artificial re- verberators,
J.-M. Jot and A. Chaigne, “Digital delay networks for designing artificial re- verberators,” in Preprints AES 90th Conv., (Paris, France), 1991. AES Preprint 3030
1991
-
[47]
Circulant and elliptic feedback delay networks for artificial reverberation,
D. Rocchesso and J. O. Smith, “Circulant and elliptic feedback delay networks for artificial reverberation,” IEEE Trans. Speech Audio Process., vol. 5, no. 1, pp. 51–63, 1997
1997
-
[48]
Efficient synthesis of room acoustics via scattering delay networks,
E. D. Sena et al., “Efficient synthesis of room acoustics via scattering delay networks,” IEEE/ACM Trans. Audio Speech Language Process., vol. 23, no. 9, pp. 1478–1492, 2015
2015
-
[49]
Application of BEM (boundary element method)-based acoustic holography to radiation analysis of sound sources with arbitrarily shaped ge- ometries,
M. R. Bai, “Application of BEM (boundary element method)-based acoustic holography to radiation analysis of sound sources with arbitrarily shaped ge- ometries,” J. Acoust. Soc. Am., vol. 92, no. 1, pp. 533–549, 1992
1992
-
[50]
The analysis of the acoustic field in irregularly shaped rooms by the finite element method,
T. Shuku and K. Ishihara, “The analysis of the acoustic field in irregularly shaped rooms by the finite element method,” J. Sound Vib. , vol. 29, no. 1, pp. 67–76, 1973
1973
-
[51]
Finite-difference time-domain simulation of low-frequency room acoustic problems,
D. Botteldooren, “Finite-difference time-domain simulation of low-frequency room acoustic problems,” J. Acoust. Soc. Am., vol. 98, no. 6, pp. 3302–3308, 1995
1995
-
[52]
Yu and L
D. Yu and L. Deng, Automatic speech recognition: A deep learning approach. Springer, 2015
2015
-
[53]
Reverberant speech recognition exploiting clarity index estimation,
P. P. Parada et al., “Reverberant speech recognition exploiting clarity index estimation,” EURASIP J. Adv. Signal Process. , vol. 2015, 2015. Article No. 54
2015
-
[54]
Speaker recognition based on deep learning: An overview,
Z. Bai and X.-L. Zhang, “Speaker recognition based on deep learning: An overview,”Neural Networks, vol. 140, pp. 65–99, 2021
2021
-
[55]
End-to-end speech emotion recog- nition using deep neural networks,
P. Tzirakis, J. Zhang, and B. W. Schuller, “End-to-end speech emotion recog- nition using deep neural networks,” in Proc. 2018 IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ’18), (Calgary, AB, Canada), pp. 5089–5093, 2018
2018
-
[56]
End-to-end speech emotion recognition using a novel context- stacking dilated convolution neural network,
D. Tang et al., “End-to-end speech emotion recognition using a novel context- stacking dilated convolution neural network,”EURASIP J. Audio, Speech, Mu- sic Process., vol. 2021, 2021. Article No. 18. 11th Convention of the European Acoustics Association M´alaga, Spain • 23rd –...
2021
-
[57]
Learning to estimate reverberation time in noisy and rever- berant rooms,
X. Xiao et al., “Learning to estimate reverberation time in noisy and rever- berant rooms,” in Proc. INTERSPEECH 2015, Dresden, Germany, pp. 3431– 3435, 2015
2015
-
[58]
Blind room acoustics characterization using recur- rent neural networks and modulation spectrum dynamics,
J. F. Santos and T. H. Falk, “Blind room acoustics characterization using recur- rent neural networks and modulation spectrum dynamics,” in Proc. AES 60th Conf., (Leuven, Belgium), 2016
2016
-
[59]
Scene-aware audio rendering via deep acoustic analysis,
Z. Tang et al., “Scene-aware audio rendering via deep acoustic analysis,”IEEE Trans. Vis. Comput. Graphics, vol. 26, no. 5, pp. 1991–2001, 2020
1991
-
[60]
Impulse response data augmentation and deep neural networks for blind room acoustic parameter estimation,
N. J. Bryan, “Impulse response data augmentation and deep neural networks for blind room acoustic parameter estimation,” inin Proc. 2020 IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ’20), 2020
2020
-
[61]
Evaluation of data augmentation techniques of room impulse re- sponses for improved AI-based estimations,
C. Kehling, “Evaluation of data augmentation techniques of room impulse re- sponses for improved AI-based estimations,” inProc. DAGA 2024, (Hannover, Germany), pp. 1285–1288, 2024
2024
-
[62]
Blind estimation of speech transmission index and room acoustic parameters based on the extended model of room impulse re- sponse,
S. Duangpummet et al., “Blind estimation of speech transmission index and room acoustic parameters based on the extended model of room impulse re- sponse,” Applied Acoustics, vol. 185, 2022. Article No. 108372
2022
-
[63]
AI-IoT platform for blind estimation of room acous- tic parameters based on deep neural networks,
J. Lopez-Ballester et al., “AI-IoT platform for blind estimation of room acous- tic parameters based on deep neural networks,” IEEE Internet of Things Jour- nal, vol. 10, no. 1, pp. 855–866, 2022
2022
-
[64]
Blind acoustic room parameter estimation using phase features,
C. Ick, A. Mehrabi, and W. Jin, “Blind acoustic room parameter estimation using phase features,” in Proc. 2023 IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ’23), (Rhodes, Greece), 2023
2023
-
[65]
Online blind reverberation time es- timation using CRNNs,
S. Deng, W. Mack, and E. A. P. Habets, “Online blind reverberation time es- timation using CRNNs,” in Proc. INTERSPEECH 2020 , (Shanghai, China), pp. 5061–5065, 2020
2020
-
[66]
Blind reverberation time estimation in dynamic acoustic condi- tions,
P. G ¨otz et al., “Blind reverberation time estimation in dynamic acoustic condi- tions,” inProc. 2022 IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ’22), (Singapore), pp. 581–585, 2022
2022
-
[67]
Joint blind room acoustic characterization from speech and music signals using convolutional recurrent neural networks,
P. Callens and M. Cernak, “Joint blind room acoustic characterization from speech and music signals using convolutional recurrent neural networks,” arXiv preprint arXiv:2010.11167, 2020
2010 arXiv
-
[68]
A universal deep room acoustics estimator,
P. S. L ´opez, P. Callens, and M. Cernak, “A universal deep room acoustics estimator,” inProc. 2021 IEEE Workshop Appls. Signal Process. Audio Acoust. (WASPAA ‘21), (New Paltz, NY , USA), pp. 356–360, 2021
2021
-
[69]
Self-supervised learning of spatial acoustic representa- tion with cross-channel signal reconstruction and multi-channel conformer,
B. Yang and X. Li, “Self-supervised learning of spatial acoustic representa- tion with cross-channel signal reconstruction and multi-channel conformer,” IEEE/ACM Trans. Audio Speech Language Process., vol. 32, pp. 4211–4225, 2024
2024
-
[70]
Exploring the power of pure attention mechanisms in blind room parameter estimation,
C. Wang et al., “Exploring the power of pure attention mechanisms in blind room parameter estimation,” EURASIP J. Audio Speech Music Process. , vol. 2024, 2024. Article No. 23
2024
-
[71]
A single-channel non-intrusive C50 estimator correlated with speech recognition performance,
P. P. Parada et al., “A single-channel non-intrusive C50 estimator correlated with speech recognition performance,” IEEE/ACM Trans. Audio Speech Lan- guage Process., vol. 24, no. 4, pp. 719–732, 2016
2016
-
[72]
Brouhaha: multi-task training for voice activity detec- tion, speech-to-noise ratio, and C50 room acoustics estimation,
M. Lavechin et al., “Brouhaha: multi-task training for voice activity detec- tion, speech-to-noise ratio, and C50 room acoustics estimation,” in Proc. 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU ’23), (Taipei, Taiwan), 2023
2023
-
[73]
Attention is all you need for blind room volume estimation,
C. Wang et al., “Attention is all you need for blind room volume estimation,” in Proc. 2024 IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ’24), (Seoul, Korea), pp. 1341–1345, 2024
2024
-
[74]
Speaker distance estimation in enclosures from single-channel audio,
M. Neri et al., “Speaker distance estimation in enclosures from single-channel audio,” IEEE/ACM Trans. Audio Speech Language Process., vol. 32, pp. 2242– 2254, 2024
2024
-
[75]
End-to-end classification of rever- berant rooms using DNNs,
C. Papayiannis, C. Evers, and P. A. Naylor, “End-to-end classification of rever- berant rooms using DNNs,” IEEE/ACM Trans. Audio Speech Language Pro- cess., vol. 28, pp. 3010–3017, 2020
2020
-
[76]
Data augmentation of room classi- fiers using generative adversarial networks,
C. Papayiannis, C. Evers, and P. A. Naylor, “Data augmentation of room classi- fiers using generative adversarial networks,”arXiv preprint arXiv:1901.03257, 2019
1901 arXiv
-
[77]
Attention is all you need,
A. Vaswani et al., “Attention is all you need,” in Proc. 31st Conf. Neural Inf. Process. Syst. (NIPS ’17), (Long Beach, CA, USA), 2017
2017
-
[78]
Evaluation of a numerical method for identifying surface acoustic impedances in a reverberant room,
N. Antonello et al., “Evaluation of a numerical method for identifying surface acoustic impedances in a reverberant room,” inProc. 10th European Congress & Exposition Noise Control Eng. (EURONOISE ‘15), (Maastricht, The Nether- lands), 2015
2015
-
[79]
Mean absorption estimation from room impulse responses using virtually supervised learning,
C. Foy, A. Deleforge, and D. D. Carlo, “Mean absorption estimation from room impulse responses using virtually supervised learning,” J. Acoust. Soc. Am., vol. 150, no. 2, pp. 1286–1299, 2021
2021
-
[80]
Room acoustical parameter estimation from room impulse responses using deep neural networks,
W. Yu and W. B. Kleijn, “Room acoustical parameter estimation from room impulse responses using deep neural networks,” IEEE/ACM Trans. Audio Speech Language Process., vol. 29, pp. 436–447, 2020
2020
-
[81]
Detecting sound-absorbing mate- rials in a room from a single impulse response using a CRNN,
C. Papayiannis, C. Evers, and P. A. Naylor, “Detecting sound-absorbing mate- rials in a room from a single impulse response using a CRNN,” arXiv preprint arXiv:1901.05852, 2019
1901 arXiv
-
[82]
Acoustic reflectors localization from stereo recordings using neural networks,
G. Bologni, R. Heusdens, and J. Martinez, “Acoustic reflectors localization from stereo recordings using neural networks,” in Proc. 2021 IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ’21), (Toronto, ON, Canada), 2021
2021
-
[83]
Reconstructing room scales with a single sound for aug- mented reality displays,
B. S. Liang et al., “Reconstructing room scales with a single sound for aug- mented reality displays,” J. Inf. Display, vol. 24, no. 1, pp. 1–12, 2023
2023
-
[84]
Room geometry estimation from higher-order ambisonics signals using convolutional recurrent neural networks,
N. Poschadel et al., “Room geometry estimation from higher-order ambisonics signals using convolutional recurrent neural networks,” inPreprints AES 150th Conv., 2021
2021
-
[85]
Data-driven 3D room geometry inference with a linear loud- speaker array and a single microphone,
C. Tuna et al., “Data-driven 3D room geometry inference with a linear loud- speaker array and a single microphone,” in Proc. Forum Acusticum 2023 , (Turin, Italy), 2023
2023
-
[86]
Data-driven joint detection and localization of acoustic reflectors,
H. N. Bicer et al. , “Data-driven joint detection and localization of acoustic reflectors,” in Proc. 2024 IEEE Int. Conf. Acoust., Speech, Signal Process. Workshops (ICASSPW ’24), (Seoul, Korea), pp. 745–749, 2024
2024
-
[87]
EchoScan: scanning complex room geometries via acous- tic echoes,
I. Yeon et al. , “EchoScan: scanning complex room geometries via acous- tic echoes,” IEEE/ACM Trans. Audio Speech Language Process. , vol. 32, pp. 4768–4782, 2024
2024
-
[88]
Towards a data-driven plane wave decomposition from multichannel room impulse responses,
D. Schindler, F. Schultz, and S. Spors, “Towards a data-driven plane wave decomposition from multichannel room impulse responses,” in Proc. DAGA 2023, (Hamburg, Germany), pp. 1671–1674, 2023
2023
-
[89]
A survey of sound source localization with deep learn- ing methods,
P. A. Grumiaux et al., “A survey of sound source localization with deep learn- ing methods,” J. Acoust. Soc. Am., vol. 152, no. 1, pp. 107–151, 2022
2022
-
[90]
Robust sound source tracking using SRP-PHAT and 3D convolutional neural networks,
D. Diaz-Guerra, A. Miguel, and J. R. Beltran, “Robust sound source tracking using SRP-PHAT and 3D convolutional neural networks,” IEEE/ACM Trans. Audio Speech Language Process., vol. 29, pp. 300–311, 2020
2020
-
[91]
DoA estimation of room reflections using NN-based MUSIC algorithm,
H. Li, W. Zhang, and L. Zhang, “DoA estimation of room reflections using NN-based MUSIC algorithm,” in in Proc. 2023 Asia Pacific Signal Inf. Pro- cess. Assoc. Annual Summit Conf. (APSIPA ASC ’23), (Taipei, Taiwan), 2023
2023
-
[92]
Double-talk-robust prediction error identifica- tion algorithms for acoustic echo cancellation,
T. van Waterschoot et al. , “Double-talk-robust prediction error identifica- tion algorithms for acoustic echo cancellation,” IEEE Trans. Signal Process., vol. 55, no. 3, pp. 846–858, 2007
2007
-
[93]
Deep learning for acoustic echo cancellation in noisy and double-talk scenarios,
H. Zhang and D. Wang, “Deep learning for acoustic echo cancellation in noisy and double-talk scenarios,” inProc. INTERSPEECH 2018, (Hyderabad, India), pp. 3239–3243, 2018
2018
-
[94]
A robust and cascaded acoustic echo cancellation based on deep learning,
C. Zhang and X. Zhang, “A robust and cascaded acoustic echo cancellation based on deep learning,” in Proc. INTERSPEECH 2020 , (Shanghai, China), pp. 5061–5065, 2020
2020
-
[95]
Acoustic echo cancellation with the dual- signal transformation LSTM network,
N. L. Westhausen and B. T. Meyer, “Acoustic echo cancellation with the dual- signal transformation LSTM network,” in Proc. 2021 IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ’21), (Toronto, ON, Canada), 2021
2021
-
[96]
Deep multitask acoustic echo cancella- tion,
A. Fazel, M. El-Khamy, and J. Lee, “Deep multitask acoustic echo cancella- tion,” in Proc. INTERSPEECH 2019, (Graz, Austria), pp. 4250–4254, 2019
2019
-
[97]
Deep learning for joint acoustic echo and acous- tic howling suppression in hybrid meetings,
H. Zhang, M. Yu, and D. Yu, “Deep learning for joint acoustic echo and acous- tic howling suppression in hybrid meetings,” in Proc. 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU ’23) , (Taipei, Tai- wan), 2023
2023
-
[98]
CAD-AEC: context-aware deep acous- tic echo cancellation,
A. Fazel, M. El-Khamy, and J. Lee, “CAD-AEC: context-aware deep acous- tic echo cancellation,” in Proc. 2020 IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ’20), 2020
2020
-
[99]
A complex spectral mapping with inplace convolution recurrent neural networks for acoustic echo cancellation,
C. Zhang, J. Liu, and X. Zhang, “A complex spectral mapping with inplace convolution recurrent neural networks for acoustic echo cancellation,” inProc. 2022 IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ’22) , (Singa- pore), pp. 751–755, 2022
2022
-
[100]
Deep learning for joint acoustic echo and noise cancellation with nonlinear distortions,
H. Zhang, K. Tan, and D. Wang, “Deep learning for joint acoustic echo and noise cancellation with nonlinear distortions,” in Proc. INTERSPEECH 2019, (Graz, Austria), pp. 4255–4259, 2019
2019
-
[101]
Deep learning-based stereophonic acoustic echo suppression without decorrelation,
L. Cheng et al., “Deep learning-based stereophonic acoustic echo suppression without decorrelation,”J. Acoust. Soc. Am., vol. 150, no. 2, pp. 816–829, 2021
2021
-
[102]
A deep learning approach to multi-channel and multi- microphone acoustic echo cancellation,
H. Zhang and D. Wang, “A deep learning approach to multi-channel and multi- microphone acoustic echo cancellation,” inProc. INTERSPEECH 2021, (Brno, Czech Republic), pp. 1139–1143, 2021
2021
-
[103]
Multi-channel and multi-microphone acoustic echo cancellation using a deep learning based approach,
H. Zhang and D. Wang, “Multi-channel and multi-microphone acoustic echo cancellation using a deep learning based approach,” arXiv preprint arXiv:2103.02552, 2021
2021 arXiv
-
[104]
A deep hierarchical fusion network for fullband acoustic echo cancellation,
H. Zhao et al., “A deep hierarchical fusion network for fullband acoustic echo cancellation,” in Proc. 2022 IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ’22), (Singapore), pp. 9112–9116, 2022
2022
-
[105]
Deep learning-based acoustic echo cancellation for surround sound systems,
G. Li et al. , “Deep learning-based acoustic echo cancellation for surround sound systems,” Appl. Sci., vol. 13, no. 3, 2023. Article No. 1266. 11th Convention of the European Acoustics Association M´alaga, Spain • 23rd – 26th June 2025 •
2023
-
[106]
A deep hybrid model for stereophonic acoustic echo control,
Y . Liu et al., “A deep hybrid model for stereophonic acoustic echo control,” Circuits Syst. Signal Process., 2024
2024
-
[107]
Fifty years of acoustic feedback control: state of the art and future challenges,
T. van Waterschoot and M. Moonen, “Fifty years of acoustic feedback control: state of the art and future challenges,”Proc. IEEE, vol. 99, no. 2, pp. 288–327, 2011
2011
-
[108]
A deep learning solution to the marginal stability problems of acoustic feedback systems for hearing aids,
C. Zheng et al., “A deep learning solution to the marginal stability problems of acoustic feedback systems for hearing aids,” J. Acoust. Soc. Am., vol. 152, no. 6, pp. 3616–3634, 2022
2022
-
[109]
Deep AHS: A deep learning approach to acoustic howling suppression,
H. Zhang, M. Yu, and D. Yu, “Deep AHS: A deep learning approach to acoustic howling suppression,” in Proc. 2023 IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ’23), (Rhodes, Greece), 2023
2023
-
[110]
Enhanced acoustic howling suppression via hybrid Kalman filter and deep learning models,
H. Zhang et al., “Enhanced acoustic howling suppression via hybrid Kalman filter and deep learning models,” IEEE/ACM Trans. Audio Speech Language Process., vol. 32, pp. 2828–2840, 2024
2024
-
[111]
P. A. Naylor and N. D. Gaubitch, eds., Speech Dereverberation. Springer, 2010
2010
-
[112]
Two-stage deep learning for noisy- reverberant speech enhancement,
Y . Zhao, Z.-Q. Wang, and D. Wang, “Two-stage deep learning for noisy- reverberant speech enhancement,” IEEE/ACM Trans. Audio Speech Language Process., vol. 27, no. 1, pp. 53–62, 2018
2018
-
[113]
Single-channel dereverberation using direct MMSE optimiza- tion and bidirectional LSTM networks,
W. Mack et al., “Single-channel dereverberation using direct MMSE optimiza- tion and bidirectional LSTM networks,” in Proc. INTERSPEECH 2018, (Hy- derabad, India), pp. 1314–1318, 2018
2018
-
[114]
Speech dereverberation constrained on room impulse re- sponse characteristics,
L. Bahrman et al., “Speech dereverberation constrained on room impulse re- sponse characteristics,” in Proc. INTERSPEECH 2024, (Kos Island, Greece), pp. 622–626, 2024
2024
-
[115]
On phase recovery and preserving early reflections for deep-learning speech dereverberation,
X. Luo, Y . Ke, X. Li, and C. Zheng, “On phase recovery and preserving early reflections for deep-learning speech dereverberation,” J. Acoust. Soc. Am. , vol. 155, no. 1, pp. 436–451, 2024
2024
-
[116]
Monaural speech dereverberation using temporal convolutional networks with self attention,
Y . Zhao et al., “Monaural speech dereverberation using temporal convolutional networks with self attention,” IEEE/ACM Trans. Audio Speech Language Pro- cess., vol. 28, pp. 1598–1607, 2020
2020
-
[117]
SkipConvGAN: Monaural speech dere- verberation using generative adversarial networks via complex time-frequency masking,
V . Kothapally and J. H. L. Hansen, “SkipConvGAN: Monaural speech dere- verberation using generative adversarial networks via complex time-frequency masking,” IEEE/ACM Trans. Audio Speech Language Process. , vol. 30, pp. 1600–1613, 2022
2022
-
[118]
DARE-Net: Speech dereverberation and room impulse response estimation,
J. Donley and P. Calamia, “DARE-Net: Speech dereverberation and room impulse response estimation,” tech. rep., Stanford University, Stanford, CA, USA, 2022
2022
-
[119]
Vincent, T
E. Vincent, T. Virtanen, and S. Gannot, eds., Audio source separation and speech enhancement. Wiley, 2018
2018
-
[120]
Deep learning for talker-dependent reverberant speaker separation: An empirical study,
M. Delfarah and D. Wang, “Deep learning for talker-dependent reverberant speaker separation: An empirical study,”IEEE/ACM Trans. Audio Speech Lan- guage Process., vol. 27, no. 11, pp. 1839–1848, 2019
2019
-
[121]
A causal and talker-independent speaker separa- tion/dereverberation deep learning algorithm: Cost associated with conversion to real-time capable operation,
E. W. Healy et al. , “A causal and talker-independent speaker separa- tion/dereverberation deep learning algorithm: Cost associated with conversion to real-time capable operation,” J. Acoust. Soc. Am., vol. 150, no. 5, pp. 3976– 3986, 2021
2021
-
[122]
An end-to-end deep learning approach to simultaneous speech dereverberation and acoustic modeling for robust speech recognition,
B. Wu et al., “An end-to-end deep learning approach to simultaneous speech dereverberation and acoustic modeling for robust speech recognition,”IEEE J. Select. Topics Signal Process., vol. 11, no. 8, pp. 1289–1300, 2017
2017
-
[123]
Zermini, Deep learning for speech separation
A. Zermini, Deep learning for speech separation . PhD thesis, University of Surrey, UK, 2020
2020
-
[124]
DBnet: DOA-driven beamforming network for end- to-end reverberant sound source separation,
A. Aroudi and S. Braun, “DBnet: DOA-driven beamforming network for end- to-end reverberant sound source separation,” in Proc. 2021 IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ’21), (Toronto, ON, Canada), 2021
2021
-
[125]
Direction specific ambisonics source separation with end-to- end deep learning,
F. Llu ´ıs et al., “Direction specific ambisonics source separation with end-to- end deep learning,” Acta Acustica, vol. 7, 2023. Article No. 29
2023
-
[126]
Spatially selective speaker separation using a DNN with a location dependent feature extraction,
A. Bohlender et al., “Spatially selective speaker separation using a DNN with a location dependent feature extraction,”IEEE/ACM Trans. Audio Speech Lan- guage Process., vol. 32, pp. 930–945, 2024
2024
-
[127]
SpatialNet: Extensively learning spatial information for multichannel joint speech separation, denoising and dereverberation,
C. Quan and X. Li, “SpatialNet: Extensively learning spatial information for multichannel joint speech separation, denoising and dereverberation,” IEEE/ACM Trans. Audio Speech Language Process., vol. 32, pp. 1310–1323, 2024
2024
-
[128]
Active noise control: a tutorial review,
S. M. Kuo and D. R. Morgan, “Active noise control: a tutorial review,” Proc. IEEE, vol. 87, no. 6, pp. 943–973, 1999
1999
-
[129]
Deep learning-assisted active noise control in a time-varying environment,
S. Im et al. , “Deep learning-assisted active noise control in a time-varying environment,” J. Mech. Sci. Technol., vol. 37, no. 3, pp. 1189–1196, 2023
2023
-
[130]
Deep ANC: A deep learning approach to active noise control,
H. Zhang and D. Wang, “Deep ANC: A deep learning approach to active noise control,” Neural Networks, vol. 141, pp. 1–10, 2021
2021
-
[131]
Deep MCANC: A deep learning approach to multi- channel active noise control,
H. Zhang and D. Wang, “Deep MCANC: A deep learning approach to multi- channel active noise control,” Neural Networks, vol. 158, 318-327, 2023
2023
-
[132]
DNoiseNet: Deep learning-based feedback active noise control in various noisy environments,
Y .-J. Cha, A. Mostafavi, and S. S. Benipal, “DNoiseNet: Deep learning-based feedback active noise control in various noisy environments,”Eng. Appl. Arti- ficial Intell., vol. 121, Article No. 105971, 2023
2023
-
[133]
Acoustic matching by embedding impulse responses,
J. Su, Z. Jin, and A. Finkelstein, “Acoustic matching by embedding impulse responses,” in Proc. 2020 IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ’20), 2020
2020
-
[134]
Visual acoustic matching,
C. Chen et al., “Visual acoustic matching,” in Proc. 2022 IEEE/CVF Conf. Comput. Vision Pattern Recognition (CVPR ’22) , (New Orleans, LA, USA), pp. 18836–18846, 2022
2022
-
[135]
N. D. Gaubitch, Blind identification of acoustic systems and enhancement of reverberant speech. PhD thesis, Imperial College London, UK, 2007
2007
-
[136]
Filtered noise shaping for time domain room impulse response estimation from reverberant speech,
C. J. Steinmetz, V . K. Ithapu, and P. Calamia, “Filtered noise shaping for time domain room impulse response estimation from reverberant speech,” in Proc. 2021 IEEE Workshop Appls. Signal Process. Audio Acoust. (WASPAA ‘21) , (New Paltz, NY , USA), pp. 221–225, 2021
2021
-
[137]
A V-RIR: Audio-visual room impulse response estima- tion,
A. Ratnarajah et al., “A V-RIR: Audio-visual room impulse response estima- tion,” in Proc. 2024 IEEE/CVF Conf. Comput. Vision Pattern Recognition (CVPR ’24), (Seattle, W A, USA), pp. 27164–27175, 2024
2024
-
[138]
Blind estimation of room impulse response from monaural re- verberant speech with segmental generative neural network,
Z. Liao et al., “Blind estimation of room impulse response from monaural re- verberant speech with segmental generative neural network,” in Proc. INTER- SPEECH 2023, (Dublin, Ireland), pp. 2723–2727, 2023
2023
-
[139]
Towards improved room impulse response estimation for speech recognition,
A. Ratnarajah et al., “Towards improved room impulse response estimation for speech recognition,” in Proc. 2023 IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ’23), (Rhodes, Greece), 2023
2023
-
[140]
Room response equalization – a review,
S. Cecchi, A. Carini, and S. Spors, “Room response equalization – a review,” Appl. Sci., vol. 8, no. 1, 2017. Article No. 16
2017
-
[141]
Data-driven local average room transfer function estimation for multi-point equalization,
C. Tuna et al., “Data-driven local average room transfer function estimation for multi-point equalization,” J. Acoust. Soc. Am., vol. 152, no. 6, pp. 3635–3647, 2022
2022
-
[142]
Reconstruction of the sound field in a room using compressive sensing,
S. A. Verburg and E. Fernandez-Grande, “Reconstruction of the sound field in a room using compressive sensing,” J. Acoust. Soc. Am. , vol. 143, no. 6, pp. 3770–3779, 2018
2018
-
[143]
Deep sound field reconstruction in real rooms: in- troducing the isobel sound field dataset,
M. S. Kristoffersen et al., “Deep sound field reconstruction in real rooms: in- troducing the isobel sound field dataset,” arXiv preprint arXiv:2102.06455 , 2021
2021 arXiv
-
[144]
Deep prior approach for room impulse response reconstruc- tion,
M. Pezzoli et al., “Deep prior approach for room impulse response reconstruc- tion,” Sensors, vol. 22, no. 7, 2022. Article No. 2710
2022
-
[145]
Efficient sound field reconstruction with conditional invertible neural networks,
X. Karakonstantis, E. Fernandez-Grande, and P. Gerstoft, “Efficient sound field reconstruction with conditional invertible neural networks,” arXiv preprint arXiv:2404.06928, 2024
2024 arXiv
-
[146]
Transformer-based virtual microphone estimator,
Z. Qiu et al., “Transformer-based virtual microphone estimator,” inProc. 2024 Workshop Hands-Free Speech Commun. Microphone Arrays (HSCMA ’24) , (Seoul, Korea), 2024
2024
-
[147]
A differentiable neural network approach to param- eter estimation of reverberation,
S. V . Lyster and C. Erkut, “A differentiable neural network approach to param- eter estimation of reverberation,” in Proc. 19th Sound Music Comput. Conf. (SMC ’22), (Saint- ´Etienne, France), pp. 358–364, 2022
2022
-
[148]
Real-time impulse response: a methodology based on machine learning approaches for a rapid impulse response generation for real-time acoustic virtual reality systems,
D. A. Sanaguano-Moreno et al., “Real-time impulse response: a methodology based on machine learning approaches for a rapid impulse response generation for real-time acoustic virtual reality systems,”Intell. Syst. Appl., vol. 21, 2024. Article No. 200306
2024
-
[149]
A frequency-domain approach to multichannel upmix,
C. Avendano and J. M. Jot, “A frequency-domain approach to multichannel upmix,” J. Audio Eng. Soc., vol. 52, no. 7/8, pp. 740–749, 2004
2004
-
[150]
Upmix B-Format Ambisonic room impulse responses using a generative model,
J. Xia and W. Zhang, “Upmix B-Format Ambisonic room impulse responses using a generative model,”Appl. Sci., vol. 13, no. 21, 2023. Article No. 11810
2023
-
[151]
Geometric deep learning and equivariant neural networks,
J. E. Gerken et al., “Geometric deep learning and equivariant neural networks,” Artif. Intell. Rev., vol. 56, pp. 14605–14662, 2023
2023
-
[152]
Learning acoustic scattering fields for dynamic interactive sound propagation,
Z. Tang, H.-Y . Meng, and D. Manocha, “Learning acoustic scattering fields for dynamic interactive sound propagation,” in Proc. 2021 IEEE Virtual Reality 3D User Interfaces (VR ’21), (Lisbon, Portugal), pp. 835–844, 2021
2021
-
[153]
MESH2IR: Neural acoustic impulse response generator for complex 3D scenes,
A. Ratnarajah et al., “MESH2IR: Neural acoustic impulse response generator for complex 3D scenes,” in Proc. 30th ACM Int. Conf. Multimedia (MM ’22), (Lisbon, Portugal), pp. 924–933, 2022
2022
-
[154]
RIR-in-a-Box: Estimating room acoustics from 3D mesh data through shoebox approximation,
L. Kelley et al., “RIR-in-a-Box: Estimating room acoustics from 3D mesh data through shoebox approximation,” in Proc. INTERSPEECH 2024, (Kos Island, Greece), 2024
2024
-
[155]
FAST-RIR: Fast neural diffuse room impulse response generator,
A. Ratnarajah et al., “FAST-RIR: Fast neural diffuse room impulse response generator,” in Proc. 2022 IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ’22), (Singapore), pp. 571–575, 2022
2022
-
[156]
Synthesis of room impulse responses by means of deep learning,
I. M. Salinas, J. A. B. Rodr ´ıguez, and G. P. Sip´an, “Synthesis of room impulse responses by means of deep learning,” in in Proc. 53o Congreso Espa ˜nol de Ac´ustica and XII Congreso Ib ´erico de Ac ´ustica (Tecniac´ustica ’22), (Elche, Spain), 2022. 11th Convention of the E...
2022
-
[157]
Predicting room impulse responses through encoder-decoder convolutional neural networks,
I. Martin et al., “Predicting room impulse responses through encoder-decoder convolutional neural networks,” inin Proc. 33rd IEEE Int. Workshop Machine Learning Signal Process. (MLSP ’23), (Rome, Italy), 2023
2023
-
[158]
Echo-aware room impulse response gener- ation,
S. Kim, J. h. Yoo, and J.-W. Choi, “Echo-aware room impulse response gener- ation,” J. Acoust. Soc. Am., vol. 156, no. 1, pp. 623–637, 2024
2024
-
[159]
Deep room impulse response completion,
J. Lin, G. G ¨otz, and S. J. Schlecht, “Deep room impulse response completion,” arXiv preprint arXiv:2402.00859, 2024
2024
-
[160]
Deep neural room acoustics primitive,
Y . He et al., “Deep neural room acoustics primitive,” in Proc. 41st Int. Conf. Machine Learning (ICML ’24), (Vienna, Austria), 2024
2024
-
[161]
Sound propagation in realistic interactive 3D scenes with parameterized sources using deep neural operators,
N. Borrel-Jensen et al., “Sound propagation in realistic interactive 3D scenes with parameterized sources using deep neural operators,” Proc. Natl. Acad. Sci., vol. 121, no. 2, 2024. Article No. e2312159120
2024
-
[162]
Room transfer function reconstruction using complex- valued neural networks and irregularly distributed microphones,
F. Ronchini et al. , “Room transfer function reconstruction using complex- valued neural networks and irregularly distributed microphones,” inProc. 32nd European Signal Process. Conf. (EUSIPCO ‘24), (Lyon, France), pp. 441–445, 2024
2024
-
[163]
Sound field reconstruction us- ing neural processes with dynamic kernels,
Z. Liang, W. Zhang, and T. D. Abhayapala, “Sound field reconstruction us- ing neural processes with dynamic kernels,” EURASIP J. Audio Speech Music Process., vol. 2024, 2024. Article No. 13
2024
-
[164]
Learning neural acoustic fields,
A. Luo et al., “Learning neural acoustic fields,” in Adv. Neural Inf. Process. Syst. 35 (NeurIPS ’22), pp. 3165–3177, 2022
2022
-
[165]
INRAS: Implicit neural representa- tion for audio scenes,
K. Su, M. Chen, and E. Shlizerman, “INRAS: Implicit neural representa- tion for audio scenes,” in Adv. Neural Inf. Process. Syst. 35 (NeurIPS ’22) , pp. 8144–8158, 2022
2022
-
[166]
Neural acoustic context field: Rendering realistic room im- pulse response with neural fields,
S. Liang et al., “Neural acoustic context field: Rendering realistic room im- pulse response with neural fields,” arXiv preprint arXiv:2309.15977, 2023
2023 arXiv
-
[167]
Few-shot audio-visual learning of environment acoustics,
S. Majumder et al., “Few-shot audio-visual learning of environment acoustics,” in Adv. Neural Inf. Process. Syst. 35 (NeurIPS ’22), pp. 2522–2536, 2022
2022
-
[168]
NeRAF: 3D scene infused neural radiance and acoustic fields,
A. Brunetto, S. Hornauer, and F. Moutarde, “NeRAF: 3D scene infused neural radiance and acoustic fields,” inProc. 13th Int. Conf. Learning Representations (ICLR ’25), (Singapore), 2025, to appear
2025
-
[169]
A V-NeRF: Learning neural fields for real-world audio-visual scene synthesis,
S. Liang et al., “A V-NeRF: Learning neural fields for real-world audio-visual scene synthesis,” in Adv. Neural Inf. Process. Syst. 36 (NeurIPS ’23), 2023
2023
-
[170]
SOAF: Scene occlusion-aware neural acoustic field,
H. Gao et al. , “SOAF: Scene occlusion-aware neural acoustic field,” arXiv preprint arXiv:2407.02264, 2024
2024 arXiv
-
[171]
Novel-view acoustic synthesis,
C. Chen et al. , “Novel-view acoustic synthesis,” in Proc. 2023 IEEE/CVF Conf. Comput. Vision Pattern Recognition (CVPR ’23) , (Vancouver, BC, Canada), pp. 6409–6419, 2023
2023
-
[172]
Novel-view acoustic synthesis from 3D reconstructed rooms,
B. Ahn et al., “Novel-view acoustic synthesis from 3D reconstructed rooms,” in Proc. INTERSPEECH 2024, (Kos Island, Greece), pp. 3260–3264, 2024
2024
-
[173]
Source localization using distributed microphones in reverberant environments based on deep learning and ray space transform,
L. Comanducci et al., “Source localization using distributed microphones in reverberant environments based on deep learning and ray space transform,” IEEE/ACM Trans. Audio Speech Language Process., vol. 28, pp. 2238–2251, 2020
2020
-
[174]
Extending GCC-PHAT using shift equivariant neural net- works,
A. Berg et al. , “Extending GCC-PHAT using shift equivariant neural net- works,” in Proc. INTERSPEECH 2022 , (Incheon, Korea), pp. 1791–1795, 2022
2022
-
[175]
D. D.-G. Aparicio, A geometric deep learning approach to sound source local- ization and tracking. PhD thesis, Universidad de Zaragoza, Spain, 2023
2023
-
[176]
Toward learning robust con- trastive embeddings for binaural sound source localization,
D. Tang, M. Taseska, and T. van Waterschoot, “Toward learning robust con- trastive embeddings for binaural sound source localization,” Front. Neuroin- form., vol. 16, 2022. Article No. 942978
2022
-
[177]
The Neural-SRP method for universal robust multi-source tracking,
E. Grinstein et al., “The Neural-SRP method for universal robust multi-source tracking,” IEEE Open J. Signal Process., vol. 5, pp. 19–28, 2023
2023
-
[178]
Machine learning- based room acoustics using flow maps and physics-informed neural networks,
N. Borrel-Jensen, A. P. Engsig-Karup, and C.-H. Jeong, “Machine learning- based room acoustics using flow maps and physics-informed neural networks,” J. Acoust. Soc. Am., vol. 151, no. 4, pp. A232–A233, 2022
2022
-
[179]
Towards reconstruction of acoustic fields via physics- informed neural networks,
K. Niebler et al. , “Towards reconstruction of acoustic fields via physics- informed neural networks,” in Proc. 51st Int. Congress & Exposition Noise Control Eng. (INTER-NOISE ’22), vol. 265, (Glasgow, Scotland, UK), 2022
2022
-
[180]
Room impulse response reconstruction with physics- informed deep learning,
X. Karakonstantis et al., “Room impulse response reconstruction with physics- informed deep learning,” J. Acoust. Soc. Am., vol. 155, no. 2, pp. 1048–1059, 2024
2024
-
[181]
Implicit neural representation with physics-informed neural networks for the reconstruction of the early part of room impulse responses,
M. Pezzoli, F. Antonacci, and A. Sarti, “Implicit neural representation with physics-informed neural networks for the reconstruction of the early part of room impulse responses,” inProc. Forum Acusticum 2023, (Turin, Italy), 2023
2023
-
[182]
Spatial extrapolation of early room impulse responses with noise-robust physics-informed neural network,
I. Tsunokini et al., “Spatial extrapolation of early room impulse responses with noise-robust physics-informed neural network,” IEICE Trans. Fundam. Elec- tron. Commun. Comput. Sci., Article No. 2024EAL2015, 2024
2024
-
[183]
Physics-informed neural network for volumetric sound field reconstruction of speech signals,
M. Olivieri et al., “Physics-informed neural network for volumetric sound field reconstruction of speech signals,” EURASIP J. Audio Speech Music Process., vol. 2024, 2024. Article No. 42
2024
-
[184]
Sound field reconstruction using a com- pact acoustics-informed neural network,
F. Ma, S. Zhao, and I. S. Burnett, “Sound field reconstruction using a com- pact acoustics-informed neural network,” J. Acoust. Soc. Am., vol. 156, no. 3, pp. 2009–2021, 2024
2009
-
[185]
Sound field estimation using deep kernel learning regularized by the wave equation,
D. Sundstr ¨om, S. Koyama, and A. Jakobsson, “Sound field estimation using deep kernel learning regularized by the wave equation,” in Proc. 2024 Int. Workshop Acoustic Signal Enhancement (IWAENC ’24), (Aalborg, Denmark), pp. 319–323, 2024
2024
-
[186]
Generative adversarial networks with physical sound field priors,
X. Karakonstantis and E. Fernandez-Grande, “Generative adversarial networks with physical sound field priors,”J. Acoust. Soc. Am., vol. 154, no. 2, pp. 1226– 1238, 2023
2023
-
[187]
Generative models for sound field reconstruc- tion,
E. Fernandez-Grande et al., “Generative models for sound field reconstruc- tion,” J. Acoust. Soc. Am., vol. 153, no. 2, pp. 1179–1190, 2023
2023
-
[188]
Building and evaluation of a real room impulse response dataset,
I. Szoke et al. , “Building and evaluation of a real room impulse response dataset,” IEEE J. Select. Topics Signal Process. , vol. 13, no. 4, pp. 863–876, 2019
2019
-
[189]
MIRaGe: Multichannel database of room impulse responses measured on high-resolution cube-shaped grid,
J. ˇCmejla et al., “MIRaGe: Multichannel database of room impulse responses measured on high-resolution cube-shaped grid,” inProc. 28th European Signal Process. Conf. (EUSIPCO ‘20) , (Amsterdam, The Netherlands), pp. 56–60, 2020
2020
-
[190]
dEchorate: a calibrated room impulse response dataset for echo-aware signal processing,
D. D. Carlo et al., “dEchorate: a calibrated room impulse response dataset for echo-aware signal processing,” EURASIP J. Audio Speech Music Process., vol. 2021, 2021. Article No. 39
2021
-
[191]
MeshRIR: A dataset of room impulse responses on meshed grid points for evaluating sound field analysis and synthesis methods,
S. Koyama et al., “MeshRIR: A dataset of room impulse responses on meshed grid points for evaluating sound field analysis and synthesis methods,” inProc. 2021 IEEE Workshop Appls. Signal Process. Audio Acoust. (WASPAA ‘21) , (New Paltz, NY , USA), pp. 1–5, 2021
2021
-
[192]
A room impulse response database for multizone sound field reproduction,
S. Zhao et al., “A room impulse response database for multizone sound field reproduction,” J. Acoust. Soc. Am., vol. 152, no. 4, pp. 2505–2512, 2022
2022
-
[193]
MYRiAD: a multi-array room acoustic database,
T. Dietzen et al., “MYRiAD: a multi-array room acoustic database,”EURASIP J. Audio Speech Music Process., vol. 2023, 2023. Article No. 17
2023
-
[194]
BRUDEX database: Binaural room impulse responses with uniformly distributed external microphones,
D. Fejgin, W. Middelberg, and S. Doclo, “BRUDEX database: Binaural room impulse responses with uniformly distributed external microphones,” in Proc. 15th ITG Conf. Speech Commun., (Aachen, Germany), pp. 126–130, 2023
2023
-
[195]
Real acoustic fields: An audio-visual room acoustics dataset and benchmark,
Z. Chen et al., “Real acoustic fields: An audio-visual room acoustics dataset and benchmark,” in Proc. 2024 IEEE/CVF Conf. Comput. Vision Pattern Recognition (CVPR ’24), (Seattle, W A, USA), pp. 21886–21896, 2024
2024
-
[196]
Room impulse response dataset of a recording studio with variable wall paneling measured using a 32- channel spherical microphone array and a B-Format microphone array,
G. Chesworth, A. Bastine, and T. Abhayapala, “Room impulse response dataset of a recording studio with variable wall paneling measured using a 32- channel spherical microphone array and a B-Format microphone array,” Appl. Sci., vol. 14, no. 5, 2024. Article No. 2095
2024
-
[197]
RealMAN: A real-recorded and annotated microphone array dataset for dynamic speech enhancement and localization,
B. Yang et al., “RealMAN: A real-recorded and annotated microphone array dataset for dynamic speech enhancement and localization,” in Adv. Neural Inf. Process. Syst. 37 (NeurIPS ’24), pp. 105997–106019, 2024
2024
-
[198]
MIRACLE – a microphone array impulse response dataset for acoustic learning,
A. Kujawski, A. J. Pelling, and E. Sarradj, “MIRACLE – a microphone array impulse response dataset for acoustic learning,” EURASIP J. Audio Speech Music Process., vol. 2024, 2024. Article No. 32
2024
-
[200]
Dataset of directional room impulse responses for realistic speech data,
S. Fragner et al., “Dataset of directional room impulse responses for realistic speech data,” Data in Brief, vol. 53, 2024. Article No. 110229
2024
-
[201]
Spatial room impulse response dataset: a robot’s journey through coupled rooms of a reverberant university building,
G. Stolz et al. , “Spatial room impulse response dataset: a robot’s journey through coupled rooms of a reverberant university building,” in Proc. DAGA 2024, (Hannover, Germany), pp. 245–247, 2024
2024
-
[202]
The tRIRjectory database: room acoustic recordings along a trajectory of moving microphones,
S. Damiano, K. MacWilliam, et al., “The tRIRjectory database: room acoustic recordings along a trajectory of moving microphones,” tech. rep., KU Leuven, Leuven, Belgium, 2025
2025
-
[203]
Sound field reconstruction using physics-informed boundary integral networks,
S. Damiano and T. van Waterschoot, “Sound field reconstruction using physics-informed boundary integral networks,” in Proc. 33rd European Sig- nal Process. Conf. (EUSIPCO ‘25) , (Palermo, Sicily, Italy), 2025, submitted for publication
2025
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.