REVIEW 3 major objections 4 minor 1 cited by
ASAudio: A Survey of Advanced Spatial Audio Research
T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This survey claims to organize all of spatial audio research into a single taxonomy by input-output representation and by generation versus understanding tasks, complete with datasets, metrics, and benchmarks.
desk verdict A plausible survey skeleton that may fill a real gap, but the abstract alone can't support the comprehensiveness claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing taxonomy itself is the central mechanism: a two-axis classification of spatial audio research by input-output representation (e.g., mono, stereo, binaural, Ambisonics, object-based, scene-based audio) and by task type (generation versus understanding). This scheme is what lets the survey systematically arrange otherwise scattered papers, datasets, and metrics into a coherent landscape.
What would settle it
Take a random sample of recent spatial audio papers from major audio and graphics venues and check whether each fits cleanly into the survey's taxonomy and appears in its references; if a substantial fraction do not, or if widely used benchmarks (e.g., standard binaural or Ambisonics datasets) are missing from the compilation, the comprehensiveness claim is refuted.
Extended reading notes
Core claim
The paper's central claim is that a comprehensive, systematically organized survey of spatial audio is now possible, and it presents one. The organization uses two axes: the input-output representation of the audio signal and the split between generation tasks (synthesizing or rendering spatial audio) and understanding tasks (analyzing or extracting information from it). Alongside the literature review, the survey catalogs datasets, evaluation metrics, and benchmarks, offering a training-perspective and an evaluation-perspective view of the field. The authors maintain a public repository of related materials.
Load-bearing premise
The survey's value rests on the assumption that its chosen taxonomy and selected literature set comprehensively cover the field; if major research lines are omitted or the categorization misrepresents the work, the survey's central purpose fails.
Editorial extensions
If this is right
- Researchers entering spatial audio can use the taxonomy as a roadmap to locate relevant methods and open problems.
- The compiled datasets and metrics provide a standardized basis for comparing future spatial audio systems.
- The generation-versus-understanding split clarifies where the field has concentrated effort and where it has not.
- The repository of related materials becomes a living index that can be updated as the field evolves.
Reading between the lines
- The representation-based axis implies that under-explored input-output pairs (for example, object-based input to binaural output) are natural targets for new synthesis pipelines, a direction the survey does not explicitly single out.
- Since generation and understanding tasks are evaluated with separate metrics, a unified metric that scores both spatial fidelity and semantic content would be a testable next step that the survey's own structure makes visible.
- If this taxonomy becomes canonical, it could serve as a shared citation framework for the field, but maintaining that status would depend on the repository being continuously updated with new work.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submission, as provided, consists solely of an abstract announcing a comprehensive survey of spatial audio. The abstract states that the paper provides a comprehensive overview, systematically reviews recent literature, organizes studies by input-output representations and by generation versus understanding tasks, and also reviews related datasets, evaluation metrics, and benchmarks. A GitHub repository is listed as a source of related materials. No full text, sections, equations, tables, or references are included in the manuscript available for review. The central claim of the paper—that it is a comprehensive and systematic survey—can therefore be neither confirmed nor refuted from the submitted content.
Significance. If the full paper delivers what the abstract promises, it would fill a notable gap: the abstract correctly notes a lack of recent comprehensive surveys in spatial audio despite rapid developments in AR/VR. The proposed organizing principle of input-output representations and generation versus understanding tasks is a reasonable and potentially useful structure, and the inclusion of datasets, metrics, and benchmarks would add practical value. The availability of a GitHub materials page is also a strength. However, because only the abstract is present, the significance is conditional: the actual breadth, accuracy, and systematic quality of the survey cannot be assessed. The abstract alone provides no evidence that the taxonomy partitions the field or that the literature selection is representative.
major comments (3)
- [Abstract] The central claim that this is a 'comprehensive overview' and a 'systematic review' of spatial audio is unsupported by the submitted text. The abstract gives no scope boundaries, inclusion/exclusion criteria, or outline of the research areas covered. A reader cannot determine whether major lines such as wave field synthesis, HRTF personalization, object-based audio, or spatial audio coding are included. This makes the central claim unfalsifiable from the manuscript, and because comprehensiveness is the primary value of a survey, this is a load-bearing issue.
- [Abstract] The categorization scheme is announced but not defined. The abstract says work is categorized by 'input-output representations' and 'generation and understanding tasks' but does not state what these categories are or how they partition the space of spatial audio research. Without such definitions, the reader cannot judge whether the taxonomy is complete or whether it overlaps or omits important subfields. This omission directly affects the survey's systematic contribution.
- [Abstract] The abstract promises a review of 'datasets, evaluation metrics, and benchmarks,' but no specific dataset, metric, or benchmark is mentioned in the available text. The GitHub repository link is provided, but the repository itself is not part of the manuscript and cannot be evaluated. As submitted, this contribution is unverifiable, and no details allow a reader to assess its scope or depth.
minor comments (4)
- [Abstract] In the sentence 'we chronologically outlining existing work,' the verb should be 'outline' to match the parallel structure with 'categorize.'
- [Abstract] The phrase 'AR, VR, and other scenarios' is vague; specifying application domains (e.g., gaming, teleconferencing, autonomous driving) would clarify the relevance of spatial audio.
- [Abstract] The GitHub URL is given without a version, commit hash, or access date. For a survey paper, including a stable reference (e.g., DOI or specific release) would improve reproducibility.
- [Abstract] The term 'input-output representations' could be more precise, perhaps 'input/output representations,' to avoid ambiguity about whether multiple representations are used for each of input and output.
Circularity Check
Survey paper with no derivation chain; no circularity found.
full rationale
ASAudio is a literature survey. Its stated contribution is to organize and review existing spatial audio research by taxonomy (input-output representations, generation vs. understanding tasks) and to compile datasets, metrics, and benchmarks. There is no fitted parameter, no predictive claim derived from an equation, and no uniqueness theorem or formal derivation whose conclusion is equivalent to its premises. The taxonomy is a categorization scheme, not a mathematical result; its completeness is an editorial judgment subject to coverage risk, not circular reasoning. No self-citation is load-bearing in the excerpt, and no result is renamed as a prediction. Under the hard rules, absence of a derivation chain means the appropriate finding is no significant circularity, score 0.
Assumptions & free parameters
assumptions (1)
- domain assumption Spatial audio research can be meaningfully categorized by input-output representations and by generation vs. understanding tasks.
Cite this review
Pith. "Pith review of ASAudio: A Survey of Advanced Spatial Audio Research." pith.science (2026). https://pith.science/paper/GTIVJN3Z
@misc{pith2026250810924,
author = {Pith},
title = {Pith review of: ASAudio: A Survey of Advanced Spatial Audio Research},
year = {2026},
howpublished = {\url{https://pith.science/paper/GTIVJN3Z}},
note = {Machine review of arXiv:2508.10924}
}
read the original abstract
With the rapid development of spatial audio technologies today, applications in AR, VR, and other scenarios have garnered extensive attention. Unlike traditional mono sound, spatial audio offers a more realistic and immersive auditory experience. Despite notable progress in the field, there remains a lack of comprehensive surveys that systematically organize and analyze these methods and their underlying technologies. In this paper, we provide a comprehensive overview of spatial audio and systematically review recent literature in the area. To address this, we chronologically outlining existing work related to spatial audio and categorize these studies based on input-output representations, as well as generation and understanding tasks, thereby summarizing various research aspects of spatial audio. In addition, we review related datasets, evaluation metrics, and benchmarks, offering insights from both training and evaluation perspectives. Related materials are available at https://github.com/dieKarotte/ASAudio.
Forward citations
Cited by 1 Pith paper
-
Dual-BEATs: Unlocking Zero-Shot Stereo Audio Perception in Audio Large Language Models via Dithering
Uncorrelated dither noise lets dual frozen BEATs encoders preserve inter-channel amplitude differences across LLM normalizers, yielding up to 97% left/center/right accuracy and zero-shot spatial generalization.
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023
arXiv 2023
-
[3]
Sound event localization and detection of overlapping sources using convolutional recurrent neural networks
Adavanne, S., Politis, A., Nikunen, J., and Virtanen, T. Sound event localization and detection of overlapping sources using convolutional recurrent neural networks. IEEE Journal of Selected Topics in Signal Processing, 13 0 (1): 0 34--48, 2018 a
2018
-
[4]
Direction of arrival estimation for multiple sound sources using convolutional recurrent neural network
Adavanne, S., Politis, A., and Virtanen, T. Direction of arrival estimation for multiple sound sources using convolutional recurrent neural network. In 2018 26th European Signal Processing Conference (EUSIPCO), pp.\ 1462--1466. IEEE, 2018 b
2018
-
[5]
Beamforming techniques for multichannel audio signal separation
Adel, H., Souad, M., Alaqeeli, A., and Hamid, A. Beamforming techniques for multichannel audio signal separation. arXiv preprint arXiv:1212.6080, 2012
arXiv 2012
-
[6]
The Conversation: Deep Audio-Visual Speech Enhancement
Afouras, T., Chung, J. S., and Zisserman, A. The conversation: Deep audio-visual speech enhancement. arXiv preprint arXiv:1804.04121, 2018
work page Pith review arXiv 2018
-
[7]
Ahn, B., Yang, K., Hamilton, B., Sheaffer, J., Ranjan, A., Sarabia, M., Tuzel, O., and Chang, J.-H. R. Novel-view acoustic synthesis from 3d reconstructed rooms. arXiv preprint arXiv:2310.15130, 2023
work page Pith review arXiv 2023
-
[8]
L., and Rafaely, B
Arbel, L., Ananthabhotla, I., Ben-Hur, Z., Alon, D. L., and Rafaely, B. On hrtf notch frequency prediction using anthropometric features and neural networks. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 816--820. IEEE, 2024
2024
Show all 204 references
-
[9]
Urban sound classification: striving towards a fair comparison
Arnault, A., Hanssens, B., and Riche, N. Urban sound classification: striving towards a fair comparison. arXiv preprint arXiv:2010.11805, 2020
2010 arXiv
-
[10]
and Schechner, Y
Barzelay, Z. and Schechner, Y. Y. Harmony in motion. In 2007 IEEE Conference on Computer Vision and Pattern Recognition, pp.\ 1--8. IEEE, 2007
2007
-
[11]
A french corpus for distant-microphone speech processing in real homes
Bertin, N., Camberlein, E., Vincent, E., Lebarbenchon, R., Peillon, S., Lamand \'e , \'E ., Sivasankaran, S., Bimbot, F., Illina, I., Tom, A., et al. A french corpus for distant-microphone speech processing in real homes. In Interspeech 2016, 2016
2016
-
[12]
J., Gannot, S., and Gerstoft, P
Bianco, M. J., Gannot, S., and Gerstoft, P. Semi-supervised source localization with deep generative modeling. In 2020 IEEE 30th International Workshop on Machine Learning for Signal Processing (MLSP), pp.\ 1--6. IEEE, 2020
2020
-
[13]
J., Gannot, S., Fernandez-Grande, E., and Gerstoft, P
Bianco, M. J., Gannot, S., Fernandez-Grande, E., and Gerstoft, P. Semi-supervised source localization in reverberant environments with deep generative modeling. IEEE Access, 9: 0 84956--84970, 2021
2021
-
[14]
On the authenticity of individual dynamic binaural synthesis
Brinkmann, F., Lindau, A., and Weinzierl, S. On the authenticity of individual dynamic binaural synthesis. The Journal of the Acoustical Society of America, 142 0 (4): 0 1784--1795, 2017
2017
-
[15]
The importance of spatial audio in modern games and virtual environments
Broderick, J., Duggan, J., and Redfern, S. The importance of spatial audio in modern games and virtual environments. In 2018 IEEE games, entertainment, media conference (GEM), pp.\ 1--9. IEEE, 2018
2018
-
[16]
Secl-umons database for sound event classification and localization
Brousmiche, M., Rouat, J., and Dupont, S. Secl-umons database for sound event classification and localization. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 756--760. IEEE, 2020
2020
-
[17]
Brown, C. P. and Duda, R. O. A structural model for binaural sound synthesis. IEEE transactions on speech and audio processing, 6 0 (5): 0 476--488, 1998
1998
-
[18]
Bryan, N. J. Impulse response data augmentation and deep neural networks for blind room acoustic parameter estimation. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1--5. IEEE, 2020
2020
-
[19]
Cao, Y., Kong, Q., Iqbal, T., An, F., Wang, W., and Plumbley, M. D. Polyphonic sound event detection and localization using a two-stage strategy. CoRR, abs/1905.00268, 2019. URL http://arxiv.org/abs/1905.00268
1905 arXiv
-
[20]
Cao, Y., Iqbal, T., Kong, Q., An, F., Wang, W., and Plumbley, M. D. An improved event-independent network for polyphonic sound event localization and detection. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 885--889...
2021
-
[21]
Matterport3d: Learning from rgb-d data in indoor environments
Chang, A., Dai, A., Funkhouser, T., Halber, M., Niessner, M., Savva, M., Song, S., Zeng, A., and Zhang, Y. Matterport3d: Learning from rgb-d data in indoor environments. arXiv preprint arXiv:1709.06158, 2017
2017 arXiv
-
[22]
Chen, C., Jain, U., Schissler, C., Gari, S. V. A., Al-Halah, Z., Ithapu, V. K., Robinson, P., and Grauman, K. Soundspaces: Audio-visual navigation in 3d environments. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part V...
2020
-
[23]
Visual acoustic matching
Chen, C., Gao, R., Calamia, P., and Grauman, K. Visual acoustic matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 18858--18868, 2022
2022
-
[24]
Spatial video streaming on apple vision pro xr headset
Chen, G., Wang, S., Chakareski, J., Koutsonikolas, D., and Dasari, M. Spatial video streaming on apple vision pro xr headset. In Proceedings of the 26th International Workshop on Mobile Computing Systems and Applications, pp.\ 115--120, 2025
2025
-
[25]
Vggsound: A large-scale audio-visual dataset
Chen, H., Xie, W., Vedaldi, A., and Zisserman, A. Vggsound: A large-scale audio-visual dataset. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 721--725. IEEE, 2020 b
2020
-
[26]
iquery: Instruments as queries for audio-visual sound separation
Chen, J., Zhang, R., Lian, D., Yang, J., Zeng, Z., and Shi, J. iquery: Instruments as queries for audio-visual sound separation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 14675--14686, 2023 a
2023
-
[27]
and Shlizerman, E
Chen, M. and Shlizerman, E. Av-cloud: Spatial audio rendering through audio-visual cloud splatting. Advances in Neural Information Processing Systems, 37: 0 141021--141044, 2024
2024
-
[28]
Chen, X., Ma, F., Zhang, Y., Bastine, A., and Samarasinghe, P. N. Head-related transfer function interpolation with a spherical cnn. arXiv preprint arXiv:2309.08290, 2023 b
2023 arXiv
-
[29]
D., Richardt, C., Kumar, A., Laney, W., Owens, A., and Richard, A
Chen, Z., Gebru, I. D., Richardt, C., Kumar, A., Laney, W., Owens, A., and Richard, A. Real acoustic fields: An audio-visual room acoustics dataset and benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 21886--21896, 2024
2024
-
[30]
Omnisep: Unified omni-modality sound separation with query-mixup
Cheng, X., Zheng, S., Wang, Z., Fang, M., Zhang, Z., Huang, R., Ma, Z., Ji, S., Zuo, J., Jin, T., et al. Omnisep: Unified omni-modality sound separation with query-mixup. arXiv preprint arXiv:2410.21269, 2024
2024 arXiv
-
[31]
Multi-channel mosra: Mean opinion score and room acoustics estimation using simulated data and a teacher model
Coldenhoff, J., Harper, A., Kendrick, P., Stojkovic, T., and Cernak, M. Multi-channel mosra: Mean opinion score and room acoustics estimation using simulated data and a teacher model. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing ...
2024
-
[32]
Simple and controllable music generation
Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., and D \'e fossez, A. Simple and controllable music generation. Advances in Neural Information Processing Systems, 36: 0 47704--47720, 2023
2023
-
[33]
See-2-sound: Zero-shot spatial environment-to-spatial sound
Dagli, R., Prakash, S., Wu, R., and Khosravani, H. See-2-sound: Zero-shot spatial environment-to-spatial sound. arXiv preprint arXiv:2406.06612, 2024
2024 arXiv
-
[34]
Fma: A dataset for music analysis
Defferrard, M., Benzi, K., Vandergheynst, P., and Bresson, X. Fma: A dataset for music analysis. arXiv preprint arXiv:1612.01840, 2016
2016 arXiv
-
[35]
Demucs: Deep extractor for music sources with extra unlabeled data remixed
D \'e fossez, A., Usunier, N., Bottou, L., and Bach, F. Demucs: Deep extractor for music sources with extra unlabeled data remixed. arXiv preprint arXiv:1909.01174, 2019
1909 arXiv
-
[36]
and Horaud, R
Deleforge, A. and Horaud, R. The cocktail party robot: Sound source separation and localisation with an active binaural head. In Proceedings of the seventh annual ACM/IEEE international conference on Human-Robot Interaction, pp.\ 431--438, 2012
2012
-
[37]
Learning spatially-aware language and audio embeddings
Devnani, B., Seto, S., Aldeneh, Z., Toso, A., Menyaylenko, E., Theobald, B.-J., Sheaffer, J., and Sarabia, M. Learning spatially-aware language and audio embeddings. Advances in Neural Information Processing Systems, 37: 0 33505--33537, 2024
2024
-
[38]
dechorate: a calibrated room impulse response database for echo-aware signal processing
Di Carlo, D., Tandeitnik, P., Foy, C., Deleforge, A., Bertin, N., and Gannot, S. dechorate: a calibrated room impulse response database for echo-aware signal processing. arXiv preprint arXiv:2104.13168, 2021
2021 arXiv
-
[39]
K., and Mehra, R
Donley, J., Tourbabin, V., Lee, J.-S., Broyles, M., Jiang, H., Shen, J., Pantic, M., Ithapu, V. K., and Mehra, R. Easycom: An augmented reality dataset to support algorithms for easy communication in noisy environments. arXiv preprint arXiv:2107.04174, 2021
2021 arXiv
-
[40]
T., and Rubinstein, M
Ephrat, A., Mosseri, I., Lang, O., Dekel, T., Wilson, K., Hassidim, A., Freeman, W. T., and Rubinstein, M. Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation. arXiv preprint arXiv:1804.03619, 2018
2018 arXiv
-
[41]
H., and Pons, J
Evans, Z., Carr, C., Taylor, J., Hawley, S. H., and Pons, J. Fast timing-conditioned latent audio diffusion. In Forty-first International Conference on Machine Learning, 2024 a
2024
-
[42]
D., Carr, C., Zukowski, Z., Taylor, J., and Pons, J
Evans, Z., Parker, J. D., Carr, C., Zukowski, Z., Taylor, J., and Pons, J. Long-form music generation with latent diffusion. arXiv preprint arXiv:2404.10301, 2024 b
2024 arXiv
-
[43]
W., Darrell, T., Freeman, W., and Viola, P
Fisher III, J. W., Darrell, T., Freeman, W., and Viola, P. Learning joint statistical models for audio-visual fusion and segregation. Advances in neural information processing systems, 13, 2000
2000
-
[44]
Fsd50k: an open dataset of human-labeled sound events
Fonseca, E., Favory, X., Pons, J., Font, F., and Serra, X. Fsd50k: an open dataset of human-labeled sound events. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30: 0 829--852, 2021
2021
-
[45]
Freesound technical demo
Font, F., Roma, G., and Serra, X. Freesound technical demo. In Proceedings of the 21st ACM international conference on Multimedia, pp.\ 411--412, 2013
2013
-
[46]
and Scholz, M
Forman, G. and Scholz, M. Apples-to-apples in cross-validation studies: pitfalls in classifier performance measurement. Acm Sigkdd Explorations Newsletter, 12 0 (1): 0 49--57, 2010
2010
-
[47]
Self-supervised moving vehicle tracking with stereo sound
Gan, C., Zhao, H., Chen, P., Cox, D., and Torralba, A. Self-supervised moving vehicle tracking with stereo sound. In Proceedings of the IEEE/CVF international conference on computer vision, pp.\ 7053--7062, 2019
2019
-
[48]
and Grauman, K
Gao, R. and Grauman, K. 2.5 d visual sound. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 324--333, 2019
2019
-
[49]
and Grauman, K
Gao, R. and Grauman, K. Visualvoice: Audio-visual speech separation with cross-modal consistency. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.\ 15490--15500. IEEE, 2021
2021
-
[50]
Learning to separate object sounds by watching unlabeled video
Gao, R., Feris, R., and Grauman, K. Learning to separate object sounds by watching unlabeled video. In Proceedings of the European conference on computer vision (ECCV), pp.\ 35--53, 2018
2018
-
[51]
A., Politis, A., Mesaros, A., Guti \'e rrez-Arriola, J
Garc \' a-Barrios, G., Krause, D. A., Politis, A., Mesaros, A., Guti \'e rrez-Arriola, J. M., and Fraile, R. Binaural source localization using deep learning and head rotation information. In 2022 30th European Signal Processing Conference (EUSIPCO), pp.\ 36--40. IEEE, 2022
2022
-
[52]
Geometry-aware multi-task learning for binaural audio generation from video
Garg, R., Gao, R., and Grauman, K. Geometry-aware multi-task learning for binaural audio generation from video. arXiv preprint arXiv:2111.10882, 2021
2021 arXiv
-
[53]
Visually-guided audio spatialization in video with geometry-aware multi-task learning
Garg, R., Gao, R., and Grauman, K. Visually-guided audio spatialization in video with geometry-aware multi-task learning. International Journal of Computer Vision, 131 0 (10): 0 2723--2737, 2023
2023
-
[54]
Bird: Big impulse response dataset
Grondin, F., Lauzon, J.-S., Michaud, S., Ravanelli, M., and Michaud, F. Bird: Big impulse response dataset. arXiv preprint arXiv:2010.09930, 2020
2010 arXiv
-
[55]
and Seguier, R
Guezenoc, C. and Seguier, R. Hrtf individualization: A survey. arXiv preprint arXiv:2003.06183, 2020
2003 arXiv
-
[56]
Psychoacoustic cues in room size perception
Hameed, S., Pakarinen, J., Valde, K., and Pulkki, V. Psychoacoustic cues in room size perception. In Audio Engineering Society Convention 116. Audio Engineering Society, 2004
2004
-
[57]
Real-time binaural speech separation with preserved spatial cues
Han, C., Luo, Y., and Mesgarani, N. Real-time binaural speech separation with preserved spatial cues. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 6404--6408. IEEE, 2020
2020
-
[58]
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp.\ 770--778, 2016
2016
-
[59]
Immersediffusion: A generative spatial audio latent diffusion model
Heydari, M., Souden, M., Conejo, B., and Atkins, J. Immersediffusion: A generative spatial audio latent diffusion model. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1--5. IEEE, 2025
2025
-
[60]
Hinton, G. E. and Salakhutdinov, R. R. Reducing the dimensionality of data with neural networks. science, 313 0 (5786): 0 504--507, 2006
2006
-
[61]
Classification of spatial audio location and content using convolutional neural networks
Hirvonen, T. Classification of spatial audio location and content using convolutional neural networks. In Audio Engineering Society Convention 138. Audio Engineering Society, 2015
2015
-
[62]
O., Jenkins, M., Liu, H., Squires, I., Cooper, S
Hogg, A. O., Jenkins, M., Liu, H., Squires, I., Cooper, S. J., and Picinali, L. Hrtf upsampling with a generative adversarial network using a gnomonic equiangular projection. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024
2024
-
[63]
C., Markovic, D., Richard, A., Gebru, I
Huang, W. C., Markovic, D., Richard, A., Gebru, I. D., and Menon, A. End-to-end binaural speech synthesis. arXiv preprint arXiv:2207.03697, 2022
2022 arXiv
-
[64]
A time-domain unsupervised learning based sound source localization method
Huang, Y., Wu, X., and Qu, T. A time-domain unsupervised learning based sound source localization method. In 2020 IEEE 3rd International Conference on Information Communication and Signal Processing (ICICSP), pp.\ 26--32. IEEE, 2020
2020
-
[65]
acoustics—measurement of room acoustic parameters—part 1: Performance spaces,
ISO, E. 3382-1, 2009,“acoustics—measurement of room acoustic parameters—part 1: Performance spaces,”. International Organization for Standardization, Brussels, Belgium, 69, 2009
2009
-
[66]
Head-related transfer function interpolation from spatially sparse measurements using autoencoder with source position conditioning
Ito, Y., Nakamura, T., Koyama, S., and Saruwatari, H. Head-related transfer function interpolation from spatially sparse measurements using autoencoder with source position conditioning. In 2022 International Workshop on Acoustic Signal Enhancement (IWAENC), pp.\ 1--5. IEEE, 2022
2022
-
[67]
A binaural room impulse response database for the evaluation of dereverberation algorithms
Jeub, M., Schafer, M., and Vary, P. A binaural room impulse response database for the evaluation of dereverberation algorithms. In 2009 16th international conference on digital signal processing, pp.\ 1--5. IEEE, 2009
2009
-
[68]
Exploring audio-visual information fusion for sound event localization and detection in low-resource realistic scenarios
Jiang, Y., Wang, Q., Du, J., Hu, M., Hu, P., Liu, Z., Cheng, S., Nian, Z., Dong, Y., Cai, M., et al. Exploring audio-visual information fusion for sound event localization and detection in low-resource realistic scenarios. In 2024 IEEE International Conference on Multimedia an...
2024
-
[69]
Modeling individual head-related transfer functions from sparse measurements using a convolutional neural network
Jiang, Z., Sang, J., Zheng, C., Li, A., and Li, X. Modeling individual head-related transfer functions from sparse measurements using a convolutional neural network. The Journal of the Acoustical Society of America, 153 0 (1): 0 248--259, 2023
2023
-
[70]
L., Gan, W.-S., et al
Jianjun, H., Tan, E. L., Gan, W.-S., et al. Natural sound rendering for headphones: integration of signal processing techniques. IEEE Signal Processing Magazine, 32 0 (2): 0 100--113, 2015
2015
-
[71]
J., and Hilton, A
Kim, H., Remaggi, L., Jackson, P. J., and Hilton, A. Immersive spatial audio reproduction for vr/ar using room acoustic modelling from 360 images. In 2019 IEEE Conference on Virtual Reality and 3D User Interfaces (VR), pp.\ 120--126. IEEE, 2019
2019
-
[72]
Visage: Video-to-spatial audio generation
Kim, J., Yun, H., and Kim, G. Visage: Video-to-spatial audio generation. In The Thirteenth International Conference on Learning Representations, 2025
2025
-
[73]
S., Park, H
Kim, J. S., Park, H. J., Shin, W., and Han, S. W. Ad-yolo: You look only once in training multiple sound event localization and detection. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1--5. IEEE, 2023
2023
-
[74]
Kingma, D. P. and Welling, M. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013
2013 arXiv
-
[75]
Habets, E
Kinoshita, K., Delcroix, M., Gannot, S., P. Habets, E. A., Haeb-Umbach, R., Kellermann, W., Leutnant, V., Maas, R., Nakatani, T., Raj, B., et al. A summary of the reverb challenge: state-of-the-art and remaining challenges in reverberant speech processing research. EURASIP Jou...
2016
-
[76]
Panoptic segmentation
Kirillov, A., He, K., Girshick, R., Rother, C., and Doll \'a r, P. Panoptic segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 9404--9413, 2019
2019
-
[77]
Efficient training of audio transformers with patchout
Koutini, K., Schl \"u ter, J., Eghbal-Zadeh, H., and Widmer, G. Efficient training of audio transformers with patchout. arXiv preprint arXiv:2110.05069, 2021
2021 arXiv
-
[78]
Meshrir: A dataset of room impulse responses on meshed grid points for evaluating sound field analysis and synthesis methods
Koyama, S., Nishida, T., Kimura, K., Abe, T., Ueno, N., and Brunnstr \"o m, J. Meshrir: A dataset of room impulse responses on meshed grid points for evaluating sound field analysis and synthesis methods. In 2021 IEEE workshop on applications of signal processing to audio and ...
2021
-
[79]
A., Garc \' a-Barrios, G., Politis, A., and Mesaros, A
Krause, D. A., Garc \' a-Barrios, G., Politis, A., and Mesaros, A. Binaural sound source distance estimation and localization for a moving listener. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023
2023
-
[80]
Audiogen: Textually guided audio generation
Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., D \'e fossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y. Audiogen: Textually guided audio generation. arXiv preprint arXiv:2209.15352, 2022
2022 arXiv
-
[81]
Bast: Binaural audio spectrogram transformer for binaural sound localization
Kuang, S., Shi, J., van der Heijden, K., and Mehrkanoon, S. Bast: Binaural audio spectrogram transformer for binaural sound localization. arXiv preprint arXiv:2207.03927, 2022
2022 arXiv
-
[82]
S., Ma, J., Thomas, M
Kushwaha, S. S., Ma, J., Thomas, M. R., Tian, Y., and Bruni, A. Diff-sage: End-to-end spatial audio generation using diffusion models. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1--5. IEEE, 2025
2025
-
[83]
Sound event detection in synthetic audio: Analysis of the dcase 2016 task results
Lafay, G., Benetos, E., and Lagrange, M. Sound event detection in synthetic audio: Analysis of the dcase 2016 task results. In 2017 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), pp.\ 11--15. IEEE, 2017
2016
-
[84]
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86 0 (11): 0 2278--2324, 1998
1998
-
[85]
Context-based evaluation of the opus audio codec for spatial audio content in virtual reality
Lee, B., Rudzki, T., Skoglund, J., and Kearney, G. Context-based evaluation of the opus audio codec for spatial audio content in virtual reality. Journal of the Audio Engineering Society, 71: 0 145--154, 04 2023 a . doi:10.17743/jaes.2022.0068
2023
-
[86]
Looking into your speech: Learning cross-modal affinity for audio-visual speech separation
Lee, J., Chung, S.-W., Kim, S., Kang, H.-G., and Sohn, K. Looking into your speech: Learning cross-modal affinity for audio-visual speech separation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 1336--1345, 2021
2021
-
[87]
Lee, J. W. and Lee, K. Neural fourier shift for binaural speech rendering. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1--5. IEEE, 2023
2023
-
[88]
W., Lee, S., and Lee, K
Lee, J. W., Lee, S., and Lee, K. Global hrtf interpolation via learned affine transformation of hyper-conditioned features. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1--5. IEEE, 2023 b
2023
-
[89]
Binauralgrad: A two-stage conditional diffusion probabilistic model for binaural audio synthesis
Leng, Y., Chen, Z., Guo, J., Liu, H., Chen, J., Tan, X., Mandic, D., He, L., Li, X., Qin, T., et al. Binauralgrad: A two-stage conditional diffusion probabilistic model for binaural audio synthesis. Advances in Neural Information Processing Systems, 35: 0 23689--23700, 2022
2022
-
[90]
See and listen: Score-informed association of sound tracks to players in chamber music performance videos
Li, B., Dinesh, K., Duan, Z., and Sharma, G. See and listen: Score-informed association of sound tracks to players in chamber music performance videos. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 2906--2910. IEEE, 2017
2017
-
[91]
Sonicsim: A customizable simulation platform for speech processing in moving sound source scenarios
Li, K., Sang, W., Zeng, C., Yang, R., Chen, G., and Hu, X. Sonicsim: A customizable simulation platform for speech processing in moving sound source scenarios. arXiv preprint arXiv:2410.01481, 2024 a
2024 arXiv
-
[92]
Binaural audio generation via multi-task learning
Li, S., Liu, S., and Manocha, D. Binaural audio generation via multi-task learning. ACM Transactions on Graphics (TOG), 40 0 (6): 0 1--13, 2021
2021
-
[93]
Binauralmusic: A diverse dataset for improving cross-modal binaural audio generation
Li, Y., Liu, S., Cheng, H., and Ye, L. Binauralmusic: A diverse dataset for improving cross-modal binaural audio generation. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 7990--7994. IEEE, 2024 b
2024
-
[94]
Cross-modal generative model for visual-guided binaural stereo generation
Li, Z., Zhao, B., and Yuan, Y. Cross-modal generative model for visual-guided binaural stereo generation. Knowledge-Based Systems, 296: 0 111814, 2024 c
2024
-
[95]
Av-nerf: Learning neural fields for real-world audio-visual scene synthesis
Liang, S., Huang, C., Tian, Y., Kumar, A., and Xu, C. Av-nerf: Learning neural fields for real-world audio-visual scene synthesis. Advances in Neural Information Processing Systems, 36: 0 37472--37490, 2023
2023
-
[96]
and Nam, J
Lim, W. and Nam, J. Enhancing spatial audio generation with source separation and channel panning loss. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 8321--8325. IEEE, 2024
2024
-
[97]
T., Ben-Hamu, H., Nickel, M., and Le, M
Lipman, Y., Chen, R. T., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, 2022
2022
-
[98]
Liu, H., Chen, Z., Yuan, Y., Mei, X., Liu, X., Mandic, D., Wang, W., and Plumbley, M. D. Audioldm: Text-to-audio generation with latent diffusion models. arXiv preprint arXiv:2301.12503, 2023
2023 arXiv
-
[99]
Omniaudio: Generating spatial audio from 360-degree video
Liu, H., Luo, T., Jiang, Q., Luo, K., Sun, P., Wan, J., Huang, R., Chen, Q., Wang, W., Li, X., et al. Omniaudio: Generating spatial audio from 360-degree video. arXiv preprint arXiv:2504.14906, 2025 a
2025 arXiv
-
[100]
Dopplerbas: Binaural audio synthesis addressing doppler effect
Liu, J., Ye, Z., Chen, Q., Zheng, S., Wang, W., Zhang, Q., and Zhao, Z. Dopplerbas: Binaural audio synthesis addressing doppler effect. arXiv preprint arXiv:2212.07000, 2022
2022 arXiv
-
[101]
Visually guided binaural audio generation with cross-modal consistency
Liu, M., Wang, J., Qian, X., and Xie, X. Visually guided binaural audio generation with cross-modal consistency. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 7980--7984. IEEE, 2024
2024
-
[102]
Visual-based spatial audio generation system for multi-speaker environments
Liu, X., Gurelli, O., Wang, Y., and Reiss, J. Visual-based spatial audio generation system for multi-speaker environments. arXiv preprint arXiv:2502.07538, 2025 b
2025 arXiv
-
[103]
Learning neural acoustic fields
Luo, A., Du, Y., Tarr, M., Tenenbaum, J., Torralba, A., and Gan, C. Learning neural acoustic fields. Advances in Neural Information Processing Systems, 35: 0 3165--3177, 2022
2022
-
[104]
and Mesgarani, N
Luo, Y. and Mesgarani, N. Conv-tasnet: Surpassing ideal time--frequency magnitude masking for speech separation. IEEE/ACM transactions on audio, speech, and language processing, 27 0 (8): 0 1256--1266, 2019
2019
-
[105]
D., Samarasinghe, P
Ma, F., Abhayapala, T. D., Samarasinghe, P. N., and Chen, X. Spatial upsampling of head-related transfer functions using a physics-informed neural network. arXiv preprint arXiv:2307.14650, 2023
2023 arXiv
-
[106]
Few-shot audio-visual learning of environment acoustics
Majumder, S., Chen, C., Al-Halah, Z., and Grauman, K. Few-shot audio-visual learning of environment acoustics. Advances in Neural Information Processing Systems, 35: 0 2522--2536, 2022
2022
-
[107]
Hrtf recommendation based on the predicted binaural colouration model
Marggraf-Turley, N., Lovedee-Turner, M., and De Sena, E. Hrtf recommendation based on the predicted binaural colouration model. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1106--1110. IEEE, 2024
2024
-
[108]
G., Pan, Z., Khurana, S., Hori, C., and Le Roux, J
Masuyama, Y., Wichern, G., Germain, F. G., Pan, Z., Khurana, S., Hori, C., and Le Roux, J. Niirf: Neural iir filter field for hrtf upsampling and personalization. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1016--...
2024
-
[109]
Hidden markov model training with contaminated speech material for distant-talking speech recognition
Matassoni, M., Omologo, M., Giuliani, D., and Svaizer, P. Hidden markov model training with contaminated speech material for distant-talking speech recognition. Computer Speech & Language, 16 0 (2): 0 205--223, 2002
2002
-
[110]
A probabilistic model for robust localization based on a binaural auditory front-end
May, T., Van De Par, S., and Kohlrausch, A. A probabilistic model for robust localization based on a binaural auditory front-end. IEEE Transactions on audio, speech, and language processing, 19 0 (1): 0 1--13, 2010
2010
-
[111]
Dataset of spatial room impulse responses in a variable acoustics room for six degrees-of-freedom rendering and analysis
McKenzie, T., McCormack, L., and Hold, C. Dataset of spatial room impulse responses in a variable acoustics room for six degrees-of-freedom rendering and analysis. arXiv preprint arXiv:2111.11882, 2021 a
2021 arXiv
-
[112]
J., and Pulkki, V
McKenzie, T., Schlecht, S. J., and Pulkki, V. Acoustic analysis and dataset of transitions between coupled rooms. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 481--485. IEEE, 2021 b
2021
-
[113]
Metrics for polyphonic sound event detection
Mesaros, A., Heittola, T., and Virtanen, T. Metrics for polyphonic sound event detection. Applied Sciences, 6 0 (6): 0 162, 2016
2016
-
[114]
Joint measurement of localization and detection of sound events
Mesaros, A., Adavanne, S., Politis, A., Heittola, T., and Virtanen, T. Joint measurement of localization and detection of sound events. In 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), pp.\ 333--337. IEEE, 2019
2019
-
[115]
P., Tancik, M., Barron, J
Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., and Ng, R. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65 0 (1): 0 99--106, 2021
2021
-
[116]
Dataset of binaural room impulse responses at multiple recording positions, source positions, and orientations in a real room
Mittag, C., B \"o hme, M., Werner, S., and Klein, S. Dataset of binaural room impulse responses at multiple recording positions, source positions, and orientations in a real room. In in Proc. of the 43rd annual convention for acoustics, DAGA, Germany, 2017
2017
-
[117]
Fundamentals of binaural technology
M ller, H. Fundamentals of binaural technology. Applied acoustics, 36 0 (3-4): 0 171--218, 1992
1992
-
[118]
Self-supervised generation of spatial audio for 360 video
Morgado, P., Nvasconcelos, N., Langlois, T., and Wang, O. Self-supervised generation of spatial audio for 360 video. Advances in neural information processing systems, 31, 2018
2018
-
[119]
Learning representations from audio-visual spatial alignment
Morgado, P., Li, Y., and Nvasconcelos, N. Learning representations from audio-visual spatial alignment. Advances in Neural Information Processing Systems, 33: 0 4733--4744, 2020
2020
-
[120]
and Neff, F
Murphy, D. and Neff, F. Spatial sound for computer games and virtual reality. In Game sound technology and player interaction: Concepts and developments, pp.\ 287--312. IGI Global Scientific Publishing, 2011
2011
-
[121]
Murphy, D. T. and Shelley, S. Openair: An interactive auralization web resource and database. In Audio Engineering Society Convention 129. Audio Engineering Society, 2010
2010
-
[122]
Wearable seld dataset: Dataset for sound event localization and detection using wearable devices around head
Nagatomo, K., Yasuda, M., Yatabe, K., Saito, S., and Oikawa, Y. Wearable seld dataset: Dataset for sound event localization and detection using wearable devices around head. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), ...
2022
-
[123]
Blind and neural network-guided convolutional beamformer for joint denoising, dereverberation, and source separation
Nakatani, T., Ikeshita, R., Kinoshita, K., Sawada, H., and Araki, S. Blind and neural network-guided convolutional beamformer for joint denoising, dereverberation, and source separation. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processi...
2021
-
[124]
Sound event localization and detection using squeeze-excitation residual cnns
Naranjo-Alcazar, J., Perez-Castanos, S., Ferrandis, J., Zuccarello, P., and Cobos, M. Sound event localization and detection using squeeze-excitation residual cnns. arXiv preprint arXiv:2006.14436, 2020
2006 arXiv
-
[125]
Nguyen, T. N. T., Jones, D. L., and Gan, W.-S. A sequence matching network for polyphonic sound event localization and detection. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 71--75. IEEE, 2020
2020
-
[126]
Nguyen, T. N. T., Watcharasupat, K. N., Nguyen, N. K., Jones, D. L., and Gan, W.-S. Salsa: Spatial cue-augmented log-spectrogram features for polyphonic sound event localization and detection. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30: 0 1749--1762, 2022
2022
-
[127]
A., and Prakash, R
Nik Khah, A., Htun, C. A., and Prakash, R. Unsupervised bayesian surprise detection in spatial audio with convolutional variational autoencoder and lstm model. In Proceedings of the 2024 ACM International Conference on Interactive Media Experiences Workshops, pp.\ 116--121, 2024
2024
-
[128]
A., Liutkus, A., and Vincent, E
Nugraha, A. A., Liutkus, A., and Vincent, E. Multichannel audio source separation with deep neural networks. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 24 0 (9): 0 1652--1664, 2016
2016
-
[129]
and Efros, A
Owens, A. and Efros, A. A. Audio-visual scene analysis with self-supervised multisensory features. In Proceedings of the European conference on computer vision (ECCV), pp.\ 631--648, 2018
2018
-
[130]
Librispeech: an asr corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S. Librispeech: an asr corpus based on public domain audio books. In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), pp.\ 5206--5210. IEEE, 2015
2015
-
[131]
Q., P \'e rez, P., and Richard, G
Parekh, S., Essid, S., Ozerov, A., Duong, N. Q., P \'e rez, P., and Richard, G. Motion informed audio source separation. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 6--10. IEEE, 2017
2017
-
[132]
3d localization of multiple sound sources with intensity vector estimates in single source zones
Pavlidi, D., Delikaris-Manias, S., Pulkki, V., and Mouchtaris, A. 3d localization of multiple sound sources with intensity vector estimates in single source zones. In 2015 23rd European Signal Processing Conference (EUSIPCO), pp.\ 1556--1560. IEEE, 2015
2015
-
[133]
Integration of spatial sound in immersive virtual environments an experimental study on effects of spatial sound on presence
Poeschl, S., Wall, K., and Doering, N. Integration of spatial sound in immersive virtual environments an experimental study on effects of spatial sound on presence. In 2013 IEEE Virtual Reality (VR), pp.\ 129--130. IEEE, 2013
2013
-
[134]
A dataset of reverberant spatial sound scenes with moving sources for sound event localization and detection
Politis, A., Adavanne, S., and Virtanen, T. A dataset of reverberant spatial sound scenes with moving sources for sound event localization and detection. arXiv preprint arXiv:2006.01919, 2020
2006 arXiv
-
[135]
A dataset of dynamic reverberant sound scenes with directional interferers for sound event localization and detection
Politis, A., Adavanne, S., Krause, D., Deleforge, A., Srivastava, P., and Virtanen, T. A dataset of dynamic reverberant sound scenes with directional interferers for sound event localization and detection. arXiv preprint arXiv:2106.06999, 2021
2021 arXiv
-
[136]
Starss22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events
Politis, A., Shimada, K., Sudarsanam, P., Adavanne, S., Krause, D., Koyama, Y., Takahashi, N., Takahashi, S., Mitsufuji, Y., and Virtanen, T. Starss22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events. arXiv preprint arXiv:2206.01948, 2022
2022 arXiv
-
[137]
Audio-visual object localization and separation using low-rank and sparsity
Pu, J., Panagakis, Y., Petridis, S., and Pantic, M. Audio-visual object localization and separation using low-rank and sparsity. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 2901--2905. IEEE, 2017
2017
-
[138]
Audio-visual cross-attention network for robotic speaker tracking
Qian, X., Wang, Z., Wang, J., Guan, G., and Li, H. Audio-visual cross-attention network for robotic speaker tracking. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31: 0 550--562, 2022
2022
-
[139]
K., Sundaresha, V., Rajagopalan, A., et al
Rachavarapu, K. K., Sundaresha, V., Rajagopalan, A., et al. Localize to binauralize: Audio spatialization from visual sound source localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp.\ 1930--1939, 2021
1930
-
[140]
Towards generating ambisonics using audio-visual cue for virtual reality
Rana, A., Ozcinar, C., and Smolic, A. Towards generating ambisonics using audio-visual cue for virtual reality. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 2012--2016. IEEE, 2019
2019
-
[141]
and Manocha, D
Ratnarajah, A. and Manocha, D. Listen2scene: Interactive material-aware binaural sound propagation for reconstructed 3d scenes. In 2024 IEEE Conference Virtual Reality and 3D User Interfaces (VR), pp.\ 254--264. IEEE, 2024
2024
-
[142]
Mesh2ir: Neural acoustic impulse response generator for complex 3d scenes
Ratnarajah, A., Tang, Z., Aralikatti, R., and Manocha, D. Mesh2ir: Neural acoustic impulse response generator for complex 3d scenes. In Proceedings of the 30th ACM International Conference on Multimedia, pp.\ 924--933, 2022
2022
-
[143]
Av-rir: Audio-visual room impulse response estimation
Ratnarajah, A., Ghosh, S., Kumar, S., Chiniya, P., and Manocha, D. Av-rir: Audio-visual room impulse response estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 27164--27175, 2024
2024
-
[144]
Impulse response estimation for robust speech recognition in a reverberant environment
Ravanelli, M., Sosi, A., Svaizer, P., and Omologo, M. Impulse response estimation for robust speech recognition in a reverberant environment. In 2012 Proceedings of the 20th European Signal Processing Conference (EUSIPCO), pp.\ 1668--1672. IEEE, 2012
2012
-
[145]
The dirha-english corpus and related tasks for distant-speech recognition in domestic environments
Ravanelli, M., Cristoforetti, L., Gretter, R., Pellin, M., Sosi, A., and Omologo, M. The dirha-english corpus and related tasks for distant-speech recognition in domestic environments. In 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), pp.\ 275--28...
2015
-
[146]
D., Krenn, S., Butler, G
Richard, A., Markovic, D., Gebru, I. D., Krenn, S., Butler, G. A., Torre, F., and Sheikh, Y. Neural synthesis of binaural speech from mono audio. In International Conference on Learning Representations, 2021
2021
-
[147]
R., Ick, C., Ding, S., Roman, A
Roman, I. R., Ick, C., Ding, S., Roman, A. S., McFee, B., and Bello, J. P. Spatial scaper: a library to simulate and augment soundscapes for sound event localization and detection in realistic rooms. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Si...
2024
-
[148]
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Medical image computing and computer-assisted intervention--MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18, ...
2015
-
[149]
E., Hinton, G
Rumelhart, D. E., Hinton, G. E., and Williams, R. J. Learning representations by back-propagating errors. nature, 323 0 (6088): 0 533--536, 1986
1986
-
[150]
w2v-seld: A sound event localization and detection framework for self-supervised spatial audio pre-training
Santos, O., Rosero, K., Masiero, B., and de Alencar Lotufo, R. w2v-seld: A sound event localization and detection framework for self-supervised spatial audio pre-training. IEEE Access, 2024
2024
-
[151]
Spatial librispeech: An augmented dataset for spatial audio learning
Sarabia, M., Menyaylenko, E., Toso, A., Seto, S., Aldeneh, Z., Pirhosseinloo, S., Zappella, L., Theobald, B.-J., Apostoloff, N., and Sheaffer, J. Spatial librispeech: An augmented dataset for spatial audio learning. arXiv preprint arXiv:2308.09514, 2023
2023 arXiv
-
[152]
and Svensson, U
Savioja, L. and Svensson, U. P. Overview of geometrical room acoustic modeling techniques. The Journal of the Acoustical Society of America, 138 0 (2): 0 708--730, 2015
2015
-
[153]
Habitat: A platform for embodied ai research
Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., et al. Habitat: A platform for embodied ai research. In Proceedings of the IEEE/CVF international conference on computer vision, pp.\ 9339--9347, 2019
2019
-
[154]
Mo \^ usai: Text-to-music generation with long-context latent diffusion
Schneider, F., Kamal, O., Jin, Z., and Sch \"o lkopf, B. Mo \^ usai: Text-to-music generation with long-context latent diffusion. arXiv preprint arXiv:2301.11757, 2023
2023 arXiv
-
[155]
Two multimodal approaches for single microphone source separation
Sedighin, F., Babaie-Zadeh, M., Rivet, B., and Jutten, C. Two multimodal approaches for single microphone source separation. In 2016 24th European Signal Processing Conference (EUSIPCO), pp.\ 110--114. IEEE, 2016
2016
-
[156]
Senocak, A., Oh, T.-H., Kim, J., Yang, M.-H., and Kweon, I. S. Learning to localize sound source in visual scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp.\ 4358--4366, 2018
2018
-
[157]
Accdoa: Activity-coupled cartesian direction of arrival representation for sound event localization and detection
Shimada, K., Koyama, Y., Takahashi, N., Takahashi, S., and Mitsufuji, Y. Accdoa: Activity-coupled cartesian direction of arrival representation for sound event localization and detection. In ICASSP 2021-2021 IEEE international conference on acoustics, speech and signal process...
2021
-
[158]
Multi-accdoa: Localizing and detecting overlapping sounds from the same class with auxiliary duplicating permutation invariant training
Shimada, K., Koyama, Y., Takahashi, S., Takahashi, N., Tsunoo, E., and Mitsufuji, Y. Multi-accdoa: Localizing and detecting overlapping sounds from the same class with auxiliary duplicating permutation invariant training. In ICASSP 2022-2022 IEEE international conference on ac...
2022
-
[159]
A., Uchida, K., Adavanne, S., Hakala, A., Koyama, Y., Takahashi, N., Takahashi, S., et al
Shimada, K., Politis, A., Sudarsanam, P., Krause, D. A., Uchida, K., Adavanne, S., Hakala, A., Koyama, Y., Takahashi, N., Takahashi, S., et al. Starss23: An audio-visual dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events. Advances in n...
2023
-
[160]
and Casey, M
Smaragdis, P. and Casey, M. Audio/visual independent components. In Proc. ICA, pp.\ 709--714, 2003
2003
-
[161]
Blind room parameter estimation using multiple multichannel speech recordings
Srivastava, P., Deleforge, A., and Vincent, E. Blind room parameter estimation using multiple multichannel speech recordings. In 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), pp.\ 226--230. IEEE, 2021
2021
-
[162]
Both ears wide open: Towards language-driven spatial audio generation
Sun, P., Cheng, S., Li, X., Ye, Z., Liu, H., Zhang, H., Xue, W., and Guo, Y. Both ears wide open: Towards language-driven spatial audio generation. arXiv preprint arXiv:2410.10676, 2024
2024 arXiv
-
[163]
Learning audio-visual source localization via false negative aware contrastive learning
Sun, W., Zhang, J., Wang, J., Liu, Z., Zhong, Y., Feng, T., Guo, Y., Zhang, Y., and Barnes, N. Learning audio-visual source localization via false negative aware contrastive learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), ...
2023
-
[164]
Sonicmotion: Dynamic spatial audio soundscapes with latent diffusion models
Templin, C., Zhu, Y., and Wang, H. Sonicmotion: Dynamic spatial audio soundscapes with latent diffusion models. arXiv preprint arXiv:2507.07318, 2025
2025
-
[165]
T., and V \"a lim \"a ki, V
Thuillier, E., Jin, C. T., and V \"a lim \"a ki, V. Hrtf interpolation using a spherical neural process meta-learner. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 32: 0 1790--1802, 2024
2024
-
[166]
Audio-visual event localization in unconstrained videos
Tian, Y., Shi, J., Li, B., Duan, Z., and Xu, C. Audio-visual event localization in unconstrained videos. In Proceedings of the European conference on computer vision (ECCV), pp.\ 247--263, 2018
2018
-
[167]
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023
2023 arXiv
-
[168]
The nigens general sound events database
Trowitzsch, I., Taghia, J., Kashef, Y., and Obermayer, K. The nigens general sound events database. arXiv preprint arXiv:1902.08314, 2019
1902 arXiv
-
[169]
P., and Hershey, J
Tzinis, E., Wisdom, S., Jansen, A., Hershey, S., Remez, T., Ellis, D. P., and Hershey, J. R. Into the wild with audioscope: Unsupervised audio-visual separation of on-screen sounds. arXiv preprint arXiv:2011.01143, 2020
2011 arXiv
-
[170]
Tzinis, E., Wisdom, S., Remez, T., and Hershey, J. R. Audioscopev2: Audio-visual attention architectures for calibrated open-domain on-screen sound separation. In European Conference on Computer Vision, pp.\ 368--385. Springer, 2022
2022
-
[171]
The sweet-home speech and multimodal corpus for home automation interaction
Vacher, M., Lecouteux, B., Chahuara, P., Portet, F., Meillon, B., and Bonnefond, N. The sweet-home speech and multimodal corpus for home automation interaction. In The 9th edition of the Language Resources and Evaluation Conference (LREC), pp.\ 4499--4506, 2014
2014
-
[172]
Heavenly mathematics: The forgotten art of spherical trigonometry
Van Brummelen, G. Heavenly mathematics: The forgotten art of spherical trigonometry. Princeton University Press, 2012
2012
-
[173]
and Sambath, M
Varghese, R. and Sambath, M. Yolov8: A novel object detection algorithm with enhanced performance and robustness. In 2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS), pp.\ 1--6. IEEE, 2024
2024
-
[174]
N., Kaiser, ., and Polosukhin, I
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, ., and Polosukhin, I. Attention is all you need. In Advances in neural information processing systems, pp.\ 5998--6008, 2017
2017
-
[175]
and Wang, D
Wang, Z.-Q. and Wang, D. Combining spectral and spatial features for deep learning based blind speaker separation. IEEE/ACM Transactions on audio, speech, and language processing, 27 0 (2): 0 457--468, 2018
2018
-
[176]
Wang, Z.-Q., Le Roux, J., and Hershey, J. R. Multi-channel deep clustering: Discriminative spectral and spatial embeddings for speaker-independent speech separation. In 2018 IEEE International conference on acoustics, speech and signal processing (ICASSP), pp.\ 1--5. IEEE, 2018
2018
-
[177]
Warnecke, M., Jamison, S., Prepelita, S., Calamia, P., and Ithapu, V. K. Hrtf personalization based on ear morphology. In Audio Engineering Society Conference: 2022 AES International Conference on Audio for Virtual and Augmented Reality. Audio Engineering Society, 2022
2022
-
[178]
J., Mandel, M
Weiss, R. J., Mandel, M. I., and Ellis, D. P. Source separation based on binaural cues and source model constraints. 2009
2009
-
[179]
J., Bharadia, D., and Gerstoft, P
Wu, Y., Ayyalasomayajula, R., Bianco, M. J., Bharadia, D., and Gerstoft, P. Sslide: Sound source localization for indoors based on deep learning. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 4680--4684. IEEE, 2021
2021
-
[180]
and Moreira Kares, E
Wuolio, L. and Moreira Kares, E. On the potential of spatial audio in enhancing virtual user experiences. 2023
2023
-
[181]
Sonic4d: Spatial audio generation for immersive 4d scene exploration
Xie, S., Zhu, H., He, T., Li, X., and Chen, Z. Sonic4d: Spatial audio generation for immersive 4d scene exploration. arXiv preprint arXiv:2506.15759, 2025
2025
-
[182]
Visually informed binaural audio generation without binaural audios
Xu, X., Zhou, H., Liu, Z., Dai, B., Wang, X., and Lin, D. Visually informed binaural audio generation without binaural audios. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 15485--15494, 2021
2021
-
[183]
Realman: A real-recorded and annotated microphone array dataset for dynamic speech enhancement and localization
Yang, B., Quan, C., Wang, Y., Wang, P., Yang, Y., Fang, Y., Shao, N., Bu, H., Xu, X., and Li, X. Realman: A real-recorded and annotated microphone array dataset for dynamic speech enhancement and localization. Advances in Neural Information Processing Systems, 37: 0 105997--10...
2024
-
[184]
Upmixing via style transfer: a variational autoencoder for disentangling spatial images and musical content
Yang, H., Wager, S., Russell, S., Luo, M., Kim, M., and Kim, W. Upmixing via style transfer: a variational autoencoder for disentangling spatial images and musical content. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p...
2022
-
[185]
Telling left from right: Learning spatial correspondence of sight and sound
Yang, K., Russell, B., and Salamon, J. Telling left from right: Learning spatial correspondence of sight and sound. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 9932--9941, 2020
2020
-
[186]
Depth anything: Unleashing the power of large-scale unlabeled data
Yang, L., Kang, B., Huang, Z., Xu, X., Feng, J., and Zhao, H. Depth anything: Unleashing the power of large-scale unlabeled data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 10371--10381, 2024 b
2024
-
[187]
and Zheng, Y
Yang, Q. and Zheng, Y. Deepear: Sound localization with binaural microphones. IEEE Transactions on Mobile Computing, 23 0 (1): 0 359--375, 2022
2022
-
[188]
Sound event localization based on sound intensity vector refined by dnn-based denoising and source separation
Yasuda, M., Koizumi, Y., Saito, S., Uematsu, H., and Imoto, K. Sound event localization based on sound intensity vector refined by dnn-based denoising and source separation. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), ...
2020
-
[189]
Lavss: Location-guided audio-visual spatial audio separation
Ye, Y., Yang, W., and Tian, Y. Lavss: Location-guided audio-visual spatial audio separation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp.\ 5508--5519, 2024
2024
-
[190]
Multi-microphone neural speech separation for far-field multi-talker speech recognition
Yoshioka, T., Erdogan, H., Chen, Z., and Alleva, F. Multi-microphone neural speech separation for far-field multi-talker speech recognition. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 5739--5743. IEEE, 2018
2018
-
[191]
Ambisonizer: Neural upmixing as spherical harmonics generation
Zang, Y., Wang, Y., and Lee, M. Ambisonizer: Neural upmixing as spherical harmonics generation. arXiv preprint arXiv:2405.13428, 2024
2024 arXiv
-
[192]
and Shao, J
Zhang, W. and Shao, J. Multi-attention audio-visual fusion network for audio spatialization. In Proceedings of the 2021 International Conference on Multimedia Retrieval, pp.\ 394--401, 2021
2021
-
[193]
N., Chen, H., and Abhayapala, T
Zhang, W., Samarasinghe, P. N., Chen, H., and Abhayapala, T. D. Surround by sound: A review of spatial audio recording and reproduction. Applied Sciences, 7 0 (5): 0 532, 2017
2017
-
[194]
and Wang, D
Zhang, X. and Wang, D. Deep learning based binaural speech separation in reverberant environments. IEEE/ACM transactions on audio, speech, and language processing, 25 0 (5): 0 1075--1084, 2017
2017
-
[195]
Hrtf field: Unifying measured hrtf magnitude representation with neural fields
Zhang, Y., Wang, Y., and Duan, Z. Hrtf field: Unifying measured hrtf magnitude representation with neural fields. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1--5. IEEE, 2023
2023
-
[196]
Isdrama: Immersive spatial drama generation through multimodal prompting
Zhang, Y., Guo, W., Pan, C., Zhu, Z., Jin, T., and Zhao, Z. Isdrama: Immersive spatial drama generation through multimodal prompting. arXiv preprint arXiv:2504.20630, 2025
2025
-
[197]
The sound of pixels
Zhao, H., Gan, C., Rouditchenko, A., Vondrick, C., McDermott, J., and Torralba, A. The sound of pixels. In Proceedings of the European conference on computer vision (ECCV), pp.\ 570--586, 2018
2018
-
[198]
Dualspec: Text-to-spatial-audio generation via dual-spectrogram guided diffusion model
Zhao, L., Chen, S., Feng, L., Zhang, X.-L., and Li, X. Dualspec: Text-to-spatial-audio generation via dual-spectrogram guided diffusion model. arXiv preprint arXiv:2502.18952, 2025
2025 arXiv
-
[199]
Magnitude modeling of personalized hrtf based on ear images and anthropometric measurements
Zhao, M., Sheng, Z., and Fang, Y. Magnitude modeling of personalized hrtf based on ear images and anthropometric measurements. Applied Sciences, 12 0 (16): 0 8155, 2022
2022
-
[200]
Interpretable binaural ratio for visually guided binaural audio generation
Zheng, T., Verma, S., and Liu, W. Interpretable binaural ratio for visually guided binaural audio generation. In 2022 International Joint Conference on Neural Networks (IJCNN), pp.\ 1--8. IEEE, 2022
2022
-
[201]
Bat: Learning to reason about spatial sounds with large language models
Zheng, Z., Peng, P., Ma, Z., Chen, X., Choi, E., and Harwath, D. Bat: Learning to reason about spatial sounds with large language models. arXiv preprint arXiv:2402.01591, 2024
2024 arXiv
-
[202]
Sep-stereo: Visually guided stereophonic audio generation by associating source separation
Zhou, H., Xu, X., Lin, D., Wang, X., and Liu, Z. Sep-stereo: Visually guided stereophonic audio generation by associating source separation. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part XII 16, pp.\ 52--69. Springer, 2020
2020
-
[203]
Zhou, Y., Wang, Z., Fang, C., Bui, T., and Berg, T. L. Visual to sound: Generating natural sound for videos in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp.\ 3550--3558, 2018
2018
-
[204]
Zmolikova, K., Delcroix, M., Burget, L., Nakatani, T., and C ernocky, J. H. Integration of variational autoencoder and spatial clustering for adaptive multi-channel neural speech separation. In 2021 IEEE Spoken Language Technology Workshop (SLT), pp.\ 889--896. IEEE, 2021
2021
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.