Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

ASAudio: A Survey of Advanced Spatial Audio Research

T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This survey claims to organize all of spatial audio research into a single taxonomy by input-output representation and by generation versus understanding tasks, complete with datasets, metrics, and benchmarks.

desk verdict A plausible survey skeleton that may fill a real gap, but the abstract alone can't support the comprehensiveness claim. read the letter →

arxiv 2508.10924 v2 pith:GTIVJN3Z submitted 2025-08-08 eess.AS cs.SD

classification eess.AScs.SD
keywords spatialaudiosurveyinput-outputrepresentationgenerationunderstandingdatasetsevaluationmetricsAR/VR
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper attempts to fill a gap: spatial audio is advancing rapidly for AR, VR, and immersive media, but no comprehensive survey organizes the field. It reviews recent literature systematically, grouping work first by the input and output representations of the audio (such as mono, binaural, or Ambisonic formats) and then by whether the task is to generate or to understand spatial audio. It also compiles related datasets, evaluation metrics, and benchmarks. If successful, the survey gives researchers a shared map of the field and a starting point for comparing methods.

What carries the argument

The organizing taxonomy itself is the central mechanism: a two-axis classification of spatial audio research by input-output representation (e.g., mono, stereo, binaural, Ambisonics, object-based, scene-based audio) and by task type (generation versus understanding). This scheme is what lets the survey systematically arrange otherwise scattered papers, datasets, and metrics into a coherent landscape.

What would settle it

Take a random sample of recent spatial audio papers from major audio and graphics venues and check whether each fits cleanly into the survey's taxonomy and appears in its references; if a substantial fraction do not, or if widely used benchmarks (e.g., standard binaural or Ambisonics datasets) are missing from the compilation, the comprehensiveness claim is refuted.

Watch

Extended reading notes

Core claim

The paper's central claim is that a comprehensive, systematically organized survey of spatial audio is now possible, and it presents one. The organization uses two axes: the input-output representation of the audio signal and the split between generation tasks (synthesizing or rendering spatial audio) and understanding tasks (analyzing or extracting information from it). Alongside the literature review, the survey catalogs datasets, evaluation metrics, and benchmarks, offering a training-perspective and an evaluation-perspective view of the field. The authors maintain a public repository of related materials.

Load-bearing premise

The survey's value rests on the assumption that its chosen taxonomy and selected literature set comprehensively cover the field; if major research lines are omitted or the categorization misrepresents the work, the survey's central purpose fails.

Editorial extensions

If this is right

  • Researchers entering spatial audio can use the taxonomy as a roadmap to locate relevant methods and open problems.
  • The compiled datasets and metrics provide a standardized basis for comparing future spatial audio systems.
  • The generation-versus-understanding split clarifies where the field has concentrated effort and where it has not.
  • The repository of related materials becomes a living index that can be updated as the field evolves.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The representation-based axis implies that under-explored input-output pairs (for example, object-based input to binaural output) are natural targets for new synthesis pipelines, a direction the survey does not explicitly single out.
  • Since generation and understanding tasks are evaluated with separate metrics, a unified metric that scores both spatial fidelity and semantic content would be a testable next step that the survey's own structure makes visible.
  • If this taxonomy becomes canonical, it could serve as a shared citation framework for the field, but maintaining that status would depend on the repository being continuously updated with new work.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The submission, as provided, consists solely of an abstract announcing a comprehensive survey of spatial audio. The abstract states that the paper provides a comprehensive overview, systematically reviews recent literature, organizes studies by input-output representations and by generation versus understanding tasks, and also reviews related datasets, evaluation metrics, and benchmarks. A GitHub repository is listed as a source of related materials. No full text, sections, equations, tables, or references are included in the manuscript available for review. The central claim of the paper—that it is a comprehensive and systematic survey—can therefore be neither confirmed nor refuted from the submitted content.

Significance. If the full paper delivers what the abstract promises, it would fill a notable gap: the abstract correctly notes a lack of recent comprehensive surveys in spatial audio despite rapid developments in AR/VR. The proposed organizing principle of input-output representations and generation versus understanding tasks is a reasonable and potentially useful structure, and the inclusion of datasets, metrics, and benchmarks would add practical value. The availability of a GitHub materials page is also a strength. However, because only the abstract is present, the significance is conditional: the actual breadth, accuracy, and systematic quality of the survey cannot be assessed. The abstract alone provides no evidence that the taxonomy partitions the field or that the literature selection is representative.

major comments (3)
  1. [Abstract] The central claim that this is a 'comprehensive overview' and a 'systematic review' of spatial audio is unsupported by the submitted text. The abstract gives no scope boundaries, inclusion/exclusion criteria, or outline of the research areas covered. A reader cannot determine whether major lines such as wave field synthesis, HRTF personalization, object-based audio, or spatial audio coding are included. This makes the central claim unfalsifiable from the manuscript, and because comprehensiveness is the primary value of a survey, this is a load-bearing issue.
  2. [Abstract] The categorization scheme is announced but not defined. The abstract says work is categorized by 'input-output representations' and 'generation and understanding tasks' but does not state what these categories are or how they partition the space of spatial audio research. Without such definitions, the reader cannot judge whether the taxonomy is complete or whether it overlaps or omits important subfields. This omission directly affects the survey's systematic contribution.
  3. [Abstract] The abstract promises a review of 'datasets, evaluation metrics, and benchmarks,' but no specific dataset, metric, or benchmark is mentioned in the available text. The GitHub repository link is provided, but the repository itself is not part of the manuscript and cannot be evaluated. As submitted, this contribution is unverifiable, and no details allow a reader to assess its scope or depth.
minor comments (4)
  1. [Abstract] In the sentence 'we chronologically outlining existing work,' the verb should be 'outline' to match the parallel structure with 'categorize.'
  2. [Abstract] The phrase 'AR, VR, and other scenarios' is vague; specifying application domains (e.g., gaming, teleconferencing, autonomous driving) would clarify the relevance of spatial audio.
  3. [Abstract] The GitHub URL is given without a version, commit hash, or access date. For a survey paper, including a stable reference (e.g., DOI or specific release) would improve reproducibility.
  4. [Abstract] The term 'input-output representations' could be more precise, perhaps 'input/output representations,' to avoid ambiguity about whether multiple representations are used for each of input and output.

Circularity Check

0 steps flagged · score 0.0 of 10

Survey paper with no derivation chain; no circularity found.

full rationale

ASAudio is a literature survey. Its stated contribution is to organize and review existing spatial audio research by taxonomy (input-output representations, generation vs. understanding tasks) and to compile datasets, metrics, and benchmarks. There is no fitted parameter, no predictive claim derived from an equation, and no uniqueness theorem or formal derivation whose conclusion is equivalent to its premises. The taxonomy is a categorization scheme, not a mathematical result; its completeness is an editorial judgment subject to coverage risk, not circular reasoning. No self-citation is load-bearing in the excerpt, and no result is renamed as a prediction. Under the hard rules, absence of a derivation chain means the appropriate finding is no significant circularity, score 0.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

No derivations or new entities are introduced; the paper is a survey and only assumes that its organizational categories are appropriate.

assumptions (1)
  • domain assumption Spatial audio research can be meaningfully categorized by input-output representations and by generation vs. understanding tasks.
    The survey relies on this taxonomy being a useful and complete organizing framework; if the categories distort the field, the survey's value drops.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ASAudio: A Survey of Advanced Spatial Audio Research." pith.science (2026). https://pith.science/paper/GTIVJN3Z

@misc{pith2026250810924,
  author       = {Pith},
  title        = {Pith review of: ASAudio: A Survey of Advanced Spatial Audio Research},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GTIVJN3Z}},
  note         = {Machine review of arXiv:2508.10924}
}
read the original abstract

With the rapid development of spatial audio technologies today, applications in AR, VR, and other scenarios have garnered extensive attention. Unlike traditional mono sound, spatial audio offers a more realistic and immersive auditory experience. Despite notable progress in the field, there remains a lack of comprehensive surveys that systematically organize and analyze these methods and their underlying technologies. In this paper, we provide a comprehensive overview of spatial audio and systematically review recent literature in the area. To address this, we chronologically outlining existing work related to spatial audio and categorize these studies based on input-output representations, as well as generation and understanding tasks, thereby summarizing various research aspects of spatial audio. In addition, we review related datasets, evaluation metrics, and benchmarks, offering insights from both training and evaluation perspectives. Related materials are available at https://github.com/dieKarotte/ASAudio.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Dual-BEATs: Unlocking Zero-Shot Stereo Audio Perception in Audio Large Language Models via Dithering

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Uncorrelated dither noise lets dual frozen BEATs encoders preserve inter-channel amplitude differences across LLM normalizers, yielding up to 97% left/center/right accuracy and zero-shot spatial generalization.

Reference graph

Works this paper leans on

204 extracted references · 49 canonical work pages · cited by 1 Pith paper

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al

    Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023

  3. [3]

    Sound event localization and detection of overlapping sources using convolutional recurrent neural networks

    Adavanne, S., Politis, A., Nikunen, J., and Virtanen, T. Sound event localization and detection of overlapping sources using convolutional recurrent neural networks. IEEE Journal of Selected Topics in Signal Processing, 13 0 (1): 0 34--48, 2018 a

  4. [4]

    Direction of arrival estimation for multiple sound sources using convolutional recurrent neural network

    Adavanne, S., Politis, A., and Virtanen, T. Direction of arrival estimation for multiple sound sources using convolutional recurrent neural network. In 2018 26th European Signal Processing Conference (EUSIPCO), pp.\ 1462--1466. IEEE, 2018 b

  5. [5]

    Beamforming techniques for multichannel audio signal separation

    Adel, H., Souad, M., Alaqeeli, A., and Hamid, A. Beamforming techniques for multichannel audio signal separation. arXiv preprint arXiv:1212.6080, 2012

  6. [6]

    The Conversation: Deep Audio-Visual Speech Enhancement

    Afouras, T., Chung, J. S., and Zisserman, A. The conversation: Deep audio-visual speech enhancement. arXiv preprint arXiv:1804.04121, 2018

  7. [7]

    Ahn, B., Yang, K., Hamilton, B., Sheaffer, J., Ranjan, A., Sarabia, M., Tuzel, O., and Chang, J.-H. R. Novel-view acoustic synthesis from 3d reconstructed rooms. arXiv preprint arXiv:2310.15130, 2023

  8. [8]

    L., and Rafaely, B

    Arbel, L., Ananthabhotla, I., Ben-Hur, Z., Alon, D. L., and Rafaely, B. On hrtf notch frequency prediction using anthropometric features and neural networks. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 816--820. IEEE, 2024

Show all 204 references
  1. [9]

    Urban sound classification: striving towards a fair comparison

    Arnault, A., Hanssens, B., and Riche, N. Urban sound classification: striving towards a fair comparison. arXiv preprint arXiv:2010.11805, 2020

  2. [10]

    and Schechner, Y

    Barzelay, Z. and Schechner, Y. Y. Harmony in motion. In 2007 IEEE Conference on Computer Vision and Pattern Recognition, pp.\ 1--8. IEEE, 2007

  3. [11]

    A french corpus for distant-microphone speech processing in real homes

    Bertin, N., Camberlein, E., Vincent, E., Lebarbenchon, R., Peillon, S., Lamand \'e , \'E ., Sivasankaran, S., Bimbot, F., Illina, I., Tom, A., et al. A french corpus for distant-microphone speech processing in real homes. In Interspeech 2016, 2016

  4. [12]

    J., Gannot, S., and Gerstoft, P

    Bianco, M. J., Gannot, S., and Gerstoft, P. Semi-supervised source localization with deep generative modeling. In 2020 IEEE 30th International Workshop on Machine Learning for Signal Processing (MLSP), pp.\ 1--6. IEEE, 2020

  5. [13]

    J., Gannot, S., Fernandez-Grande, E., and Gerstoft, P

    Bianco, M. J., Gannot, S., Fernandez-Grande, E., and Gerstoft, P. Semi-supervised source localization in reverberant environments with deep generative modeling. IEEE Access, 9: 0 84956--84970, 2021

  6. [14]

    On the authenticity of individual dynamic binaural synthesis

    Brinkmann, F., Lindau, A., and Weinzierl, S. On the authenticity of individual dynamic binaural synthesis. The Journal of the Acoustical Society of America, 142 0 (4): 0 1784--1795, 2017

  7. [15]

    The importance of spatial audio in modern games and virtual environments

    Broderick, J., Duggan, J., and Redfern, S. The importance of spatial audio in modern games and virtual environments. In 2018 IEEE games, entertainment, media conference (GEM), pp.\ 1--9. IEEE, 2018

  8. [16]

    Secl-umons database for sound event classification and localization

    Brousmiche, M., Rouat, J., and Dupont, S. Secl-umons database for sound event classification and localization. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 756--760. IEEE, 2020

  9. [17]

    Brown, C. P. and Duda, R. O. A structural model for binaural sound synthesis. IEEE transactions on speech and audio processing, 6 0 (5): 0 476--488, 1998

  10. [18]

    Bryan, N. J. Impulse response data augmentation and deep neural networks for blind room acoustic parameter estimation. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1--5. IEEE, 2020

  11. [19]

    Cao, Y., Kong, Q., Iqbal, T., An, F., Wang, W., and Plumbley, M. D. Polyphonic sound event detection and localization using a two-stage strategy. CoRR, abs/1905.00268, 2019. URL http://arxiv.org/abs/1905.00268

  12. [20]

    Cao, Y., Iqbal, T., Kong, Q., An, F., Wang, W., and Plumbley, M. D. An improved event-independent network for polyphonic sound event localization and detection. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 885--889...

  13. [21]

    Matterport3d: Learning from rgb-d data in indoor environments

    Chang, A., Dai, A., Funkhouser, T., Halber, M., Niessner, M., Savva, M., Song, S., Zeng, A., and Zhang, Y. Matterport3d: Learning from rgb-d data in indoor environments. arXiv preprint arXiv:1709.06158, 2017

  14. [22]

    Chen, C., Jain, U., Schissler, C., Gari, S. V. A., Al-Halah, Z., Ithapu, V. K., Robinson, P., and Grauman, K. Soundspaces: Audio-visual navigation in 3d environments. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part V...

  15. [23]

    Visual acoustic matching

    Chen, C., Gao, R., Calamia, P., and Grauman, K. Visual acoustic matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 18858--18868, 2022

  16. [24]

    Spatial video streaming on apple vision pro xr headset

    Chen, G., Wang, S., Chakareski, J., Koutsonikolas, D., and Dasari, M. Spatial video streaming on apple vision pro xr headset. In Proceedings of the 26th International Workshop on Mobile Computing Systems and Applications, pp.\ 115--120, 2025

  17. [25]

    Vggsound: A large-scale audio-visual dataset

    Chen, H., Xie, W., Vedaldi, A., and Zisserman, A. Vggsound: A large-scale audio-visual dataset. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 721--725. IEEE, 2020 b

  18. [26]

    iquery: Instruments as queries for audio-visual sound separation

    Chen, J., Zhang, R., Lian, D., Yang, J., Zeng, Z., and Shi, J. iquery: Instruments as queries for audio-visual sound separation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 14675--14686, 2023 a

  19. [27]

    and Shlizerman, E

    Chen, M. and Shlizerman, E. Av-cloud: Spatial audio rendering through audio-visual cloud splatting. Advances in Neural Information Processing Systems, 37: 0 141021--141044, 2024

  20. [28]

    Chen, X., Ma, F., Zhang, Y., Bastine, A., and Samarasinghe, P. N. Head-related transfer function interpolation with a spherical cnn. arXiv preprint arXiv:2309.08290, 2023 b

  21. [29]

    D., Richardt, C., Kumar, A., Laney, W., Owens, A., and Richard, A

    Chen, Z., Gebru, I. D., Richardt, C., Kumar, A., Laney, W., Owens, A., and Richard, A. Real acoustic fields: An audio-visual room acoustics dataset and benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 21886--21896, 2024

  22. [30]

    Omnisep: Unified omni-modality sound separation with query-mixup

    Cheng, X., Zheng, S., Wang, Z., Fang, M., Zhang, Z., Huang, R., Ma, Z., Ji, S., Zuo, J., Jin, T., et al. Omnisep: Unified omni-modality sound separation with query-mixup. arXiv preprint arXiv:2410.21269, 2024

  23. [31]

    Multi-channel mosra: Mean opinion score and room acoustics estimation using simulated data and a teacher model

    Coldenhoff, J., Harper, A., Kendrick, P., Stojkovic, T., and Cernak, M. Multi-channel mosra: Mean opinion score and room acoustics estimation using simulated data and a teacher model. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing ...

  24. [32]

    Simple and controllable music generation

    Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., and D \'e fossez, A. Simple and controllable music generation. Advances in Neural Information Processing Systems, 36: 0 47704--47720, 2023

  25. [33]

    See-2-sound: Zero-shot spatial environment-to-spatial sound

    Dagli, R., Prakash, S., Wu, R., and Khosravani, H. See-2-sound: Zero-shot spatial environment-to-spatial sound. arXiv preprint arXiv:2406.06612, 2024

  26. [34]

    Fma: A dataset for music analysis

    Defferrard, M., Benzi, K., Vandergheynst, P., and Bresson, X. Fma: A dataset for music analysis. arXiv preprint arXiv:1612.01840, 2016

  27. [35]

    Demucs: Deep extractor for music sources with extra unlabeled data remixed

    D \'e fossez, A., Usunier, N., Bottou, L., and Bach, F. Demucs: Deep extractor for music sources with extra unlabeled data remixed. arXiv preprint arXiv:1909.01174, 2019

  28. [36]

    and Horaud, R

    Deleforge, A. and Horaud, R. The cocktail party robot: Sound source separation and localisation with an active binaural head. In Proceedings of the seventh annual ACM/IEEE international conference on Human-Robot Interaction, pp.\ 431--438, 2012

  29. [37]

    Learning spatially-aware language and audio embeddings

    Devnani, B., Seto, S., Aldeneh, Z., Toso, A., Menyaylenko, E., Theobald, B.-J., Sheaffer, J., and Sarabia, M. Learning spatially-aware language and audio embeddings. Advances in Neural Information Processing Systems, 37: 0 33505--33537, 2024

  30. [38]

    dechorate: a calibrated room impulse response database for echo-aware signal processing

    Di Carlo, D., Tandeitnik, P., Foy, C., Deleforge, A., Bertin, N., and Gannot, S. dechorate: a calibrated room impulse response database for echo-aware signal processing. arXiv preprint arXiv:2104.13168, 2021

  31. [39]

    K., and Mehra, R

    Donley, J., Tourbabin, V., Lee, J.-S., Broyles, M., Jiang, H., Shen, J., Pantic, M., Ithapu, V. K., and Mehra, R. Easycom: An augmented reality dataset to support algorithms for easy communication in noisy environments. arXiv preprint arXiv:2107.04174, 2021

  32. [40]

    T., and Rubinstein, M

    Ephrat, A., Mosseri, I., Lang, O., Dekel, T., Wilson, K., Hassidim, A., Freeman, W. T., and Rubinstein, M. Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation. arXiv preprint arXiv:1804.03619, 2018

  33. [41]

    H., and Pons, J

    Evans, Z., Carr, C., Taylor, J., Hawley, S. H., and Pons, J. Fast timing-conditioned latent audio diffusion. In Forty-first International Conference on Machine Learning, 2024 a

  34. [42]

    D., Carr, C., Zukowski, Z., Taylor, J., and Pons, J

    Evans, Z., Parker, J. D., Carr, C., Zukowski, Z., Taylor, J., and Pons, J. Long-form music generation with latent diffusion. arXiv preprint arXiv:2404.10301, 2024 b

  35. [43]

    W., Darrell, T., Freeman, W., and Viola, P

    Fisher III, J. W., Darrell, T., Freeman, W., and Viola, P. Learning joint statistical models for audio-visual fusion and segregation. Advances in neural information processing systems, 13, 2000

  36. [44]

    Fsd50k: an open dataset of human-labeled sound events

    Fonseca, E., Favory, X., Pons, J., Font, F., and Serra, X. Fsd50k: an open dataset of human-labeled sound events. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30: 0 829--852, 2021

  37. [45]

    Freesound technical demo

    Font, F., Roma, G., and Serra, X. Freesound technical demo. In Proceedings of the 21st ACM international conference on Multimedia, pp.\ 411--412, 2013

  38. [46]

    and Scholz, M

    Forman, G. and Scholz, M. Apples-to-apples in cross-validation studies: pitfalls in classifier performance measurement. Acm Sigkdd Explorations Newsletter, 12 0 (1): 0 49--57, 2010

  39. [47]

    Self-supervised moving vehicle tracking with stereo sound

    Gan, C., Zhao, H., Chen, P., Cox, D., and Torralba, A. Self-supervised moving vehicle tracking with stereo sound. In Proceedings of the IEEE/CVF international conference on computer vision, pp.\ 7053--7062, 2019

  40. [48]

    and Grauman, K

    Gao, R. and Grauman, K. 2.5 d visual sound. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 324--333, 2019

  41. [49]

    and Grauman, K

    Gao, R. and Grauman, K. Visualvoice: Audio-visual speech separation with cross-modal consistency. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.\ 15490--15500. IEEE, 2021

  42. [50]

    Learning to separate object sounds by watching unlabeled video

    Gao, R., Feris, R., and Grauman, K. Learning to separate object sounds by watching unlabeled video. In Proceedings of the European conference on computer vision (ECCV), pp.\ 35--53, 2018

  43. [51]

    A., Politis, A., Mesaros, A., Guti \'e rrez-Arriola, J

    Garc \' a-Barrios, G., Krause, D. A., Politis, A., Mesaros, A., Guti \'e rrez-Arriola, J. M., and Fraile, R. Binaural source localization using deep learning and head rotation information. In 2022 30th European Signal Processing Conference (EUSIPCO), pp.\ 36--40. IEEE, 2022

  44. [52]

    Geometry-aware multi-task learning for binaural audio generation from video

    Garg, R., Gao, R., and Grauman, K. Geometry-aware multi-task learning for binaural audio generation from video. arXiv preprint arXiv:2111.10882, 2021

  45. [53]

    Visually-guided audio spatialization in video with geometry-aware multi-task learning

    Garg, R., Gao, R., and Grauman, K. Visually-guided audio spatialization in video with geometry-aware multi-task learning. International Journal of Computer Vision, 131 0 (10): 0 2723--2737, 2023

  46. [54]

    Bird: Big impulse response dataset

    Grondin, F., Lauzon, J.-S., Michaud, S., Ravanelli, M., and Michaud, F. Bird: Big impulse response dataset. arXiv preprint arXiv:2010.09930, 2020

  47. [55]

    and Seguier, R

    Guezenoc, C. and Seguier, R. Hrtf individualization: A survey. arXiv preprint arXiv:2003.06183, 2020

  48. [56]

    Psychoacoustic cues in room size perception

    Hameed, S., Pakarinen, J., Valde, K., and Pulkki, V. Psychoacoustic cues in room size perception. In Audio Engineering Society Convention 116. Audio Engineering Society, 2004

  49. [57]

    Real-time binaural speech separation with preserved spatial cues

    Han, C., Luo, Y., and Mesgarani, N. Real-time binaural speech separation with preserved spatial cues. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 6404--6408. IEEE, 2020

  50. [58]

    Deep residual learning for image recognition

    He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp.\ 770--778, 2016

  51. [59]

    Immersediffusion: A generative spatial audio latent diffusion model

    Heydari, M., Souden, M., Conejo, B., and Atkins, J. Immersediffusion: A generative spatial audio latent diffusion model. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1--5. IEEE, 2025

  52. [60]

    Hinton, G. E. and Salakhutdinov, R. R. Reducing the dimensionality of data with neural networks. science, 313 0 (5786): 0 504--507, 2006

  53. [61]

    Classification of spatial audio location and content using convolutional neural networks

    Hirvonen, T. Classification of spatial audio location and content using convolutional neural networks. In Audio Engineering Society Convention 138. Audio Engineering Society, 2015

  54. [62]

    O., Jenkins, M., Liu, H., Squires, I., Cooper, S

    Hogg, A. O., Jenkins, M., Liu, H., Squires, I., Cooper, S. J., and Picinali, L. Hrtf upsampling with a generative adversarial network using a gnomonic equiangular projection. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024

  55. [63]

    C., Markovic, D., Richard, A., Gebru, I

    Huang, W. C., Markovic, D., Richard, A., Gebru, I. D., and Menon, A. End-to-end binaural speech synthesis. arXiv preprint arXiv:2207.03697, 2022

  56. [64]

    A time-domain unsupervised learning based sound source localization method

    Huang, Y., Wu, X., and Qu, T. A time-domain unsupervised learning based sound source localization method. In 2020 IEEE 3rd International Conference on Information Communication and Signal Processing (ICICSP), pp.\ 26--32. IEEE, 2020

  57. [65]

    acoustics—measurement of room acoustic parameters—part 1: Performance spaces,

    ISO, E. 3382-1, 2009,“acoustics—measurement of room acoustic parameters—part 1: Performance spaces,”. International Organization for Standardization, Brussels, Belgium, 69, 2009

  58. [66]

    Head-related transfer function interpolation from spatially sparse measurements using autoencoder with source position conditioning

    Ito, Y., Nakamura, T., Koyama, S., and Saruwatari, H. Head-related transfer function interpolation from spatially sparse measurements using autoencoder with source position conditioning. In 2022 International Workshop on Acoustic Signal Enhancement (IWAENC), pp.\ 1--5. IEEE, 2022

  59. [67]

    A binaural room impulse response database for the evaluation of dereverberation algorithms

    Jeub, M., Schafer, M., and Vary, P. A binaural room impulse response database for the evaluation of dereverberation algorithms. In 2009 16th international conference on digital signal processing, pp.\ 1--5. IEEE, 2009

  60. [68]

    Exploring audio-visual information fusion for sound event localization and detection in low-resource realistic scenarios

    Jiang, Y., Wang, Q., Du, J., Hu, M., Hu, P., Liu, Z., Cheng, S., Nian, Z., Dong, Y., Cai, M., et al. Exploring audio-visual information fusion for sound event localization and detection in low-resource realistic scenarios. In 2024 IEEE International Conference on Multimedia an...

  61. [69]

    Modeling individual head-related transfer functions from sparse measurements using a convolutional neural network

    Jiang, Z., Sang, J., Zheng, C., Li, A., and Li, X. Modeling individual head-related transfer functions from sparse measurements using a convolutional neural network. The Journal of the Acoustical Society of America, 153 0 (1): 0 248--259, 2023

  62. [70]

    L., Gan, W.-S., et al

    Jianjun, H., Tan, E. L., Gan, W.-S., et al. Natural sound rendering for headphones: integration of signal processing techniques. IEEE Signal Processing Magazine, 32 0 (2): 0 100--113, 2015

  63. [71]

    J., and Hilton, A

    Kim, H., Remaggi, L., Jackson, P. J., and Hilton, A. Immersive spatial audio reproduction for vr/ar using room acoustic modelling from 360 images. In 2019 IEEE Conference on Virtual Reality and 3D User Interfaces (VR), pp.\ 120--126. IEEE, 2019

  64. [72]

    Visage: Video-to-spatial audio generation

    Kim, J., Yun, H., and Kim, G. Visage: Video-to-spatial audio generation. In The Thirteenth International Conference on Learning Representations, 2025

  65. [73]

    S., Park, H

    Kim, J. S., Park, H. J., Shin, W., and Han, S. W. Ad-yolo: You look only once in training multiple sound event localization and detection. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1--5. IEEE, 2023

  66. [74]

    Kingma, D. P. and Welling, M. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013

  67. [75]

    Habets, E

    Kinoshita, K., Delcroix, M., Gannot, S., P. Habets, E. A., Haeb-Umbach, R., Kellermann, W., Leutnant, V., Maas, R., Nakatani, T., Raj, B., et al. A summary of the reverb challenge: state-of-the-art and remaining challenges in reverberant speech processing research. EURASIP Jou...

  68. [76]

    Panoptic segmentation

    Kirillov, A., He, K., Girshick, R., Rother, C., and Doll \'a r, P. Panoptic segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 9404--9413, 2019

  69. [77]

    Efficient training of audio transformers with patchout

    Koutini, K., Schl \"u ter, J., Eghbal-Zadeh, H., and Widmer, G. Efficient training of audio transformers with patchout. arXiv preprint arXiv:2110.05069, 2021

  70. [78]

    Meshrir: A dataset of room impulse responses on meshed grid points for evaluating sound field analysis and synthesis methods

    Koyama, S., Nishida, T., Kimura, K., Abe, T., Ueno, N., and Brunnstr \"o m, J. Meshrir: A dataset of room impulse responses on meshed grid points for evaluating sound field analysis and synthesis methods. In 2021 IEEE workshop on applications of signal processing to audio and ...

  71. [79]

    A., Garc \' a-Barrios, G., Politis, A., and Mesaros, A

    Krause, D. A., Garc \' a-Barrios, G., Politis, A., and Mesaros, A. Binaural sound source distance estimation and localization for a moving listener. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023

  72. [80]

    Audiogen: Textually guided audio generation

    Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., D \'e fossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y. Audiogen: Textually guided audio generation. arXiv preprint arXiv:2209.15352, 2022

  73. [81]

    Bast: Binaural audio spectrogram transformer for binaural sound localization

    Kuang, S., Shi, J., van der Heijden, K., and Mehrkanoon, S. Bast: Binaural audio spectrogram transformer for binaural sound localization. arXiv preprint arXiv:2207.03927, 2022

  74. [82]

    S., Ma, J., Thomas, M

    Kushwaha, S. S., Ma, J., Thomas, M. R., Tian, Y., and Bruni, A. Diff-sage: End-to-end spatial audio generation using diffusion models. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1--5. IEEE, 2025

  75. [83]

    Sound event detection in synthetic audio: Analysis of the dcase 2016 task results

    Lafay, G., Benetos, E., and Lagrange, M. Sound event detection in synthetic audio: Analysis of the dcase 2016 task results. In 2017 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), pp.\ 11--15. IEEE, 2017

  76. [84]

    Gradient-based learning applied to document recognition

    LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86 0 (11): 0 2278--2324, 1998

  77. [85]

    Context-based evaluation of the opus audio codec for spatial audio content in virtual reality

    Lee, B., Rudzki, T., Skoglund, J., and Kearney, G. Context-based evaluation of the opus audio codec for spatial audio content in virtual reality. Journal of the Audio Engineering Society, 71: 0 145--154, 04 2023 a . doi:10.17743/jaes.2022.0068

  78. [86]

    Looking into your speech: Learning cross-modal affinity for audio-visual speech separation

    Lee, J., Chung, S.-W., Kim, S., Kang, H.-G., and Sohn, K. Looking into your speech: Learning cross-modal affinity for audio-visual speech separation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 1336--1345, 2021

  79. [87]

    Lee, J. W. and Lee, K. Neural fourier shift for binaural speech rendering. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1--5. IEEE, 2023

  80. [88]

    W., Lee, S., and Lee, K

    Lee, J. W., Lee, S., and Lee, K. Global hrtf interpolation via learned affine transformation of hyper-conditioned features. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1--5. IEEE, 2023 b

  81. [89]

    Binauralgrad: A two-stage conditional diffusion probabilistic model for binaural audio synthesis

    Leng, Y., Chen, Z., Guo, J., Liu, H., Chen, J., Tan, X., Mandic, D., He, L., Li, X., Qin, T., et al. Binauralgrad: A two-stage conditional diffusion probabilistic model for binaural audio synthesis. Advances in Neural Information Processing Systems, 35: 0 23689--23700, 2022

  82. [90]

    See and listen: Score-informed association of sound tracks to players in chamber music performance videos

    Li, B., Dinesh, K., Duan, Z., and Sharma, G. See and listen: Score-informed association of sound tracks to players in chamber music performance videos. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 2906--2910. IEEE, 2017

  83. [91]

    Sonicsim: A customizable simulation platform for speech processing in moving sound source scenarios

    Li, K., Sang, W., Zeng, C., Yang, R., Chen, G., and Hu, X. Sonicsim: A customizable simulation platform for speech processing in moving sound source scenarios. arXiv preprint arXiv:2410.01481, 2024 a

  84. [92]

    Binaural audio generation via multi-task learning

    Li, S., Liu, S., and Manocha, D. Binaural audio generation via multi-task learning. ACM Transactions on Graphics (TOG), 40 0 (6): 0 1--13, 2021

  85. [93]

    Binauralmusic: A diverse dataset for improving cross-modal binaural audio generation

    Li, Y., Liu, S., Cheng, H., and Ye, L. Binauralmusic: A diverse dataset for improving cross-modal binaural audio generation. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 7990--7994. IEEE, 2024 b

  86. [94]

    Cross-modal generative model for visual-guided binaural stereo generation

    Li, Z., Zhao, B., and Yuan, Y. Cross-modal generative model for visual-guided binaural stereo generation. Knowledge-Based Systems, 296: 0 111814, 2024 c

  87. [95]

    Av-nerf: Learning neural fields for real-world audio-visual scene synthesis

    Liang, S., Huang, C., Tian, Y., Kumar, A., and Xu, C. Av-nerf: Learning neural fields for real-world audio-visual scene synthesis. Advances in Neural Information Processing Systems, 36: 0 37472--37490, 2023

  88. [96]

    and Nam, J

    Lim, W. and Nam, J. Enhancing spatial audio generation with source separation and channel panning loss. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 8321--8325. IEEE, 2024

  89. [97]

    T., Ben-Hamu, H., Nickel, M., and Le, M

    Lipman, Y., Chen, R. T., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, 2022

  90. [98]

    Liu, H., Chen, Z., Yuan, Y., Mei, X., Liu, X., Mandic, D., Wang, W., and Plumbley, M. D. Audioldm: Text-to-audio generation with latent diffusion models. arXiv preprint arXiv:2301.12503, 2023

  91. [99]

    Omniaudio: Generating spatial audio from 360-degree video

    Liu, H., Luo, T., Jiang, Q., Luo, K., Sun, P., Wan, J., Huang, R., Chen, Q., Wang, W., Li, X., et al. Omniaudio: Generating spatial audio from 360-degree video. arXiv preprint arXiv:2504.14906, 2025 a

  92. [100]

    Dopplerbas: Binaural audio synthesis addressing doppler effect

    Liu, J., Ye, Z., Chen, Q., Zheng, S., Wang, W., Zhang, Q., and Zhao, Z. Dopplerbas: Binaural audio synthesis addressing doppler effect. arXiv preprint arXiv:2212.07000, 2022

  93. [101]

    Visually guided binaural audio generation with cross-modal consistency

    Liu, M., Wang, J., Qian, X., and Xie, X. Visually guided binaural audio generation with cross-modal consistency. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 7980--7984. IEEE, 2024

  94. [102]

    Visual-based spatial audio generation system for multi-speaker environments

    Liu, X., Gurelli, O., Wang, Y., and Reiss, J. Visual-based spatial audio generation system for multi-speaker environments. arXiv preprint arXiv:2502.07538, 2025 b

  95. [103]

    Learning neural acoustic fields

    Luo, A., Du, Y., Tarr, M., Tenenbaum, J., Torralba, A., and Gan, C. Learning neural acoustic fields. Advances in Neural Information Processing Systems, 35: 0 3165--3177, 2022

  96. [104]

    and Mesgarani, N

    Luo, Y. and Mesgarani, N. Conv-tasnet: Surpassing ideal time--frequency magnitude masking for speech separation. IEEE/ACM transactions on audio, speech, and language processing, 27 0 (8): 0 1256--1266, 2019

  97. [105]

    D., Samarasinghe, P

    Ma, F., Abhayapala, T. D., Samarasinghe, P. N., and Chen, X. Spatial upsampling of head-related transfer functions using a physics-informed neural network. arXiv preprint arXiv:2307.14650, 2023

  98. [106]

    Few-shot audio-visual learning of environment acoustics

    Majumder, S., Chen, C., Al-Halah, Z., and Grauman, K. Few-shot audio-visual learning of environment acoustics. Advances in Neural Information Processing Systems, 35: 0 2522--2536, 2022

  99. [107]

    Hrtf recommendation based on the predicted binaural colouration model

    Marggraf-Turley, N., Lovedee-Turner, M., and De Sena, E. Hrtf recommendation based on the predicted binaural colouration model. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1106--1110. IEEE, 2024

  100. [108]

    G., Pan, Z., Khurana, S., Hori, C., and Le Roux, J

    Masuyama, Y., Wichern, G., Germain, F. G., Pan, Z., Khurana, S., Hori, C., and Le Roux, J. Niirf: Neural iir filter field for hrtf upsampling and personalization. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1016--...

  101. [109]

    Hidden markov model training with contaminated speech material for distant-talking speech recognition

    Matassoni, M., Omologo, M., Giuliani, D., and Svaizer, P. Hidden markov model training with contaminated speech material for distant-talking speech recognition. Computer Speech & Language, 16 0 (2): 0 205--223, 2002

  102. [110]

    A probabilistic model for robust localization based on a binaural auditory front-end

    May, T., Van De Par, S., and Kohlrausch, A. A probabilistic model for robust localization based on a binaural auditory front-end. IEEE Transactions on audio, speech, and language processing, 19 0 (1): 0 1--13, 2010

  103. [111]

    Dataset of spatial room impulse responses in a variable acoustics room for six degrees-of-freedom rendering and analysis

    McKenzie, T., McCormack, L., and Hold, C. Dataset of spatial room impulse responses in a variable acoustics room for six degrees-of-freedom rendering and analysis. arXiv preprint arXiv:2111.11882, 2021 a

  104. [112]

    J., and Pulkki, V

    McKenzie, T., Schlecht, S. J., and Pulkki, V. Acoustic analysis and dataset of transitions between coupled rooms. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 481--485. IEEE, 2021 b

  105. [113]

    Metrics for polyphonic sound event detection

    Mesaros, A., Heittola, T., and Virtanen, T. Metrics for polyphonic sound event detection. Applied Sciences, 6 0 (6): 0 162, 2016

  106. [114]

    Joint measurement of localization and detection of sound events

    Mesaros, A., Adavanne, S., Politis, A., Heittola, T., and Virtanen, T. Joint measurement of localization and detection of sound events. In 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), pp.\ 333--337. IEEE, 2019

  107. [115]

    P., Tancik, M., Barron, J

    Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., and Ng, R. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65 0 (1): 0 99--106, 2021

  108. [116]

    Dataset of binaural room impulse responses at multiple recording positions, source positions, and orientations in a real room

    Mittag, C., B \"o hme, M., Werner, S., and Klein, S. Dataset of binaural room impulse responses at multiple recording positions, source positions, and orientations in a real room. In in Proc. of the 43rd annual convention for acoustics, DAGA, Germany, 2017

  109. [117]

    Fundamentals of binaural technology

    M ller, H. Fundamentals of binaural technology. Applied acoustics, 36 0 (3-4): 0 171--218, 1992

  110. [118]

    Self-supervised generation of spatial audio for 360 video

    Morgado, P., Nvasconcelos, N., Langlois, T., and Wang, O. Self-supervised generation of spatial audio for 360 video. Advances in neural information processing systems, 31, 2018

  111. [119]

    Learning representations from audio-visual spatial alignment

    Morgado, P., Li, Y., and Nvasconcelos, N. Learning representations from audio-visual spatial alignment. Advances in Neural Information Processing Systems, 33: 0 4733--4744, 2020

  112. [120]

    and Neff, F

    Murphy, D. and Neff, F. Spatial sound for computer games and virtual reality. In Game sound technology and player interaction: Concepts and developments, pp.\ 287--312. IGI Global Scientific Publishing, 2011

  113. [121]

    Murphy, D. T. and Shelley, S. Openair: An interactive auralization web resource and database. In Audio Engineering Society Convention 129. Audio Engineering Society, 2010

  114. [122]

    Wearable seld dataset: Dataset for sound event localization and detection using wearable devices around head

    Nagatomo, K., Yasuda, M., Yatabe, K., Saito, S., and Oikawa, Y. Wearable seld dataset: Dataset for sound event localization and detection using wearable devices around head. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), ...

  115. [123]

    Blind and neural network-guided convolutional beamformer for joint denoising, dereverberation, and source separation

    Nakatani, T., Ikeshita, R., Kinoshita, K., Sawada, H., and Araki, S. Blind and neural network-guided convolutional beamformer for joint denoising, dereverberation, and source separation. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processi...

  116. [124]

    Sound event localization and detection using squeeze-excitation residual cnns

    Naranjo-Alcazar, J., Perez-Castanos, S., Ferrandis, J., Zuccarello, P., and Cobos, M. Sound event localization and detection using squeeze-excitation residual cnns. arXiv preprint arXiv:2006.14436, 2020

  117. [125]

    Nguyen, T. N. T., Jones, D. L., and Gan, W.-S. A sequence matching network for polyphonic sound event localization and detection. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 71--75. IEEE, 2020

  118. [126]

    Nguyen, T. N. T., Watcharasupat, K. N., Nguyen, N. K., Jones, D. L., and Gan, W.-S. Salsa: Spatial cue-augmented log-spectrogram features for polyphonic sound event localization and detection. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30: 0 1749--1762, 2022

  119. [127]

    A., and Prakash, R

    Nik Khah, A., Htun, C. A., and Prakash, R. Unsupervised bayesian surprise detection in spatial audio with convolutional variational autoencoder and lstm model. In Proceedings of the 2024 ACM International Conference on Interactive Media Experiences Workshops, pp.\ 116--121, 2024

  120. [128]

    A., Liutkus, A., and Vincent, E

    Nugraha, A. A., Liutkus, A., and Vincent, E. Multichannel audio source separation with deep neural networks. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 24 0 (9): 0 1652--1664, 2016

  121. [129]

    and Efros, A

    Owens, A. and Efros, A. A. Audio-visual scene analysis with self-supervised multisensory features. In Proceedings of the European conference on computer vision (ECCV), pp.\ 631--648, 2018

  122. [130]

    Librispeech: an asr corpus based on public domain audio books

    Panayotov, V., Chen, G., Povey, D., and Khudanpur, S. Librispeech: an asr corpus based on public domain audio books. In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), pp.\ 5206--5210. IEEE, 2015

  123. [131]

    Q., P \'e rez, P., and Richard, G

    Parekh, S., Essid, S., Ozerov, A., Duong, N. Q., P \'e rez, P., and Richard, G. Motion informed audio source separation. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 6--10. IEEE, 2017

  124. [132]

    3d localization of multiple sound sources with intensity vector estimates in single source zones

    Pavlidi, D., Delikaris-Manias, S., Pulkki, V., and Mouchtaris, A. 3d localization of multiple sound sources with intensity vector estimates in single source zones. In 2015 23rd European Signal Processing Conference (EUSIPCO), pp.\ 1556--1560. IEEE, 2015

  125. [133]

    Integration of spatial sound in immersive virtual environments an experimental study on effects of spatial sound on presence

    Poeschl, S., Wall, K., and Doering, N. Integration of spatial sound in immersive virtual environments an experimental study on effects of spatial sound on presence. In 2013 IEEE Virtual Reality (VR), pp.\ 129--130. IEEE, 2013

  126. [134]

    A dataset of reverberant spatial sound scenes with moving sources for sound event localization and detection

    Politis, A., Adavanne, S., and Virtanen, T. A dataset of reverberant spatial sound scenes with moving sources for sound event localization and detection. arXiv preprint arXiv:2006.01919, 2020

  127. [135]

    A dataset of dynamic reverberant sound scenes with directional interferers for sound event localization and detection

    Politis, A., Adavanne, S., Krause, D., Deleforge, A., Srivastava, P., and Virtanen, T. A dataset of dynamic reverberant sound scenes with directional interferers for sound event localization and detection. arXiv preprint arXiv:2106.06999, 2021

  128. [136]

    Starss22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events

    Politis, A., Shimada, K., Sudarsanam, P., Adavanne, S., Krause, D., Koyama, Y., Takahashi, N., Takahashi, S., Mitsufuji, Y., and Virtanen, T. Starss22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events. arXiv preprint arXiv:2206.01948, 2022

  129. [137]

    Audio-visual object localization and separation using low-rank and sparsity

    Pu, J., Panagakis, Y., Petridis, S., and Pantic, M. Audio-visual object localization and separation using low-rank and sparsity. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 2901--2905. IEEE, 2017

  130. [138]

    Audio-visual cross-attention network for robotic speaker tracking

    Qian, X., Wang, Z., Wang, J., Guan, G., and Li, H. Audio-visual cross-attention network for robotic speaker tracking. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31: 0 550--562, 2022

  131. [139]

    K., Sundaresha, V., Rajagopalan, A., et al

    Rachavarapu, K. K., Sundaresha, V., Rajagopalan, A., et al. Localize to binauralize: Audio spatialization from visual sound source localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp.\ 1930--1939, 2021

  132. [140]

    Towards generating ambisonics using audio-visual cue for virtual reality

    Rana, A., Ozcinar, C., and Smolic, A. Towards generating ambisonics using audio-visual cue for virtual reality. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 2012--2016. IEEE, 2019

  133. [141]

    and Manocha, D

    Ratnarajah, A. and Manocha, D. Listen2scene: Interactive material-aware binaural sound propagation for reconstructed 3d scenes. In 2024 IEEE Conference Virtual Reality and 3D User Interfaces (VR), pp.\ 254--264. IEEE, 2024

  134. [142]

    Mesh2ir: Neural acoustic impulse response generator for complex 3d scenes

    Ratnarajah, A., Tang, Z., Aralikatti, R., and Manocha, D. Mesh2ir: Neural acoustic impulse response generator for complex 3d scenes. In Proceedings of the 30th ACM International Conference on Multimedia, pp.\ 924--933, 2022

  135. [143]

    Av-rir: Audio-visual room impulse response estimation

    Ratnarajah, A., Ghosh, S., Kumar, S., Chiniya, P., and Manocha, D. Av-rir: Audio-visual room impulse response estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 27164--27175, 2024

  136. [144]

    Impulse response estimation for robust speech recognition in a reverberant environment

    Ravanelli, M., Sosi, A., Svaizer, P., and Omologo, M. Impulse response estimation for robust speech recognition in a reverberant environment. In 2012 Proceedings of the 20th European Signal Processing Conference (EUSIPCO), pp.\ 1668--1672. IEEE, 2012

  137. [145]

    The dirha-english corpus and related tasks for distant-speech recognition in domestic environments

    Ravanelli, M., Cristoforetti, L., Gretter, R., Pellin, M., Sosi, A., and Omologo, M. The dirha-english corpus and related tasks for distant-speech recognition in domestic environments. In 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), pp.\ 275--28...

  138. [146]

    D., Krenn, S., Butler, G

    Richard, A., Markovic, D., Gebru, I. D., Krenn, S., Butler, G. A., Torre, F., and Sheikh, Y. Neural synthesis of binaural speech from mono audio. In International Conference on Learning Representations, 2021

  139. [147]

    R., Ick, C., Ding, S., Roman, A

    Roman, I. R., Ick, C., Ding, S., Roman, A. S., McFee, B., and Bello, J. P. Spatial scaper: a library to simulate and augment soundscapes for sound event localization and detection in realistic rooms. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Si...

  140. [148]

    U-net: Convolutional networks for biomedical image segmentation

    Ronneberger, O., Fischer, P., and Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Medical image computing and computer-assisted intervention--MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18, ...

  141. [149]

    E., Hinton, G

    Rumelhart, D. E., Hinton, G. E., and Williams, R. J. Learning representations by back-propagating errors. nature, 323 0 (6088): 0 533--536, 1986

  142. [150]

    w2v-seld: A sound event localization and detection framework for self-supervised spatial audio pre-training

    Santos, O., Rosero, K., Masiero, B., and de Alencar Lotufo, R. w2v-seld: A sound event localization and detection framework for self-supervised spatial audio pre-training. IEEE Access, 2024

  143. [151]

    Spatial librispeech: An augmented dataset for spatial audio learning

    Sarabia, M., Menyaylenko, E., Toso, A., Seto, S., Aldeneh, Z., Pirhosseinloo, S., Zappella, L., Theobald, B.-J., Apostoloff, N., and Sheaffer, J. Spatial librispeech: An augmented dataset for spatial audio learning. arXiv preprint arXiv:2308.09514, 2023

  144. [152]

    and Svensson, U

    Savioja, L. and Svensson, U. P. Overview of geometrical room acoustic modeling techniques. The Journal of the Acoustical Society of America, 138 0 (2): 0 708--730, 2015

  145. [153]

    Habitat: A platform for embodied ai research

    Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., et al. Habitat: A platform for embodied ai research. In Proceedings of the IEEE/CVF international conference on computer vision, pp.\ 9339--9347, 2019

  146. [154]

    Mo \^ usai: Text-to-music generation with long-context latent diffusion

    Schneider, F., Kamal, O., Jin, Z., and Sch \"o lkopf, B. Mo \^ usai: Text-to-music generation with long-context latent diffusion. arXiv preprint arXiv:2301.11757, 2023

  147. [155]

    Two multimodal approaches for single microphone source separation

    Sedighin, F., Babaie-Zadeh, M., Rivet, B., and Jutten, C. Two multimodal approaches for single microphone source separation. In 2016 24th European Signal Processing Conference (EUSIPCO), pp.\ 110--114. IEEE, 2016

  148. [156]

    Senocak, A., Oh, T.-H., Kim, J., Yang, M.-H., and Kweon, I. S. Learning to localize sound source in visual scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp.\ 4358--4366, 2018

  149. [157]

    Accdoa: Activity-coupled cartesian direction of arrival representation for sound event localization and detection

    Shimada, K., Koyama, Y., Takahashi, N., Takahashi, S., and Mitsufuji, Y. Accdoa: Activity-coupled cartesian direction of arrival representation for sound event localization and detection. In ICASSP 2021-2021 IEEE international conference on acoustics, speech and signal process...

  150. [158]

    Multi-accdoa: Localizing and detecting overlapping sounds from the same class with auxiliary duplicating permutation invariant training

    Shimada, K., Koyama, Y., Takahashi, S., Takahashi, N., Tsunoo, E., and Mitsufuji, Y. Multi-accdoa: Localizing and detecting overlapping sounds from the same class with auxiliary duplicating permutation invariant training. In ICASSP 2022-2022 IEEE international conference on ac...

  151. [159]

    A., Uchida, K., Adavanne, S., Hakala, A., Koyama, Y., Takahashi, N., Takahashi, S., et al

    Shimada, K., Politis, A., Sudarsanam, P., Krause, D. A., Uchida, K., Adavanne, S., Hakala, A., Koyama, Y., Takahashi, N., Takahashi, S., et al. Starss23: An audio-visual dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events. Advances in n...

  152. [160]

    and Casey, M

    Smaragdis, P. and Casey, M. Audio/visual independent components. In Proc. ICA, pp.\ 709--714, 2003

  153. [161]

    Blind room parameter estimation using multiple multichannel speech recordings

    Srivastava, P., Deleforge, A., and Vincent, E. Blind room parameter estimation using multiple multichannel speech recordings. In 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), pp.\ 226--230. IEEE, 2021

  154. [162]

    Both ears wide open: Towards language-driven spatial audio generation

    Sun, P., Cheng, S., Li, X., Ye, Z., Liu, H., Zhang, H., Xue, W., and Guo, Y. Both ears wide open: Towards language-driven spatial audio generation. arXiv preprint arXiv:2410.10676, 2024

  155. [163]

    Learning audio-visual source localization via false negative aware contrastive learning

    Sun, W., Zhang, J., Wang, J., Liu, Z., Zhong, Y., Feng, T., Guo, Y., Zhang, Y., and Barnes, N. Learning audio-visual source localization via false negative aware contrastive learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), ...

  156. [164]

    Sonicmotion: Dynamic spatial audio soundscapes with latent diffusion models

    Templin, C., Zhu, Y., and Wang, H. Sonicmotion: Dynamic spatial audio soundscapes with latent diffusion models. arXiv preprint arXiv:2507.07318, 2025

  157. [165]

    T., and V \"a lim \"a ki, V

    Thuillier, E., Jin, C. T., and V \"a lim \"a ki, V. Hrtf interpolation using a spherical neural process meta-learner. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 32: 0 1790--1802, 2024

  158. [166]

    Audio-visual event localization in unconstrained videos

    Tian, Y., Shi, J., Li, B., Duan, Z., and Xu, C. Audio-visual event localization in unconstrained videos. In Proceedings of the European conference on computer vision (ECCV), pp.\ 247--263, 2018

  159. [167]

    Llama 2: Open foundation and fine-tuned chat models

    Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023

  160. [168]

    The nigens general sound events database

    Trowitzsch, I., Taghia, J., Kashef, Y., and Obermayer, K. The nigens general sound events database. arXiv preprint arXiv:1902.08314, 2019

  161. [169]

    P., and Hershey, J

    Tzinis, E., Wisdom, S., Jansen, A., Hershey, S., Remez, T., Ellis, D. P., and Hershey, J. R. Into the wild with audioscope: Unsupervised audio-visual separation of on-screen sounds. arXiv preprint arXiv:2011.01143, 2020

  162. [170]

    Tzinis, E., Wisdom, S., Remez, T., and Hershey, J. R. Audioscopev2: Audio-visual attention architectures for calibrated open-domain on-screen sound separation. In European Conference on Computer Vision, pp.\ 368--385. Springer, 2022

  163. [171]

    The sweet-home speech and multimodal corpus for home automation interaction

    Vacher, M., Lecouteux, B., Chahuara, P., Portet, F., Meillon, B., and Bonnefond, N. The sweet-home speech and multimodal corpus for home automation interaction. In The 9th edition of the Language Resources and Evaluation Conference (LREC), pp.\ 4499--4506, 2014

  164. [172]

    Heavenly mathematics: The forgotten art of spherical trigonometry

    Van Brummelen, G. Heavenly mathematics: The forgotten art of spherical trigonometry. Princeton University Press, 2012

  165. [173]

    and Sambath, M

    Varghese, R. and Sambath, M. Yolov8: A novel object detection algorithm with enhanced performance and robustness. In 2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS), pp.\ 1--6. IEEE, 2024

  166. [174]

    N., Kaiser, ., and Polosukhin, I

    Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, ., and Polosukhin, I. Attention is all you need. In Advances in neural information processing systems, pp.\ 5998--6008, 2017

  167. [175]

    and Wang, D

    Wang, Z.-Q. and Wang, D. Combining spectral and spatial features for deep learning based blind speaker separation. IEEE/ACM Transactions on audio, speech, and language processing, 27 0 (2): 0 457--468, 2018

  168. [176]

    Wang, Z.-Q., Le Roux, J., and Hershey, J. R. Multi-channel deep clustering: Discriminative spectral and spatial embeddings for speaker-independent speech separation. In 2018 IEEE International conference on acoustics, speech and signal processing (ICASSP), pp.\ 1--5. IEEE, 2018

  169. [177]

    Warnecke, M., Jamison, S., Prepelita, S., Calamia, P., and Ithapu, V. K. Hrtf personalization based on ear morphology. In Audio Engineering Society Conference: 2022 AES International Conference on Audio for Virtual and Augmented Reality. Audio Engineering Society, 2022

  170. [178]

    J., Mandel, M

    Weiss, R. J., Mandel, M. I., and Ellis, D. P. Source separation based on binaural cues and source model constraints. 2009

  171. [179]

    J., Bharadia, D., and Gerstoft, P

    Wu, Y., Ayyalasomayajula, R., Bianco, M. J., Bharadia, D., and Gerstoft, P. Sslide: Sound source localization for indoors based on deep learning. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 4680--4684. IEEE, 2021

  172. [180]

    and Moreira Kares, E

    Wuolio, L. and Moreira Kares, E. On the potential of spatial audio in enhancing virtual user experiences. 2023

  173. [181]

    Sonic4d: Spatial audio generation for immersive 4d scene exploration

    Xie, S., Zhu, H., He, T., Li, X., and Chen, Z. Sonic4d: Spatial audio generation for immersive 4d scene exploration. arXiv preprint arXiv:2506.15759, 2025

  174. [182]

    Visually informed binaural audio generation without binaural audios

    Xu, X., Zhou, H., Liu, Z., Dai, B., Wang, X., and Lin, D. Visually informed binaural audio generation without binaural audios. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 15485--15494, 2021

  175. [183]

    Realman: A real-recorded and annotated microphone array dataset for dynamic speech enhancement and localization

    Yang, B., Quan, C., Wang, Y., Wang, P., Yang, Y., Fang, Y., Shao, N., Bu, H., Xu, X., and Li, X. Realman: A real-recorded and annotated microphone array dataset for dynamic speech enhancement and localization. Advances in Neural Information Processing Systems, 37: 0 105997--10...

  176. [184]

    Upmixing via style transfer: a variational autoencoder for disentangling spatial images and musical content

    Yang, H., Wager, S., Russell, S., Luo, M., Kim, M., and Kim, W. Upmixing via style transfer: a variational autoencoder for disentangling spatial images and musical content. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p...

  177. [185]

    Telling left from right: Learning spatial correspondence of sight and sound

    Yang, K., Russell, B., and Salamon, J. Telling left from right: Learning spatial correspondence of sight and sound. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 9932--9941, 2020

  178. [186]

    Depth anything: Unleashing the power of large-scale unlabeled data

    Yang, L., Kang, B., Huang, Z., Xu, X., Feng, J., and Zhao, H. Depth anything: Unleashing the power of large-scale unlabeled data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 10371--10381, 2024 b

  179. [187]

    and Zheng, Y

    Yang, Q. and Zheng, Y. Deepear: Sound localization with binaural microphones. IEEE Transactions on Mobile Computing, 23 0 (1): 0 359--375, 2022

  180. [188]

    Sound event localization based on sound intensity vector refined by dnn-based denoising and source separation

    Yasuda, M., Koizumi, Y., Saito, S., Uematsu, H., and Imoto, K. Sound event localization based on sound intensity vector refined by dnn-based denoising and source separation. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), ...

  181. [189]

    Lavss: Location-guided audio-visual spatial audio separation

    Ye, Y., Yang, W., and Tian, Y. Lavss: Location-guided audio-visual spatial audio separation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp.\ 5508--5519, 2024

  182. [190]

    Multi-microphone neural speech separation for far-field multi-talker speech recognition

    Yoshioka, T., Erdogan, H., Chen, Z., and Alleva, F. Multi-microphone neural speech separation for far-field multi-talker speech recognition. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 5739--5743. IEEE, 2018

  183. [191]

    Ambisonizer: Neural upmixing as spherical harmonics generation

    Zang, Y., Wang, Y., and Lee, M. Ambisonizer: Neural upmixing as spherical harmonics generation. arXiv preprint arXiv:2405.13428, 2024

  184. [192]

    and Shao, J

    Zhang, W. and Shao, J. Multi-attention audio-visual fusion network for audio spatialization. In Proceedings of the 2021 International Conference on Multimedia Retrieval, pp.\ 394--401, 2021

  185. [193]

    N., Chen, H., and Abhayapala, T

    Zhang, W., Samarasinghe, P. N., Chen, H., and Abhayapala, T. D. Surround by sound: A review of spatial audio recording and reproduction. Applied Sciences, 7 0 (5): 0 532, 2017

  186. [194]

    and Wang, D

    Zhang, X. and Wang, D. Deep learning based binaural speech separation in reverberant environments. IEEE/ACM transactions on audio, speech, and language processing, 25 0 (5): 0 1075--1084, 2017

  187. [195]

    Hrtf field: Unifying measured hrtf magnitude representation with neural fields

    Zhang, Y., Wang, Y., and Duan, Z. Hrtf field: Unifying measured hrtf magnitude representation with neural fields. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 1--5. IEEE, 2023

  188. [196]

    Isdrama: Immersive spatial drama generation through multimodal prompting

    Zhang, Y., Guo, W., Pan, C., Zhu, Z., Jin, T., and Zhao, Z. Isdrama: Immersive spatial drama generation through multimodal prompting. arXiv preprint arXiv:2504.20630, 2025

  189. [197]

    The sound of pixels

    Zhao, H., Gan, C., Rouditchenko, A., Vondrick, C., McDermott, J., and Torralba, A. The sound of pixels. In Proceedings of the European conference on computer vision (ECCV), pp.\ 570--586, 2018

  190. [198]

    Dualspec: Text-to-spatial-audio generation via dual-spectrogram guided diffusion model

    Zhao, L., Chen, S., Feng, L., Zhang, X.-L., and Li, X. Dualspec: Text-to-spatial-audio generation via dual-spectrogram guided diffusion model. arXiv preprint arXiv:2502.18952, 2025

  191. [199]

    Magnitude modeling of personalized hrtf based on ear images and anthropometric measurements

    Zhao, M., Sheng, Z., and Fang, Y. Magnitude modeling of personalized hrtf based on ear images and anthropometric measurements. Applied Sciences, 12 0 (16): 0 8155, 2022

  192. [200]

    Interpretable binaural ratio for visually guided binaural audio generation

    Zheng, T., Verma, S., and Liu, W. Interpretable binaural ratio for visually guided binaural audio generation. In 2022 International Joint Conference on Neural Networks (IJCNN), pp.\ 1--8. IEEE, 2022

  193. [201]

    Bat: Learning to reason about spatial sounds with large language models

    Zheng, Z., Peng, P., Ma, Z., Chen, X., Choi, E., and Harwath, D. Bat: Learning to reason about spatial sounds with large language models. arXiv preprint arXiv:2402.01591, 2024

  194. [202]

    Sep-stereo: Visually guided stereophonic audio generation by associating source separation

    Zhou, H., Xu, X., Lin, D., Wang, X., and Liu, Z. Sep-stereo: Visually guided stereophonic audio generation by associating source separation. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part XII 16, pp.\ 52--69. Springer, 2020

  195. [203]

    Zhou, Y., Wang, Z., Fang, C., Bui, T., and Berg, T. L. Visual to sound: Generating natural sound for videos in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp.\ 3550--3558, 2018

  196. [204]

    Zmolikova, K., Delcroix, M., Burget, L., Nakatani, T., and C ernocky, J. H. Integration of variational autoencoder and spatial clustering for adaptive multi-channel neural speech separation. In 2021 IEEE Spoken Language Technology Workshop (SLT), pp.\ 889--896. IEEE, 2021

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.