REVIEW 4 major objections 4 minor 1 cited by
HARP: A Large-Scale Higher-Order Ambisonic Room Impulse Response Dataset
T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper introduces HARP, a dataset of 100,000 simulated 7th-order Ambisonic room impulse responses, claimed to be the first large-scale HOA-RIR dataset.
desk verdict A large-scale 7th-order HOA RIR dataset would be a real contribution, but this version doesn't yet show that the 64-microphone encoding works, so it's a promising announcement rather than a usable dataset. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the 64-microphone superposition array. Each capsule's directivity is one real spherical harmonic $Y_{n,m}(\theta,\phi)$ derived from the complex spherical harmonics $Y_n^m(\theta,\phi)$ with N3D normalization; the signal at position $\mathbf{r}_m$ is written as $$p(\mathbf{r}_m,t)=\sum_{n=0}^{7}\sum_{m=-n}^{n}c_n^m(t)Y_n^m(\theta_m,\phi_m),$$ so the weights $c_n^m(t)$ are the HOA coefficients up to order 7. Because a 7th-order field has $(7+1)^2=64$ coefficients, 64 suitably arranged capsules match the degrees of freedom. These directivities are inserted into an image-source room simulator that computes, for each room configuration, impulse responses to all 64 microphones; the 0th-order response doubles as the omnidirectional channel for RT60 analysis.
What would settle it
Re-simulate one room configuration from HARP in an independent image-source or boundary-element solver, place the same source and receiver, sample the sound field with the described 64-microphone array, and compare the resulting 7th-order coefficients against a reference computed by analytic projection of the same field onto spherical harmonics; if the error exceeds the sampling discretization error of a 7th-order grid, the dataset's core encoding claim fails.
Extended reading notes
Core claim
The paper's central claim is that a 64-microphone superposition arrangement, each microphone having a directivity given by one real spherical harmonic up to order 7, can sample a simulated room sound field directly in the spherical-harmonics domain, and that applying this configuration inside an image-source room simulator yields 100,000 valid 7th-order Ambisonic RIRs with varied room geometries, materials, and source-receiver positions. It reports that these are, to its knowledge, the first large-scale HOA-RIRs, stored in AmbiX (ACN/SN3D) format with metadata, and it argues that this scale and order is what distinguishes HARP from earlier datasets with only a few hundred samples or lower-order encoding.
Load-bearing premise
The whole dataset stands on the unstated and unvalidated assumption that the 64-microphone superposition array recovers the true 7th-order spherical-harmonic coefficients of each simulated sound field.
Editorial extensions
If this is right
- HARP gives machine-learning researchers 100,000 labelled HOA-RIRs to train or pre-train models for room parameter estimation, dereverberation, source localization, and spatial upsampling.
- The AmbiX/SN3D format lets the responses feed directly into standard Ambisonic renderers without conversion.
- The RT60 distribution concentrated between 0.4 and 0.8 seconds means most samples match typical living rooms and offices, with edge cases included for more extreme acoustics.
- The inclusion of multiple room geometries and material lookup tables supports experiments on generalization across acoustic environments.
- If the encoding works as claimed, this is the first large-scale HOA-RIR dataset, providing a reference scale for other synthetic and measured spatial-audio datasets.
Reading between the lines
- If the array is validated, the same construction can be carried to even higher orders or to other simulation engines, since the count 64 only reflects the $(N+1)^2$ harmonic count for $N=7$.
- A natural test the paper does not report is comparing the dataset's 0th-order RIRs against measured RIRs in similar rooms; a systematic mismatch would quantify how much of the real-world gap comes from the image-source approximation.
- The lack of diffraction and scattering in the image-source method suggests models trained on HARP may need fine-tuning on measured data before use in rooms with furniture or other obstacles.
- The claim of being the first large-scale HOA-RIR dataset rests on the comparison table; if a comparable 7th-order dataset appears, the contribution would shift from firstness to scale and methodology.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces HARP, a proposed dataset of 100,000 simulated 7th-order Higher-Order Ambisonic Room Impulse Responses (HOA-RIRs) generated with the Image Source Method using the Pyroomacoustics library. The authors describe a 64-microphone configuration intended to capture RIRs directly in the spherical harmonics domain, a room-simulation setup with randomized room geometries, absorption materials from a lookup table, and source-receiver positions, and they compare HARP with existing ambisonic RIR datasets. They also note a planned contribution to Pyroomacoustics through a new SphericalHarmonicDirectivity class.
Significance. If the generation method is correct and the dataset is released, HARP would be a valuable large-scale resource for spatial audio and machine-learning research, offering far more 7th-order HOA-RIRs than existing datasets. The paper also claims a first large-scale HOA-RIR dataset and a new directivity implementation in a widely used simulation library. However, the central technical claim—that the 64-microphone configuration directly captures 7th-order HOA coefficients—is not derived, specified, or validated, and the paper currently provides no dataset link. The practical significance therefore depends entirely on unverified assertions.
major comments (4)
- [Section III-A, Eqs. (1)-(4)] The manuscript never specifies whether the 64 microphones are coincident directional sensors or spatially distributed transducers, nor does it give their positions, polar patterns, or the combination weights that map the 64 outputs to the HOA coefficients c_n^m(t). Eq. (3) is the spherical harmonic expansion of the pressure at a point, but the output of a directional microphone is a weighted integral of the incident field against the directivity pattern, not the point pressure. If the microphones are spatially separated, a 64x64 linear system or a spherical sampling scheme such as a t-design is required, and none is described. The central claim that the configuration 'captures RIRs directly in the Spherical Harmonics domain' is therefore unsupported as written.
- [Section III-A vs. Section III-D] Section III-A states that N3D normalization is used for the spherical harmonic coefficients, while Section III-D states that each RIR is stored in the AmbiX format (ACN/SN3D). N3D and SN3D are different normalizations, and no conversion factors between them are provided. Without this conversion, the stored coefficients may be scaled incorrectly relative to the simulated microphone outputs, corrupting the dataset for downstream Ambisonics rendering and analysis.
- [Section IV, Figure 4] The only quantitative validation presented is the RT60 distribution computed from the 0th-order (omnidirectional) channel. This cannot detect errors in the higher-order channels (orders 1-7), which are the entire point of the dataset. Moreover, the caption states 'Histogram to be updated in camera-ready version,' indicating that the figure is a placeholder. There is no analytic free-field benchmark, no comparison to measured HOA-RIRs, no comparison to an independently computed sound field, and no error metric for the spherical harmonic coefficients. The dataset's central premise is therefore unverified.
- [Dataset availability] The paper claims to introduce a dataset and to provide it as a resource, but it gives no URL, repository, or access instructions. For a dataset paper, the availability of the data and metadata is a load-bearing component; without it, the contribution cannot be evaluated or used by the community. A revision should include a working link and a data availability statement.
minor comments (4)
- [Throughout] The text contains several typographical errors, such as 'FOr instance,' 'capturedd,' and 'f ¨ur'; these should be corrected in a revision.
- [Section III-B] The phrase 'a very image source order (40)' should read 'a very high image source order (40)' or similar.
- [Figure 2] The 'Free field measurement' shown in Figure 2 is not described quantitatively in the text; the authors should specify what is plotted and how it verifies the directivity of the proposed configuration.
- [Section IV, Table I] The comparison table would be more useful if it included the number of RIRs for ARNI SRIR (currently listed as '-') and the exact room configurations for HARP, rather than the qualitative 'Wide variety.'
Circularity Check
No significant circularity: the HOA coefficients are generated by directivity-defined sensors rather than fitted or predicted, though the method is under-specified and unvalidated.
full rationale
Walking the derivation chain: Section III-A defines 64 microphone directivities from the spherical-harmonic basis of order <= 7 (Eqs. 1 and 2). Because 64 = (7+1)^2, the channels correspond one-to-one to the HOA coefficients c_n^m(t); a simulated sensor whose directivity is Y_n^m outputs the projection of the incident field onto that basis function. That is the standard definition of ideal HOA encoding, not a fitted parameter renamed as a prediction. Eq. (3) is only the SH expansion statement, and no parameter is optimized against a target or held-out quantity. The sole quantitative check, the RT60 distribution in Figure 4, uses the 0th-order channel and therefore validates the room simulation's reverberation characteristics but not the higher-order coefficients; this is a validation gap, not circularity. The one self-citation [2] appears in a list of downstream applications and is not load-bearing. The paper is under-specified (no microphone coordinates or combination weights, no N3D-to-SN3D conversion factor) and unvalidated against external benchmarks, but no load-bearing step reduces to its own input by construction. Hence no significant circularity is found.
Assumptions & free parameters
free parameters (3)
- maximum image source order =
40
- number of source-receiver combinations per room =
20
- lookup table absorption coefficients =
material-dependent values from standard references
assumptions (4)
- domain assumption Linear acoustics and superposition hold for the simulated rooms.
- domain assumption The image source method with 40th-order reflections produces RIRs representative of real rooms.
- ad hoc to paper A set of 64 spherical harmonic directional microphones can directly measure HOA coefficients up to 7th order.
- domain assumption N3D normalized real spherical harmonics are orthonormal and correctly implemented in the modified Pyroomacoustics directivity class.
Cite this review
Pith. "Pith review of HARP: A Large-Scale Higher-Order Ambisonic Room Impulse Response Dataset." pith.science (2026). https://pith.science/paper/VHV2WNAU
@misc{pith2026241114207,
author = {Pith},
title = {Pith review of: HARP: A Large-Scale Higher-Order Ambisonic Room Impulse Response Dataset},
year = {2026},
howpublished = {\url{https://pith.science/paper/VHV2WNAU}},
note = {Machine review of arXiv:2411.14207}
}
read the original abstract
This contribution introduces a dataset of 7th-order Ambisonic Room Impulse Responses (HOA-RIRs), created using the Image Source Method. By employing higher-order Ambisonics, our dataset enables precise spatial audio reproduction, a critical requirement for realistic immersive audio applications. Leveraging the virtual simulation, we present a unique microphone configuration, based on the superposition principle, designed to optimize sound field coverage while addressing the limitations of traditional microphone arrays. The presented 64-microphone configuration allows us to capture RIRs directly in the Spherical Harmonics domain. The dataset features a wide range of room configurations, encompassing variations in room geometry, acoustic absorption materials, and source-receiver distances. A detailed description of the simulation setup is provided alongside for an accurate reproduction. The dataset serves as a vital resource for researchers working on spatial audio, particularly in applications involving machine learning to improve room acoustics modeling and sound field synthesis. It further provides a very high level of spatial resolution and realism crucial for tasks such as source localization, reverberation prediction, and immersive sound reproduction.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
Physics-Informed Direction-Aware Neural Acoustic Fields
A physics-informed neural network enforces momentum and continuity equations on all four Ambisonic channels, outperforming data-only and W-channel-only baselines for room impulse response interpolation.
Reference graph
Works this paper leans on
-
[1]
S. Moreau, J. Daniel, and S. Bertet, “3d sound field recording with higher order ambisonics–objective measurements and validation of a 4th order spherical microphone,” in 120th Convention of the AES, pp. 20–23, 2006
work page 2006
-
[2]
Blind room acoustic parameters estimation using mobile audio transformer,
S. Saini and J. Peissig, “Blind room acoustic parameters estimation using mobile audio transformer,” in 2023 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2023
work page 2023
-
[3]
N. J. Bryan, “Impulse response data augmentation and deep neural networks for blind room acoustic parameter estimation,” in 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020
work page 2020
-
[4]
Online Blind Reverberation Time Estimation Using CRNNs,
S. Deng, W. Mack, and E. A. Habets, “Online Blind Reverberation Time Estimation Using CRNNs,” in Proc. Interspeech 2020 , pp. 5061–5065, 2020
work page 2020
-
[5]
Online reverberation time and clarity estimation in dynamic acoustic conditions,
P. G ¨otz, C. Tuna, A. Walther, and E. A. P. Habets, “Online reverberation time and clarity estimation in dynamic acoustic conditions,” The Journal of the Acoustical Society of America , vol. 153, pp. 3532–3542, 06 2023
work page 2023
-
[6]
Single-and multi-microphone speech dereverberation using spectral enhancement,
E. A. P. Habets, “Single-and multi-microphone speech dereverberation using spectral enhancement,” 2007
work page 2007
-
[7]
Metamgc: a music generation framework for concerts in metaverse,
C. Jin, F. Wu, J. Wang, Y . Liu, Z. Guan, and Z. Han, “Metamgc: a music generation framework for concerts in metaverse,” EURASIP journal on audio, speech, and music processing , vol. 2022, no. 1, p. 31, 2022. Submitted to ICASSP 2025 Workshop
work page 2022
-
[8]
A parametric spatial audio compression codec for higher-order ambisonics,
C. Hold, “A parametric spatial audio compression codec for higher-order ambisonics,” 2024
work page 2024
Show all 24 references
-
[9]
Parametric late rever- beration from broadband directional estimates,
N. Meyer-Kahlen, S. J. Schlecht, and T. Lokki, “Parametric late rever- beration from broadband directional estimates,” in 2021 Immersive and 3D Audio: from Architecture to Automotive (I3DA) , pp. 1–10, 2021
2021
-
[10]
Upmix b-format ambisonic room impulse responses using a generative model,
J. Xia and W. Zhang, “Upmix b-format ambisonic room impulse responses using a generative model,” Applied Sciences, vol. 13, no. 21, p. 11810, 2023
2023
-
[11]
Overview and evaluation of sound event localization and detection in dcase 2019,
A. Politis, A. Mesaros, S. Adavanne, T. Heittola, and T. Virtanen, “Overview and evaluation of sound event localization and detection in dcase 2019,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 684–698, 2020
2019
-
[12]
Pyroomacoustics: A python package for audio room simulation and array processing algorithms,
R. Scheibler, E. Bezzam, and I. Dokmani ´c, “Pyroomacoustics: A python package for audio room simulation and array processing algorithms,” in 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP), pp. 351–355, IEEE, 2018
2018
-
[13]
Zotter and M
F. Zotter and M. Frank, Ambisonics: A practical 3D audio theory for recording, studio production, sound reinforcement, and virtual reality . Springer Nature, 2019
2019
-
[14]
How to make ambisonics sound good,
M. Frank, “How to make ambisonics sound good,” in Forum Acus- ticum,(Krakow), 2014
2014
-
[15]
Investigation on localisation accuracy for first and higher order ambisonics reproduced sound sources,
S. Bertet, J. Daniel, E. Parizet, and O. Warusfel, “Investigation on localisation accuracy for first and higher order ambisonics reproduced sound sources,” Acta Acustica united with Acustica , vol. 99, no. 4, pp. 642–657, 2013
2013
-
[16]
Ambisonics encoding of other audio formats for multiple listening conditions,
J. Daniel, J.-B. Rault, and J.-D. Polack, “Ambisonics encoding of other audio formats for multiple listening conditions,” in Audio Engineering Society Convention 105 , Audio Engineering Society, 1998
1998
-
[17]
A comparative study of 3-d audio encoding and rendering techniques,
J.-M. Jot, V . Larcher, and J.-M. Pernaux, “A comparative study of 3-d audio encoding and rendering techniques,” in Audio engineering society conference: 16th international conference: Spatial sound reproduction , Audio Engineering Society, 1999
1999
-
[18]
Image method for efficiently simulating small-room acoustics,
J. B. Allen and D. A. Berkley, “Image method for efficiently simulating small-room acoustics,” The Journal of the Acoustical Society of America, vol. 65, no. 4, pp. 943–950, 1979
1979
-
[19]
Database of omnidirectional and b-format room impulse responses,
R. Stewart and M. Sandler, “Database of omnidirectional and b-format room impulse responses,” in 2010 IEEE International Conference on Acoustics, Speech and Signal Processing , pp. 165–168, 2010
2010
-
[20]
A dataset of reverberant spatial sound scenes with moving sources for sound event localization and detection,
A. Politis, S. Adavanne, and T. Virtanen, “A dataset of reverberant spatial sound scenes with moving sources for sound event localization and detection,” in Detection and Classification of Acoustic Scenes and Events Workshop, 2020
2020
-
[21]
A dataset of higher-order ambisonic room impulse responses and 3d models measured in a room with varying furniture,
G. G ¨otz, S. J. Schlecht, and V . Pulkki, “A dataset of higher-order ambisonic room impulse responses and 3d models measured in a room with varying furniture,” in 2021 Immersive and 3D Audio: from Architecture to Automotive (I3DA) , pp. 1–8, IEEE, 2021
2021
-
[22]
Openair: An interactive auralization web resource and database,
D. T. Murphy and S. Shelley, “Openair: An interactive auralization web resource and database,” in Audio Engineering Society Convention 129 , Audio Engineering Society, 2010
2010
-
[23]
Dataset of spatial room impulse responses in a variable acoustics room for six degrees-of- freedom rendering and analysis,
T. McKenzie, L. McCormack, and C. Hold, “Dataset of spatial room impulse responses in a variable acoustics room for six degrees-of- freedom rendering and analysis,” 2021
2021
-
[24]
Homula-rir: A room impulse response dataset for teleconferencing and spatial audio applications acquired through higher- order microphones and uniform linear microphone arrays,
F. Miotello, P. Ostan, M. Pezzoli, L. Comanducci, A. Bernardini, F. An- tonacci, and A. Sarti, “Homula-rir: A room impulse response dataset for teleconferencing and spatial audio applications acquired through higher- order microphones and uniform linear microphone arrays,” in ...
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.