{"id":"f2865feb-e8fc-43c1-a7e3-fa8fcb60d22d","arxiv_id":"2411.14207","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"HARP is a synthetic dataset of 100,000 seventh-order Ambisonic room impulse responses generated with a 64-microphone spherical harmonic configuration in Pyroomacoustics.","lead":"This paper describes HARP, a simulated dataset of 100,000 seventh-order Ambisonic room impulse responses created with the image source method. It proposes a 64-microphone spherical-harmonic configuration, but provides no dataset link and no validation of the recordings.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 64-microphone superposition capture of 7th-order HOA coefficients is never specified or validated; the paper's central dataset claim rests on it.","rationale":"The reader's weakest assumption is also the most load-bearing concern: the 64-microphone configuration must recover genuine 7th-order spherical-harmonic coefficients, otherwise every RIR in the dataset is invalid. I agree with the reader's REJECT verdict. The manuscript provides no microphone coordinates, no combination weights, no inversion procedure, and no comparison to analytic or measured sound fields. The only quantitative evaluation shown, the RT60 histogram, uses the 0th-order channel and says nothing about the correctness of orders 1-7. The additional N3D-to-SN3D normalization ambiguity compounds the problem, as the data format claim could be inconsistent with the equations. The implementation of a SphericalHarmonicDirectivity class in Pyroomacoustics is a plausible step, but it is not by itself evidence of correct HOA encoding. The proposed free-field analytic test would settle the concern: if the 64 normalized outputs match the spherical harmonic basis evaluated at the source direction, then the encoding pipeline is doing what is claimed; if not, the dataset's central premise fails. As written, the paper lacks the dataset link, code, and validation needed to support its central claim, so the reader's REJECT is appropriate and no verdict adjustment is needed.","tokens_in":6021,"tokens_out":6362,"duration_ms":65426,"concrete_test":"Run an analytic free-field validation of the complete pipeline. Place a single point source in free space (no reflective walls) at known spherical coordinates (r0, theta0, phi0) and generate the 64-channel output with the exact HARP configuration and code. At the direct-path delay, normalize each channel by the (0,0) omnidirectional channel. The resulting normalized output for channel (n,m) should equal the analytic ratio Y_n^m(theta0, phi0) / Y_0^0 for all n<=7, up to a common scale and time delay. Repeat with a second source direction and with two simultaneous sources to verify linearity and superposition. Additionally, feed the stored AmbiX (ACN/SN3D) files into a standard SN3D decoder and confirm that the rendered direction matches the known source direction. If these comparisons deviate beyond numerical precision, the claimed direct SH capture is invalid.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that HARP provides 100,000 valid 7th-order HOA RIRs depends on the assertion in Section III-A that a 64-microphone superposition configuration \"captures RIRs directly in the Spherical Harmonics domain.\" The manuscript never states whether the 64 microphones are coincident directional sensors with SH patterns or pressure sensors at distinct positions, and it gives no coordinates, weights, inversion, or validation. Eq. (3) models pressure p(r_m,t) as a sum of SH functions, but for a directional microphone the output is a weighted integral of the field against the pattern, not the pressure at a point; the paper does not show how the 64 outputs map to coefficients c_n^m. If the microphones are spatially separated, a 64x64 linear system (or a spherical t-design sampling scheme) is needed, and none is described. A second, independent problem appears between Section III-A (N3D normalization, Eq. 1) and Section III-D (storage in ACN/SN3D format): no conversion factor is given, so the stored coefficients may be scaled by the wrong normalization. The only quantitative check, the RT60 histogram in Figure 4, uses the 0th-order channel only and cannot detect errors in orders 1-7. Consequently, the fundamental premise of the dataset is unverified, not merely unoptimized.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces HARP, a proposed dataset of 100,000 simulated 7th-order Higher-Order Ambisonic Room Impulse Responses (HOA-RIRs) generated with the Image Source Method using the Pyroomacoustics library. The authors describe a 64-microphone configuration intended to capture RIRs directly in the spherical harmonics domain, a room-simulation setup with randomized room geometries, absorption materials from a lookup table, and source-receiver positions, and they compare HARP with existing ambisonic RIR datasets. They also note a planned contribution to Pyroomacoustics through a new SphericalHarmonicDirectivity class.","tokens_in":6326,"tokens_out":3279,"duration_ms":32220,"significance":"If the generation method is correct and the dataset is released, HARP would be a valuable large-scale resource for spatial audio and machine-learning research, offering far more 7th-order HOA-RIRs than existing datasets. The paper also claims a first large-scale HOA-RIR dataset and a new directivity implementation in a widely used simulation library. However, the central technical claim—that the 64-microphone configuration directly captures 7th-order HOA coefficients—is not derived, specified, or validated, and the paper currently provides no dataset link. The practical significance therefore depends entirely on unverified assertions.","major_comments":[{"comment":"The manuscript never specifies whether the 64 microphones are coincident directional sensors or spatially distributed transducers, nor does it give their positions, polar patterns, or the combination weights that map the 64 outputs to the HOA coefficients c_n^m(t). Eq. (3) is the spherical harmonic expansion of the pressure at a point, but the output of a directional microphone is a weighted integral of the incident field against the directivity pattern, not the point pressure. If the microphones are spatially separated, a 64x64 linear system or a spherical sampling scheme such as a t-design is required, and none is described. The central claim that the configuration 'captures RIRs directly in the Spherical Harmonics domain' is therefore unsupported as written.","section":"Section III-A, Eqs. (1)-(4)"},{"comment":"Section III-A states that N3D normalization is used for the spherical harmonic coefficients, while Section III-D states that each RIR is stored in the AmbiX format (ACN/SN3D). N3D and SN3D are different normalizations, and no conversion factors between them are provided. Without this conversion, the stored coefficients may be scaled incorrectly relative to the simulated microphone outputs, corrupting the dataset for downstream Ambisonics rendering and analysis.","section":"Section III-A vs. Section III-D"},{"comment":"The only quantitative validation presented is the RT60 distribution computed from the 0th-order (omnidirectional) channel. This cannot detect errors in the higher-order channels (orders 1-7), which are the entire point of the dataset. Moreover, the caption states 'Histogram to be updated in camera-ready version,' indicating that the figure is a placeholder. There is no analytic free-field benchmark, no comparison to measured HOA-RIRs, no comparison to an independently computed sound field, and no error metric for the spherical harmonic coefficients. The dataset's central premise is therefore unverified.","section":"Section IV, Figure 4"},{"comment":"The paper claims to introduce a dataset and to provide it as a resource, but it gives no URL, repository, or access instructions. For a dataset paper, the availability of the data and metadata is a load-bearing component; without it, the contribution cannot be evaluated or used by the community. A revision should include a working link and a data availability statement.","section":"Dataset availability"}],"minor_comments":[{"comment":"The text contains several typographical errors, such as 'FOr instance,' 'capturedd,' and 'f ¨ur'; these should be corrected in a revision.","section":"Throughout"},{"comment":"The phrase 'a very image source order (40)' should read 'a very high image source order (40)' or similar.","section":"Section III-B"},{"comment":"The 'Free field measurement' shown in Figure 2 is not described quantitatively in the text; the authors should specify what is plotted and how it verifies the directivity of the proposed configuration.","section":"Figure 2"},{"comment":"The comparison table would be more useful if it included the number of RIRs for ARNI SRIR (currently listed as '-') and the exact room configurations for HARP, rather than the qualitative 'Wide variety.'","section":"Section IV, Table I"}],"recommendation":"major_revision","confidential_remarks":"The paper is closer to a workshop report than a complete dataset publication. The central technical gap—the missing specification and validation of the 64-microphone SH capturing scheme—is potentially fixable if the authors can clarify the sensor geometry and provide an independent validation of the higher-order coefficients. However, the placeholder RT60 figure and the absence of a dataset link are atypical for a serious journal submission and should be addressed decisively. I would not reject outright because the underlying idea is plausible, but the revision must supply the missing derivation, normalization details, and validation before the paper can be considered for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline is simple: if HARP is valid, it's the first large-scale HOA-RIR dataset at 7th order, and that matters for the ML side of spatial audio. Existing datasets top out at 4th order and a few hundred samples (TAU-SRIR, MOTUS), so 100,000 simulated 7th-order RIRs is a genuine step up in scale. The idea of using spherical-harmonic directivity patterns for the microphones inside Pyroomacoustics is also reasonable, and extending that library is a useful, reproducible contribution if the code is released.\n\nBut the paper doesn't yet support its central claim. The 64-microphone configuration that supposedly captures HOA coefficients directly is never fully specified. I don't need psychoanalysis here, just the math: the paper gives equations for spherical harmonics and says the signal is a weighted sum, but it never states whether the microphones are coincident with different SH patterns or spread out, and it doesn't give coordinates, weights, or an inversion. That's a load-bearing gap. If the microphones are simply pressure sensors at different locations, you need a 64x64 linear system to get coefficients. If they are directional SH patterns, you need to explain how the integral in the directivity maps to the coefficient in Eq. (3). Neither is done.\n\nThere's also a normalization inconsistency: Eq. (1) uses N3D, but the stored format is ACN/SN3D, and no conversion factor is provided. That could silently scale every coefficient. The only validation, the RT60 histogram from the 0th-order channel, can't catch errors in orders 1–7. And there is no dataset link, no free-field benchmark, no comparison against analytic or measured fields.\n\nThe paper's own conclusion is honest about known limitations (ISM misses diffraction etc.), which is good. But the missing validation is not a minor omission; it's the difference between a dataset paper and a dataset announcement.\n\nFor whom is this useful? Researchers who want a large synthetic HOA training resource. After a revision that spells out the microphone setup, validates the encoding (even just a plane-wave test), fixes the normalization, and releases the data, this could become a standard resource. Right now I wouldn't cite it or train on it.\n\nMy recommendation: send it to peer review, because the artifact could be important and the flaws are fixable, but the referee should demand the missing details and a validation section before acceptance.","headline":"A large-scale 7th-order HOA RIR dataset would be a real contribution, but this version doesn't yet show that the 64-microphone encoding works, so it's a promising announcement rather than a usable dataset.","tokens_in":6793,"tokens_out":2716,"would_cite":false,"duration_ms":28616,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces HARP, a dataset of 100,000 simulated 7th-order Ambisonic room impulse responses, claimed to be the first large-scale HOA-RIR dataset.","keywords":["Higher-Order Ambisonics","Room Impulse Response dataset","Image Source Method","Spherical Harmonics","Spatial audio","Machine learning acoustics","Reverberation","HOA-RIR"],"falsifier":"Re-simulate one room configuration from HARP in an independent image-source or boundary-element solver, place the same source and receiver, sample the sound field with the described 64-microphone array, and compare the resulting 7th-order coefficients against a reference computed by analytic projection of the same field onto spherical harmonics; if the error exceeds the sampling discretization error of a 7th-order grid, the dataset's core encoding claim fails.","tokens_in":1604,"feed_emoji":"🎧","tokens_out":7402,"duration_ms":101258,"temperature":0.7,"pith_summary":"This paper introduces HARP, a simulated dataset of 100,000 room impulse responses encoded as 7th-order Ambisonics (HOA-RIRs). It is offered as the first large-scale resource of its kind, filling a gap left by existing datasets that stop at lower orders or cover only a few rooms. The responses are generated by the Image Source Method across a variety of room geometries and materials, using a 64-microphone array whose directivities are the real spherical harmonics, so the array captures HOA coefficients directly. If the encoding is correct, HARP supplies machine-learning researchers with the scale and spatial resolution needed for tasks such as room parameter estimation, dereverberation, and immersive rendering.","feed_headline":"First large-scale 7th-order Ambisonic room response dataset","feed_subtitle":"HARP spans varied rooms and materials, built for spatial-audio and machine-learning research.","key_machinery":"The load-bearing object is the 64-microphone superposition array. Each capsule's directivity is one real spherical harmonic $Y_{n,m}(\\theta,\\phi)$ derived from the complex spherical harmonics $Y_n^m(\\theta,\\phi)$ with N3D normalization; the signal at position $\\mathbf{r}_m$ is written as $$p(\\mathbf{r}_m,t)=\\sum_{n=0}^{7}\\sum_{m=-n}^{n}c_n^m(t)Y_n^m(\\theta_m,\\phi_m),$$ so the weights $c_n^m(t)$ are the HOA coefficients up to order 7. Because a 7th-order field has $(7+1)^2=64$ coefficients, 64 suitably arranged capsules match the degrees of freedom. These directivities are inserted into an image-source room simulator that computes, for each room configuration, impulse responses to all 64 microphones; the 0th-order response doubles as the omnidirectional channel for RT60 analysis.","core_discovery":"The paper's central claim is that a 64-microphone superposition arrangement, each microphone having a directivity given by one real spherical harmonic up to order 7, can sample a simulated room sound field directly in the spherical-harmonics domain, and that applying this configuration inside an image-source room simulator yields 100,000 valid 7th-order Ambisonic RIRs with varied room geometries, materials, and source-receiver positions. It reports that these are, to its knowledge, the first large-scale HOA-RIRs, stored in AmbiX (ACN/SN3D) format with metadata, and it argues that this scale and order is what distinguishes HARP from earlier datasets with only a few hundred samples or lower-order encoding.","pith_inferences":["If the array is validated, the same construction can be carried to even higher orders or to other simulation engines, since the count 64 only reflects the $(N+1)^2$ harmonic count for $N=7$.","A natural test the paper does not report is comparing the dataset's 0th-order RIRs against measured RIRs in similar rooms; a systematic mismatch would quantify how much of the real-world gap comes from the image-source approximation.","The lack of diffraction and scattering in the image-source method suggests models trained on HARP may need fine-tuning on measured data before use in rooms with furniture or other obstacles.","The claim of being the first large-scale HOA-RIR dataset rests on the comparison table; if a comparable 7th-order dataset appears, the contribution would shift from firstness to scale and methodology."],"forward_implications":["HARP gives machine-learning researchers 100,000 labelled HOA-RIRs to train or pre-train models for room parameter estimation, dereverberation, source localization, and spatial upsampling.","The AmbiX/SN3D format lets the responses feed directly into standard Ambisonic renderers without conversion.","The RT60 distribution concentrated between 0.4 and 0.8 seconds means most samples match typical living rooms and offices, with edge cases included for more extreme acoustics.","The inclusion of multiple room geometries and material lookup tables supports experiments on generalization across acoustic environments.","If the encoding works as claimed, this is the first large-scale HOA-RIR dataset, providing a reference scale for other synthetic and measured spatial-audio datasets."],"supporting_citations":[{"why":"Supplies the room-simulation framework in which the 64-microphone spherical-harmonic directivities are implemented and the 100,000 RIRs are computed.","marker":"[12]"},{"why":"Supplies the Image Source Method used to model all reflections in the simulated rooms.","marker":"[18]"},{"why":"Provides the higher-order Ambisonics background and validation approach that the proposed 7th-order encoding extends.","marker":"[1]"},{"why":"Serves as a comparison baseline in the dataset survey, representing a first-order dataset with 700 RIRs.","marker":"[19]"},{"why":"Serves as a comparison baseline, representing a 4th-order dataset with 114 RIRs.","marker":"[20]"},{"why":"Serves as a comparison baseline, representing a 3rd-order dataset with 3,320 RIRs.","marker":"[21]"},{"why":"Serves as a comparison baseline, representing a first-order dataset with fewer than 50 responses.","marker":"[22]"},{"why":"Serves as a comparison baseline, representing a 4th-order six-degrees-of-freedom dataset.","marker":"[23]"},{"why":"Serves as a comparison baseline, representing a 2nd-order dataset with 25 responses in a seminar room.","marker":"[24]"}],"fun_headline_variants":["HARP delivers 100k 7th-order Ambisonic room responses","First 7th-order Ambisonic RIR dataset with 100k samples","64-microphone superposition captures room acoustics in SH domain","HARP: 100k high-order Ambisonic room impulse responses"],"cache_read_input_tokens":8960,"weakest_assumption_plain":"The whole dataset stands on the unstated and unvalidated assumption that the 64-microphone superposition array recovers the true 7th-order spherical-harmonic coefficients of each simulated sound field.","fun_headline_variants_meta":{"raw":{"variants":["HARP delivers 100k 7th-order Ambisonic room responses","First 7th-order Ambisonic RIR dataset with 100k samples","64-microphone superposition captures room acoustics in SH domain","HARP: 100k high-order Ambisonic room impulse responses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000486,"raw_usage":{"total_tokens":2363,"prompt_tokens":879,"completion_tokens":1484,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":1405}},"tokens_in":495,"tokens_out":1484,"duration_ms":11509,"temperature":1.0,"reasoning_tokens":1405,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:24:59.163867+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-simulate one room configuration from HARP in an independent image-source or boundary-element solver, place the same source and receiver, sample the sound field with the described 64-microphone array, and compare the resulting 7th-order coefficients against a reference computed by analytic projection of the same field onto spherical harmonics; if the error exceeds the sampling discretization error of a 7th-order grid, the dataset's core encoding claim fails.","supporting_citations":[{"cited_title":"Pyroomacoustics: A python package for audio room simulation and array processing algorithms,","cited_arxiv_id":null,"evidence_quote":"Supplies the room-simulation framework in which the 64-microphone spherical-harmonic directivities are implemented and the 100,000 RIRs are computed."},{"cited_title":"3d sound field recording with higher order ambisonics–objective measurements and validation of a 4th order spherical microphone,","cited_arxiv_id":null,"evidence_quote":"Provides the higher-order Ambisonics background and validation approach that the proposed 7th-order encoding extends."},{"cited_title":"Database of omnidirectional and b-format room impulse responses,","cited_arxiv_id":null,"evidence_quote":"Serves as a comparison baseline in the dataset survey, representing a first-order dataset with 700 RIRs."},{"cited_title":"A dataset of higher-order ambisonic room impulse responses and 3d models measured in a room with varying furniture,","cited_arxiv_id":null,"evidence_quote":"Serves as a comparison baseline, representing a 3rd-order dataset with 3,320 RIRs."},{"cited_title":"Openair: An interactive auralization web resource and database,","cited_arxiv_id":null,"evidence_quote":"Serves as a comparison baseline, representing a first-order dataset with fewer than 50 responses."},{"cited_title":"Dataset of spatial room impulse responses in a variable acoustics room for six degrees-of- freedom rendering and analysis,","cited_arxiv_id":null,"evidence_quote":"Serves as a comparison baseline, representing a 4th-order six-degrees-of-freedom dataset."},{"cited_title":"Homula-rir: A room impulse response dataset for teleconferencing and spatial audio applications acquired through higher- order microphones and uniform linear microphone arrays,","cited_arxiv_id":null,"evidence_quote":"Serves as a comparison baseline, representing a 2nd-order dataset with 25 responses in a seminar room."}],"review_version":1}