{"id":"dd0e74e9-f80f-47fe-929b-38f6f5ecbbf1","arxiv_id":"2504.21847","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"AV-DAR uses multi-view images and acoustic beam tracing in a trainable pipeline to estimate room impulse responses from far fewer measurements than prior methods.","lead":"This paper introduces a system that predicts how sound reflects inside a room using a few photos and microphone recordings. It could make spatial audio in virtual reality and telepresence cheaper to produce by cutting down the number of required acoustic measurements.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Vision-prior contribution is supported by only one room-scale ablation; without cross-scene w/o Vision results, the multimodal claim is not load-bearing.","rationale":"The reader's weakest_assumption is that visual appearance reliably encodes acoustic surface properties. My concern is adjacent but not identical: the paper provides partial evidence for that assumption (Table 3, Figure 5, and the vision ablations in Section A.2.4), but it has not demonstrated that the vision prior is load-bearing outside one room and one training scale. If the w/o Vision variant still matches or beats the strongest baselines on other rooms, the central multimodal claim would be overstated rather than false. The remaining components (beam tracing, residual field, IPE) are well ablated and internally consistent, so no fatal flaw is present. The verdict should remain CONDITIONAL, but the condition should explicitly include a cross-scene vision ablation and uncertainty estimates across seeds.","tokens_in":18748,"tokens_out":17912,"duration_ms":203558,"concrete_test":"Run the w/o Vision variant on all six rooms (RAF-Empty, RAF-Furnished, HAA Classroom, HAA Complex, HAA Dampened, HAA Hallway) at the 0.1% and 1% RAF scales and the 12-listener HAA scale, with at least 3 random seeds, and report mean and standard deviation for C50, EDT, T60, and Loudness for the full model, w/o Vision, and the best baseline. If the vision-attributable improvement is below one standard error in a majority of room-metric cells, the claim that multi-view priors drive the data-efficiency gains is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that multi-view visual priors make RIR estimation data-efficient and accurate (Section 1). The only direct ablation of this claim is Table 3 on RAF-Furnished at 0.1% data: removing vision raises C50 from 1.98 to 2.13 (+7.6%) and EDT from 80.1 ms to 98.6 ms (+23%), while T60 improves slightly from 15.2% to 14.3%. No w/o Vision ablation is reported for the other five rooms or at the 1% training scale. This matters because the headlined '10x data' comparison is not uniformly favorable even for the full model: at 0.1%, RAF-Furnished EDT is 80.1 ms versus AVR's 72.3 ms at 1%. If the mixed and modest vision contribution seen in Table 3 persists on HAA and RAF-Empty, the large reported gains would be attributable mainly to the beam-tracing plus residual structure, not to the multi-view vision prior. The title and abstract present the vision prior as the central contribution, so this evidence gap is load-bearing for the strongest_claim, even though the empirical method itself appears sound.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces AV-DAR, a differentiable room acoustic rendering framework that combines acoustic beam tracing with multi-view visual features. The RIR is decomposed into a source response, a beam-traced reflection response conditioned on visual features and integrated positional encodings, and a learned residual field for late reverberation and diffraction. The method is evaluated on the Real Acoustic Field (RAF) dataset and the Hearing Anything Anywhere (HAA) dataset, reporting improved C50, EDT, T60, and loudness errors over learning-based and physics-based baselines, including a claimed 10x data-efficiency advantage on RAF and large same-scale gains. The paper also provides ablations, qualitative wave-field visualizations, interpretable reflection-response maps, and supplementary computational-cost and failure-case analyses.","tokens_in":18996,"tokens_out":4158,"duration_ms":45429,"significance":"If the claims hold, AV-DAR would be a meaningful step toward few-shot, physics-based room impulse response estimation, and the paper contains several strengths: it is, to my knowledge, the first to integrate acoustic beam tracing into a fully differentiable end-to-end RIR renderer; it evaluates on six real-world rooms from two datasets; it includes component ablations, view-count ablations, failure cases, and a disclosed data-quality exclusion in the supplement; and the learned reflection responses are visualized and shown to be material-aware. However, the headline claims about the multi-view vision prior and about 'significantly outperforming' baselines are not yet fully supported by the evidence as presented: all comparisons are single-run, the vision ablation is confined to one scene and one training scale, and a key metric in the 10x-data comparison is not favorable to the method. The central architecture is plausible and the empirical direction is promising, but the load-bearing evidence needs strengthening before the paper's strongest claims can be accepted.","major_comments":[{"comment":"All reported results are single runs without error bars, multiple random seeds, or significance tests. The abstract and §4.2 repeatedly state that AV-DAR 'significantly outperforms' prior methods, but no statistical support is provided. This matters because several differences are small or inconsistent; for example, in Table 3 the T60 error is lower without vision (14.3) than with the full model (15.2). Please report variance over multiple training runs and, where appropriate, paired significance tests for the main comparisons.","section":"§4.2, Tables 1, 2, 3, and 6"},{"comment":"The vision-prior ablation ('w/o Vision') is reported only for RAF-Furnished at 0.1% training data. Since the multi-view vision prior is the central contribution stated in the title and introduction, the absence of w/o Vision results for RAF-Empty, the four HAA rooms, and the 1% training scale is a load-bearing gap: the large reported gains could be driven primarily by the beam-tracing plus residual structure rather than by the visual features. The text's statement that 'each component is essential' is also not uniformly supported by Table 3, where removing vision improves T60. Please add cross-scene vision ablations and discuss metric-specific effects.","section":"§4.2, Table 3 and ablation paragraph"},{"comment":"For the HAA dataset, the multi-view 'images' are rendered from Polycam reconstructions rather than real photographs. The paper discloses this, but the central claim concerns visual priors from multi-view images, and the HAA experiments therefore rest on an unvalidated assumption that rendered Polycam views are an acceptable substitute for real photographs. The manuscript should explicitly frame HAA as testing the method with rendered visual inputs, and either provide evidence that this surrogate preserves the relevant material cues or add a caveat that real-photo performance on HAA is not yet evaluated.","section":"§4.1, Implementation Details (HAA images)"},{"comment":"The statement that AV-DAR 'achieves comparable performance to models trained on 10 times more data' is not uniformly true across the reported metrics. For example, on RAF-Furnished EDT, Ours at 0.1% data reports 80.1 ms, which is worse than AVR at 1% data (72.3 ms) and NAF++ at 1% data (74.9 ms). The claim should be qualified to identify the metrics and rooms for which the 10x-data comparison holds, and the paper should avoid implying overall superiority on all metrics.","section":"§4.2, Table 1 and '10x data' claim"}],"minor_comments":[{"comment":"The word 'quires' appears in the text describing the attention mechanism; it should be 'queries.'","section":"§3.5, Eq. (15)"},{"comment":"There is a duplicated phrase 'the importance of importance of visual information' and a typo 'evluat- ing' in the supplementary; these should be corrected.","section":"§A.2.4"},{"comment":"The paper specifies that HAA images are rendered at 512x512 resolution from Polycam reconstructions, but it does not state how many rendered views are used per HAA room; please provide these counts for reproducibility.","section":"§4.1, Implementation Details"},{"comment":"The beam-tracing apex angle, the number of beams, the source directional sharpness parameter, and the frequency discretization are free parameters, but no sensitivity analysis is reported; a short sensitivity study or a justification for the chosen values would strengthen reproducibility.","section":"§3.3 and §A.1.4"}],"recommendation":"major_revision","confidential_remarks":"The main risk is that the headline claims (vision-prior contribution and statistical significance) outrun the evidence. The paper is technically sound in its architecture, and I see no circularity; the missing pieces are cross-scene vision ablations, multi-run statistics, and a more careful qualification of the 10x-data comparison. These are fixable within the manuscript's scope, hence major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know about AV-DAR is that it does something genuinely new: it puts acoustic beam tracing inside an end-to-end differentiable renderer and conditions the learned reflection responses on multi-view visual features. That integration is not in the prior literature, and the engineering is careful. The results across six real rooms on RAF and HAA are broad and mostly consistent, with large relative gains over strong baselines at equal training scale and a credible 10x data-efficiency story at 0.1% training data.\n\nThe second thing: the vision prior, which is the headline of the title and abstract, is only ablated in one room (RAF-Furnished) at one data scale (0.1%). The ablation shows vision helps EDT and C50 but slightly hurts T60. That is a real support gap for the central claim. I trust the stress-test note here: without w/o Vision results for the other five rooms, the multimodal contribution could be a modest boost to a pipeline that is mostly carried by the beam tracing plus learned residual structure. The residual component is crucial (ablating it is catastrophic), and the paper would be stronger if it framed vision as a helpful, uneven component rather than the main driver.\n\nOther soft spots are minor. All tables report single runs, so 'significantly outperforming' is not statistically supported; adding error bars or repeated runs would help. The exclusion of one invalid RIR due to speaker failure is disclosed honestly but only in the supplement; that should be in the main text.\n\nOverall this is a solid empirical contribution from a serious group. The novelty is real and the method is efficient, with fast inference and reasonable training times. The main fix is to broaden the vision ablation and tone down the multimodal claim until it is backed by more scenes. I would send it to a serious referee.","headline":"Genuinely novel integration of beam tracing with multi-view visual cues, but the vision-prior claim is only ablated in one room, so treat the multimodal headline as promising rather than proven.","tokens_in":19474,"tokens_out":2941,"would_cite":true,"duration_ms":29395,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Combining multi-view visual features with acoustic beam tracing in a differentiable renderer yields accurate room impulse response estimation from sparse real-world measurements, outperforming prior methods and matching models trained on…","keywords":["room impulse response","differentiable rendering","acoustic beam tracing","audio-visual learning","multi-view vision","material-aware reflection","spatial audio","real-world RIR estimation"],"falsifier":"Take a room whose walls wear a single uniform paint hiding different substrates (drywall on one side, concrete masonry on another), record RIRs, and compare AV-DAR with and without the visual branch: if the vision-conditioned model no longer beats the acoustic-only version, the visual-prior premise is falsified for that setting. More directly, impedance-tube measurements of two surface samples with identical RGB appearance but different absorption coefficients would yield reflection responses that a vision-only encoder cannot distinguish.","tokens_in":18548,"feed_emoji":"🎧","tokens_out":11437,"duration_ms":99414,"temperature":0.7,"pith_summary":"Realistic spatial audio requires knowing a room's impulse response—how sound bounces from a speaker to a listener—but dense measurements are costly and pure learning methods need huge datasets. The paper proposes that a room's appearance can supply most of the missing information, because visible surface materials largely determine how sound reflects. Its AV-DAR system feeds multi-view images through a vision encoder whose material-aware features condition a differentiable beam-tracing renderer, so reflection responses are learned from a handful of recorded RIRs. On two real-world datasets with six rooms, AV-DAR trained on 0.1% of the measurements matches baselines trained on 1%, and at equal scale it improves relative error by 16.6% to 50.9%. If the visual-acoustic correlation holds, this makes per-room acoustic capture cheap enough for consumer AR/VR.","feed_headline":"Room images replace most of the measurements needed to render sound","feed_subtitle":"Multi-view images as material cues let a physics-based renderer match models trained on 10× more room impulse responses.","key_machinery":"The central object is the differentiable RIR renderer built on acoustic beam tracing, with a vision-conditioned multi-scale reflection response. Beam tracing represents sound as volumetric cones rather than zero-width rays, so a listener is registered as 'hit' by a specular path without Monte Carlo oversampling; each path's frequency response is the product of reflection responses at hit points, mapped to the time domain by a minimum-phase transform and accumulated with propagation loss and delay. Because a beam's footprint on a surface is an ellipse whose size grows with travel distance, the reflection response is evaluated at an integrated positional encoding (IPE) that averages Fourier features over the elliptical region. A multi-view vision encoder supplies a material-aware feature at each surface point, aggregated across cameras by cross-attention and across neighboring samples by point-transformer fusion, so the same point can have different effective reflection responses depending on viewing context and beam scale. A residual neural field treats every surface point as a secondary source and Monte Carlo-integrates over solid angles to model diffuse reflections and late reverberation. This decomposition lets gradients flow from the RIR loss back through all components, fine-tuning the visual features themselves for acoustic prediction.","core_discovery":"AV-DAR is the claim that RIR rendering can be made both physical and data-lean by letting vision supply the surface reflection properties that acoustics alone cannot identify from sparse measurements. The paper decomposes the RIR into a learnable source response, a reflection response computed by tracing specular beams through coarse room geometry, and a residual field capturing diffuse reflections and late reverberation. Vision enters through a multi-view encoder that aggregates pixel-aligned features across cameras into a material-aware descriptor at each surface point, conditioning the frequency-dependent reflection response at every beam hit. The entire pipeline is optimized end-to-end against ground-truth RIRs, and on the Real Acoustic Field dataset the model trained on 0.1% of data matches baselines trained on 1%, with relative gains of 16.6% to 50.9% at equal scale; on four Hearing Anything Anywhere rooms trained on 12 locations it outperforms all prior physics- and learning-based baselines on nearly every metric.","pith_inferences":["If the visual-to-acoustic correlation holds across scenes, the same vision encoder could be trained across many rooms to enable zero-shot RIR prediction for never-seen spaces from images alone—an extension the paper explicitly leaves to future work.","The two-level cross-attention aggregation of multi-view features is a general mechanism: any inverse problem where image evidence constrains physical surface parameters (e.g., thermal emissivity, tactile roughness) could reuse this pattern of conditioning a differentiable physical simulator.","The current method relies on known rough geometry (e.g., a few planes); coupling AV-DAR's beam tracer with automatic image-based geometry estimation would remove the last manual step, yielding a purely vision-driven acoustic renderer.","Because the residual field is position-dependent and learned, the same decomposition could be reused to estimate physical parameters such as frequency-dependent absorption coefficients that other room acoustic solvers could consume."],"forward_implications":["A new room can be captured with a camera and a handful of microphone recordings rather than tens of thousands of source–listener pairs.","Physics-based rendering becomes fast enough for interactive use: inference stays under 70 ms for a two-second RIR in the tested rooms, an order of magnitude faster than differentiable image-source rendering.","The learned acoustic model is inspectable: optimized reflection responses align with material categories (carpet absorbs high frequencies, metal reflects them), so engineers can trace where a prediction comes from.","The approach transfers across very different real venues—offices, a classroom, a hallway, a dampened room—indicating the physics-vision combination generalizes beyond the training scenes rather than overfitting."],"supporting_citations":[{"why":"Supplies the cone-beam tracing primitive that replaces image-source enumeration and guarantees specular path detection.","marker":"[21]"},{"why":"Provides the RIR decomposition used in Eq. 4 and the differentiable image-source baseline (DiffRIR), plus the Hearing Anything Anywhere dataset and evaluation protocol.","marker":"[67]"},{"why":"Real Acoustic Field dataset: dense real-world RIRs with co-captured images, used for the scaling comparisons from 0.01% to 100% of the training data.","marker":"[13]"},{"why":"Supplies the integrated positional encoding that averages Fourier features over the beam's elliptical contact region in the multi-scale reflection response.","marker":"[4]"},{"why":"Acoustic Volume Rendering (A VR) is the volume-rendering baseline used for comparison and for the quantitative claim that beam tracing is more efficient and accurate.","marker":"[35]"},{"why":"Pre-trained vision encoder (DINOv2) that produces the per-view pixel-aligned features; ablations show it substantially outperforms ResNet18.","marker":"[49]"},{"why":"Point Transformer fusion propagates features from sampled basis points to arbitrary query points, enabling the material-aware feature F(x,Σ).","marker":"[71]"},{"why":"Minimum-phase transform converts the frequency-domain path response into a causal time-domain impulse in the reflection renderer.","marker":"[45]"}],"fun_headline_variants":["Vision priors let physics-based room acoustics match 10× data models","Multi-view images cut room acoustic data needs by 10×","Physics + vision: room acoustics with 0.1% of the data","AV-DAR: Room sound from images, not just measurements","Seeing the room makes sound rendering data-lean"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire method rests on the assumption that surface appearance reliably predicts acoustic reflection behavior—if visually similar materials reflect sound very differently, or the same material looks different under varied lighting, the vision encoder will train on a misleading signal and the data-efficiency gains vanish.","fun_headline_variants_meta":{"raw":{"variants":["Vision priors let physics-based room acoustics match 10× data models","Multi-view images cut room acoustic data needs by 10×","Physics + vision: room acoustics with 0.1% of the data","AV-DAR: Room sound from images, not just measurements","Seeing the room makes sound rendering data-lean"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000271,"raw_usage":{"total_tokens":1603,"prompt_tokens":892,"completion_tokens":711,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":621}},"tokens_in":508,"tokens_out":711,"duration_ms":6324,"temperature":1.0,"reasoning_tokens":621,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:52:28.186238+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a room whose walls wear a single uniform paint hiding different substrates (drywall on one side, concrete masonry on another), record RIRs, and compare AV-DAR with and without the visual branch: if the vision-conditioned model no longer beats the acoustic-only version, the visual-prior premise is falsified for that setting. More directly, impedance-tube measurements of two surface samples with identical RGB appearance but different absorption coefficients would yield reflection responses that a vision-only encoder cannot distinguish.","supporting_citations":[{"cited_title":"A beam tracing approach to acoustic modeling for interactive virtual environments","cited_arxiv_id":null,"evidence_quote":"Supplies the cone-beam tracing primitive that replaces image-source enumeration and guarantees specular path detection."},{"cited_title":"Hearing anything any- where","cited_arxiv_id":null,"evidence_quote":"Provides the RIR decomposition used in Eq. 4 and the differentiable image-source baseline (DiffRIR), plus the Hearing Anything Anywhere dataset and evaluation protocol."},{"cited_title":"Real acoustic fields: An audio-visual room acous- tics dataset and benchmark","cited_arxiv_id":null,"evidence_quote":"Real Acoustic Field dataset: dense real-world RIRs with co-captured images, used for the scaling comparisons from 0.01% to 100% of the training data."},{"cited_title":"Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P","cited_arxiv_id":null,"evidence_quote":"Supplies the integrated positional encoding that averages Fourier features over the beam's elliptical contact region in the multi-scale reflection response."},{"cited_title":"Acoustic volume rendering for neural impulse re- sponse fields","cited_arxiv_id":null,"evidence_quote":"Acoustic Volume Rendering (A VR) is the volume-rendering baseline used for comparison and for the quantitative claim that beam tracing is more efficient and accurate."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Pre-trained vision encoder (DINOv2) that produces the per-view pixel-aligned features; ablations show it substantially outperforms ResNet18."},{"cited_title":"Point transformer","cited_arxiv_id":null,"evidence_quote":"Point Transformer fusion propagates features from sampled basis points to arbitrary query points, enabling the material-aware feature F(x,Σ)."},{"cited_title":"Gregory McDaniel and Cory L","cited_arxiv_id":null,"evidence_quote":"Minimum-phase transform converts the frequency-domain path response into a causal time-domain impulse in the reflection renderer."}],"review_version":1}