{"id":"37737895-e0c8-47a4-a5b1-420962214eeb","arxiv_id":"2608.00463","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Scene2Sound generates object-anchored, spatially consistent soundscapes for 3D Gaussian Splatting worlds without retraining, by merging multi-view detections whose rendering Gaussian sets overlap.","lead":"This paper presents Scene2Sound, a system that automatically adds sound to 3D virtual worlds, anchoring each sound to a fixed 3D object so the audio stays consistent as a listener moves through the world. It combines off-the-shelf vision-language, segmentation, and text-to-audio models, and adds two new metrics for checking whether soundscapes respond correctly to listener motion.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ablating Gaussian set matching leaves CGC F1 unchanged (0.2593 vs. 0.2589), so the core mechanism's contribution to claimed spatial consistency is unproven; the LMC comparison uses confounded pre-adoption audio.","rationale":"The reader's weakest assumption is the GSM premise, and the paper's own ablation supports that concern: the headline grounding metric CGC F1 is numerically identical with and without Gaussian set matching. This is the most load-bearing issue because GSM is presented as the technical novelty that enables view-consistent 3D instances, and the central claim of spatial consistency rests on it. The paper's defense—that precision improves significantly—is real but does not rescue the headline claim, since the abstract and contributions do not qualify spatial consistency as a precision-only effect. The LMC row in the same table is not comparable because the ablation variant uses cached pre-adoption audio, and the subjective study was run at an earlier configuration, so neither the objective nor the perceptual evidence isolates the contribution of GSM. The hollow-structure assumption behind viewpoint selection is acknowledged by the authors as a known failure mode and is less central to the claimed contribution. The configuration mismatch in the subjective study is important but secondary, since the perceptual benefit of the overall pipeline might survive even if GSM specifically were not the cause. The threshold sweep further weakens the premise: CGC improves as merging becomes more conservative, which is inconsistent with the idea that cross-view Gaussian set overlap is a reliable object-identity cue at the adopted threshold. The appropriate verdict remains CONDITIONAL, as additional evidence (matched ablation and subjective test at the final configuration) is needed; the preprint's transparency and validation of the metrics prevent rejection. I therefore leave the reader's verdict unchanged.","tokens_in":22606,"tokens_out":4838,"duration_ms":46162,"concrete_test":"Run a matched comparison of the final Scene2Sound configuration against the w/o-instance-association variant under identical conditions: same final audio prompts, per-label energy normalization, ambient gain, and the Sec. V-F subjective protocol (19 participants, 24 scenes, all three MOS axes), and recompute LMC/CGC at the final rendering configuration for both variants. If CGC F1 remains within noise and the Spatial Congruency MOS is not significantly higher for the full system, the GSM mechanism's contribution to the claimed spatial consistency is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table V shows that removing Gaussian set matching (w/o inst. assoc.) leaves the held-out grounding F1 essentially unchanged (CGC 0.2593 vs. 0.2589), with only the precision/recall operating point shifting (0.441/0.229 to 0.302/0.316). The LMC difference cited in favor of GSM (+0.283 vs. +0.228) is not a clean ablation: per the table footnote, the variant row is scored with cached pre-adoption audio, while only the Scene2Sound row uses the adopted audio prompts and gain normalization. Thus neither spatial-consistency axis cleanly demonstrates that GSM, the paper's advertised backbone, produces the headline result. The Appendix C threshold sweep compounds this: CGC rises monotonically with J_min up to 0.50 (0.159 to 0.286), indicating that a more conservative, less-merge-heavy association—not the shared-primitive-index cue per se—improves held-out grounding. Because the abstract and contributions attribute spatial consistency to view-consistent association via GSM, the central claim is currently supported only by a precision improvement, which is not the metric headline. The subjective study (Sec. V-F) also used an earlier configuration (J_min=0.05, placement cap 3), so it cannot validate the final GSM-based system.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Scene2Sound, a training-free pipeline that takes a pretrained 3D Gaussian Splatting (3DGS) scene and produces an object-based soundscape: automatic viewpoint selection, VLM-based sound-event detection, text-to-audio synthesis, cross-view instance association via Gaussian set matching (GSM), and real-time object-based spatialization. The authors propose two spatial-consistency metrics: Listener-Motion Consistency (LMC), which checks that audio distance tracks listener displacement, and Cross-View Grounding Consistency (CGC), a held-out reprojection F1 that checks whether placed sources are supported by views not used for placement. Experiments on 24 generated scenes and 81 real-world 360-degree scenes compare Scene2Sound with eleven baselines, including a re-implemented SonoWorld. The paper reports that Scene2Sound preserves audio quality while achieving positive LMC and higher CGC than baselines, and a 19-participant listening study shows higher MOS for preference, spatial congruency, and semantic congruency.","tokens_in":22775,"tokens_out":5634,"duration_ms":48811,"significance":"The paper addresses a genuine gap: adding persistent, viewpoint-consistent sound to navigable 3DGS worlds. If the central claims hold, the training-free design and object-based representation are practically valuable, and the proposed metrics are a useful contribution. The paper deserves credit for validating LMC against two real multi-position datasets and a learning check (ViGAS moves LMC from 0 to +0.455), for a held-out CGC protocol with construct-validity corruptions, for a co-located-instance stress test (0/16,284 over-merges), and for committing to release the SoundscapePLY testbed and code. However, the current evidence does not cleanly demonstrate that GSM, the advertised backbone, causes the headline spatial-consistency gains, and several secondary claims outrun the measurements.","major_comments":[{"comment":"The ablation 'w/o inst. assoc.' leaves CGC F1 essentially unchanged (0.2593 vs 0.2589) and only shifts the operating point from precision to recall (0.441/0.229 vs 0.302/0.316). The listener-axis contrast (+0.283 vs +0.228) is not a clean ablation because, per the table footnote, the variant row is scored with cached pre-adoption audio while the Scene2Sound row uses the adopted audio prompts and gain normalization. The threshold sweep in Appendix C further shows CGC rising monotonically with J_min up to 0.50 (0.159 to 0.286), suggesting the benefit may come from conservative merging rather than the shared-primitive-index cue per se. Consequently, neither spatial-consistency axis provides clean causal support for GSM as the source of the headline result. Please run the 'w/o inst. assoc.' variant at the final adopted configuration with matched audio prompts and gain normalization, and report LMC, CGC F1, precision, and recall.","section":"Sec. V-D, Table V"},{"comment":"The 19-participant subjective study was conducted at an earlier pipeline configuration (placement cap 3, J_min=0.05, earlier audio-prompt template), not at the final configuration (J_min=0.15, no cap, adopted prompts) used for all quantitative results. Since the subjective evaluation is presented as confirmation of the proposed system's perceptual benefit, it cannot validate the final GSM-based system; at most it validates the object-based rendering architecture under an earlier association policy. Please either rerun the study at the final configuration or explicitly limit the subjective claim to the rendering architecture and state that the association stage was not perceptually validated.","section":"Sec. V-F, App. F"},{"comment":"The real-world transfer evidence is weaker than the abstract's claim that Scene2Sound 'remains spatially consistent' on 3DGS scenes from real-world 360-degree captures. On D-SAV360, LMC is +0.068 over all 81 scenes and +0.120 on the 46 scenes with grounded point sources; the 35 scenes without grounded sources score exactly 0, and the single capture position precludes the cross-view association that the method's grounding axis depends on, so CGC is not computed in that regime. Please either temper the real-world spatial-consistency claim to 'weakly above chance on the listener axis' or provide a multi-position real-capture evaluation.","section":"Sec. V-B, App. H"}],"minor_comments":[{"comment":"FAD is computed from 24 samples (one per scene) and the paper acknowledges that this is noisy, but no confidence intervals are reported for the FAD differences; given that the metric is used comparatively, bootstrap CIs around the FAD contrasts would make the 'preserves audio quality' claim easier to evaluate.","section":"Sec. IV, Audio Quality"},{"comment":"The distance weight w_d(d_ij) is only described as inverse-distance weighting; providing the explicit functional form in the main text or a specific equation number in the supplementary material would improve reproducibility.","section":"Sec. III-B, Eq. (2)"},{"comment":"The LMC values in the J_min sweep (+0.253 at the adopted column) predate the final rendering revision and differ from Table III's +0.283; the main text and appendix should be harmonized or have the discrepancy explicitly explained at one location.","section":"Table V, App. C"},{"comment":"The phrase 'cross-view association' for the three yaw-rotated panoramas from the same capture position may mislead readers; these are same-position rotations, not spatially distinct views, so the limitations of the single-position protocol should be stated at the first use of that phrase in the appendix.","section":"App. H"}],"recommendation":"major_revision","confidential_remarks":"This is a solid systems paper with useful metrics and a reproducible pipeline. The main obstacle is the gap between the abstract's attribution of spatial consistency to GSM and the ablation evidence, which currently leaves the mechanism's contribution unproven. I believe this is fixable with a clean ablation at the adopted configuration and a more measured real-world spatial-consistency claim, hence major_revision rather than reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Scene2Sound is a genuinely new system: it takes a pre-trained 3DGS world and produces an object-based soundscape with persistent 3D anchors, training-free, and it ships a more transparent evaluation than most papers in this area. The task framing (auditory grounding for navigable 3D worlds) is novel relative to per-view spatial audio and single-panorama methods, and the two consistency metrics — LMC for listener motion, CGC for held-out grounding — are a real contribution, with construct validation, external anchors on real recordings, and a learning check. The Gaussian set matching idea, using the rasterizer's tile–Gaussian sets as a cross-view correspondence index, is simple and plausibly useful. Credit also to the appendices: threshold sweeps, co-located-instance stress tests, stage-wise diagnostics, and honest statements about the hollow-structure assumption and the absence of multi-view real data.\n\nBut the paper's central claim — that GSM drives the spatial consistency gain — is not actually supported. Ablating instance association leaves CGC F1 essentially unchanged (0.2593 vs 0.2589); only precision moves (0.441 vs 0.302). And the LMC comparison that favors Scene2Sound (+0.283 vs +0.228) is not a clean ablation: the variant is scored with cached pre-adoption audio while the full system uses the adopted prompts and gain normalization. So the contribution of the backbone to the headline result is unproven. The subjective study also evaluated an earlier configuration (J_min=0.05, placement cap 3), so the perceptual benefits, while plausibly real, are not evidence for the reported final system. Real-world transfer is single-view only and does not exercise multi-view association at all.\n\nThese are real soft spots, but they are not fatal. The pipeline as a whole is coherent, the evaluation is unusually candid, and the metrics/dataset are useful independent of the GSM claim. What is missing is one clean ablation re-run at the adopted configuration and a config-matched subjective study, or at least a clear acknowledgment that the listener-axis gain is a property of the whole pipeline, not GSM specifically.\n\nThis deserves serious peer review. I would send it out with a request for a clean ablation and a config-matched subjective study, or a revision of the claims to match what is actually measured. The task, metrics, and dataset are valuable enough that the paper should be given a chance.","headline":"A transparent, well-engineered system paper with real novelty in task and metrics, but the advertised GSM backbone does not carry the headline result — claims need to be matched to the evidence.","tokens_in":23491,"tokens_out":3327,"would_cite":true,"duration_ms":28255,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Scene2Sound claims that any pre-trained 3D Gaussian Splatting world can be turned into a spatially consistent audio-visual scene, training-free, by using shared Gaussians as the anchor that keeps each sound attached to its object.","keywords":["3D Gaussian Splatting","spatial audio","soundscape generation","auditory grounding","Gaussian set matching","multi-view instance association","object-based audio","training-free pipeline"],"falsifier":"Take a thin or highly reflective object, render it from two nearly orthogonal viewpoints, compute the Jaccard similarity of the Gaussian sets behind its segmentation masks, and check whether same-object pairs fall below the 0.15 threshold: if they do, the object would be placed as two sources and held-out CGC precision would drop, which would contradict the shared-primitive identity premise.","tokens_in":22243,"feed_emoji":"🔊","tokens_out":6287,"duration_ms":50675,"temperature":0.7,"pith_summary":"Scene2Sound claims that a silent 3D Gaussian Splatting world can be given a coherent soundscape without any training or per-scene optimization, by grounding each sound in a specific object and anchoring it to a persistent 3D position. The paper's central bet is that the 3DGS representation itself supplies the correspondence cue: when two camera views see the same object, the same Gaussians render both detections, so Jaccard overlap over those Gaussian sets identifies the object across views. On 24 curated generated worlds and 81 real-world 360-degree reconstructions, the pipeline keeps audio quality competitive with per-viewpoint generators while staying spatially consistent: audio responds to listener motion (LMC +0.283 against a chance level of 0) and held-out views support the placed sources (CGC 0.259 vs 0.065 for a single-panorama baseline). A 19-participant study found the object-based rendering preferable to three spatial baselines on preference, spatial congruency, and semantic congruency. If correct, the result turns any pre-trained 3DGS scene into a navigable audio-visual world from the scene file alone.","feed_headline":"Silent 3D scenes gain sound that stays put as you move","feed_subtitle":"A training-free pipeline links multi-view sightings of the same object via shared Gaussians and spatializes each sound in real time.","key_machinery":"Gaussian set matching (GSM): Jaccard similarity $J(G_{k,m}, G_{k',m'}) = |G_{k,m}\\cap G_{k',m'}|/|G_{k,m}\\cup G_{k',m'}|$ between the sets of Gaussians that the renderer's tile metadata records as contributing to two detected regions. It does the work of cross-view instance association: along with a fixed threshold $J_{\\min}=0.15$ and Union-Find clustering, it merges observations of one physical object into a single 3D source and keeps distinct objects separate, without learned features, per-Gaussian parameters, or per-scene optimization. Source positions are weighted centroids of the merged Gaussian sets, so the anchors inherit the scene's actual geometry.","core_discovery":"The discovery is that persistent sound sources in a 3DGS world can be created by reading the rasterizer's own bookkeeping. When a segmentation model marks an object in a rendered view, the tile metadata records exactly which Gaussians contributed to those pixels; another view of the same object shares many of the same Gaussians, while a different object does not. Scene2Sound lifts per-view masks to Gaussian sets, merges sets whose Jaccard similarity exceeds 0.15 by Union-Find, and places each merged instance at the weighted centroid of its Gaussians. That gives each sound a stable 3D anchor that survives arbitrary listener motion, with audio synthesized per source by a text-to-audio model and spatialized by a standard object-based engine. The paper argues that this primitive-index shortcut is something multi-view images plus per-view depth cannot offer, because only 3DGS has a shared persistent set of primitives across views.","pith_inferences":["The Gaussian-set overlap trick is not specific to audio: it is a general way to fuse any multi-view 2D signal (labels, semantics, contact points) onto a pre-trained 3DGS world without optimization, so the same association module could serve other scene-authoring tasks.","Because LMC and CGC are defined on rendered audio and held-out views rather than on this pipeline's internals, they could become reusable checks for any generative audio-visual scene model that claims spatial consistency.","The real-scene results are limited to a single capture position, so the next evident bottleneck is data: multi-position real 360-degree captures paired with soundscape annotations would likely close the LMC gap from +0.068 toward +0.283.","A testable extension would use GSM as a change detector: static objects keep their Gaussian sets across time, so re-rendering at later states and re-matching could flag dynamic or moving sources without any tracking model."],"forward_implications":["Any pre-trained 3DGS scene, whether generated or reconstructed, can receive a soundscape in about 222 seconds on one GPU, with no retraining or per-scene optimization.","Object-based audio attached to persistent 3D anchors can be re-spatialized at arbitrary listener poses, so navigation-consistent sound no longer needs a fixed viewpoint or panorama.","Per-viewpoint and single-panorama pipelines cannot pass both consistency axes: they either freeze the waveform (LMC 0) or change it without geometric cause; Scene2Sound reports LMC +0.283 and held-out grounding CGC 0.259.","Upstream errors are contained per source: a hallucinated VLM proposal is dropped when segmentation cannot ground it, and surviving sources are unaffected, so the pipeline degrades gracefully.","The proposed metrics LMC and CGC give the task a measurable definition of spatial consistency, validated against real multi-position recordings and a trained acoustic-synthesis model."],"supporting_citations":[{"why":"Defines the 3D Gaussian splatting scene representation whose persistent primitives serve as the shared index across views.","marker":"[1]"},{"why":"Supplies the tile-based rendering metadata that records which Gaussians contribute to each image tile, the raw material for Gaussian set matching.","marker":"[48]"},{"why":"Grounds VLM grounding queries as pixel-accurate masks, the 2D observations that get lifted to Gaussian sets.","marker":"[18]"},{"why":"Identifies sound-emitting objects and assigns source IDs and audio prompts from the multi-view panoramas.","marker":"[17]"},{"why":"Synthesizes the per-source waveforms from text prompts; swapping it tests backend sensitivity.","marker":"[19]"},{"why":"The single-panorama spatial-audio baseline that Scene2Sound is compared against on spatial consistency and held-out grounding.","marker":"[14]"},{"why":"Supplies the object-based audio representation in which each source is a signal plus a persistent 3D position.","marker":"[16]"}],"fun_headline_variants":["Sound stays put in 3D worlds via Gaussian set linking","No training: consistent 3D soundscape from Gaussian anchors","Persistent 3D sound via Gaussian set overlap","Anchoring audio to Gaussians for consistent sound as you move","3D scenes get sound that sticks, thanks to Gaussian anchors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is GSM's object-identity rule: observations of the same physical object from different views are exactly those whose tile-contributed Gaussian sets have Jaccard similarity above 0.15 — if a real object seen from steep angles, or thin or reflective geometry, shares fewer Gaussians than that, its instances fragment and its sound drifts per view.","fun_headline_variants_meta":{"raw":{"variants":["Sound stays put in 3D worlds via Gaussian set linking","No training: consistent 3D soundscape from Gaussian anchors","Persistent 3D sound via Gaussian set overlap","Anchoring audio to Gaussians for consistent sound as you move","3D scenes get sound that sticks, thanks to Gaussian anchors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00084,"raw_usage":{"total_tokens":3706,"prompt_tokens":1034,"completion_tokens":2672,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":650,"completion_tokens_details":{"reasoning_tokens":2587}},"tokens_in":650,"tokens_out":2672,"duration_ms":18416,"temperature":1.0,"reasoning_tokens":2587,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:19:48.757417+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a thin or highly reflective object, render it from two nearly orthogonal viewpoints, compute the Jaccard similarity of the Gaussian sets behind its segmentation masks, and check whether same-object pairs fall below the 0.15 threshold: if they do, the object would be placed as two sources and held-out CGC precision would drop, which would contradict the shared-primitive identity premise.","supporting_citations":[{"cited_title":"3D gaussian splatting for real-time radiance field rendering,","cited_arxiv_id":null,"evidence_quote":"Defines the 3D Gaussian splatting scene representation whose persistent primitives serve as the shared index across views."},{"cited_title":"gsplat: An open-source library for gaussian splatting,","cited_arxiv_id":null,"evidence_quote":"Supplies the tile-based rendering metadata that records which Gaussians contribute to each image tile, the raw material for Gaussian set matching."},{"cited_title":"Stable audio open,","cited_arxiv_id":null,"evidence_quote":"Synthesizes the per-source waveforms from text prompts; swapping it tests backend sensitivity."},{"cited_title":"SonoWorld: From one image to a 3D audio-visual scene,","cited_arxiv_id":null,"evidence_quote":"The single-panorama spatial-audio baseline that Scene2Sound is compared against on spatial consistency and held-out grounding."},{"cited_title":"Object-based 3D audio production for virtual reality using the audio definition model,","cited_arxiv_id":null,"evidence_quote":"Supplies the object-based audio representation in which each source is a signal plus a persistent 3D position."}],"review_version":2}