{"id":"4236e86e-5aaa-4932-89ce-35166ac555ce","arxiv_id":"2608.09435","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"ST-Omni-R1, trained on the new synthetic ST-OmniQA benchmark, outperforms general audio-visual models on spatio-temporal sound event reasoning by combining FOA trajectory tokens with panoramic visual context.","lead":"This paper introduces ST-OmniQA, a synthetic benchmark of 40,000 panoramic videos with spatial audio and 400,000 questions that test whether models can recognize, localize, and track sound sources over time and bind them to visible objects. It also presents ST-Omni-R1, a model that fuses sound direction and motion cues with panoramic video, reporting 77.83% average accuracy versus 37.28% for the best general baseline.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 77.83% headline is measured entirely by DeepSeek-v4-flash, which is also the free-form reward model during RT-RL; without human agreement or a judge-free check, the numeric gap may reflect judge alignment rather than semantic skill.","rationale":"The reader's weakest-assumption identification is correct: the unvalidated DeepSeek-v4-flash judge is load-bearing for every reported accuracy. My read adds a specific aggravating factor: the same class of judge is used not only for evaluation but also as the free-form reward model during reasoning-tree RL, so the model can be explicitly optimized toward the judge's stylistic preferences. That makes the concern sharper but does not change the verdict. The architectural story and benchmark design are plausible, and the internal ablations in Table 4 are consistent with a real effect from curriculum and modality fusion. However, with no released artifacts, no error bars, and no human or exact-match validation of the judge, the strongest quantitative claims are appropriately conditional. I do not see an internal inconsistency that would justify rejection, and the transfer results are interesting enough to warrant further scrutiny rather than dismissal. Therefore the reader's CONDITIONAL verdict should stand unchanged.","tokens_in":12485,"tokens_out":4472,"duration_ms":48813,"concrete_test":"Select a stratified sample of 200 ST-OmniQA test responses (50 per level) plus, if feasible, 50 responses from each transfer benchmark. Have at least two human annotators blind to model identity label each response as semantically equivalent to the reference using the paper's definition; compute inter-annotator agreement and agreement between DeepSeek-v4-flash and the human labels, then recompute per-level accuracy and the gap versus the best baseline under human labels. If the gap remains near 40 points and judge-human agreement is high, the concern is resolved; if human-judged accuracy drops materially or differs by level, the headline needs qualification. As a cheaper auxiliary check, recompute Level A and B accuracy by exact or deterministic answer matching, where answers are unique by construction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Every reported accuracy in Tables 2 and 3 is a DeepSeek-v4-flash semantic-equivalence judgment, and the same evaluator scores free-form responses during Stage-II reward computation (Eq. 14 and the following sentence in Section \"Stage II\"). No human agreement, per-category validation, or error analysis is reported for this judge. Because reasoning-tree RL optimizes the model toward this judge's notion of semantic equivalence, a systematic stylistic preference of the judge would be amplified in the final numbers, potentially inflating the 77.83% vs. 37.28% gap and the transfer results without any genuine semantic gain. The concern is not that the judge is necessarily biased; it is that the central quantitative claim is a property of an unvalidated black box, and the training procedure gives the model direct access to that same black box during optimization. This is an evaluation-pipeline assumption rather than an architectural flaw, but it is load-bearing: changing or validating the judge could change all headline comparisons, especially for levels with free-form answers.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ST-OmniQA, a synthetic benchmark built from 40K panoramic videos with synchronized first-order Ambisonics (FOA) audio, containing 400K question-answer pairs organized into four capability levels that cover source perception, multi-source localization, trajectory relations, and audio-visual binding. The paper also proposes ST-Omni-R1, a model that combines an STA-encoder producing one semantic token and 40 trajectory tokens with panoramic visual context, and trains it with progressive curriculum learning followed by reasoning-tree reinforcement learning. The authors report 77.83% average semantic accuracy on ST-OmniQA versus 37.28% for the best evaluated baseline, and additional transfer results on TAU-NIGENS, L3DAS22, and STARSS23.","tokens_in":12736,"tokens_out":4100,"duration_ms":43181,"significance":"If the results hold, the benchmark and model fill a genuine gap: treating sound sources as persistent, spatially located, moving entities that must be bound to visible instances, rather than as clip-level acoustic events. The paper's strengths include the modality-necessity constraint in Eq. (2), which reduces unimodal shortcuts; the scene-room-level split that prevents leakage; the executable reasoning graphs behind Level C/D traces; and the transfer evaluation on three public spatial-audio benchmarks. The central weakness is that all reported accuracy numbers are produced by a single unvalidated LLM judge that also participates in the reward computation during training, so the quantitative claims are not yet established. The manuscript is worth pursuing after the evaluation pipeline is independently validated and statistical variability is reported.","major_comments":[{"comment":"All semantic-accuracy numbers in Tables 2-4 are computed by DeepSeek-v4-flash, yet no human agreement, per-category validation, or error analysis is reported. Because every headline comparison passes through this judge, the 77.83% versus 37.28% gap and the transfer numbers are properties of the judge as much as of the models. Please report human agreement on a representative sample, per-level and per-question-type accuracy, and a judge-free exact-match check for structured fields such as event class, DoA, distance, and motion state.","section":"Experiments, \"Evaluation protocol\" paragraph"},{"comment":"The same semantic-equivalence evaluator that produces the reported accuracy is used to score free-form responses inside the RT-RL reward in Eq. (14). The model is therefore optimized directly toward the judge's notion of semantic equivalence, and any stylistic preference of the judge is amplified during training; the Level D jump from 90.90 (SFT) to 94.70 (SFT+RT-RL) in Table 4 may reflect judge alignment rather than reasoning skill. Please validate the judge independently, evaluate with a different judge or human raters, and report the agreement between RT-RL reward scores and human judgments.","section":"Stage II, Eq. (14) and the following sentence"},{"comment":"No seed variance, error bars, or significance tests are reported, despite small differences such as the Level C comparison of 55.70 versus 54.40 in Table 4. Please report at least three seeds with standard deviations, or bootstrap confidence intervals, for the main comparisons so that the reader can assess whether the reported gaps are statistically meaningful.","section":"Experiments, \"Implementation details\" and Tables 2-4"},{"comment":"The baseline comparisons are not controlled: general-purpose baselines are evaluated without task-specific tuning and do not receive spatial audio, while BAT retains its own free-form prompting protocol, and on the transfer benchmarks BAT scores are near zero on several axes. The claim of superiority over the best baseline is therefore not established under a common protocol. Please add at least one spatial-audio-capable baseline under the same input and evaluation protocol, or present an explicit audio-only controlled comparison that reports both models with matched modalities.","section":"Experiments, \"Main Results\" and Tables 2-3"}],"minor_comments":[{"comment":"References for Hyun-Bin et al. 2026 and Oh et al. 2026 cite the same arXiv identifier (2606.14141) but are listed as separate works; please merge or disambiguate them.","section":"References"},{"comment":"Footnote 1 states that additional details appear in the Appendix, but no appendix is included in the manuscript; please either add the appendix or remove the pointer.","section":"Experiments, footnote 1"},{"comment":"The paper does not state whether the benchmark, evaluation prompts, or model weights will be released; given the benchmark's intended reuse, an availability statement would be valuable.","section":"Experiments, general"},{"comment":"The reference for Yang, Liu, and Li 2022 contains the typo \"Acouts.\" in the venue name; please correct it to \"Acoustics\".","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and interesting problem, and the benchmark design shows care (modality-necessity filtering, scene-room splits, executable reasoning graphs). However, the evaluation bottleneck is serious: a single unvalidated proprietary judge that also appears in the training reward, with no human agreement, no seeds, and no release plan. I would ask the authors to fix the evaluation pipeline before acceptance, since the central quantitative claims all pass through it. The duplicate reference entry for the same arXiv paper also suggests the related-work section needs careful curation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: ST-OmniQA is a solid, well-designed benchmark for a real gap — binding moving sound sources to visible objects over time — and ST-Omni-R1 shows a clear, large margin over existing baselines. But the headline numbers rest on an unvalidated LLM judge that also serves as the RL reward model, and no artifacts are released. Treat the quantitative claims as conditional; the benchmark itself is worth engaging with.\n\nWhat's actually new: 40K synthetic indoor videos (Matterport3D + SoundSpaces2.0), 400K QA pairs, four capability levels, and modality-necessity filtering (Eq. 2) that ensures Level D questions can only be answered with joint audio-visual evidence. That's a thoughtful design. The trajectory-token interface (K=40 tokens over time-varying FOA intensity) is a reasonable way to give an LLM access to source motion, and the transfer results on TAU-NIGENS, L3DAS22, and STARSS23 support the claim that the learned representations generalize beyond the synthetic benchmark.\n\nThe soft spots are mostly about evaluation. Every reported accuracy is a DeepSeek-v4-flash semantic-equivalence judgment, with no human agreement or per-category error analysis. Since the same evaluator scores free-form answers during Stage-II RL, the model is optimized toward that judge's notion of correctness; a stylistic bias would be amplified, not averaged out. This is a genuine concern, though partially mitigated by the fact that many questions have annotation-defined targets that can be checked directly, and the RL gain over SFT is only ~2.5 points. Still, the absolute gap (77.83% vs 37.28%) could shift if the judge is swapped. There are also no error bars, no seeds, and no released code or data — disappointing for a benchmark paper. The baselines are a bit uneven: general models have no spatial audio at all, so the margin is not surprising; BAT is the only spatial comparison and it's evaluated with its own protocol.\n\nThe math and architecture are fine as far as I can tell. The loss functions are standard, the curriculum is sensible, and the transfer evaluation is a good idea. The citation pattern looks okay, aside from a duplicate entry for the concurrent ST-AudioLM work.\n\nWho's this for? People building audio-visual LLMs or evaluating spatial reasoning will find the benchmark useful. I'd like to see this go through peer review, with the judge validated and artifacts released. If I were editor, I'd send it to review and ask for those in revision.\n\nMy recommendation: engage with it, but keep the numbers at arm's length until the evaluation protocol is validated.","headline":"ST-OmniQA is a well-designed benchmark for a real gap, but the headline numbers rest on an unvalidated LLM judge that also serves as the RL reward model, so treat the quantitative claims as conditional until the evaluator and artifacts are sorted out.","tokens_in":13282,"tokens_out":2655,"would_cite":true,"duration_ms":25909,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that dynamic sound scenes require tracking each source's identity, position, and motion over time, and presents a model and benchmark that do so.","keywords":["spatio-temporal audio-visual reasoning","first-order Ambisonics","sound event localization and detection","multimodal language models","trajectory tracking","audio-visual question answering","curriculum learning","reinforcement learning"],"falsifier":"Take a few hundred held-out ST-OmniQA answers from all four levels and have human annotators score semantic equivalence instead of the automated judge; if the human scores change the ranking or shrink the gap, the headline claim fails. As a second check, ablate the 40 trajectory tokens from the model's audio input; if Level C and Level D accuracy do not fall, the trajectory mechanism is not carrying the result.","tokens_in":12319,"feed_emoji":"🔊","tokens_out":11519,"duration_ms":96739,"temperature":0.7,"pith_summary":"Understanding a dynamic sound scene means deciding what emits each sound, where the source is, how it moves, and which visible object it belongs to; the paper argues this is a tracking-and-binding problem, not a clip-level classification problem. To evaluate this capability, the paper builds ST-OmniQA, a synthetic benchmark of 40,000 panoramic videos with synchronized first-order Ambisonics audio and 400,000 question-answer pairs organized into four levels, from single-source sound-event recognition to multi-source trajectory reasoning and visual-instance binding. The paper then proposes ST-Omni-R1, which represents each clip as one semantic audio token plus 40 time-ordered trajectory tokens and fuses those with panoramic visual tokens inside a language-model decoder, trained by progressive curriculum learning and reasoning-tree reinforcement learning. The reported result is 77.83% average semantic accuracy on ST-OmniQA against 37.28% for the best evaluated baseline, with transfer gains on three public spatial-audio benchmarks.","feed_headline":"AI model tracks moving sounds, beating best baseline by 40 points","feed_subtitle":"A new benchmark and model link sound events to visible objects and motion over time, scoring 77.83% vs 37.28%.","key_machinery":"The load-bearing object is the STA-encoder's split audio representation: a single semantic token for clip-level event content and $K=40$ trajectory tokens that retain time-varying direction, distance, and activity. The input is a seven-channel FOA-derived feature tensor combining channel-wise log-Mel spectrograms with Mel-projected acoustic-intensity components, processed by an Audio Spectrogram Transformer; temporal self-attention over frequency-pooled patch features produces the ordered trajectory tokens. Their initialization loss $L_{\\text{traj}}$ supervises per-bin binary activity, unit direction vectors, and log-distance, while a frozen static encoder preserves event semantics. These audio tokens are projected by a trainable connector and concatenated with video and question tokens as $C = [V_{\\text{tok}}; A_{\\text{tok}}; Q_{\\text{tok}}]$, which lets the language model bind acoustic trajectories to visible instances. The benchmark uses the constraint $|C_A|>1$, $|C_V|>1$, $|C_{AV}|=1$ so that a Level-D answer is unique only after cross-modal binding.","core_discovery":"The central claim is that time-varying source geometry is the missing ingredient in audio-language and vision-language models, and that it can be learned and expressed through FOA-derived semantic and trajectory representations. In the benchmark, each source carries a multimodal state $S_i(t) = \\{e_i, a_i(t), g_i(t), m_i(t), v_i(t), o_i, R_i(t)\\}$, where $g_i(t)$ holds azimuth, elevation, and source-listener distance and $R_i(t)$ holds relations to landmarks, occluders, and other sources. ST-Omni-R1's STA-encoder turns a four-channel FOA waveform into a semantic token $s$ and $K=40$ temporally ordered trajectory tokens $\\tau_1,\\dots,\\tau_K$, supervised during initialization by per-bin activity, unit direction, and log-distance targets; the connector inserts these into the decoder context together with video tokens. Progressive curriculum stages and reasoning-tree reinforcement learning train the model, which reaches 84.70%, 76.20%, 55.70%, and 94.70% semantic accuracy on the four levels and a 77.83% average, and outperforms the BAT baseline on the TAU-NIGENS, L3DAS22, and STARSS23 transfer sets. Semantic accuracy here means whether a generated answer conveys the same meaning as the reference.","pith_inferences":["Because the benchmark scenes are rendered from indoor 3D meshes with simulated room acoustics, real-world deployment could face a domain gap that the three transfer datasets only partially cover.","The single automated semantic-equivalence judge makes the reported gap contingent on that judge's scoring behavior; a human-validated scoring subset would be the most direct check, and the paper does not report one.","The trajectory-token interface could be reused outside question answering, for example as a generic audio-state module for embodied agents that must localize and follow a moving sound source; the paper does not make that claim.","Because each question is generated from executable reasoning graphs over structured scene states, the benchmark could be extended to harder compositional or physical-reasoning questions beyond the four levels covered."],"forward_implications":["General-purpose audio-visual and audio-only models that compress a clip into global event semantics will miss source-level geometry; ST-OmniQA puts this gap at 37.28% best baseline versus 77.83%.","The learned spatial and motion token sequence transfers to real-world audio and audio-visual spatial datasets, so the representation is not tied only to the synthetic benchmark.","Because Level-D questions are constructed so that neither modality alone determines the answer, strong performance on that level is evidence of genuine cross-modal binding rather than unimodal shortcutting.","Progressive curriculum stages and reasoning-tree RL both add capability, with Stage II lifting the hardest scene-grounded level from 90.90% to 94.70%."],"supporting_citations":[{"why":"Supplies the Matterport3D scene meshes used to instantiate the indoor environments in ST-OmniQA.","marker":"(Chang et al. 2017)"},{"why":"Supplies SoundSpaces 2.0 room-acoustics simulation that renders time-varying reverberation for moving sources.","marker":"(Chen et al. 2022)"},{"why":"Provides the AST transformer backbone from which the STA-encoder is initialized.","marker":"(Gong, Chung, and Glass 2021)"},{"why":"Provides the pretrained visual encoder and language decoder that ST-Omni-R1 builds on.","marker":"(Bai et al. 2025)"},{"why":"Supplies the GRPO group-relative policy optimization used for reasoning-tree reinforcement learning.","marker":"(Shao et al. 2024)"},{"why":"Defines the BAT spatial-audio language model used as the main spatial baseline and transfer comparison.","marker":"(Zheng et al. 2024)"},{"why":"Supplies STARSS23, the real-world audio-visual dataset used to test transfer of spatial and motion representations.","marker":"(Shimada et al. 2023)"},{"why":"Supplies TAU-NIGENS, the dynamic reverberant-scene dataset used as an audio-only transfer test.","marker":"(Politis et al. 2021)"},{"why":"Supplies L3DAS22, the real-office 3D audio dataset used as an audio-only transfer test.","marker":"(Guizzo et al. 2022)"},{"why":"Provides the automated semantic-equivalence evaluator through which all reported accuracy scores pass.","marker":"(Xu et al. 2026)"}],"fun_headline_variants":["Sound-aware AI locates and tracks sources in 360° video","Tracking moving audio sources: new benchmark and model","AI that hears and sees: sound source tracking leaps","From 37% to 78%: model reasons about sound movement","Spatial audio model tracks sources, beats baselines by 40"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Every reported accuracy number is produced by a fixed automated semantic-equivalence judge with no reported human agreement, error analysis, or per-category validation; if that judge is lenient or biased toward ST-Omni-R1's answer style, the headline gap and the transfer numbers are not established.","fun_headline_variants_meta":{"raw":{"variants":["Sound-aware AI locates and tracks sources in 360° video","Tracking moving audio sources: new benchmark and model","AI that hears and sees: sound source tracking leaps","From 37% to 78%: model reasons about sound movement","Spatial audio model tracks sources, beats baselines by 40"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000681,"raw_usage":{"total_tokens":3149,"prompt_tokens":1058,"completion_tokens":2091,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":674,"completion_tokens_details":{"reasoning_tokens":2005}},"tokens_in":674,"tokens_out":2091,"duration_ms":15216,"temperature":1.0,"reasoning_tokens":2005,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T17:26:26.377115+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a few hundred held-out ST-OmniQA answers from all four levels and have human annotators score semantic equivalence instead of the automated judge; if the human scores change the ranking or shrink the gap, the headline claim fails. As a second check, ablate the 40 trajectory tokens from the model's audio input; if Level C and Level D accuracy do not fall, the trajectory mechanism is not carrying the result.","supporting_citations":[{"cited_title":"2022 , volume =","cited_arxiv_id":null,"evidence_quote":"Supplies SoundSpaces 2.0 room-acoustics simulation that renders time-varying reverberation for moving sources."},{"cited_title":"2024 , volume =","cited_arxiv_id":null,"evidence_quote":"Defines the BAT spatial-audio language model used as the main spatial baseline and transfer comparison."},{"cited_title":"2021 , volume =","cited_arxiv_id":null,"evidence_quote":"Supplies TAU-NIGENS, the dynamic reverberant-scene dataset used as an audio-only transfer test."}],"review_version":1}