{"id":"b960bf67-4434-4dfd-84b0-e4fb83f57ef8","arxiv_id":"2607.11433","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A query-scoped evidence state that conditions planning, verification, repair, and stopping raises OmniGAIA accuracy to 45.6% and WorldSense to 58.3% versus free-form agent baselines.","lead":"Omni-Decision is a training-free agent that keeps a structured evidence ledger—what is grounded, missing, conflicting, or still needed—while answering questions over video, audio, web, and code. The shared state improves multi-step omni-modal QA by large margins over free-form tool agents on two benchmarks.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The central causal claim is still under-isolated: state gains may largely reflect higher tool budgets and longer trajectories rather than the shared evidence ledger itself.","rationale":"I agree with the reader's weakest-assumption diagnosis and do not find a deeper internal inconsistency. The systems contribution is real: shared state_digest, deterministic reducer commits, readiness from Ut/Ft/Ct, and inspectable stuck-rate localization (Appendix H) are well specified, and WorldSense transfer plus backend swaps reduce pure dataset-overfit risk. The remaining issue is causal attribution of the headline accuracy jump. Because the paper itself reports the large tool-call disparity and confines the no-state test to a 120-case subset without budget matching, CONDITIONAL remains the right verdict; the concrete full-split matched-budget ablation above is the single check that would most cleanly settle whether the evidence ledger, rather than longer seeking under strong proprietary models, is doing the work. No stronger objection (e.g., broken readiness definition or non-reproducible judging) is needed to move the verdict further.","tokens_in":24719,"tokens_out":673,"duration_ms":8175,"concrete_test":"Re-run the default gpt-5.2 + gemini-3.1-pro configuration on the full 360 OmniGAIA split with three arms: (A) full Omni-Decision; (B) no-state free-form history with the same max steps and a hard cap of 11 tool calls/query; (C) no-state with the base agent's ~2.5-call budget. If (B) recovers most of the +27 pp over base while matching (A)'s call volume, the state-interface claim weakens; if (A) still leads (B) by a large margin under matched budgets, the claim is substantially strengthened.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's strongest claim is that explicit query-scoped evidence-state control, not merely more tools or stronger backends, drives the large OmniGAIA gains (+27.3 pp over the official base agent; Table 1). The load-bearing soft spot is that this causal isolation is incomplete. Appendix F shows Omni-Decision averages 11.22 tool requests per query versus 2.45 for OmniGAIA base (4,038 vs 881 total), and the system is allowed up to 15 iterations with critic-gated repair. The no-state ablation (Table 2) is only on a fixed 120-case stratified subset, not the full 360-split, and does not match tool-call budgets or trajectory length: removing St drops subset accuracy by 9.17/10.83/3.33 pp depending on backend, but a no-state agent that still runs ~11 tool calls under the same planner/critic/finalizer could close much of that gap via extra search and verification alone. Ready(St) and can_advance also couple stopping to the state object, so 'state' and 'budgeted multi-step seeking' are partially entangled. Without a full-benchmark, budget-matched no-state (or free-form-history) control, the separable contribution of the ledger remains only partially sealed.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes Omni-Decision, a training-free agent system for omni-modal evidence-seeking QA that centers control on an explicit query-scoped evidence state St = {Et, Ct, Ft, Ut} (confirmed evidence, conflicts, fact/computation dependencies, open needs). Planning, validation, repair, finalization, and insufficient stopping all read a shared state_digest; only a deterministic reducer commits normalized tool observations and critic verdicts. On OmniGAIA the system reaches 45.56% accuracy versus 18.33% for the official base agent (+27.23 pp) and 25.51% for OmniAgent; on WorldSense it reaches 58.32% versus 28.06% for OmniAgent. Controlled planner/perception swaps, a 120-case no-state ablation (Table 2), and human progress audits on 40 Medium/Hard cases are offered as evidence that the shared state interface contributes beyond free-form trajectories.","tokens_in":25072,"tokens_out":923,"duration_ms":11769,"significance":"If the causal claim holds, the work is a useful systems contribution: it reframes long-horizon omni-modal QA as query-scoped evidence closure rather than trajectory or role engineering, and it supplies an inspectable read/write contract that other agent stacks could adopt without new training. Strengths include a clear state formalism (§3.1–3.2, Algorithm 1), deterministic commit semantics, Wilson CIs on the full 360-split (Appendix C), multi-backend diagnostics (Table 2), transfer to a closed-world MCQ setting (WorldSense), and process-level audits that separate partial evidence-chain progress from zero-progress failures (§4.3, Appendices D–E). These are concrete engineering and evaluation assets for multimodal agent research, even if absolute SOTA against native end-to-end models is not claimed.","major_comments":[{"comment":"The central claim that explicit evidence-state control—not merely longer multi-step tool use—drives the OmniGAIA gains is only partially isolated. Appendix F reports 4,038 tool requests (avg. 11.22/query) for Omni-Decision versus 881 (avg. 2.45) for OmniGAIA base, with a max of 15 iterations and critic-gated repair. Table 2’s no-state ablation is restricted to a fixed 120-case subset and does not match tool-call budget or trajectory length. A budget-matched free-form-history (or no-state) control on the full 360-split is needed to seal the separable contribution of St; without it, ready(St)/can_advance and extended seeking remain entangled.","section":"§4.2 Table 2; Appendix F"},{"comment":"Comparisons to Minimal agent (5.56%) and OmniGAIA base (18.33%) understate the risk that gains partly reflect stronger orchestration under gpt-5.2 + gemini-3.1-pro rather than the state object alone. Table 2 shows large backend sensitivity (45.56% → 33.33% → 11.39% when perception/planner degrade), while the state effect shrinks to −3.33 pp under the weakest stack. The manuscript should either (i) report a full-benchmark no-state run with the default backends and matched call caps, or (ii) substantially soften claims that attribute the +27.3 pp primarily to evidence-state control.","section":"§4.1–4.2; Table 1–2"},{"comment":"Answer readiness ready(St)=(Ut=∅)∧complete(Ft)∧(Ct=∅) is treated as an adequate stopping criterion (§3.2), yet Appendix H documents frequent stuck/forced-answer exits when the critic correctly flags unclosed slots that perception cannot fill. The paper does not quantify how often forced answers under revision/budget limits inflate or deflate official accuracy relative to principled abstention. A breakdown of accepted vs forced vs insufficient outcomes on the full OmniGAIA split would make the control claim falsifiable and clarify whether state-conditioned stopping helps or merely delays unsupported answers.","section":"§3.2; Appendix H Table 11–12"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is that this is a clean systems argument with numbers: free-form multimodal agent histories are a bad control substrate, and a query-scoped evidence ledger (confirmed atoms, conflicts, fact/compute slots, open needs) plus a deterministic reducer is a better one. They get 45.6% on OmniGAIA vs 18.3% for the official base agent, transfer to WorldSense at 58.3%, and back that with planner/perception swaps, no-state subset drops, Wilson CIs, human progress audits, and stuck-rate localization.\n\nWhat is actually new is not “agents with tools” or “critics.” They say so. The contribution is the shared read/write contract: everyone reads the same state_digest; only the reducer commits typed events; readiness and insufficient stop are derived from Ut/Ft/Ct rather than left in dialogue. Appendices ship reducer contracts, need statuses, and full prompts. That is more engineering honesty than most multimodal agent papers.\n\nThe soft spot the stress-test flags is real and load-bearing for the causal claim. Omni-Decision averages ~11 tool calls vs ~2.5 for the base agent; max 15 steps with critic-gated repair. The no-state ablation is only on a 120-case subset and is not budget-matched. So “state” and “longer, more careful seeking under strong gpt-5.2 + gemini backends” are partly entangled. I would not oversell the +27 pp as pure ledger magic. I also would not dismiss the ablations: drops appear across three backend configs, audits show many wrong answers still make partial chain progress, and stuck cases often look like recognized unclosed slots the perception backend cannot fill. That is useful diagnosis even if isolation is incomplete. No public code is a practical minus for a systems claim.\n\nMath is light by design (ready/can_advance definitions, field-wise ⊕); citations are appropriate and not circular. This is for people building long-horizon multimodal agents who care about inspectable control and abstention, not for pure perception SOTA hunting.\n\nI would send it to peer review. Ask for a full-split, budget-matched no-state (or free-form history) control and clearer cost reporting; do not desk-reject. Worth reading and, for agent-control work, worth citing.","headline":"Useful systems paper with real gains and honest audits; the state-vs-budget confound is real but does not erase the contribution.","tokens_in":25718,"tokens_out":563,"would_cite":true,"duration_ms":7021,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A shared evidence state turns omni-modal QA into query-scoped evidence closure, lifting multi-step accuracy far above free-form agent baselines.","keywords":["omni-modal QA","evidence-seeking agents","query-scoped evidence state","evidence closure","multi-step multimodal reasoning","tool-using agents","OmniGAIA","WorldSense"],"falsifier":"Re-run Omni-Decision and the no-state baseline on the full OmniGAIA split with hard-matched tool budgets, step limits, and identical backends; if the large accuracy gap disappears under equal call counts, the central claim that explicit evidence-state control is the separable driver would not hold.","tokens_in":25538,"feed_emoji":"🧭","tokens_out":697,"duration_ms":6094,"temperature":0.7,"pith_summary":"Omni-modal evidence-seeking QA asks systems to answer questions whose clues are sparsely scattered across video, audio, images, web pages, and computation. Most agent systems leave those clues in free-form histories or tool logs, so they cannot cleanly track what is grounded, what still conflicts, and when the chain is ready to answer. This paper argues that the missing piece is not another role topology or perception model, but an explicit query-scoped evidence state that records confirmed evidence, unresolved conflicts, fact and computation dependencies, and open needs. Omni-Decision is a training-free system in which planning, acquisition, validation, repair, finalization, and insufficient stopping all read the same state digest, while only a deterministic reducer commits normalized observations. On open-world OmniGAIA it reaches 45.6% accuracy (+27.3 points over the official base agent) and transfers to WorldSense at 58.3%; no-state ablations and progress audits support a separable contribution from the shared state interface.","feed_headline":"Shared evidence state lifts omni-modal agent accuracy by 27 points","feed_subtitle":"Planning, repair, and stopping read one query-scoped ledger of needs, facts, and conflicts","key_machinery":"Query-scoped evidence state St = {Et, Ct, Ft, Ut}: confirmed evidence atoms, unresolved conflicts, external fact and computation slots, and open evidence needs. A bounded state digest conditions the planner, critic, and finalizer; only the reducer commits normalized tool observations and critic verdicts through deterministic field updates, with readiness defined as empty needs, complete dependencies, and no conflicts.","core_discovery":"Long-horizon omni-modal evidence seeking is best framed as query-scoped evidence closure rather than longer context or free-form trajectories. Maintaining a structured state of confirmed evidence, conflicts, fact/computation dependencies, and open needs, and conditioning every control decision on that shared view, yields large accuracy gains on multi-step omni-modal QA without training.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Query-scoped evidence state lifts omni-modal QA 27 points","Shared evidence ledger drives 27-point omni-modal agent gain","Evidence-closure state beats free-form trajectories by 27 points","Structured needs and conflicts raise OmniGAIA accuracy 27 points","Training-free evidence state control yields 27-point QA lift"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The accuracy gains are attributed mainly to the evidence-state interface itself, even though the full system also runs much longer tool trajectories under strong proprietary planner and perception backends, and the no-state ablation is only a fixed subset diagnostic rather than a full matched-budget comparison.","fun_headline_variants_meta":{"raw":{"variants":["Query-scoped evidence state lifts omni-modal QA 27 points","Shared evidence ledger drives 27-point omni-modal agent gain","Evidence-closure state beats free-form trajectories by 27 points","Structured needs and conflicts raise OmniGAIA accuracy 27 points","Training-free evidence state control yields 27-point QA lift"]},"model":"grok-4.5","effort":"low","cost_usd":0.005378,"raw_usage":{"total_tokens":1495,"prompt_tokens":802,"num_sources_used":0,"completion_tokens":94,"cost_in_usd_ticks":53780000,"prompt_tokens_details":{"text_tokens":802,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":599,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":802,"tokens_out":94,"duration_ms":4892,"temperature":1.0,"reasoning_tokens":599,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T05:37:34.262257+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run Omni-Decision and the no-state baseline on the full OmniGAIA split with hard-matched tool budgets, step limits, and identical backends; if the large accuracy gap disappears under equal call counts, the central claim that explicit evidence-state control is the separable driver would not hold.","supporting_citations":[],"review_version":1}