{"id":"436378f1-0fa4-4ad4-802a-afcf22806b28","arxiv_id":"2606.09669","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"SpatialWorld is a new multi-simulator benchmark showing top multimodal agents achieve under 18% success on interactive spatial tasks requiring active exploration and long-horizon planning.","lead":"The paper introduces SpatialWorld, a benchmark integrating eight simulators with 760 human-annotated tasks to test multimodal agents on interactive spatial reasoning under vision-only partial observability. Smart generalists should read it because even GPT-5 succeeds on only 17.4% of tasks, exposing a clear gap in AI spatial capabilities for real-world applications.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader correctly flagged the absence of methodological detail as the binding uncertainty. No additional load-bearing flaw is detectable from the abstract text; the central claim is an empirical report whose soundness cannot be further assessed without the full protocol.","tokens_in":1788,"tokens_out":244,"duration_ms":13840,"concrete_test":"Re-run the full evaluation pipeline on a 10% random subsample of tasks using the exact agent prompts and action interface described in the methods; if the TSR for GPT-5 shifts by >5 absolute points, the headline numbers are sensitive to undocumented implementation choices.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract presents a benchmark construction and empirical result (low TSR across 15 agents) whose validity hinges on the fidelity of the 760 tasks, reference trajectories, and terminal verifiers. Without access to the full methodological sections, no internal inconsistency, hidden assumption, or unsupported inference can be isolated from the provided text alone. The reported performance gap is a direct empirical observation under the stated protocol; any deeper critique would require the missing details on protocol unification, verifier construction, or domain sampling.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces SpatialWorld, a unified benchmark for interactive spatial reasoning in multimodal agents. It integrates eight heterogeneous simulation backends under a simulator-agnostic protocol, with 760 human-annotated tasks across domains like household routines and social collaboration. Each task includes a human-validated initial state, reference trajectory, and terminal-state verifier. Evaluation of 15 agents shows low performance, with GPT-5 achieving 17.4% average TSR and Qwen-3.5 at 14.1%, exposing mismatches between success and efficiency plus domain variations.","tokens_in":1873,"tokens_out":379,"duration_ms":23726,"significance":"If the tasks and verifiers hold, the work is significant for providing the first large-scale, cross-simulator test of active spatial understanding under partial observability. The low TSR results and identified bottlenecks in exploration/planning offer concrete evidence of current MLLM limitations, positioning the benchmark as a useful testbed. The shared protocol across backends is a clear strength for generalizability.","major_comments":[{"comment":"Benchmark construction (methods section describing task creation and verifiers): The criteria for selecting the 760 tasks, potential annotation biases, and exact implementation of the terminal-state verifiers are not detailed. This is load-bearing for the central TSR claims, as the reported performance gaps (e.g., 17.4% for GPT-5) cannot be interpreted without confirming that the tasks accurately and representatively measure interactive spatial understanding.","section":"Benchmark construction"}],"minor_comments":[{"comment":"The abstract mentions 'eight heterogeneous simulation backends' but does not list them or their domains explicitly; adding this would improve clarity without altering the claims.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the positive evaluation of SpatialWorld's significance and for highlighting the need for greater transparency in benchmark construction. We agree that additional detail on task selection, annotation processes, and verifier implementation is warranted to strengthen interpretability of the TSR results and will revise the methods section accordingly.","responses":[{"response":"We acknowledge that the current manuscript provides only high-level descriptions of task creation and verifiers. In the revised version we will insert a new subsection (Methods 3.2) that explicitly details: (1) the multi-stage selection criteria used to curate the 760 tasks across the eight simulators (diversity in domain, horizon length, and required spatial operations, with explicit balancing to avoid over-representation of any single simulator); (2) the annotation protocol, including the number of annotators per task, inter-annotator agreement metrics, and steps taken to reduce selection and confirmation biases (e.g., blind review of initial states and reference trajectories); and (3) the precise implementation of each terminal-state verifier, including the predicate logic, simulator-specific APIs invoked, and example verification traces for representative tasks. These additions will directly support the validity of the reported performance gaps.","revision_made":"yes","referee_comment":"[Benchmark construction] Benchmark construction (methods section describing task creation and verifiers): The criteria for selecting the 760 tasks, potential annotation biases, and exact implementation of the terminal-state verifiers are not detailed. This is load-bearing for the central TSR claims, as the reported performance gaps (e.g., 17.4% for GPT-5) cannot be interpreted without confirming that the tasks accurately and representatively measure interactive spatial understanding."}],"tokens_in":1369,"tokens_out":365,"duration_ms":10981,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that the paper builds SpatialWorld as a single protocol running across eight different simulators, with 760 human-annotated tasks that require agents to explore under partial vision and output text actions. The headline result is that even the best model they test hits only 17.4% task success rate on average.\n\nWhat is actually new is the attempt to make the evaluation simulator-agnostic while keeping vision-only inputs and a shared action format. That moves past the usual static VQA setups or single-simulator papers. They also supply reference trajectories and terminal-state verifiers for each task, which lets them run consistent scoring.\n\nThe evaluation of 15 agents and the follow-up notes on efficiency gaps plus domain differences are straightforward and useful to see. The low numbers line up with the claim that active spatial planning is still hard.\n\nThe soft spot is that the abstract gives little on how the tasks were sampled across domains or exactly how the verifiers were built and validated. Without those steps it is harder to judge whether the 17% figure reflects agent limits or benchmark quirks. The full paper may cover this, but it is not visible here.\n\nThis is for people working on embodied or multimodal agents who need a broader testbed than current options. The construction is concrete enough and the results are reported plainly, so it deserves a serious referee even if revisions will be needed on the protocol description.","headline":"SpatialWorld gives a practical multi-simulator benchmark for interactive spatial reasoning and shows current agents top out at low success rates, but the abstract skimps on task construction details.","tokens_in":2430,"tokens_out":368,"would_cite":false,"duration_ms":16838,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A new benchmark across eight simulators shows even top multimodal agents succeed on fewer than 18 percent of interactive spatial tasks.","keywords":["spatial reasoning","multimodal agents","interactive benchmark","simulation environments","task success rate","partial observability","long-horizon planning"],"falsifier":"An agent that achieves greater than 50 percent average task success rate on the full set of 760 tasks while following the same vision-only and text-action rules would show the reported performance ceiling is not fundamental.","tokens_in":2694,"feed_emoji":"🗺️","tokens_out":675,"duration_ms":17173,"temperature":0.7,"pith_summary":"The paper presents SpatialWorld as a unified evaluation platform for how well multimodal agents handle interactive spatial reasoning when they must actively explore partially observed environments and carry out real-world tasks. It supplies 760 human-annotated tasks, reference trajectories, and terminal-state verifiers that run on eight different simulation engines through one shared protocol and text-action interface. When fifteen current agents are tested, the best performer reaches only 17.4 percent average task success while the strongest open-source model reaches 14.1 percent. The results point to clear shortfalls in active exploration and long-horizon planning. The benchmark is offered as a shared testbed that future agents must clear before claims of robust spatial competence can be accepted.","feed_headline":"Top agents reach only 17% success on interactive spatial tasks","feed_subtitle":"SpatialWorld runs 760 tasks across eight simulators under one protocol and shows exploration and planning as the main limits.","key_machinery":"SpatialWorld benchmark, a simulator-agnostic collection of tasks and verifiers that forces agents to gather egocentric visual evidence and issue decisions through a single text-based action interface.","core_discovery":"SpatialWorld integrates eight heterogeneous simulation backends under a simulator-agnostic protocol and supplies 760 tasks with human-validated initial states, reference trajectories, and terminal verifiers; under vision-only partial observability and a unified text action space, fifteen advanced agents achieve at most 17.4 percent average task success rate, exposing persistent gaps in active exploration and long-horizon planning.","pith_inferences":["Designers of future agents may need to add explicit spatial memory or mapping modules rather than relying solely on larger models.","The benchmark could be extended by adding human performance baselines on the same tasks to quantify the remaining gap.","Because the action interface is text-only, improvements in language-to-action grounding could raise scores without changing the visual pipeline."],"forward_implications":["Task success rates and execution efficiency are often mismatched, so efficiency metrics must be tracked separately.","Performance varies sharply across domains such as household routines and social collaboration, indicating domain-specific weaknesses.","Active exploration under partial observability and long-horizon planning remain the dominant bottlenecks for current agents.","A shared protocol across simulators allows direct comparison of agents without simulator-specific tuning."],"fun_headline_variants":["Spatial agents hit 17% success across 760 benchmark tasks","Top model scores 17.4% in SpatialWorld interactive tasks","Eight simulators test agents with max 17% task success rate","Planning limits agents to 17% success in SpatialWorld benchmark"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 760 tasks, reference trajectories, and verifiers across the eight simulators accurately and representatively measure interactive spatial understanding needed for real-world tasks.","fun_headline_variants_meta":{"raw":{"variants":["Spatial agents hit 17% success across 760 benchmark tasks","Top model scores 17.4% in SpatialWorld interactive tasks","Eight simulators test agents with max 17% task success rate","Planning limits agents to 17% success in SpatialWorld benchmark"]},"model":"grok-4.3","cost_usd":0.006428,"raw_usage":{"total_tokens":2951,"prompt_tokens":706,"num_sources_used":0,"completion_tokens":70,"cost_in_usd_ticks":64278000,"prompt_tokens_details":{"text_tokens":706,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2175,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":706,"tokens_out":70,"duration_ms":9862,"temperature":1.0,"reasoning_tokens":2175,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T16:33:33.645334+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An agent that achieves greater than 50 percent average task success rate on the full set of 760 tasks while following the same vision-only and text-action rules would show the reported performance ceiling is not fundamental.","supporting_citations":[],"review_version":1}