{"id":"269bd257-e842-4a53-9352-8b9dd633f18e","arxiv_id":"2506.21876","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new 23-dimension benchmark finds that even the best VLMs score near random on motion trajectory, temporal extension, and several prediction tasks, far below humans, suggesting weak internal world models.","lead":"WM-ABench decomposes world modeling into 23 atomic perception and prediction skills and evaluates 15 vision-language models on over 100,000 simulated test instances. The results show that even frontier models fail at basic spatial, temporal, and motion reasoning, often near random, indicating they lack robust internal world models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central negative claim rests on synthetic-to-real transfer; §5.4's real-world recast is too underpowered to establish it.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: simulator-generated images and questions may not measure the same world-modeling ability that appears in real-world settings. My stress-test sharpens this by showing that the paper's own real-world validation (Table 4) is too underpowered to resolve it: ~50 data points per dimension, no confidence intervals, and no significance tests. This matters because the abstract and conclusion make general claims about VLMs' world-modeling abilities, not merely about performance on synthetic renderings. The authors themselves flag the OOD risk in the Limitations section, so this is not a manufactured concern. I agree with the reader's conditional verdict: the benchmark is valuable, with strong independent support from human baselines, multiple simulators, counterfactual hard negatives, and a real-world recast that partially addresses transfer. The concern does not warrant rejection; it warrants the conditional acceptance the reader already recommended, plus a concrete scaling-up of the real-world validation. I therefore leave the verdict unchanged and agree with the reader's weakest-assumption analysis.","tokens_in":31300,"tokens_out":6307,"duration_ms":74612,"concrete_test":"Expand the real-world recast of §5.4 to at least 300 instances per dimension across the five perception and three prediction dimensions, using real video with ground-truth state annotations (e.g., robot manipulation trajectories with tracked object positions, driving videos with vehicle paths) and the same forced-choice format with hard negative options. Evaluate the same model suite, compute per-model per-dimension accuracy with 95% confidence intervals, and run a paired comparison against synthetic scores. If real-world Motion Trajectory accuracy is significantly above chance (e.g., >60%) while synthetic accuracy remains near random, the out-of-distribution concern lands and the central claim needs qualification; if it remains near chance, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that VLMs lack robust internal world models depends on the assumption that WM-ABench's simulator-rendered tasks measure the same abilities used in real-world settings. The only direct evidence for transfer is §5.4 (Table 4), which recasts ~50 real-world data points per dimension for 8 of 23 dimensions. With n≈50, a reported accuracy of 30% (near the 25% random baseline) has a 95% confidence interval of roughly 17–43%, making it statistically indistinguishable from chance. Table 4 reports no confidence intervals, significance tests, or correlation measures, yet the text concludes that 'bias from our use of simulation data is likely to be minimal.' The authors' Limitations section explicitly concedes that the synthetic images 'do not look realistic' and may be out-of-distribution. Because the abstract and conclusion make general claims about VLMs' 'basic world modeling abilities' rather than about synthetic-scene performance, this underpowered validation is the weakest load-bearing link in the argument. If real-world versions of the near-random dimensions (especially Motion Trajectory and Temporal Extension) are solvable by VLMs, the headline negative conclusions would overstate the real-world gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-stage conceptual framework, decomposing world modeling into perception (visual, spatial, temporal, quantitative, motion) and prediction (mechanistic simulation, transitive inference, compositional inference), and introduces WM-ABench, a large simulator-generated benchmark with 23 fine-grained dimensions, controlled counterfactual options, and human baselines. The authors evaluate 15 VLMs in 660 experiments, reporting that models are strong at some static visual tasks but near-random on motion trajectory and several prediction tasks, and that their representations entangle orthogonal attributes such as color and speed. A small real-world recast of 8 dimensions is used to argue that simulation results transfer.","tokens_in":31450,"tokens_out":4067,"duration_ms":47870,"significance":"If the benchmark and transfer claim hold, this is a useful diagnostic contribution: the controlled counterfactual design, the use of simulator ground truth, and the inclusion of human baselines give the central negative findings about synthetic-scene performance a solid evidentiary basis. The paper is also commendably concrete about its limitations, including the realism gap of synthetic images. However, the headline conclusions in the abstract and conclusion are stated about VLMs' 'basic world modeling abilities' in general, and the only direct evidence that the simulator results transfer to real-world settings is a small, underpowered recast in Section 5.4. The benchmark itself appears to be a genuine resource, and the framework is clearly presented, but the generalizability claims need stronger support or explicit tempering.","major_comments":[{"comment":"The real-world validation is underpowered for the strong conclusion drawn from it. Table 4 reports approximately 50 data points per dimension and no confidence intervals; for example, Motion Trajectory accuracies range from 23.3% to 40.0% against a 25% random baseline, and with n=50 the 95% binomial confidence interval half-width is roughly 14 percentage points, making several of these values statistically indistinguishable from chance. Spatial Positioning values (26.7–53.3%) are similarly wide. The text then concludes that 'bias from our use of simulation data is likely to be minimal,' but no significance tests, confidence intervals, or explicit alignment criteria are provided. Since the Limitations section concedes that the synthetic images 'do not look realistic' and may be out-of-distribution, the abstract and conclusion should either restrict their claims to simulator-rendered tasks or be accompanied by a statistically adequate real-world validation.","section":"Section 5.4 / Table 4 / Limitations"},{"comment":"Frontier-model results are based on only 100 instances per subtask and are reported without confidence intervals or significance tests. For instance, OpenAI o3 achieves 48% on Motion Trajectory in Table 1, but with n=100 the standard error for a proportion near 0.5 is 5 percentage points, so the claim that this is 'near-random' is compatible with true accuracy well above 50%. Likewise, the 'human parity' claims for static perception must be qualified by the uncertainty in both the model and human estimates. The authors should report exact binomial confidence intervals or equivalent uncertainty quantification for all asterisked entries, and should temper statements such as 'Frontier Models Reach Human Parity in Static Perception' accordingly.","section":"Section 4.4 / Tables 1 and 2"},{"comment":"The perception-filtering analysis that underpins 'VLMs struggle with intuitive physics even under accurate state perception' is underspecified. Appendix D.2 states that the protocol excluded approximately 30% of instances, but Section 5.3 describes the filter as instances 'where all models correctly answer all relevant perception questions,' without clarifying whether the filter is model-specific or global, which tasks it applies to, or how the three listed queries map onto each prediction task. Table 3 reports only accuracy differences between filtered and unfiltered inputs, with no sample sizes, no per-condition accuracies, and no confidence intervals. As a result, the quantitative conclusion that 'limited perception capability is not the only cause' is not adequately supported for the three tasks shown, and the generalization to prediction tasks generally is even less supported. The authors should report the filtered and unfiltered accuracies separately, with uncertainties, and should state the filtering rule precisely.","section":"Section 5.3 / Table 3 / Appendix D.2"}],"minor_comments":[{"comment":"There is a typo in the first sentence: 'Depite' should be 'Despite'.","section":"Section 4.4"},{"comment":"The sentence 'We use both ManiSkill framework version 2 (Gu et al., 2023) and 3 (Tao et al., 2024))' has an unmatched closing parenthesis; remove the extra parenthesis.","section":"Appendix C.2"},{"comment":"The caption uses the abbreviation 'filtered (correct state perception)' without defining the filtering procedure or the sample size; please define it in the caption or refer to a precise appendix section.","section":"Table 3"},{"comment":"The compositional inference display contains mismatched parentheses in 'S(i)t ∼ Compose(S(1)t , . . . S(n)t )'; this should be cleaned up for readability.","section":"Appendix A.3"},{"comment":"The human evaluation uses 50 samples per task with three annotators, but the reported human accuracies in Tables 1 and 2 are point estimates; adding a simple binomial confidence interval or noting the sample size alongside each human row would help the reader judge the human baselines.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The benchmark and the simulator-based measurements are a solid contribution, and the negative findings are clearly valuable for the community. The main risk is that the paper's broad, abstract-level wording overstates the external validity of results that are, on the evidence provided, primarily about synthetic-scene performance. The real-world recast in Table 4 is a reasonable attempt but is too small to carry the weight placed on it. I see this as fixable within the scope of a revision: add uncertainty quantification, clarify the filtering procedure, and either strengthen the real-world validation or explicitly scope the conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is worth taking seriously. WM-ABench is a real contribution: it decomposes world modeling into 23 atomic dimensions, generates controlled counterfactual negatives from six simulators, runs 15 VLMs, and anchors everything with human baselines. The central negative finding—that current VLMs are near-random on motion trajectory and many physical prediction tasks under these controlled conditions—is well supported by the scale and by simulator ground truth. The entanglement analysis (e.g., color biasing speed judgments) is a nice diagnostic that goes beyond a simple accuracy table.\n\nThe soft spot is the real-world transfer section, §5.4. Roughly 50 examples per dimension from repurposed real datasets is too small to establish that \"bias from our use of simulation data is likely to be minimal.\" A reported 30% accuracy on a 4-choice task with n≈50 is within noise of chance, and Table 4 has no confidence intervals or significance tests. The Limitations section honestly concedes the synthetic images look out-of-distribution. So the headline claim should be scoped to simulated environments until a proper real-world validation exists. I want to stress: this does not sink the benchmark. The synthetic results are informative on their own, and Table 4's trends are consistent with them. It only means the generalization claim is currently overclaimed.\n\nOther smaller issues: frontier models were evaluated on 100 instances per task without error bars; the Section D.2 filter removes roughly 30% of prediction instances before the Table 3 analysis, and the paper should report both filtered and unfiltered numbers; and parsing failures are skipped, which could introduce selection bias if correlated with model confusion.\n\nWho should read it: anyone building or evaluating VLMs for embodied or physical reasoning, and benchmark designers. It deserves peer review—conditional accept. The required fixes are modest: release data and code, add confidence intervals, report the unfiltered perception-prediction analysis, and either expand the real-world recast or soften the generalization wording.","headline":"Useful diagnostic benchmark with solid synthetic-world evidence of VLM blind spots, but the real-world recast in §5.4 is too thin to carry the generalization claim.","tokens_in":32106,"tokens_out":2196,"would_cite":true,"duration_ms":22448,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Vision-language models, tested on a new 23-dimension atomic benchmark, score near chance on motion trajectory, spatial reconstruction, temporal extension, and compositional physics, and so lack robust internal world models.","keywords":["world models","vision-language models","benchmark evaluation","intuitive physics","spatial reasoning","temporal perception","motion trajectory","counterfactual generation"],"falsifier":"Render the same controlled physical events and counterfactual wrong answers with photorealistic images, or with real video that has exact ground truth, and rerun the trajectory, temporal-extension, and collision tasks. If models jump to human-level scores while the physical content is unchanged, the near-random results are a rendering artifact; if they stay near random, the missing-world-model claim holds.","tokens_in":31069,"feed_emoji":"🧠","tokens_out":12597,"duration_ms":127153,"temperature":0.7,"pith_summary":"This paper asks whether vision-language models (VLMs) carry a genuine internal world model, and answers with a systematic, atomic test. The authors propose a two-stage framework, perception then prediction, and build WM-ABench: 23 fine-grained dimensions, over 100,000 instances from six simulators, with counterfactual wrong answers so models cannot win by visual shortcuts. Across 660 experiments on 15 commercial and open-source VLMs, they report that almost all models are near random when judging motion trajectories, weak on spatial and temporal judgment, poor at predicting collisions and multi-step changes, and entangled in how they represent attributes, with some models treating blue objects as faster than green ones. If the benchmark is a fair measure, current VLMs lack robust, independent internal world models, and the paper supplies a checklist for where to improve them.","feed_headline":"Vision-language models are near-random at motion and physics tests","feed_subtitle":"660 experiments across 23 dimensions find current models cannot track object motion, collisions, or multi-step change.","key_machinery":"The load-bearing object is WM-ABench itself, a benchmark that decomposes world modeling into 23 atomic tasks organized by a two-stage framework: perception (visual, spatial, temporal, quantitative, motion) and prediction (mechanistic simulation, transitive inference, compositional inference). The key mechanism that makes the scores interpretable is counterfactual option generation: every question's wrong answers are produced either by perturbing the action while fixing the true previous state, or by perturbing the previous state while fixing the action, so the false states look visually similar to the true next state. That controls spurious cues and forces a model to simulate world dynamics rather than match surface features. Hard negative generation plus single-variable control is what lets the benchmark attribute failures to atomic dimensions, and a standardized relative entanglement score measures how much changing one attribute, such as color, shifts performance on an unrelated task.","core_discovery":"The paper's central claim is that state-of-the-art VLMs do not have robust internal world models in the mechanistic sense: they can name shapes and colors and count objects, but they largely fail when asked to reconstruct 3D relations from multiple views, compare durations or speeds across episodes, trace a moving object's trajectory, or predict the next state after physical events and agent actions. The strongest evidence is the near-random motion-trajectory results, and the entanglement analysis showing that performance shifts when irrelevant attributes like color change. The authors also show the deficit is not only perceptual: filtering to cases where models answered all perceptual verification questions correctly, next-state prediction for slide and drop improves only slightly, or decreases for collide, implying missing physical knowledge rather than mere perception errors. A real-world recast of eight dimensions replicates the main pattern, giving the authors reason to think the simulator-based findings generalize.","pith_inferences":["An implication the authors leave implicit is that WM-ABench's single-variable perturbation design could double as training data: fine-tuning on the hard negatives might teach models to stop using color as a proxy for speed, a failure mode the entanglement analysis exposes.","A reader wanting to locate the bottleneck could run the same questions with text-only state descriptions; if scores stay equally low, the missing ability is in language-grounded world knowledge, whereas if scores collapse further, vision grounding is the main culprit.","The relative entanglement metric could be ported to other VLM benchmarks: computing how much color, position, or material shifts reported accuracy on unrelated tasks would turn any multiple-choice benchmark into a check for shortcut reliance."],"forward_implications":["Current VLMs should not be trusted as general-purpose world simulators in robotics, driving, or other planning systems that depend on tracking trajectories and predicting collisions, since motion-trajectory accuracy near random is a direct warning.","Improving perception alone will not be enough: once the authors filter to cases where models had correct state perception, next-state prediction for slide and drop barely improves and collision prediction worsens, pointing to missing physical knowledge.","Static perception strengths are not a proxy for world modeling: state-of-the-art models are near-perfect on color, shape, and counting while failing spatial positioning, temporal extension, and compositional prediction.","Frontier models close some gaps, such as near-human size comparison and driving navigation, but not others such as multi-view 3D reconstruction, temporal extension, and multi-step manipulation, so the deficits are selective rather than a single general limitation."],"supporting_citations":[{"why":"The simulator used to render the multi-view spatial, temporal, motion, and compositional-physics images.","marker":"Gan et al., 2021"},{"why":"The physical-scene data set that supplies the collide, slide, and drop transition tasks.","marker":"Bear et al., 2022"},{"why":"The robotics simulator used for the manipulation and agent-action prediction tasks.","marker":"Tao et al., 2024"},{"why":"The earlier robotics simulator that provides manipulation demonstrations for lift, push, and drop.","marker":"Gu et al., 2023"},{"why":"The indoor navigation simulator that produces agent navigation state pairs.","marker":"Szot et al., 2021a"},{"why":"The urban driving simulator used for vehicle navigation and trajectory prediction.","marker":"Dosovitskiy et al., 2017a"},{"why":"The Bayesian account of perception that motivates the two-stage perception-then-prediction framing.","marker":"Knill and Pouget, 2004"},{"why":"The core-knowledge account from comparative psychology that justifies the selected perception dimensions.","marker":"Spelke, 2000"},{"why":"The mental-simulation account that motivates mechanistic prediction as a distinct ability.","marker":"Hegarty, 2004"},{"why":"The locality-of-experience account that motivates transitive, step-by-step future extrapolation.","marker":"Prystawski et al., 2023"}],"fun_headline_variants":["VLMs can't predict object motion—near-random on trajectory tests","Even with perfect perception, VLMs fail at physical simulation","World model benchmark: VLMs near-random on motion, biased by color","VLMs ace colors but can't simulate physics—motion tests near-random","Study: VLMs lack internal world models, fail to predict change"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusions assume that answering these simulator-rendered, counterfactual questions fairly measures the world-modeling ability that would matter with real-world images; the paper itself cautions that the synthetic images are not realistic, so if models fail mainly because the visuals are unfamiliar, the near-random scores would overstate the real-world gap.","fun_headline_variants_meta":{"raw":{"variants":["VLMs can't predict object motion—near-random on trajectory tests","Even with perfect perception, VLMs fail at physical simulation","World model benchmark: VLMs near-random on motion, biased by color","VLMs ace colors but can't simulate physics—motion tests near-random","Study: VLMs lack internal world models, fail to predict change"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00106,"raw_usage":{"total_tokens":4451,"prompt_tokens":955,"completion_tokens":3496,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":3404}},"tokens_in":571,"tokens_out":3496,"duration_ms":27258,"temperature":1.0,"reasoning_tokens":3404,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:16:23.456619+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Render the same controlled physical events and counterfactual wrong answers with photorealistic images, or with real video that has exact ground truth, and rerun the trajectory, temporal-extension, and collision tasks. If models jump to human-level scores while the physical content is unchanged, the near-random results are a rendering artifact; if they stay near random, the missing-world-model claim holds.","supporting_citations":[],"review_version":1}