{"id":"bc324b59-ef3e-4d4d-81da-a70a3736670a","arxiv_id":"2506.04633","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"STARE is a 4K-task benchmark showing multimodal LLMs perform near random chance on multi-step spatial simulation tasks such as cube net folding and tangrams, despite strong 2D transformation results.","lead":"STARE is a new benchmark of about 4,000 visual-spatial puzzles that asks whether multimodal AI can solve problems by mentally simulating transformations. Across ten models, accuracy collapses on multi-step tasks like cube net folding and tangrams, while humans stay near perfect, showing current AI has weak spatial imagination.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Near-chance cube-net scores may reflect a 3D perception bottleneck, not a simulation deficit; the paper's own folded-face probe is at 57.4%.","rationale":"The reader's weakest assumption was ground-truth label correctness. While that is a legitimate validity concern, it is unlikely to be the most load-bearing issue: two human participants achieved near-perfect accuracy (99% on cube nets, 87.5-94% on tangrams), which would be improbable if the programmatic generators systematically mislabeled a large fraction of items. The paper's own data, however, reveal a more direct threat to the central claim. The perception probes show that GPT-4o's 3D perception of folded states is at chance (57.4%), and providing the final folded form raises cube-net accuracy to 100%. This means the failure to benefit from intermediate visual simulations could be explained by an inability to perceive the rendered 3D states, not by an inability to perform mental simulation. The paper acknowledges this in Q3 but does not incorporate it into the central conclusion, which remains 'models cannot effectively perform visual simulation.' This conflation is material: if perception is the bottleneck, the benchmark demonstrates a limitation in 3D visual perception, not in visual simulation per se. The proposed conditional analysis would settle the question. I still regard the benchmark as useful and the overall verdict of CONDITIONAL as appropriate, so the reader's verdict stands; the concern is an additional condition rather than a reason to reject.","tokens_in":39613,"tokens_out":8512,"duration_ms":102803,"concrete_test":"For the cube-net folding task with visual simulations, use the perception probes to identify items where the model correctly perceives every intermediate state (e.g., correctly answers 'has face X been folded?' for all faces in the intermediate images). Restrict the evaluation to this correctly-perceived subset and recompute F1 with and without visual simulation. If F1 improves substantially on the subset, the benefit was present but masked by perception errors; if it remains near chance, the simulation deficit is confirmed. This directly tests whether the near-chance scores and inconsistent gains are attributable to perception rather than simulation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that multimodal models cannot effectively perform visual simulation, with near-chance cube-net and tangram scores and inconsistent gains from intermediate visual simulations as evidence. But the paper's own perception probes (Section 3.3, Table 2) show GPT-4o's accuracy on the 3D perception probe ('has face 6 been folded?') is 57.4%, essentially chance, while 2D color and connectivity are near ceiling. When the final folded form is provided, cube-net accuracy jumps to 100% (Q3). This strongly suggests that the bottleneck for benefiting from visual simulations may be the inability to perceive the 3D state depicted in intermediate renderings, not a lack of internal simulation. The paper acknowledges this ('these specific perceptual errors in folding explain the limited benefits from visual simulations observed in Table 1') yet the abstract and conclusion still frame the result as 'models cannot effectively perform visual simulation.' The claim conflates visual perception with visual simulation: a model that cannot tell whether a face is folded cannot be said to have failed to simulate it. This concern is internal to the paper's own data and is more load-bearing than label correctness, since near-perfect human performance already provides strong evidence the labels are mostly correct, whereas no analysis currently separates perception from simulation.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces STARE, a benchmark of approximately 4,000 programmatically generated tasks spanning 2D and 3D geometric transformations, cube net folding, tangram puzzles, and real-world perspective and temporal reasoning. Ten multimodal large language models are evaluated in settings with and without intermediate visual simulations, alongside two human participants. The main findings are that models perform near chance on cube net folding and tangram puzzles, that gains from intermediate visual simulations are inconsistent across models and tasks, and that humans achieve near-perfect accuracy but take longer without visual guidance. The paper attributes these gaps to a lack of multi-step visual simulation ability in current MLLMs.","tokens_in":39801,"tokens_out":3714,"duration_ms":44084,"significance":"If the central claim holds, STARE would be a useful diagnostic for spatial reasoning in multimodal models, complementing existing verbal-reasoning benchmarks. The programmatic generation makes the benchmark reproducible and extensible, and the error analysis, including the 3D folded-face perception probe, is a valuable attempt to decompose model failures. However, the central claim as stated conflates perception with simulation: the paper's own probe shows GPT-4o at 57.4% on a 3D perception task, which undermines the interpretation that near-chance cube net scores reflect a pure simulation deficit. The human baseline is also too thin to support the strong comparative claims. With these issues addressed, the benchmark could be a solid contribution to evaluating spatial cognition in MLLMs.","major_comments":[{"comment":"The paper's own perception probe shows GPT-4o at 57.4% on 'has face 6 been folded?', essentially chance, while 2D color and connectivity are near ceiling; providing the final folded form raises cube net accuracy to 100%. The text acknowledges that these perceptual errors explain the limited benefits from visual simulations, yet the Abstract and Conclusion state that 'models cannot effectively perform visual simulation.' This conflates a 3D perception deficit with a simulation deficit: a model that cannot perceive whether a face is folded has failed at perception, not necessarily at simulation. The central claim needs to be reframed, or supported by a condition that isolates simulation from perception (e.g., symbolic state descriptions or non-visual intermediate representations), before the paper can claim evidence about visual simulation ability.","section":"§3.3, Table 2, Q3; Abstract; §4"},{"comment":"The human baseline consists of two undergraduates, and no error bars, repeated runs, or per-subject variability are reported for either humans or models. The strong claims about human near-perfect accuracy and response-time reductions (e.g., 28.9s down to 17.1s on tangram puzzles) rest on an essentially anecdotal sample. Report at least per-subject scores and confidence intervals, or increase the number of participants, to make the human-model comparison statistically meaningful.","section":"§3.1; Table 1"},{"comment":"The paper explicitly states that the tangram question-only set has a 'bias': models can achieve about 75% accuracy by comparing total piece areas. This means the 'without visual simulation' tangram condition in Table 1 is not a pure spatial-simulation test, and the near-chance F1 scores in that condition are not clean evidence about simulation ability. The dataset should be redesigned so that area-sum is non-diagnostic (e.g., equal-area solvable and unsolvable instances), or the analysis should exclude or condition on those instances.","section":"§2.2/E.4; §3.3 Q5, Table 4"},{"comment":"The validity of the entire benchmark depends on programmatic generators: the cube net algorithm's overlap and disconnection checks and the tangram invalid-case construction must never mislabel a puzzle. The paper describes these checks but provides no formal verification, no independent audit, and no human-validation statistics on the generated labels. Given that near-chance scores could also arise from systematic mislabeling, the authors should report a human audit of a random sample (e.g., 100 instances per task) or otherwise verify label correctness.","section":"§E.3, E.4; §2.2"}],"minor_comments":[{"comment":"The header row splits 'Temp-oral' and 'Pers-pective' awkwardly; ensure the table is readable in the final version.","section":"Table 1"},{"comment":"The claim of r≈0.88 across 11 models with p≈5e-4 is fragile; with only 11 data points and a drop to r≈0.58 when open-source models are removed, the wording 'strong correlation' overstates the evidence. Hedge the claim or use rank correlation.","section":"§3.2, Appendix G"},{"comment":"The reference list contains malformed entries (e.g., '[11] et al. Johnson, Justin' and several others with misplaced author names), which should be cleaned before publication.","section":"References"},{"comment":"The paper uses accuracy for multiple-choice tasks and F1 for binary tasks, then macro-averages across tasks; since these metrics have different ranges and chance levels, the 'overall' column is hard to interpret. State this limitation explicitly.","section":"§3.1, Evaluation Metrics"},{"comment":"The text describes the benchmark as containing ~4K instances, while Table 6 sums to 3,937; align the wording to avoid a minor inconsistency.","section":"§2.1, Table 6"}],"recommendation":"major_revision","confidential_remarks":"The perception/simulation confound is the main substantive issue. I would encourage the editor to ask for a revision that either reframes the central claim to 'models lack robust 3D spatial perception and consequently fail at visual simulation' or adds a condition that removes the perception confound. The human baseline also needs strengthening. The benchmark itself is a useful contribution if these points are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: STARE is worth having. The programmatic generation of intermediate visual states across 2D/3D transforms, cube nets, and tangrams is a real addition—existing benchmarks mostly skip explicit step-by-step simulation because it is expensive to annotate. The results are consistent across ten models: multi-step spatial tasks like cube-net folding and tangram are near chance, while simpler 2D transforms are not. The human data, though thin, gives a meaningful calibration point. If the GitHub repo ships what the paper describes, this is a reusable testbed.\n\nThe soft spots are real but not fatal. The biggest is interpretive. The paper's own folded-face probe is 57.4% for GPT-4o, and when the final folded form is presented, cube-net accuracy jumps to 100%. That is strong evidence that a large part of the cube-net failure is 3D perception, not mental simulation. The paper acknowledges this in Q3, yet the abstract and conclusion still say models 'cannot effectively perform visual simulation.' That conflation matters: a model that cannot tell whether a face is folded has not been shown to lack simulation ability. The fix is straightforward—reframe the claim as 'models fail on tasks that require visual simulation, in part because they cannot perceive the 3D states the simulation renders'—or add a control that gives the model the correct perceptual state and measures simulation proper. I do not think the stress-test concern sinks the paper, but it should change the headline.\n\nOther items are minor. No error bars or repeated runs; open-source models are run with temperature 0.7, so single-run variance could be nontrivial. The human baseline is two undergraduates. The tangram question-only set has an acknowledged area-sum shortcut, which muddies the Q+Steps comparison. The synthetic-to-real correlation drops from 0.88 to 0.58 when open-source models are excluded, so that claim should be softened. The label-correctness worry from the reader is the least of my concerns: near-perfect human performance on the same tasks is independent evidence that the generators are mostly right.\n\nVerdict: send it to review. The benchmark is a solid contribution, the analysis is mostly careful, and the perception/simulation distinction is a fixable framing problem rather than a load-bearing flaw.","headline":"A genuinely useful benchmark whose headline claim runs ahead of its own perception probes.","tokens_in":40379,"tokens_out":2700,"would_cite":true,"duration_ms":32861,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that current multimodal large language models cannot effectively perform multi-step visual simulation: they score close to random chance on cube net folding and tangram puzzles, and intermediate visual steps help only…","keywords":["spatial cognition","visual simulation","multimodal large language models","benchmark","cube net folding","tangram puzzles","mental imagery","visual reasoning"],"falsifier":"Run an independent audit of the generated labels: brute-force fold every cube net (or enumerate all 11 cube nets) and run an exact-cover solver on every tangram puzzle. If the mislabel rate is above a couple of percent, the near-chance model scores could reflect broken ground truth rather than missing simulation; if it is zero, the claim stands on the generated data.","tokens_in":39399,"feed_emoji":"🧩","tokens_out":6022,"duration_ms":67380,"temperature":0.7,"pith_summary":"STARE is a new benchmark of about 4,000 spatial tasks built to test whether multimodal AI models can do what humans do naturally: run a step-by-step mental picture of a transformation rather than just talking through it. The paper finds that models handle simple 2D transformations well but score close to random chance on cube net folding and tangram puzzles, which require several mental moves in 3D. Showing models intermediate visual steps helps on some tasks and hurts on others, and the best model tested still lands below 60% overall, while humans reach the mid-90s. The authors conclude that current models cannot effectively perform visual simulation, and they use the benchmark to locate the failure in 3D perception and multi-step integration rather than in basic color or flat-shape recognition.","feed_headline":"AI models score near chance on multi-step spatial puzzles","feed_subtitle":"A 4,000-task benchmark finds near-random results on cube net folding and tangrams, where humans reach 99% accuracy.","key_machinery":"The carrying mechanism is STARE's programmatic generator, which synthesizes each task with an explicit ground-truth simulation: a folding algorithm that rotates cube-net faces 90 degrees about shared edges and checks for overlaps and disconnections, a recursive segmentation-and-scramble routine for tangram puzzles, and Matplotlib and Blender renderings of every intermediate state. That generator lets the authors produce matched without-visual-simulation and with-visual-simulation versions of the same item, turning the benchmark into a controlled intervention where the only difference between the two conditions is the intermediate imagery. The explicit step structure of each task is what makes the claim about visual simulation testable.","core_discovery":"On its own terms, the paper establishes that state-of-the-art multimodal large language models do not perform multi-step visual simulation the way humans do. Across the benchmark's 2D and 3D transformations, cube net folding, tangram puzzles, temporal frame reasoning, and perspective reasoning, models stay within a few points of random chance on the multi-step spatial tasks, even when they are given explicit intermediate visualizations. The one systematic exception is straightforward 2D transformation, where accuracy reaches the 80-90% range. Human testers scored 87.5-99% on the same items, and their response times dropped by 7.5 seconds on average when intermediate steps were shown, evidence that the tasks genuinely are solved by running mental simulations. Because models improve inconsistently with those same visual aids, and because a probing test shows they fail specifically at judging whether a face has been folded into depth, the paper concludes the bottleneck is the capacity to simulate spatial change step by step, not the ability to see or describe the shapes.","pith_inferences":["A direct test of whether the deficit is architectural: fine-tune a model on pairs of initial and intermediate visual states, then measure cube-net accuracy held out; if it stays near chance, missing simulation is a hard architectural limit rather than an experience gap.","Because models can reach about 75% on tangram question-only items by comparing piece areas, the near-chance numbers on the full tangram set may understate the failure: on items where the area heuristic cannot work, models likely do even worse.","The human response-time data suggest a new evaluation signal: compare a model's time-to-answer with and without intermediate visuals; a model that answers about as fast without simulation is plausibly pattern-matching rather than simulating.","The same generator-controlled design of matched with and without intermediate states extends naturally to deformable bodies, articulated mechanisms, and physical-prediction tasks, where the intermediate states are equally well-defined."],"forward_implications":["Success on 2D transformation tasks should not be read as general spatial competence: the same models drop to near-chance on multi-step 3D tasks, so benchmarks that stop at 2D overstate ability.","Intermediate visual simulation is not a reliable assist: models that improve on some tasks and decline on others (GPT-4o and o1 on tangrams, Claude and Gemini Flash on cube nets) have not internalized the visual steps.","The strong correlation (r≈0.88) between synthetic-task performance and real-world task performance implies that gains on these abstract simulation tasks should transfer to practical settings such as navigation and assembly.","Because explicit verbal reasoning steps do not help cube net folding and actively hurt tangram performance, chain-of-thought prompting cannot substitute for visual simulation.","A specific deficit in 3D perception, with GPT-4o identifying a folded face at only 57.4% accuracy while being perfect on color, locates the bottleneck in depth-aware perception rather than in 2D vision."],"supporting_citations":[{"why":"Establishes the mental-rotation paradigm that motivates the multi-step simulation account; supplies the cognitive premise that humans solve spatial tasks by analog simulation.","marker":"[6]"},{"why":"Introduces mental animation, the step-by-step simulation of mechanical systems that STARE's intermediate-visualization design operationalizes.","marker":"[7]"},{"why":"A parallel spatial cognition evaluation cited as the most relevant prior benchmark, which STARE extends by adding explicit intermediate visual simulations.","marker":"[16]"},{"why":"The perceptual benchmark showing that models can see but not perceive; STARE contrasts its own multi-step simulation tasks against it.","marker":"[29]"},{"why":"The closest prior mental-imagery benchmark, which STARE distinguishes by focusing on explicit step-by-step simulation rather than spatial memory from video.","marker":"[33]"},{"why":"The tool used to programmatically generate the 2D transformation, cube-net, and tangram stimuli and their ground-truth states.","marker":"[34]"},{"why":"The 3D mesh-construction approach STARE follows for generating abstract 3D shapes.","marker":"[35]"},{"why":"The source of the temporally consistent video frames used to build the temporal frame reasoning task.","marker":"[37]"},{"why":"The real-scene dataset that supplies the indoor environments for the perspective reasoning task.","marker":"[38]"}],"fun_headline_variants":["AI flunks 3D spatial puzzles that humans ace","Multimodal models near chance on tangrams and cube folding","Visual aids can't fix AI spatial reasoning failures","Benchmark: AI lags humans on multi-step spatial sims","Models random on complex spatial tasks humans solve"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that STARE's ground-truth labels are correct: the cube-net folding algorithm's overlap and disconnection checks and the tangram invalid-case construction must never mislabel a puzzle, since near-chance scores only mean 'no visual simulation' if the labels are right.","fun_headline_variants_meta":{"raw":{"variants":["AI flunks 3D spatial puzzles that humans ace","Multimodal models near chance on tangrams and cube folding","Visual aids can't fix AI spatial reasoning failures","Benchmark: AI lags humans on multi-step spatial sims","Models random on complex spatial tasks humans solve"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1472,"prompt_tokens":1016,"completion_tokens":456,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":632,"completion_tokens_details":{"reasoning_tokens":377}},"tokens_in":632,"tokens_out":456,"duration_ms":4977,"temperature":1.0,"reasoning_tokens":377,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:36:23.658461+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run an independent audit of the generated labels: brute-force fold every cube net (or enumerate all 11 cube nets) and run an exact-cover solver on every tangram puzzle. If the mislabel rate is above a couple of percent, the near-chance model scores could reflect broken ground truth rather than missing simulation; if it is zero, the claim stands on the generated data.","supporting_citations":[{"cited_title":"Mental rotation of three-dimensional objects.Science, 171 (3972):701–703, 1971","cited_arxiv_id":null,"evidence_quote":"Establishes the mental-rotation paradigm that motivates the multi-step simulation account; supplies the cognitive premise that humans solve spatial tasks by analog simulation."},{"cited_title":"Mental animation: Inferring motion from static displays of mechanical systems.Journal of Experimental Psychology: Learning, Memory, and Cognition, 18(5):1084–1102, 1992","cited_arxiv_id":null,"evidence_quote":"Introduces mental animation, the step-by-step simulation of mechanical systems that STARE's intermediate-visualization design operationalizes."},{"cited_title":"Matplotlib: Visualization with python.https://matplotlib.org/, 2012","cited_arxiv_id":null,"evidence_quote":"The tool used to programmatically generate the 2D transformation, cube-net, and tangram stimuli and their ground-truth states."},{"cited_title":"Lawrence Zitnick, and Ross Girshick","cited_arxiv_id":null,"evidence_quote":"The 3D mesh-construction approach STARE follows for generating abstract 3D shapes."},{"cited_title":"Objectron: A large scale dataset of object-centric videos in the wild with pose annotations.Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021","cited_arxiv_id":null,"evidence_quote":"The source of the temporally consistent video frames used to build the temporal frame reasoning task."}],"review_version":1}