{"id":"cf8b9935-f5af-4658-9317-976f2a36a782","arxiv_id":"2506.05523","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 500-video benchmark with six reasoning categories shows state-of-the-art vision-language models scoring below 25%, with particularly low performance on planning and abstract reasoning.","lead":"MORSE-500 is a new benchmark of 500 programmatically generated videos that test AI reasoning in six areas, from math to planning. On it, even the strongest vision-language models score below 25%, and the reported human baseline is 55.4%, suggesting large room for improvement.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Below-chance planning scores (0–5% vs. 12–33% random) suggest the reported planning deficit is an evaluation artifact; the central 'planning bottleneck' claim needs re-scoring before it can stand.","rationale":"The reader's conditional verdict and weakest assumption focus on the 2fps/32-frame proxy for the strongest models. My concern is closely related but sharper: the planning scores are not merely possibly depressed by frame sampling — they are below the multiple-choice chance floor, which is internally inconsistent with a correct evaluation pipeline. This is a correctness risk, not just a missing detail. The released code and intermediate evaluation logs make the concern directly checkable, which is a real strength of the paper. If re-scoring confirms an artifact, the benchmark remains useful but the headline conclusion about planning must be retracted or softened. If re-scoring shows the low scores are genuine, the paper needs to explain why models systematically underperform random guessing. Either way, the verdict remains conditional pending this check.","tokens_in":62079,"tokens_out":6396,"duration_ms":87403,"concrete_test":"Using the released evaluation logs, inspect the raw responses for all planning items from o3, Gemini 2.5 Pro, and Gemini 2.5 Flash. (1) Count how often the correct option letter or sequence appears anywhere in the raw output versus the extracted score; (2) compute a random-guess baseline using the actual number of options per item; (3) re-score the same outputs with a more tolerant LLM extractor. If the raw correct-present rate is at or above chance, or if extractor disagreement is large, the Table 2 planning numbers are invalid and the Section 5 planning claim must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 2 reports planning accuracy of 0–5% for every model, with human performance at 56.0%. Several planning tasks are explicitly multiple-choice (Appendix B.4 shows a 5-option Frozen Lake question; Table 3's maze and rope-knot tasks also use choices), so random guessing should yield roughly 12.5–33% or higher. A model that scores consistently below chance is not showing 'near-random performance' — it is showing evidence of systematic measurement error, whether from answer extraction, letter/option mapping, unreadable or truncated question text under the 512px/2fps/32-frame protocol, or some other pipeline bug. The paper provides no random baseline, no per-task breakdown for the 100 planning items, and no error analysis. Since Section 5's headline conclusion singles out planning as the 'particularly concerning' deficit and uses it to argue for 'fundamental architectural limitations,' these numbers are load-bearing. An artifact here would not necessarily invalidate the overall human–model gap, but it would invalidate the paper's strongest and most emphasized result.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MORSE-500, a 500-clip video benchmark spanning six reasoning categories (abstract, mathematical, physical, planning, spatial, temporal), with all videos generated or curated through deterministic Python scripts, generative video models, and real footage. The design goals are temporal-first, vision-centric evaluation (questions appear inside the video), a cognitive-science-grounded taxonomy, and programmatic difficulty scaling for future \"living\" benchmark versions. The authors evaluate 19 closed- and open-source models, all promoted with the minimal instruction \"Answer the question in this video,\" and report a large human-model performance gap (55.4% human vs. 23.6% best model), with particularly low scores in planning (0-5%) and abstract reasoning. The dataset, generation scripts, and evaluation harness are publicly released.","tokens_in":62290,"tokens_out":4420,"duration_ms":55669,"significance":"If the reported results are taken at face value, MORSE-500 is a useful diagnostic resource: it is one of the few video benchmarks that spans six reasoning types under controlled generation, and the release of deterministic scripts, ground-truth annotations, and an evaluation harness is a genuine contribution to reproducible benchmark development. The programmatic generation pipeline with explicit difficulty parameters is a strong design choice for a benchmark meant to evolve. The headline finding that frontier models lag human performance by roughly thirty points across categories, and especially on planning, would be an important stress-test signal for the multimodal reasoning community. However, the paper's most emphasized quantitative conclusions currently rest on evaluation choices that are either undocumented or not yet shown to be artifact-free, so the significance of the stated planning bottleneck in particular cannot be established from the manuscript as written.","major_comments":[{"comment":"The central human-vs-model comparison is confounded by input modality. Human performance (55.4%) is reported on native video, while the strongest API models (o3, o4-mini, Gemini 2.5 Pro, Gemini 2.5 Flash, etc.) are evaluated on 2 fps, 32-frame stills as stated in §3.1. The only frame-sampling ablation, Table 4, is computed on Gemini 2.5 Flash and shows overall scores between 18.4% and 19.2% across settings; it does not quantify the native-video loss for o3 or Gemini 2.5 Pro. Since §5's claim of \"fundamental architectural limitations rather than scaling issues\" depends directly on the size of the 55.4%-vs-23.6% gap, the authors need to either evaluate the strongest models on native video where the API supports it or run a carefully matched subset experiment to bound the frame-sampling cost. As written, part of the reported gap could be an evaluation-protocol artifact rather than a reasoning deficit.","section":"§3.1, §3.2, Table 2"},{"comment":"The paper's strongest and most emphasized result, the planning bottleneck, is undermined by the fact that every model scores 0-5% on the planning category, which is below chance for representative items. Appendix B.4 shows a Frozen Lake planning question with five options, implying a 20% random baseline; the 0-5% numbers across all models therefore suggest a systematic measurement artifact, such as letter/option mapping errors, answer-extraction failures, unreadable arrow symbols at the 512px/2fps/32-frame protocol, or a bug in the planning evaluation pipeline. The paper reports no random or majority baselines, no per-planning-subcategory breakdown (mazes vs. rope knots vs. robot manipulation), and no error analysis. Calling this performance \"near-random\" is inaccurate. The authors must add baselines, a per-task breakdown, and a re-scoring pass before the claim that planning reveals a fundamental architectural limitation can be accepted.","section":"§3.2, Table 2 (Planning column), Appendix B.4, §5"},{"comment":"The human baseline is insufficiently specified and the model results have no uncertainty estimates. Section 2.3.2 states that \"paper authors\" evaluated \"randomly sampled data slices\" but gives no number of evaluators, number of items per evaluator, whether evaluators watched native video, question-format constraints, or inter-annotator agreement. In Table 2, category sizes are as small as 64 items (abstract and physical), so the reported differences between closely ranked models (e.g., 17.4% vs. 16.8% overall) are within plausible sampling noise; no confidence intervals or significance tests are provided. For a benchmark whose purpose is diagnostic comparison, the human protocol and statistical uncertainty are part of the measurement itself and should be reported.","section":"§2.3.2, Table 2"}],"minor_comments":[{"comment":"The text refers to \"Table 6\" for qualitative examples, but the displayed artifact is labeled as a figure (Figure 6); the table/figure numbering should be made consistent.","section":"§3.3"},{"comment":"The phrase \"full video input (as a mp3 file format)\" should read \"video input (as an mp4 file)\"; audio is not part of this experiment.","section":"§3.4"},{"comment":"The protocol says \"we provided detailed instructions on the output formatting in the video\" while also stating \"No few-shot examples or format-specific guidance were provided\"; clarify that the on-screen format instructions are the only guidance and are not considered few-shot examples.","section":"§3.1"},{"comment":"The model name is sometimes written as Qwen2-235B-A22B and elsewhere as Qwen3-235B-A22B; please use the correct name consistently.","section":"Appendix C.2"},{"comment":"The phrase \"various Gemini 2.5 Pro and OpenAI o3\" is imprecise; list model names exactly (e.g., \"Gemini 2.5 Pro, OpenAI o3, and related variants\").","section":"Abstract"},{"comment":"The difficulty calibration (\"20% current-model capabilities, 50% moderate extensions, 30% stress tests\") is asserted without a measured operationalization; consider defining the calibration procedure or reporting model pass rates per difficulty band.","section":"§2.2"}],"recommendation":"major_revision","confidential_remarks":"The central planning result, which is the paper's most striking claim, appears potentially artifact-driven (sub-chance accuracy on multiple-choice items), and the strongest models are evaluated on frame stacks rather than video. Both issues are fixable within the manuscript's scope, so I do not recommend rejection, but the claims in their current form outrun the evidence. If the authors can add random baselines, a planning subcategory breakdown, a controlled native-video comparison for at least one frontier model, and proper uncertainty/CI reporting, the benchmark would be a solid contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: MORSE-500 is a worthwhile benchmark, well-built and honestly documented, but the paper's most striking number—planning accuracy at 0–5%—is probably an evaluation artifact, and the 'architectural limitation' conclusion built on it needs to be softened or fixed.\n\nWhat's genuinely new is the integration: six reasoning categories in video form, generated by deterministic scripts with explicit difficulty controls. That combination doesn't exist in the cited benchmarks. The release is real: code, data, harness. They also disclose the frame-sampling protocol for API-bound models and include a small ablation showing 2fps/32-frame is a reasonable operating point. That's more transparency than most benchmark papers.\n\nThe below-chance planning scores are the biggest problem. Many planning tasks are multiple-choice with 5 options (see Appendix B.4), so random is 20%; every model scores 0–5%. That's not 'near-random', that's systematic measurement error—likely answer extraction or option-mapping failure under the 512px/32-frame protocol. The paper provides no random baseline, no per-task breakdown, and no error analysis. Section 5 singles out planning as evidence of fundamental limits; that specific claim is unsupported.\n\nSecond, the human baseline is produced by the authors with no protocol details. Third, no confidence intervals anywhere, which matters for per-category comparisons. The frame-sampling concern for Gemini/o3 is real but partially mitigated by their ablation; I would not expect it to close the overall human-model gap, but it could move category-level numbers.\n\nAlso, the 'truly vision-centric' claim is a bit overstated: questions appear as text overlays that models can OCR. Not fatal, but worth noting.\n\nOverall, the benchmark deserves a serious referee. The right outcome is conditional acceptance after the planning numbers are re-scored and re-reported, CIs added, and human protocol documented. The overall human-vs-model gap is large enough that it will probably survive, but the paper's sharpest conclusion about planning does not.","headline":"A solid, well-released video benchmark with a real evaluation bug in its headline planning numbers.","tokens_in":62828,"tokens_out":2459,"would_cite":false,"duration_ms":30094,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims MORSE-500, a 500-clip video benchmark with questions embedded in footage, measures the strongest vision-language models at 23.6% against 55.4% for humans, with planning near random — evidence of architectural limits, not…","keywords":["video benchmark","multimodal reasoning","vision-language models","programmatic generation","temporal reasoning","planning","abstract reasoning","stress test"],"falsifier":"Run the strongest models on the 100 planning clips with native video input once the APIs support it (or with 8 fps and 128 frames as an intermediate check), and separately have human raters answer the same clips from only 2 fps, 32-frame stills. If model planning accuracy rises well above 5%, or human accuracy under the frame cap falls toward model levels, the claimed architectural limit is at least partly a stimulus-sampling artifact.","tokens_in":61921,"feed_emoji":"🎬","tokens_out":8860,"duration_ms":102919,"temperature":0.7,"pith_summary":"MORSE-500 is a video benchmark built to test whether vision-language models can reason over time: 500 fully scripted clips, each with a question rendered into the footage itself, spanning abstract, mathematical, physical, planning, spatial, and temporal reasoning. The paper's central claim is that the strongest available models fail on a large scale — OpenAI o3 tops out at 23.6% overall against 55.4% for human raters, and planning questions draw 0–5% accuracy across every model tested, which the authors take as evidence of fundamental architectural limits in temporal integration and multi-step reasoning rather than a scaling problem. The design that makes the measurement durable is programmatic generation: every clip comes from deterministic scripts whose complexity parameters can be increased without limit, so the benchmark can keep making harder instances as models improve. A reader should care because MORSE-500 turns the vague worry that models do not really reason about videos into a concrete, reproducible, and expandable measurement that is currently very low.","feed_headline":"AI models score 23.6% where humans score 55.4% on video reasoning","feed_subtitle":"New 500-clip scripted benchmark shows near-random planning results (0–5%) for every model tested.","key_machinery":"The load-bearing mechanism is the programmatic generation pipeline: deterministic Python scripts built on Manim, Matplotlib, and MoviePy, plus generative video models and curated real footage, that render each reasoning task as an animated clip and derive the ground-truth answer from the same parameter values that set the scene. Because entity count, reasoning depth, distractor density, temporal dynamics, and visual complexity are script parameters, difficulty is controllable and reproducible, and the question-in-the-video design forces models to extract both the question and the evidence from pixels rather than from a separate prompt. The evaluation harness borrows the answer-extraction and string-matching protocol of MathVista for scoring.","core_discovery":"On its own terms, the paper reports two discoveries. The first is the benchmark itself: 500 videos built so that answering requires genuine temporal understanding, with questions embedded as on-screen text inside the clip, categories grounded in a cognitive taxonomy, and roughly a third of instances deliberately beyond current model capabilities. The second is the measured state of the field: every evaluated model, from Gemini 2.5 Pro and OpenAI o3 down to 3B open-weight models, scores between 5.0% and 23.6% overall, with abstract reasoning below 25% and planning at 0–5% — near-random given the multiple-choice structure — while humans reach 55.4%. The paper interprets the uniformity of the collapse across scales and training paradigms as evidence that current multimodal systems are doing elaborate pattern matching rather than genuine multi-step reasoning over dynamic scenes, and it releases the corpus, scripts, and harness so the measurement can be repeated and extended.","pith_inferences":["Editorial inference: if the planning failure is a bag-of-frames problem, frame-shuffling a planning clip should barely change model accuracy while destroying human accuracy; this is a cheap test the paper does not run.","Editorial inference: the real-versus-generated physics task converts video-generator quality into benchmark difficulty, so it will get harder automatically as generators improve — a free difficulty ladder the paper does not highlight.","Editorial inference: humans were not tested under the same 2 fps, 32-frame constraint applied to the strongest models; a matched human ablation would separate frames lost from reasoning lost in the headline 31.8-point gap."],"forward_implications":["If the measured gap holds, no current vision-language model is close to human-level multimodal reasoning on dynamic tasks, so claims about deploying such models in autonomous, video-driven decision systems need to be tempered.","The benchmark's generation scripts let difficulty be scaled along entity count, reasoning depth, distractor density, and temporal dynamics, so it can keep producing harder instances instead of saturating.","The paper's ablation shows model accuracy drops as the same content is given as a single image, then multiple images, then video — so distributing information over time is itself a current failure mode to target.","The planning category (0–5% across all models) is the clearest diagnostic: multi-step, goal-directed reasoning over video is essentially unachieved and should be the first thing tested in new models."],"supporting_citations":[{"why":"The top closed-source baseline; Gemini 2.5 Pro's 21.8% overall score anchors the human-versus-model comparison, and its frame-sampled input protocol motivates the paper's weakest assumption.","marker":"[Google, 2025]"},{"why":"The strongest evaluated model family (o3 at 23.6% overall) sets the upper bound of model performance on the benchmark.","marker":"[OpenAI, 2025]"},{"why":"MathVista supplies the LLM-based answer extraction and string-matching evaluation protocol, plus inspiration for the mathematical reasoning tasks.","marker":"[Lu et al., 2024]"},{"why":"ARC-AGI-2 pattern-induction puzzles are adapted into the animated abstract reasoning questions, the category where the paper reports the largest deficits.","marker":"[Chollet et al., 2025]"},{"why":"The Physics IQ benchmark provides the real footage used in the real-versus-generated physical reasoning discrimination task.","marker":"[Motamed et al., 2025]"},{"why":"OpenAI Gym's FrozenLake environment is adapted into the fogged maze planning tasks that yield the near-chance planning scores.","marker":"[Brockman et al., 2016]"},{"why":"MimicPlay robot manipulation clips are used for the action-sequence ordering questions in the planning category.","marker":"[Wang et al., 2023]"}],"fun_headline_variants":["AI video reasoning scores 23.6% vs human 55.4%","New stress-test: AI planning near random on 500 clips","MORSE-500: scripted videos expose AI reasoning gaps","Abstract reasoning below 25% for all AI models tested","Benchmark evolves: 500 clips, AI scores under 25% overall"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The strongest models (Gemini 2.5 Pro and the OpenAI o-series) received the videos as 2 frames-per-second stills capped at 32 frames because their APIs do not yet accept video input, so the headline human-versus-model gap could shrink if native video input helps them.","fun_headline_variants_meta":{"raw":{"variants":["AI video reasoning scores 23.6% vs human 55.4%","New stress-test: AI planning near random on 500 clips","MORSE-500: scripted videos expose AI reasoning gaps","Abstract reasoning below 25% for all AI models tested","Benchmark evolves: 500 clips, AI scores under 25% overall"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000846,"raw_usage":{"total_tokens":3729,"prompt_tokens":1040,"completion_tokens":2689,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":656,"completion_tokens_details":{"reasoning_tokens":2596}},"tokens_in":656,"tokens_out":2689,"duration_ms":21508,"temperature":1.0,"reasoning_tokens":2596,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:20:16.194049+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the strongest models on the 100 planning clips with native video input once the APIs support it (or with 8 fps and 128 frames as an intermediate check), and separately have human raters answer the same clips from only 2 fps, 32-frame stills. If model planning accuracy rises well above 5%, or human accuracy under the frame cap falls toward model levels, the claimed architectural limit is at least partly a stimulus-sampling artifact.","supporting_citations":[],"review_version":1}