{"id":"c2dc472a-781a-4668-95f3-7080f369df69","arxiv_id":"2508.02095","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"VLM4D benchmarks spatiotemporal reasoning in VLMs and finds large gaps versus humans, with proposed methods showing partial improvement.","lead":"This paper introduces VLM4D, a benchmark with real and synthetic videos plus question-answer pairs to test whether vision language models can reason about object motion, rotation, and perspective. It reports that state-of-the-art VLMs fall well short of human performance, and that reconstruction and fine-tuning techniques narrow the gap.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The benchmark-validity assumption is the key load-bearing point: without static-frame and language-only ablations, the reported VLM-human gaps do not establish fundamental spatiotemporal deficiencies.","rationale":"The reader's UNVERDICTED verdict stems from lack of full text. My stress-test identifies the same weakest assumption but makes it more operational: the benchmark's temporal necessity is unverified. This is not an accusation of data fabrication; it is an ordinary construct-validity concern. Since the abstract alone cannot resolve it, the verdict should remain UNVERDICTED pending release of ablations and data. If the controls pass, the claim would be substantially strengthened. If they fail, the claim of fundamental deficiency would need to be downgraded. The recommended verdict is therefore unchanged, and the concrete test provides a path to confirmation.","tokens_in":627,"tokens_out":2401,"duration_ms":29493,"concrete_test":"Run three control conditions on VLM4D with the same models: (1) random static frame instead of the full video, (2) shuffled frame order, and (3) question-only text with no visual input. Compare accuracy against the full-video condition. If (1) or (3) yield accuracy close to full-video accuracy, then the benchmark does not isolate spatiotemporal reasoning. Additionally, report the exact human evaluation protocol (number of replays allowed, time per question) and inter-annotator agreement per question category.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that state-of-the-art VLMs exhibit fundamental deficiencies in spatiotemporal reasoning—depends on VLM4D's question-answer pairs actually requiring dynamic, multi-frame understanding. The abstract reports significant performance gaps but provides no internal evidence that correct answers cannot be derived from a single frame, from language priors, or from annotation patterns. For example, a question about 'rotation' might be answerable from a static image of an object's orientation, and a question about 'motion continuity' might be answerable by common-sense trajectories of typical scenes. Without control conditions that ablate temporal information (random static frames, frame order shuffled, question-only text), the observed gaps are uninterpretable as a measure of temporal reasoning. The human baseline also needs scrutiny: if humans could replay videos or had additional time, the comparison is confounded by viewing protocol rather than model capability. Thus the benchmark's construct validity is the weakest link in the argument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces VLM4D, a benchmark designed to evaluate spatiotemporal reasoning in vision-language models (VLMs) using real-world and synthetic videos with curated question-answer pairs. The authors report that state-of-the-art open and closed-source VLMs perform significantly worse than humans, especially on tasks requiring integration of multiple visual cues and temporal coherence. They also propose two directions to improve spatiotemporal comprehension: leveraging 4D feature field reconstruction and targeted spatiotemporal supervised fine-tuning. The claims are based on the abstract only, as the full text is not available in the submitted manuscript.","tokens_in":922,"tokens_out":3821,"duration_ms":45538,"significance":"If the benchmark is valid and the reported gaps are reproducible, VLM4D would fill a genuine gap in dynamic scene understanding evaluation and could catalyze research on temporal grounding in VLMs. The proposed improvement directions are plausible and of interest to the community. However, the significance of the work rests entirely on the construct validity of the benchmark and the rigor of the evaluation, neither of which can be assessed from the abstract alone. The paper's central claim of 'fundamental deficiencies' is potentially impactful but is currently unsupported by the visible evidence.","major_comments":[{"comment":"The abstract reports 'significant performance gaps' between VLMs and humans, but it provides no evidence that the question-answer pairs actually require dynamic, multi-frame understanding. Without ablations such as static-frame input, shuffled frame order, or question-only inference, the observed gaps could arise from single-frame visual cues or language priors rather than deficiencies in spatiotemporal reasoning. This is the central load-bearing point of the paper, so it must be addressed with concrete control conditions or ablation studies.","section":"Abstract"},{"comment":"The human baseline is described only as 'human baselines,' with no specification of the viewing protocol, such as whether participants were allowed to replay videos, the time limits imposed, the number of attempts, or whether answers were open-ended or multiple-choice. A confounded protocol could explain or exaggerate the reported gaps, so the methods for collecting human performance must be fully described and justified.","section":"Abstract"},{"comment":"The claimed effectiveness of '4D feature field reconstruction and targeted spatiotemporal supervised fine-tuning' is stated without any experimental details, including dataset splits, hyperparameters, or evaluation protocols. There is also a risk of circularity if the supervised fine-tuning data overlaps with the VLM4D test set; the paper must explicitly state how such overlap is avoided or measured.","section":"Abstract"},{"comment":"The manuscript made available to the referee contains only the abstract; the main text with the benchmark construction, dataset statistics, model list, evaluation protocol, and full result tables is absent. Consequently, none of the quantitative claims or analyses can be checked, and the paper is not in a reviewable state. This is a blocking incompleteness that must be resolved before any substantive assessment is possible.","section":"Full text"}],"minor_comments":[{"comment":"The description of VLM4D as 'the first benchmark specifically designed to evaluate the spatiotemporal reasoning capabilities of VLMs' should be supported by a brief comparison with existing video question-answering benchmarks such as NExT-QA, EgoSchema, or TGIF-QA to clarify the claimed novelty.","section":"Abstract"},{"comment":"The term '4D feature field reconstruction' is introduced without definition or citation; the abstract should either explain the concept briefly or reference the approach so that readers can understand the proposed direction.","section":"Abstract"},{"comment":"The phrase 'fundamental deficiencies in existing models' is a strong generalization; unless the benchmark demonstrably covers a broad and representative set of spatiotemporal tasks, the authors may wish to qualify it as 'limitations on VLM4D tasks' to avoid overclaiming.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The submission as provided is incomplete, consisting only of an abstract. I cannot verify any of the technical claims. The stress-test concern about benchmark validity is well-founded and should be resolved with ablations and a detailed human-protocol description. I recommend requesting the full manuscript before making a final editorial decision; the novelty claim of 'first benchmark' should also be checked against the literature during revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a benchmark paper with a plausible and useful niche — VLM4D aims to be the first benchmark specifically for spatiotemporal reasoning in VLMs, and the abstract reports a clear gap between humans and state-of-the-art open and closed models. That alone makes it worth a look if you work on VLM evaluation.\n\nWhat's genuinely new: the benchmark construction with real and synthetic videos, and QA pairs targeting translation, rotation, perspective, and motion continuity. The abstract also gives a useful diagnostic: models fall down when they have to integrate multiple visual cues and maintain temporal coherence. That's a more specific finding than the usual \"VLMs are bad at video.\" The two exploration directions — 4D feature fields and spatiotemporal SFT — are reasonable bets, though the abstract only says they \"demonstrate effectiveness\" without numbers.\n\nWhere I'd pump the brakes: we only got the abstract; the full text in our package is a placeholder. So I can't check the QA construction, the human protocol, or the evaluation pipeline. The biggest soft spot is construct validity. The headline claim — fundamental deficiencies in spatiotemporal reasoning — depends on the questions actually requiring dynamic multi-frame understanding. The abstract doesn't report static-frame baselines, shuffled-frame controls, or question-only checks. Without those, a large human-model gap might just mean the questions correlate with language priors or single-image cues. The human baseline protocol is also unstated; if humans could replay or pause, the comparison is unfair in the other direction. These aren't proven flaws, but they're exactly the evidence the paper needs to include.\n\nThe citation pattern and framing seem fine from the abstract; no red flags there.\n\nVerdict: this deserves a serious referee. The first-benchmark claim is checkable, and if the controls are in the full text, it could be a solid contribution. If they're not, the authors should add them. I'd send it out, but with the expectation that the review asks for those ablations. I wouldn't cite it myself until I've seen the full methodology.","headline":"A promising first-benchmark claim for VLM spatiotemporal reasoning, but the abstract alone doesn't carry the load: the construct-validity controls are missing.","tokens_in":1252,"tokens_out":2539,"would_cite":false,"duration_ms":29020,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces VLM4D, a benchmark of real and synthetic videos with question-answer pairs designed to test whether vision-language models can track objects through space and time, and reports that current models fall well short of…","keywords":["VLM4D","spatiotemporal reasoning","vision language models","video understanding benchmark","4D feature fields","motion continuity","perspective awareness"],"falsifier":"A single controlled experiment would settle the benchmark's validity: paraphrase every question pair without changing the underlying spatiotemporal inference, and regrade the same models. If the human-model gap shrinks or reverses under paraphrasing, the gap is a wording artifact; if it persists across paraphrases and across a different set of videos, the deficiency claim is supported.","tokens_in":485,"feed_emoji":"🎥","tokens_out":6091,"duration_ms":69006,"temperature":0.7,"pith_summary":"This paper introduces VLM4D, a benchmark of real and synthetic videos with question-answer pairs designed to test spatiotemporal awareness in vision-language models (VLMs). The authors' central aim is to establish that current state-of-the-art open and closed-source VLMs, despite strong static-image performance, fall well short of human baselines when they must track how objects move, rotate, and shift in perspective over time. Evaluations across the benchmark show the largest failures come from integrating multiple visual cues at once and keeping a coherent story across frames. The paper also presents evidence that two interventions—reconstructing a 4D feature field of the scene and fine-tuning on spatiotemporal question-answer data—improve these abilities. If the benchmark is a fair measure, then dynamic scene understanding is a distinct and still-unsolved capability that standard video benchmarks do not expose.","feed_headline":"New benchmark shows AI vision models lag on moving 3D scenes","feed_subtitle":"VLM4D tests whether vision-language models track objects moving and rotating through space; most fail at keeping time and motion coherent.","key_machinery":"The central object is VLM4D, a benchmark whose video-question pairs are designed to separate spatiotemporal reasoning into measurable components: translating objects, rotating objects, viewpoint and perspective changes, and continuity of motion across frames. The curated pairing is what carries the argument, because each question is constructed to require a specific spatial-temporal inference rather than a static-image shortcut. The paper also uses a 4D feature field reconstruction as an intervention: it builds a representation in which every spatial location can be queried over time, giving the model an explicit handle on motion, and combines it with supervised fine-tuning on spatiotemporal question-answer pairs. Human baselines on the same questions provide the reference point against which model gaps are defined.","core_discovery":"The central discovery is that modern VLMs do not genuinely reason about space and time together: on VLM4D they show significant gaps against human baselines, and error analysis identifies two specific failure modes—models have trouble combining multiple visual cues in one inference, and they lose temporal coherence as motion unfolds. The benchmark itself is the load-bearing artifact: each video is paired with curated questions that isolate translational motion, rotational motion, perspective awareness, and motion continuity, so that a score can be attributed to a particular spatiotemporal skill. The paper further reports that feeding the model a reconstructed 4D feature field—a representation that lets a network query any point in space at any time—plus targeted spatiotemporal supervised fine-tuning, measurably improves comprehension on these tasks. In the authors' framing, the gap is not a matter of scale but a fundamental deficiency in how existing models encode dynamic structure.","pith_inferences":["Beyond the paper: the benchmark could be turned into a diagnostic by measuring each model's score per subtask (translation, rotation, perspective, continuity) to map exactly which spatiotemporal skill is missing, rather than reporting one aggregate gap.","A testable consequence the authors do not state: if QA paraphrase sets produce the same ranking of models, the conclusion would be robust; if not, the benchmark's wording would need to be treated as part of what is being measured.","The 4D feature field result points toward a larger design principle: future VLMs may need an explicit world model queried in space and time, rather than frame-wise attention over a video, to close the gap."],"forward_implications":["If VLM4D is a valid measure, then state-of-the-art VLMs should be re-evaluated for dynamic scene understanding before being deployed in settings where motion matters, such as robotics and autonomous driving.","The reported human-model gap implies that progress on static image benchmarks does not transfer to spatiotemporal reasoning; new evaluation and training signals are needed.","The success of 4D feature field reconstruction suggests that explicit time-varying scene representations can supply the structural cues that frame-wise attention misses.","Targeted spatiotemporal fine-tuning, as demonstrated, offers a practical route to improve these abilities in existing VLMs without changing the architecture."],"supporting_citations":[],"fun_headline_variants":["AI vision models fail at moving 3D reasoning, new benchmark finds","VLMs can't track motion or time in new spatiotemporal benchmark","New benchmark shows VLMs lose temporal coherence on moving scenes","4D feature fields help VLMs but humans still far ahead on motion","VLM4D: VLMs struggle with rotating objects and perspective shifts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that VLM4D's curated videos and question pairs measure spatiotemporal reasoning well enough that the models' low scores reflect a genuine inability to reason about dynamic scenes, not quirks of how the questions are phrased or which videos were chosen.","fun_headline_variants_meta":{"raw":{"variants":["AI vision models fail at moving 3D reasoning, new benchmark finds","VLMs can't track motion or time in new spatiotemporal benchmark","New benchmark shows VLMs lose temporal coherence on moving scenes","4D feature fields help VLMs but humans still far ahead on motion","VLM4D: VLMs struggle with rotating objects and perspective shifts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000233,"raw_usage":{"total_tokens":1484,"prompt_tokens":926,"completion_tokens":558,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":480}},"tokens_in":542,"tokens_out":558,"duration_ms":7155,"temperature":1.0,"reasoning_tokens":480,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T05:08:28.833929+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A single controlled experiment would settle the benchmark's validity: paraphrase every question pair without changing the underlying spatiotemporal inference, and regrade the same models. If the human-model gap shrinks or reverses under paraphrasing, the gap is a wording artifact; if it persists across paraphrases and across a different set of videos, the deficiency claim is supported.","supporting_citations":[],"review_version":1}