{"id":"d4def7ec-80f9-42a6-9b44-9668b956eee9","arxiv_id":"2507.09815","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The new VRU-Accident benchmark (1K videos, 6K QA pairs, 1K dense captions) shows the best evaluated MLLM reaches 66.9% on VRU-accident VQA versus 94.7% for human experts, with the weakest performance on causal and preventive reasoning.","lead":"VRU-Accident is a new benchmark of 1,000 real-world dashcam videos showing accidents between vehicles and pedestrians or cyclists, with 6,000 multiple-choice questions and dense written descriptions. It shows that today's multimodal AI models handle basic visual attributes well but fall far short of human experts on reasoning about accident causes, types, and preventability.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reuse of MM-AU/DoTA videos may inflate the reported zero-shot accuracies; a source-split re-evaluation is needed before the benchmark numbers can be taken at face value.","rationale":"I agree with the reader that dataset contamination from MM-AU/DoTA reuse is the most load-bearing threat to the paper's central evaluation claim. The paper is otherwise internally consistent: the benchmark construction is described in sufficient detail, the VQA and dense captioning protocols are reasonable as a first pass, and the conclusion that MLLMs are weaker on causal and preventive categories is plausible given the reported category-wise pattern. However, the 'zero-shot on unseen scenarios' claim is an explicit part of Section 4.1, and Section 2 admits that a large fraction of the videos come from existing public datasets. Because contamination would affect the magnitude but probably not the direction of the findings, the appropriate disposition is the same conditional verdict the reader reached: require contamination diagnostics before accepting the quantitative results. If the proposed source-split check shows no accuracy gap, the paper can be upgraded without further changes.","tokens_in":19372,"tokens_out":8473,"duration_ms":99262,"concrete_test":"Ask the authors to release the source identifier for each of the 1,000 videos (which are from MM-AU, which from DoTA, and which are newly collected web videos), then re-run the Table 4 VQA evaluation separately on the public-source subset and the newly collected subset using identical prompts and decoding settings. If the newly collected subset shows comparable or higher accuracy on every category, contamination from MM-AU/DoTA reuse is not inflating the reported numbers; if accuracy on the newly collected subset drops materially, especially on Weather & Light or Traffic Environment, the zero-shot results should be re-reported on a contamination-free split before the headline claims are trusted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claim is that 17 MLLMs tested 'zero-shot' reach only 66.9% average accuracy versus 94.7% for human experts, with especially poor results on accident cause and prevention (Section 4.2, Table 4). Section 2 explicitly says VRU-Accident 'combines the MM-AU and DoTA samples while adding more samples reaching 1,000 videos.' MM-AU and DoTA are public dashcam datasets assembled from web-scraped videos, and MLLM video pretraining corpora are also largely web-scraped. Consequently, exact or near-duplicate videos may already appear in the training data of the evaluated models. If so, the zero-shot protocol in Section 4.1 is not evaluating 'unseen VRU-related accident scenarios,' and the reported accuracies confound genuine video reasoning with memorization of specific footage. This does not by itself overturn the claim that MLLMs struggle on causal and preventive reasoning; decontaminated scores would likely be lower, not higher. Rather, it undermines the specific benchmark numbers and the claim that models 'perform reasonably well on visually grounded attributes,' both of which are load-bearing for the benchmark's utility as a fair zero-shot evaluation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"VRU-Accident introduces a benchmark of 1,000 real-world dashcam videos of accidents or near-accidents involving vulnerable road users, annotated with 6,000 multiple-choice VQA pairs across six categories (24,000 candidate options, 3,400+ unique answers) and 1,000 dense scene descriptions. The curation pipeline is semi-automatic: human experts provide correct VQA answers, GPT-4o generates counterfactual distractors and initial dense captions, and human experts verify the results. The paper evaluates 17 MLLMs (15 open-source, 2 closed-source) in a zero-shot setting on VQA and dense captioning, reporting best average VQA accuracy of 66.9% (Gemini 1.5-flash) versus 94.7% for human experts, with the largest gaps on Accident Cause and Prevention Measure. The central conclusion is that current MLLMs handle visually grounded attributes reasonably well but struggle with causal, preventive, and fine-grained accident reasoning.","tokens_in":19607,"tokens_out":5443,"duration_ms":59993,"significance":"If the benchmark numbers are taken at face value, VRU-Accident is a useful and timely resource: it targets an under-served safety-critical niche, combines VQA and dense captioning in one benchmark, includes human verification in the annotation loop, covers diverse VRU accident scenarios, and ships inference scripts, model outputs, and evaluation code for reproducibility. The paper's qualitative finding that all evaluated models fall far below human performance on causal and preventive reasoning is plausible and worth publishing, provided the validity concerns below are resolved. The benchmark's potential impact for MLLM evaluation in autonomous driving justifies serious attention, but the headline numbers currently rest on assumptions about video contamination and reference independence that are not verified.","major_comments":[{"comment":"The paper states in Section 2 that VRU-Accident 'combines the MM-AU and DoTA samples while adding more samples reaching 1,000 videos.' MM-AU and DoTA are public web-scraped dashcam datasets, and the evaluated MLLMs were pretrained on large web-scale video corpora. The zero-shot protocol described in Section 4.1 therefore does not guarantee that the videos are unseen, and the Table 4 accuracies may conflate genuine reasoning with memorization of specific footage. This is load-bearing for the central claim that models 'perform reasonably well on visually grounded attributes' yet fail on causal reasoning. The authors should provide a source-split analysis (e.g., report results separately for newly collected videos versus reused MM-AU/DoTA videos) and a decontamination check, or explicitly weaken the zero-shot claim. Decontaminated scores would likely be lower, not higher, so the qualitative conclusion about causal reasoning may survive, but the benchmark's quantitative claims cannot be accepted without this analysis.","section":"Section 2, Table 1; Section 4.1"},{"comment":"The dense captioning ground truth is generated by GPT-4o and then human-verified, rather than written independently by human annotators. The paper appropriately excludes GPT-4 from the dense captioning leaderboard in Table 5, but all other models are still scored against references produced by the same model family and in the same narrative style. As a result, the reported SPICE, METEOR, COMET, and ROUGE scores reward stylistic and content choices of GPT-4o, not necessarily accident-understanding quality. This is load-bearing for the dense captioning conclusions. The authors should either collect a human-written reference subset, add human preference or correctness evaluation, or report robustness of model rankings across alternative reference sets.","section":"Section 3.3, Supplementary D, Table 5"},{"comment":"The human expert row in Table 4 reports 94.7% average accuracy, but the manuscript does not state the number of annotators, their qualifications, the annotation protocol, or inter-annotator agreement. Without this information, the human baseline and the label reliability of the VQA ground truth are not established. This matters because the paper's central comparison is between MLLM performance and human expert performance, and because the benchmark's utility depends on the correctness and consistency of the ground-truth answers. The authors should report annotation statistics and agreement metrics, and describe how borderline cases (e.g., near-miss events discussed in Supplementary B.3) were adjudicated.","section":"Section 4.2, Table 4"},{"comment":"The VQA distractors are generated by GPT-4o from the human-labeled ground truth, and the analysis in Supplementary Figure 8 shows that distractors in several categories have question-to-answer cosine similarity comparable to or higher than the ground truth. While this suggests the distractors are semantically plausible, it also raises the question of whether some multiple-choice items are ambiguous or have more than one defensible answer. The paper does not report human agreement on the final candidate sets or any validation that the three distractors are unambiguously incorrect. Since the benchmark's discriminative power depends on distractor quality, the authors should provide human validation statistics or a discussion of ambiguity resolution.","section":"Section 3.2, Supplementary D"}],"minor_comments":[{"comment":"There are numerous typos and formatting artifacts, including 'traffic accidentes' in the Introduction, 'One the other hand' in Section 2, 'Unlikely accident classification tasks' in Section 3.4 (should be 'Unlike'), 'Configiguration' in Table 3, and inconsistent rendering of 'LLaVA' as 'LLaV A' throughout the text and tables.","section":"Throughout"},{"comment":"The model lists in Table 4 and Figure 4 are inconsistent: Figure 4 includes 'Mobile-VideoGPT(0.5B)', which does not appear in Table 4, while Table 4 lists 'Mobile-VideoGPT(1.5B)'. The authors should align these lists and describe which model versions were actually evaluated.","section":"Table 4 and Figure 4"},{"comment":"The METEOR formula as written, METEOR = (Precision × Recall) / (Precision + α × Recall + (1 − α)), does not match the standard METEOR formulation in the cited reference [19]. Please clarify the formula or use the standard definition with the appropriate penalty term.","section":"Supplementary B.2, Eq. (4)"},{"comment":"The main text refers to a 'Dense Captioning generator G_DC' without identifying it as GPT-4o; the identity is only revealed in the supplementary material. This should be stated explicitly in the main text, since it directly affects interpretation of the dense captioning evaluation.","section":"Section 3.3 and Supplementary D"},{"comment":"The paper says BLEU is omitted because scores are below 0.1, but BLEU is a precision-oriented metric that is known to be harsh on video captions; a one-sentence justification in the main text would help readers interpret the metric choice.","section":"Section 4.1"},{"comment":"Since the benchmark reuses videos from MM-AU and DoTA, the paper should document the licenses and any privacy or consent considerations for the dashcam footage, especially because the dataset is publicly released and contains identifiable people in accident situations.","section":"Section 2 and Data Release"}],"recommendation":"major_revision","confidential_remarks":"This is a solid benchmark contribution with a clear evaluation gap, but the two main validity issues—possible train/video overlap for the zero-shot VQA numbers and GPT-4-generated dense caption references—are load-bearing for the headline claims. If the authors can supply a source-split or contamination analysis, human-written caption references (at least on a subset), and annotation reliability statistics, the paper would be much stronger; without those, the quantitative claims should be substantially weakened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a useful benchmark that should be reviewed, but the central quantitative claim is undercut by a plausible contamination problem. The paper's qualitative conclusion—MLLMs are much better at visually grounded attributes than at causal and preventive reasoning—is likely robust, and probably understated if anything, but the specific 66.9% vs 94.7% comparison needs a source-split re-evaluation before it can be cited as a fair zero-shot result.\n\nWhat's new and good: it's the first VRU-focused video VQA + dense captioning benchmark from real dashcam accidents. The annotation design is thoughtful: six categories including cause, type, and prevention; 6K questions with large unique answer sets; 1K dense captions. The semi-automatic pipeline with human verification is practical, and they evaluate 17 models with both closed and open weights. The finding that models fall below 50% on cause and prevention across the board is consistent across model families and is a useful signal for the field. They also commit to releasing scripts, outputs, and evaluation results, which is creditable.\n\nSoft spots, in proportion: the biggest one is overlapping source videos. Section 2 is explicit that the benchmark combines MM-AU and DoTA samples, and both are web-scraped dashcam collections that likely appear in MLLM pretraining corpora. Without a contamination analysis, the zero-shot protocol in Section 4.1 is not really evaluating 'unseen' scenarios. Decontaminated scores would almost certainly be lower, so the qualitative gap widens rather than shrinks, but the benchmark's numerical claims are compromised. They should report scores split by source (original vs. reused) or at least do video-level near-duplicate removal.\n\nSecond, dense caption references are generated by GPT-4 and then human-verified. Excluding GPT-4 from the leaderboard is the right call, but the references are still written in GPT-4's narrative style, which biases METEOR and ROUGE toward models that imitate that style. Independent human-written references for a random subset would calibrate this.\n\nThird, no random baseline. Given that options are not uniformly distributed across categories, reporting a simple chance baseline per category would help readers interpret the low scores.\n\nMinor: the human expert baseline is probably inflated because the same annotators who wrote the ground truth took the test, but that's common and not a real problem.\n\nBottom line: the benchmark is a genuine resource and the main finding about causal reasoning is real. Send it to peer review, but require source-split analysis and a calibration sample of human-written captions. If those come back clean, this is a solid standard testbed.","headline":"A useful VRU accident benchmark that should get peer review, but the zero-shot numbers need a source-split re-evaluation before the performance gap can be taken at face value.","tokens_in":20115,"tokens_out":2929,"would_cite":true,"duration_ms":34213,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI models lag humans by 28 points on pedestrian-crash reasoning","keywords":["video question answering","dense captioning","vulnerable road users","multimodal large language models","accident scene understanding","autonomous driving safety","counterfactual distractors","benchmark"],"falsifier":"Run the same 17 models on a newly recorded set of 1,000 VRU accident and near-accident videos that were never posted publicly and postdate every model's training cutoff, using the identical VQA and dense-caption protocol. If average accuracy approaches the 94.7% human level, the claimed reasoning deficit is contamination, not a genuine capability gap.","tokens_in":19199,"feed_emoji":"🚗","tokens_out":7342,"duration_ms":74708,"temperature":0.7,"pith_summary":"VRU-Accident is a large-scale vision-language benchmark built from 1,000 real-world dashcam videos of accidents and near-accidents involving pedestrians and cyclists, annotated with 6,000 multiple-choice questions across six safety-critical categories and 1,000 dense scene descriptions. The paper's central claim is that current multimodal large language models (MLLMs) can answer visually grounded questions about weather, lighting, and traffic environment but cannot reliably reason about accident cause, type, and prevention. In zero-shot evaluation of 17 models, the best model reaches 66.9% average accuracy against 94.7% for human experts, and dense captions frequently omit collision dynamics or hallucinate road-user motion. The benchmark is offered as a standardized testbed for diagnosing and improving accident-centric reasoning in MLLMs for autonomous driving safety.","feed_headline":"AI models lag humans by 28 points on pedestrian-crash reasoning","feed_subtitle":"A 1,000-video benchmark finds multimodal models can see scenes but cannot explain why crashes happen.","key_machinery":"The load-bearing mechanism is the benchmark's six-category VQA design paired with counterfactual distractor generation. For each video $V_i$ and each category $j$, human experts supply a ground-truth answer $y_j$; a VQA generator $G_{VQA}$ produces three contextually plausible but incorrect answers $y^*_j$, so the candidate set $\\hat{Y}_j = \\{y_j, y^{*(1)}_j, y^{*(2)}_j, y^{*(3)}_j\\}$ forces models to reason over dynamically constructed, semantically diverse answer spaces rather than choose from a fixed label set. A second generator $G_{DC}$ produces dense captions $C_i$ describing weather, environment, road-user appearance, kinematics, spatial relations, and collision sequence, and human experts verify all annotations. The three causal categories — Accident Type, Accident Cause, and Prevention Measure — carry the diagnostic weight: they separate visually grounded attribute recognition from the causal and counterfactual reasoning that the paper claims current MLLMs lack.","core_discovery":"The paper introduces VRU-Accident as the first large-scale vision-language benchmark aimed specifically at evaluating MLLMs in safety-critical scenarios involving vulnerable road users. Each of its 1,000 ego-view dashcam videos carries six VQA pairs, one per category — Weather & Light, Traffic Environment, Road Configuration, Accident Type, Accident Cause, and Accident Prevention Measure — with one correct answer and three counterfactual distractors per question, yielding 24,000 candidate options and over 3.4K unique answers; each video also has a dense, temporally grounded caption. Annotations are produced semi-automatically: human experts label ground truth, GPT-4o generates plausible distractors and initial dense captions, and human experts verify. Evaluated zero-shot on 17 MLLMs (15 open-source and 2 closed-source), the benchmark shows that models do well on visually grounded attributes but drop sharply on causal and preventive reasoning: the best model, Gemini 1.5-flash, averages 66.9% versus 94.7% for human experts, and dense captions score low on SPICE, METEOR, COMET, and ROUGE while often misdescribing collision dynamics. The authors conclude that VRU-Accident provides a diagnostic and generative testbed that exposes the gap between visual grounding and true accident reasoning in current MLLMs.","pith_inferences":["Beyond the paper, because the videos are drawn from public datasets MM-AU and DoTA, an implicit assumption is that the evaluated MLLMs did not see these exact clips during pretraining; a clean held-out test with newly collected videos would be a stronger test of the claimed reasoning gap.","Beyond the paper, the counterfactual distractor pairs could be reused as training data for contrastive accident reasoning, teaching models to prefer causally correct explanations over visually plausible ones.","Beyond the paper, reference captions were initially generated by GPT-4o, so reported captioning scores may understate models whose narrative style differs from GPT-4o; rewriting references by human experts would harden the dense-caption leaderboard.","Beyond the paper, the six categories could be extended to near-miss severity, VRU intention prediction, and liability attribution, which would make the benchmark directly useful for autonomous vehicle decision-making rather than only for diagnosis."],"forward_implications":["Model rankings on VRU-Accident give a standardized zero-shot measure of accident-scene reasoning, so progress in safety-critical MLLMs can be tracked against a fixed human baseline of 94.7%.","Because visually grounded categories are easy while Accident Cause and Prevention Measure are hard, benchmarks that only test scene attributes will overstate MLLM readiness for autonomous driving perception.","The counterfactual distractor design means a model must distinguish the correct cause from plausible alternatives rather than memorize label distributions, so high accuracy requires genuine causal reasoning.","Low dense-caption scores and hallucinated pedestrian motion show that current MLLMs lack temporally grounded collision understanding, not just vocabulary for describing scenes.","Open-source models such as Qwen2-VL(7B) can serve as practical baselines for accident captioning, but all evaluated models fall short of human-level narrative accuracy."],"supporting_citations":[{"why":"MM-AU supplies 510 VRU accident videos reused in VRU-Accident and provides prior accident captions that the benchmark extends.","marker":"[11, 12]"},{"why":"DoTA supplies dashcam accident videos and serves as a source of additional VRU clips.","marker":"[50]"},{"why":"DADA-2000 contributes 223 VRU videos and prior driver-attention accident analysis.","marker":"[10]"},{"why":"SUTD-TrafficQA defines the traffic VQA benchmark task that VRU-Accident adapts to ego-view VRU accidents.","marker":"[47]"},{"why":"GPT-4 (GPT-4o) generates the three counterfactual distractors per question and the initial dense captions.","marker":"[30]"},{"why":"The Traffic Accident Benchmark for Causality Recognition motivates the accident cause and prevention categories.","marker":"[53]"},{"why":"Gemini 1.5-flash is the best evaluated model, setting the 66.9% average accuracy that defines the current ceiling.","marker":"[39]"},{"why":"SPICE scores dense captions on scene-graph tuples, grounding the claim that captions miss causal and spatiotemporal content.","marker":"[6]"},{"why":"COMET is a neural metric used to score dense captions for semantic adequacy.","marker":"[34]"}],"fun_headline_variants":["AI lags humans by 28 points on pedestrian crash reasoning","New benchmark: AI can see accidents but not explain why","First VRU accident benchmark shows AI's causal reasoning gap","Why did that pedestrian crash? AI models fail the test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion that today's MLLMs lack VRU accident reasoning assumes the reused public dashcam videos were not part of the models' pretraining data; if the models have already seen these clips, the reported zero-shot accuracy gaps are inflated.","fun_headline_variants_meta":{"raw":{"variants":["AI lags humans by 28 points on pedestrian crash reasoning","New benchmark: AI can see accidents but not explain why","First VRU accident benchmark shows AI's causal reasoning gap","Why did that pedestrian crash? AI models fail the test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000353,"raw_usage":{"total_tokens":1992,"prompt_tokens":1089,"completion_tokens":903,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":705,"completion_tokens_details":{"reasoning_tokens":835}},"tokens_in":705,"tokens_out":903,"duration_ms":10099,"temperature":1.0,"reasoning_tokens":835,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:46:27.383520+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 17 models on a newly recorded set of 1,000 VRU accident and near-accident videos that were never posted publicly and postdate every model's training cutoff, using the identical VQA and dense-caption protocol. If average accuracy approaches the 94.7% human level, the claimed reasoning deficit is contamination, not a genuine capability gap.","supporting_citations":[{"cited_title":"Crandall","cited_arxiv_id":null,"evidence_quote":"DoTA supplies dashcam accident videos and serves as a source of additional VRU clips."},{"cited_title":"Dada-2000: Can driving accident be pre- dicted by driver attention? analyzed by a benchmark, 2019","cited_arxiv_id":null,"evidence_quote":"DADA-2000 contributes 223 VRU videos and prior driver-attention accident analysis."},{"cited_title":"Sutd-trafficqa: A question answering benchmark and an efficient network for video rea- soning over traffic events, 2021","cited_arxiv_id":null,"evidence_quote":"SUTD-TrafficQA defines the traffic VQA benchmark task that VRU-Accident adapts to ego-view VRU accidents."},{"cited_title":"Gpt-4 technical report, 2024","cited_arxiv_id":null,"evidence_quote":"GPT-4 (GPT-4o) generates the three counterfactual distractors per question and the initial dense captions."},{"cited_title":"Traffic Accident Bench- mark for Causality Recognition","cited_arxiv_id":null,"evidence_quote":"The Traffic Accident Benchmark for Causality Recognition motivates the accident cause and prevention categories."},{"cited_title":"Gemini: A family of highly capable multi- modal models, 2025","cited_arxiv_id":null,"evidence_quote":"Gemini 1.5-flash is the best evaluated model, setting the 66.9% average accuracy that defines the current ceiling."},{"cited_title":"Spice: Semantic propositional image cap- tion evaluation","cited_arxiv_id":null,"evidence_quote":"SPICE scores dense captions on scene-graph tuples, grounding the claim that captions miss causal and spatiotemporal content."},{"cited_title":"Comet: A neural framework for mt evaluation, 2020","cited_arxiv_id":null,"evidence_quote":"COMET is a neural metric used to score dense captions for semantic adequacy."}],"review_version":1}