{"id":"6cfdecb9-798d-418c-b131-891b6e070a64","arxiv_id":"2608.05560","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Current multimodal AI models detect sports hazards with high sensitivity but low causal accuracy and high prompt-induced false alarms, according to the SPRINT benchmark.","lead":"This paper introduces SPRINT, a benchmark of 2,888 real-world sports videos with fine-grained labels for when and why accidents happen. Tests of seven multimodal AI models show they often signal danger but rarely name the cause, and they raise false alarms when prompts mention danger explicitly.","discovery_kind":"new_method","skeptic_critique":null,"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces SPRINT, a benchmark of 2,888 real-world sports videos (2,440 accident, 448 safe controls) with fine-grained temporal annotations (earliest cue T1, most obvious moment T2) and hierarchical cause labels (macro factors H1, direct causes H2). The authors evaluate seven MLLMs under multiple prompts and temporal truncation windows along three progressive dimensions: hazard detection (D1), factor coverage (D2), and cause identification (D3). The central empirical finding is a large gap between D1 and D3, along with strong prompt-induced false alarms on safe videos. A fine-tuning experiment on Qwen3-VL-8B-Instruct shows that SPRINT annotations can improve performance on most metrics. The paper concludes that current MLLMs exhibit only superficial proactive safety, lacking stable, cause-grounded early warning.","tokens_in":27548,"tokens_out":6223,"duration_ms":118687,"significance":"If the measurement is valid, SPRINT addresses a real gap: proactive physical hazard anticipation in dynamic video, which existing safety benchmarks do not systematically cover. The use of real-world footage, manual verification of safe controls, multiple prompt conditions, and a diagnostic false-alarm protocol are genuine design strengths, and the paper ships a substantial benchmark with detailed tables. The primary claims, especially the sharp D1-versus-D3 gap and the prompt-sensitivity of early warning, are qualitatively robust across models and settings. However, the validity of the D3 measurement—the load-bearing evidence for the 'superficial proactive safety' conclusion—rests on an automatic evaluator that is itself one of the evaluated models, with only 83% human agreement on D3. The dataset curation also contains an unexplained count discrepancy. These issues do not undermine the qualitative direction of the findings but need to be addressed before the benchmark and its headline numbers can be taken at face value.","major_comments":[{"comment":"The D3 automatic evaluation uses Gemini 3 Flash as the judge, and Gemini 3 Flash is also one of the seven models being scored. Human agreement with this judge is only 83% on D3, and no per-model agreement or self-preference analysis is reported. Since the central claim that models 'fall below 50% on cause identification' depends entirely on these D3 scores, the authors should either (a) report human-evaluated D3 for all models, (b) use a second, non-evaluated judge and show agreement, or (c) at minimum provide evidence that the judge does not systematically favor or penalize particular models or response styles (e.g., concise versus hedging answers). Without this, the exact D3 numbers, and hence the 'superficial proactive safety' conclusion, remain uncertain.","section":"§4.3 and Table 9"},{"comment":"The dataset size arithmetic is inconsistent. Section 3.2 states that after feature-based deduplication the accident pipeline yields 2,630 valid videos, while Section 3.1 and Table 6 report 2,440 accident videos. The difference of 190 videos is never explained, and adding the 448 safe videos to 2,630 gives 3,078 rather than the stated total of 2,888. Since the scale of the benchmark is a primary contribution, the authors must clarify whether 2,630 is a typo, whether additional filtering was applied, or whether the reported totals are correct.","section":"§3.1 and §3.2"},{"comment":"The headline 'the best model exceeds 95% in signaling hazards yet falls below 50% in identifying their causes' is prompt-dependent. Table 9 shows that Doubao-seed-1.8 reaches D3 = 0.53 under Prompt2 and 0.64 under Prompt3, while the below-50% value occurs under Prompt1 (D3 = 0.42). The paper does not specify that the headline refers to the descriptive prompt, and Figure 3 is presented as an aggregate view. The authors should either qualify the claim by prompt or report the D3 range across prompts, so that readers do not take 'below 50%' as the universal best-case result.","section":"Abstract and §5.1"},{"comment":"No measures of variance or statistical significance are reported for any of the evaluation metrics. Several comparisons that support specific conclusions—for example, the claim in Section 5.2 that 'GPT-5 shows stronger prompt robustness'—rest on differences of a few points in Table 2 that could easily be within sampling noise. Moreover, the false-alarm rates in Table 11 are computed on only 448 safe videos, and some gaps (e.g., Gemini-3-Flash P1 0.73 vs P2 0.12) are large, but others (e.g., GPT-5 P2 0.28 vs Doubao P2 0.13) lack error bars. Reporting confidence intervals or conducting significance tests across sampled videos or repeated evaluations would substantially strengthen the benchmark's quantitative claims.","section":"§4.3 and Tables 2, 9, 11"}],"minor_comments":[{"comment":"The fine-tuning experiments use Qwen3-VL-8B-Instruct, whereas the main evaluation tables (Tables 9–11) list 'Qwen3-VL-8B' without the Instruct suffix. It should be stated explicitly whether the 'Base' rows in Tables 3 and 5 refer to the Instruct variant, and if so, why the same variant is not included in the main evaluation.","section":"§6.1 and Tables 3, 5"},{"comment":"The diagnostic test is described as truncating safe videos at the 'moment of maximum motion intensity,' but Figure 5 labels the condition 'Video Context (Up to T1),' which is misleading because safe videos have no T1 annotation. Please use consistent terminology throughout.","section":"§3.3 and Figure 5"},{"comment":"The prompt names contain formatting artifacts such as 'P rompt1 1' and 'P rompt2 1'. These should be cleaned up in the final version.","section":"§4.2"},{"comment":"The fine-tuning training data is generated by GPT-5 from the same annotations that define the test set. The paper should discuss whether the D3 gains could partly reflect learning the annotation format or the generator's phrasing, rather than improved visual causal reasoning, and ideally include a small human-annotated or out-of-distribution test set.","section":"§6.1 and §6.3"},{"comment":"The automatic evaluator prompts request JSON for 'Level1', 'Level2', and 'Level3' responses, but the main text describes only three evaluation dimensions (D1, D2, D3). It would help readers to clarify how the evaluator's output fields map to the D1/D2/D3 metrics.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid benchmark contribution, and the qualitative findings are likely to survive the fixes requested above. The two issues I would weigh most heavily as an editor are the self-evaluator confound for D3 and the dataset count inconsistency; both are fixable but need to be addressed before publication. The related-work positioning relative to PaSBench-Video should also be checked for novelty overlap, since SPRINT and PaSBench-Video are both video benchmarks for proactive safety warning."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"SPRINT is worth a serious look. It is the first video benchmark I know of that tests proactive physical hazard inference at scale on fully real footage: 2,888 clips, 14 sports, with T1/T2 temporal annotations, two-level causal labels, and 448 safe controls specifically for false-alarm diagnosis. The prompt-sensitivity diagnostic is the strongest piece: the difference between explicit danger prompts and neutral descriptions is large, and the static first-frame test convincingly shows that a lot of the apparent 'early warning' is keyword-triggered. The fine-tuning section shows the annotations are usable as training signal, and the appendix documenting an abandoned self-filmed data collection—because models could tell the footage was staged—is the kind of negative result that makes me trust the authors' judgment.\n\nThe central claim, that current MLLMs can signal hazards but cannot ground them in causes and over-warn under explicit prompts, is probably right. The D1-to-D3 gap is consistent across seven models and several prompt sets, and the false-alarm numbers on safe videos are stark. The reader's conditional verdict is fair.\n\nThe soft spots are real but not fatal. The abstract's 'best model exceeds 95%... falls below 50%' is not the worst-case summary of the full tables—some models do score above 50% on D3 in several conditions; it depends on which prompt and window you pick. They should quote a specific configuration or use an average. Second, the D3 numbers rest on Gemini 3 Flash as the judge, and Gemini 3 Flash is also one of the evaluated models. Human agreement on D3 is only 83%, and the rubric's 'shotgun approach' penalty is strict. If the judge penalizes certain answer styles, D3 could be systematically low. The D1-to-D3 gap is so large that it would likely survive, but the paper should at least report a second judge or a small human-scored subset for D3. Third, no confidence intervals on any point estimate; with 2,440 videos, some of the finer sport-level comparisons in Appendix E may be noise. Minor: the data availability statement says both 'open-sourced upon acceptance' and 'annotations are available at GitHub URL,' which should be reconciled.\n\nWho should read this: people building safety evaluations for video MLLMs and anyone working on temporal action anticipation or causal video reasoning. It deserves a serious referee round before acceptance; the issues are fixable with reporting changes, not new experiments.","headline":"Solid new benchmark with a robust central finding; referee it with one eye on the auto-evaluator.","tokens_in":28126,"tokens_out":2568,"would_cite":true,"duration_ms":426085,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Current multimodal AI can flag an imminent sports hazard more than 95% of the time yet names its actual cause less than half the time, and explicit danger prompts trigger frequent false alarms on hazard-free videos.","keywords":["proactive risk inference","multimodal large language models","sports accident videos","hazard detection","cause identification","early warning","benchmark","false alarms"],"falsifier":"Re-score a stratified sample of the best models' full-video responses on D3 with an independent judge and expert panels using a structured cause rubric; if the best model's D3 exceeds 50%, or if a rewrite of the prompt removing the word 'danger' eliminates the D1–D3 gap and the false-alarm jump on safe videos, the claim of superficial proactive safety would be substantially weakened.","tokens_in":27438,"feed_emoji":"⚠️","tokens_out":7340,"duration_ms":53159,"temperature":0.7,"pith_summary":"This paper introduces SPRINT, a benchmark of 2,888 real-world sports accident and safe-control videos, and uses it to test whether multimodal large language models can give cause-grounded early warnings of physical hazards. The central finding is a sharp split: the best model flags imminent danger in more than 95% of accident videos, yet identifies the accident's actual cause in fewer than 50%. Diagnostic experiments on hazard-free videos show that simply asking 'Is there any danger?' pushes false-alarm rates up sharply, which the authors read as evidence that current warnings are driven more by prompt-induced bias than by grounded visual understanding. If the conclusion holds, it matters for any safety application—autonomous driving, fall detection, workplace monitoring—where an early warning is only useful if it says why something is about to go wrong.","feed_headline":"AI models signal danger 95% of the time but name causes under 50%","feed_subtitle":"Sports-accident benchmark shows early warnings are prompt-driven, not grounded in physical understanding.","key_machinery":"The load-bearing object is SPRINT itself: 2,888 real-world videos (2,440 accidents, 448 safe controls) spanning 14 sports and 3 settings, with two timestamps per accident—the earliest cue (T1) and the most obvious moment (T2)—and two levels of cause annotation: macro inducing factors (H1, aligned with the host–agent–environment dimensions of the Haddon Matrix) and free-text direct cause descriptions (H2). The evaluation protocol defines three hierarchical binary metrics—D1 (hazard mentioned), D2 (factor coverage), D3 (cause specificity)—so that a model only scores on D3 if it has scored on D1 and D2. Safe videos annotated with a moment of maximum motion intensity act as negative control probes, isolating prompt-induced false alarms from genuine hazard perception.","core_discovery":"The paper's claim is that current MLLMs exhibit only superficial proactive safety: they can signal that a hazard is present or imminent, but they lack stable, cause-grounded early warning. The evidence comes from three progressive evaluation dimensions on SPRINT—hazard detection (D1), factor coverage (D2), and direct cause identification (D3). Under full-video evaluation, the strongest closed-source model exceeds 95% on D1 but falls below 50% on D3; open-source models fall below 60% on D1 unless explicitly prompted to look for danger. In the temporal-window experiment, shifting from an explicit danger inquiry to a neutral description drops the best early-window D1 score from 88% to 59%, and on the 448 safe control videos explicit danger prompts raise false-positive rates to as high as 0.88. Fine-tuning Qwen3-VL-8B on the SPRINT annotations more than triples D3 (from 0.16 to 0.52) and sharpens early-window detection, which the paper offers as evidence that the benchmark captures teachable skill rather than an unfixable limitation.","pith_inferences":["The D1–D3 split is probably a signature of anomaly detection rather than physical reasoning: the models latch onto salient motion changes (a fall, a collision) and fail on hazards with subtle kinematics, which predicts the observed near-floor performance on pole vault and interpersonal collisions.","The prompt-bias result suggests that instruction-tuned safety behavior is partly lexical—models associate the word 'danger' with a warning response—so an untested extension would be measuring whether removing risk vocabulary from prompts degrades detection in real driving or fall-monitoring settings.","Because the D3 auto-judge is itself a tested model and human agreement on D3 is 83%, the absolute D3 numbers are likely lower bounds; the comparative ordering across models and the D1–D3 gap are the sturdier conclusions.","A concrete extension of the benchmark idea would be a structured H2 taxonomy (e.g., kinematic cause categories) to make cause-identification scoring independent of free-text matching, and a domain-transfer test that fine-tunes on SPRINT and evaluates on driving or fall videos."],"forward_implications":["Safety evaluations that only check whether a model refuses or flags harmful content miss the proactive dimension; SPRINT shows a model can 'pass' detection while lacking the causal understanding needed for intervention.","A high D1/low D3 profile means a deployed early-warning system would trigger alerts without being able to tell a driver, clinician, or worker what is going wrong and what to correct.","Prompt sensitivity becomes a measurable axis: the same model can swing from 88% to 59% detection merely by rewording the question, so a single-prompt evaluation overstates proactive capability.","The fine-tuning results indicate that temporally and causally annotated video data can partially close the gap, suggesting the deficit is not purely architectural.","Because the hazard taxonomy mirrors the Haddon Matrix, scores on SPRINT offer a proxy for cause-grounded reasoning in other physical-safety domains such as driving and fall prevention."],"supporting_citations":[{"why":"Supplies the injury-prevention framework whose host–agent–environment dimensions justify sports accidents as a proxy for broader physical-safety reasoning.","marker":"Haddon Jr, 1968"},{"why":"PaSBench, the proactive risk-awareness benchmark that SPRINT extends from images to real-world sports video.","marker":"Yuan et al., 2025"},{"why":"PaSBench-Video, the closest prior streaming-video proactive benchmark, whose temporal calibration and false-positive control challenges SPRINT builds on.","marker":"Zhao et al., 2026"},{"why":"CLIP features used for feature-based video deduplication during dataset curation.","marker":"Radford et al., 2021"},{"why":"Driving-accident anticipation benchmark cited to ground the claim that sports pre-accident cues share reasoning with autonomous driving.","marker":"Fang et al., 2019"},{"why":"Video-capture evidence for fall detection, supporting the domain-transfer argument for elderly fall prevention.","marker":"Robinovitch et al., 2022"}],"fun_headline_variants":["MLLMs flag 95% of hazards, explain under 50%","Proactive safety gap: signal > understanding in MLLMs","SPRINT: MLLMs detect hazards but miss causes","95% hazard detection, under 50% cause ID in MLLMs","Prompt-driven false alarms expose superficial MLLM safety"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion that cause identification is below 50% rests on the assumption that Gemini 3 Flash's automatic scoring of free-text cause descriptions is a valid measure of true cause identification, an assumption the paper itself qualifies with an 83% human agreement rate on D3 and with the fact that the judge is one of the seven models under test.","fun_headline_variants_meta":{"raw":{"variants":["MLLMs flag 95% of hazards, explain under 50%","Proactive safety gap: signal > understanding in MLLMs","SPRINT: MLLMs detect hazards but miss causes","95% hazard detection, under 50% cause ID in MLLMs","Prompt-driven false alarms expose superficial MLLM safety"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000671,"raw_usage":{"total_tokens":3094,"prompt_tokens":1019,"completion_tokens":2075,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":635,"completion_tokens_details":{"reasoning_tokens":1985}},"tokens_in":635,"tokens_out":2075,"duration_ms":12436,"temperature":1.0,"reasoning_tokens":1985,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T10:31:53.004211+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-score a stratified sample of the best models' full-video responses on D3 with an independent judge and expert panels using a structured cause rubric; if the best model's D3 exceeds 50%, or if a rewrite of the prompt removing the word 'danger' eliminates the D1–D3 gap and the false-alarm jump on safe videos, the claim of superficial proactive safety would be substantially weakened.","supporting_citations":[],"review_version":1}