{"id":"828ac82e-51fc-4e32-96c0-6fb36bc9047c","arxiv_id":"2507.02074","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured survey of 2023-2025 LLM and VLM methods for crash detection in video, with notable internal inconsistencies in reported numbers.","lead":"This paper reviews how large language models are being applied to detect car crashes in video, organizing recent methods into a taxonomy and comparing datasets and benchmarks. It is a useful starting map of a young field, but internal contradictions and unverifiable references limit its reliability.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Survey's empirical map is not self-consistent: DAD size, CrashLLM metrics, and unsourced failure-rate numbers in §VII-F conflict inside the paper, so the central deployment-gap claim rests on unverified transcriptions until corrected.","rationale":"The reader's weakest assumption is exactly the one I find most load-bearing: the survey's value is as a faithful synthesis, and its central claims about methods, readiness, and gaps stand or fall on whether the numbers it reports from primary sources are accurate and self-consistent. The visible internal contradictions—DAD size, CrashLLM's headline metric, the unverifiable TrafficLens citation, and the unsourced failure rates in §VII-F—confirm that this assumption is currently violated. I do not see a separate, more fundamental flaw in the taxonomy or argument structure: the high-level claim that LLM/VLM methods add contextual reasoning and explanation at significant computational cost can survive corrections to individual rows. The right remedy is the same one the reader proposes: verify the disputed entries against primary sources and correct the survey before it is used as a reference. Therefore the reader's CONDITIONAL verdict remains appropriate, and I set verdict_should_be to UNCHANGED rather than moving it.","tokens_in":28348,"tokens_out":7003,"duration_ms":84723,"concrete_test":"Perform a primary-source audit of the flagged entries: (1) count videos in the original DAD paper (Chan et al., ICPR 2016); (2) check CrashLLM's original paper (arXiv:2402.16682) to determine whether 'over 92%' refers to a different task or subset than the 53.8% macro-F1; (3) search for a NEC Labs 'TrafficLens' technical report and, if it does not exist, replace the §II-E2 example; (4) locate any source for the §VII-F failure rates; and (5) recompute the affected Table II, IV, and V rows with verified values. If DAD should be ~620 videos or CrashLLM's 92% figure cannot be confirmed, correct the text and revise any deployment conclusions that depend on those numbers.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is a map of LLM-based crash detection and a deployment gap, and that map is built from transcribed numbers. The assumption that those transcriptions are accurate and self-consistent is load-bearing, and it is visibly violated. Table II lists DAD as '1,500+ videos', §VII-A repeats 'DAD: 1,500 videos', but §X-A says DAD contains '~620 dashcam videos (with positive/negative splits)'. §II-D and §II-E attribute 'over 92% accuracy' to CrashLLM, while §IV-B and Table IV report macro-F1 34.9%→53.8% on CrashEvent; the paper never reconciles these as different metrics or tasks. §II-E2's dynamic-template example relies on TrafficLens [59], whose reference has 'Unknown' as the author and no URL, making the citation unverifiable. §VII-F asserts 12–16% OOD drops, 40% occlusion misses, 25% weather degradation, 30% lighting reduction, and 60% adversarial false negatives without citing any source. These are not cosmetic typos: they are the quantitative content of the survey's comparative and deployment analysis. If any flagged row is wrong, readers will draw incorrect conclusions about which methods are deployment-ready and which datasets support which claims. The 'Note on Computational Estimates' in §VI-D further concedes that latency and deployment classes are estimates, weakening the empirical basis of Table V. As written, the paper does not yet provide a reliable synthesis, though it is repairable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript surveys recent (2023–2025) LLM/VLM approaches to video-based crash detection. It proposes a taxonomy along fusion level, prompt strategy, and LLM role; reviews key datasets; describes three architectural families; discusses evaluation metrics; compares recent methods in Table IV; analyzes deployment readiness in Table V; and outlines challenges and future directions. The paper's central claims are that LLM-based methods add contextual understanding and explanation generation, that the field is shifting from reactive pixel-based detection to context-aware event interpretation, and that a substantial deployment gap remains due to latency, compute, and robustness issues.","tokens_in":28671,"tokens_out":5113,"duration_ms":55869,"significance":"If its factual basis were reliable, this survey would be a useful entry point for researchers, with a clear taxonomy and broad coverage of a fast-moving area. The manuscript also deserves credit for explicitly acknowledging in §VI-D that computational classifications and latency figures are approximate. However, the survey's value as a synthesis depends on accurate transcription of primary sources, and the visible internal contradictions in dataset sizes, model metrics, architecture descriptions, and unsourced failure statistics undermine confidence in the comparison tables and in the deployment-gap argument. The paper is repairable, but in its current form it is not yet a reliable map of the field.","major_comments":[{"comment":"The DAD dataset size is reported inconsistently: Table II and §VII-A say 1,500+ videos, while §X-A states that DAD contains '~620 dashcam videos (with positive/negative splits)'. Since DAD is used as a benchmark for VERA, Video-LLaMA, and CRASH, this discrepancy affects the dataset map and the data-bottleneck argument. Please verify the count against the primary source [61] and state one consistent number, or clearly explain the version/split difference.","section":"Table II, §VII-A, §X-A"},{"comment":"CrashLLM's reported performance is contradictory: §II-D2 and §II-E1 credit it with 'over 92% accuracy', while §IV-B and Table IV report macro-F1 improving from 34.9% to 53.8% on CrashEvent. The paper never reconciles these as different metrics or tasks. In addition, §IV-B describes CrashLLM's visual backbone as a Swin Transformer, while §VI-B describes it as ResNet-50. Since CrashLLM is a flagship example in the taxonomy and comparison, these inconsistencies must be resolved with precise metric definitions and a single architecture description.","section":"§II-D2, §II-E1, §IV-B, Table IV"},{"comment":"The robustness numbers in §VII-F — 12–16% out-of-distribution drops, 40% occlusion misses, 25% weather degradation, 30% lighting reduction, and 60% adversarial false negatives — are presented as empirical findings without any citation. Section VII-G repeats the '12–16% on average' figure without a source. These figures are load-bearing for the deployment-gap claim, so they must either be traced to specific sources with the evaluation protocol described, or be removed and rephrased as unsupported estimates.","section":"§VII-F, §VII-G"},{"comment":"Video-LLaMA's role in crash detection is described inconsistently. Table IV states 'N/A (no crash benchmarks)' and §VI-B says the original paper does not report crash-specific benchmarks, but §IV-A says Video-LLaMA detects 'crash anomalies on datasets like UCF-Crime', §II-C presents it as a step toward temporal reasoning in crash detection, and §III-A says DAD is used to train VLMs like Video-LLaMA. Please distinguish the original model's reported benchmarks from downstream applications that use it as a backbone.","section":"§IV-A, §II-C, §III-A, Table IV"},{"comment":"The deployment-readiness conclusions rest partly on unmeasured values. The 'Note on Computational Estimates' in §VI-D concedes that computational classifications and latency estimates are approximate, yet Table V labels methods 'Deployment Ready' or 'No' based on those columns and on reported metrics. Please either obtain latency and memory measurements from the primary sources, or clearly mark all estimate-derived cells in Table V and soften the readiness labels accordingly.","section":"§VI-D, Table V"}],"minor_comments":[{"comment":"Reference [59] (TrafficLens) lists 'Unknown' as the author and provides no URL or venue, making the citation unverifiable; since this example anchors the 'Dynamic Event Templates' taxonomy category, it should be replaced with a traceable source or removed.","section":"Reference [59], §II-E2"},{"comment":"Reference [64] attributes UCF-Crime to 'W. Soomro, A. R. Zamir, and M. Shah'; the actual authors are Sultani, Chen, and Shah (as correctly listed in reference [10]). Please correct the citation.","section":"Reference [64], §III-A"},{"comment":"The word 'demonstaring' in §IV-B should be 'demonstrating'.","section":"§IV-B"},{"comment":"Section V-B refers to 'anticipation-focused datasets such as CRASH', but CRASH is presented in Table IV and §VI-B as a method/model. Please clarify whether CRASH is also a dataset and avoid mixing these roles.","section":"§V-B, Table IV"},{"comment":"References [44] and [67] refer to the same CRASH system with different author lists and one incomplete entry ('Y. Liao et al.'). Please consolidate them into a single complete citation.","section":"References [44] and [67]"},{"comment":"The timeline in Figure 3 lists 'LLaV A-1.5' with an extra space; this should read 'LLaVA-1.5' for consistency with the text.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the survey cites the authors' own works [57] and [86] as evidence of field progress, and one of them ([57]) is a workshop paper. This is not disqualifying, but I recommend asking the authors to ensure those entries are independently verifiable and that claims drawn from them are not overstated. The contradictions in Tables II and IV should be checked against primary sources before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful organizational survey of LLM/VLM crash detection, but its comparative tables and deployment-gap argument are built on a set of transcribed numbers that conflict with each other in visible places. It deserves a serious referee, but only with a fixing pass attached.\n\nWhat's good: the taxonomy in §II-E — fusion level, prompt strategy, LLM role, input granularity, detection focus — is a sensible way to carve up the recent literature. The paper correctly identifies the central tension (reasoning quality vs. deployability), and the discussion of architecture families in §IV gives a newcomer a fair sense of the design space. Table V attempts an honest deployment-readiness assessment, and the 'Note on Computational Estimates' at least concedes that the latency figures are approximate.\n\nThe soft spots are real and match the stress-test note. DAD is '1,500+ videos' in Table II and §VII-A but '~620 dashcam videos' in §X-A. CrashLLM is credited with 'over 92% accuracy' in §II-D and §II-E1 while Table IV and §VI-A report macro-F1 34.9%→53.8% on CrashEvent; the paper never reconciles these as different metrics. Video-LLaMA is said in §IV-A to detect crash anomalies on UCF-Crime, but §VI-B explicitly says the original paper reports no crash-specific benchmarks. Section VII-F asserts specific failure rates — 40% occlusion misses, 25% weather degradation, 30% lighting reduction, 60% adversarial false negatives — with no source; only the 12–16% OOD drop carries a citation. Reference [59] has 'Unknown' as its author and no URL, making the dynamic-template example unverifiable.\n\nNone of these is a fatal conceptual flaw; the taxonomy and the deployment-gap argument survive the scrubbing. But the survey's value is precisely as a quantitative map, and if any of those rows are wrong, the reader is misled. The 'first comprehensive analysis' claim is also stronger than the evidence; there is no systematic search protocol, and the review leans heavily on two works by the authors themselves ([57], [86]) as evidence of the field's progress. That is not a sin, but it should be tempered.\n\nThe paper is for newcomers who want a readable map of the field and for practitioners scoping whether LLM-based crash detection is worth building on. They should verify primary sources before quoting any number. I would send it to peer review with a major-revision recommendation focused on fixing the internal contradictions and sourcing the failure statistics.","headline":"Useful taxonomy, but the transcribed numbers don't hang together; needs a fixing pass before it can be trusted as a reference.","tokens_in":29181,"tokens_out":3729,"would_cite":false,"duration_ms":35290,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM-based video crash detection has shifted from pixel anomaly detection to context-aware event interpretation, but it still faces a deployment gap.","keywords":["crash detection","large language models","vision-language models","video understanding","multimodal learning","autonomous driving","video anomaly detection","survey"],"falsifier":"A reader could go to the original DAD and CrashLLM papers and check the transcribed numbers: if DAD's true size is about 620 videos rather than 1,500+, or if CrashLLM's reported performance is macro-F1 53.8% rather than 'over 92% accuracy', then the comparison tables and the deployment-readiness labels derived from them would need revision.","tokens_in":28137,"feed_emoji":"🚗","tokens_out":5716,"duration_ms":65603,"temperature":0.7,"pith_summary":"This survey argues that crash detection from video has entered a new phase: instead of just flagging pixels that look anomalous, systems built on large language models and vision-language models interpret crashes as events, generate explanations, and reason about causes. The paper's central claim is that this shift brings real capabilities but also a hard deployment gap, because the models with the richest reasoning are too slow, too memory-hungry, and too prone to hallucination for real-time safety-critical use. It organizes the field into a taxonomy of fusion strategies, LLM roles, and architectures, maps the available datasets, and compares methods. If the survey's reading is right, the next research step is not higher detection accuracy but standardized latency reporting, causal-annotation datasets, and efficiency-first architectures.","feed_headline":"LLM crash detection can explain, not yet deploy","feed_subtitle":"Reasoning-rich crash models still miss the sub-100ms and low-memory bar for real cars, the survey finds.","key_machinery":"The carrying machinery is a three-axis taxonomy: fusion level (early token concatenation, late fusion with separate heads, or cross-attention), prompt strategy (static prompts, dynamic event templates, learned visual questions), and LLM role (passive captioning, active QA/inference, generative reasoning). Around this taxonomy the survey builds an architecture classification (visual encoder plus LLM decoder, frozen LLM plus learned adapter, joint vision-language pretraining) and an accuracy-versus-latency trade-off diagram. The taxonomy does the argument's load-bearing work: it converts scattered recent systems into a design space, then reads the deployment gap off that space.","core_discovery":"The survey's central finding is that current LLM/VLM crash detection systems cluster around a few design choices, and that both their capabilities and their deployment barriers follow from those choices. Systems like VERA and Holmes-VAD can verbalize why a crash happened, not merely that one occurred, but they pay for this with latencies in the hundreds of milliseconds and multi-gigabyte memory footprints. By contrast, the only systems the survey labels deployment-ready in real time, such as LA V AD and CRASH, are training-free or lightweight and give up most of the contextual reasoning that makes LLMs valuable. The paper also claims that the biggest data bottleneck is the lack of large-scale datasets with explicit causal chains linking fine-grained events to crash outcomes, which blocks progress from correlation to true causal reasoning.","pith_inferences":["If the survey's accuracy-versus-latency picture holds, a hierarchical pipeline—a cheap frame-level detector that flags candidate clips, with LLM reasoning applied only over those clips—is the most direct way to close the deployment gap; the paper lists this as a solution direction but does not test it.","The survey's critique of heterogeneous metrics implies a concrete next step: a standard crash-detection leaderboard reporting latency, accuracy, and explanation quality on fixed hardware, which would let the field quantify its trade-offs instead of asserting them.","Because the survey relies on transcribed numbers, re-verifying each method's reported metric against its original paper is a cheap first test of the survey's map; the visible internal conflicts suggest this re-check is needed before the tables are used to guide decisions."],"forward_implications":["If the survey's map is correct, no current 7-billion-parameter crash detection LLM can run within the sub-100ms latency budgets of safety-critical vehicle systems; only lightweight training-free models are currently labeled deployment-ready.","The reported absence of large-scale datasets linking fine-grained temporal events to crash outcomes means current models cannot move from correlation to causal reasoning, and this is the field's main data bottleneck.","The taxonomy implies that choosing a fusion strategy is choosing a point on the accuracy-latency-explainability trade-off, so future systems should be compared along all three axes rather than by accuracy alone.","The survey's call to report latency alongside accuracy, and to evaluate across datasets, would become standard practice if its assessment is accepted."],"supporting_citations":[{"why":"Supplies the crash-as-language-task baseline with the macro-F1 improvement (34.9% to 53.8%) that anchors the heavyweight class and the deployment critique.","marker":"[53]"},{"why":"Provides VERA, the verbalized anomaly-detection method whose 86.55% AUC on DAD and UCF-Crime anchors the medium-scale reasoning end of the trade-off.","marker":"[54]"},{"why":"Introduces Video-LLaMA, the early-fusion video-language architecture that the taxonomy repeatedly uses as an example and that lacks crash-specific benchmarks.","marker":"[38]"},{"why":"Provides LA V AD, the training-free lightweight method whose 85.0% AUC and deployment-ready label anchor the efficiency end of the spectrum.","marker":"[60]"},{"why":"DAD is one of the two primary crash datasets used for training and evaluation; its size figures appear in the comparison tables and in the conclusion.","marker":"[61]"},{"why":"UCF-Crime is the anomaly benchmark used to compare VERA, Holmes-VAD, LA V AD, and Video-LLaMA, and its road-accident subset underpins several reported AUC numbers.","marker":"[64]"},{"why":"ScVLM and SHRP 2 NDS provide the hybrid supervised-contrastive method and naturalistic driving dataset used to argue for moving beyond binary crash detection.","marker":"[65]"},{"why":"Holmes-VAD supplies the explainable heavyweight method with 86.5% AUC on UCF-Crime, used to illustrate the memory-wall limitation.","marker":"[85]"}],"fun_headline_variants":["Crash AI: Explainable models too slow for cars","LLM crash detection: reason vs react tradeoff","Survey: Crash reasoning AI misses deployment bar","Crash AI survey: lack of causal data blocks progress","Why LLM crash AI can't yet hit real-time"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's comparisons rest on the accuracy, latency, and dataset-size numbers it transcribes from the papers it cites, and those transcriptions are not internally consistent: DAD appears as 1,500+ videos in Table II but about 620 in the conclusion, and CrashLLM appears as over 92% accuracy in Section II-D but macro-F1 53.8% in Table IV.","fun_headline_variants_meta":{"raw":{"variants":["Crash AI: Explainable models too slow for cars","LLM crash detection: reason vs react tradeoff","Survey: Crash reasoning AI misses deployment bar","Crash AI survey: lack of causal data blocks progress","Why LLM crash AI can't yet hit real-time"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000355,"raw_usage":{"total_tokens":1849,"prompt_tokens":788,"completion_tokens":1061,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":404,"completion_tokens_details":{"reasoning_tokens":984}},"tokens_in":404,"tokens_out":1061,"duration_ms":12394,"temperature":1.0,"reasoning_tokens":984,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:38:18.112988+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could go to the original DAD and CrashLLM papers and check the transcribed numbers: if DAD's true size is about 620 videos rather than 1,500+, or if CrashLLM's reported performance is macro-F1 53.8% rather than 'over 92% accuracy', then the comparison tables and the deployment-readiness labels derived from them would need revision.","supporting_citations":[{"cited_title":"Pentagon relation and Biedenharn-Elliott identity","cited_arxiv_id":"2402.16682","evidence_quote":"Supplies the crash-as-language-task baseline with the macro-F1 improvement (34.9% to 53.8%) that anchors the heavyweight class and the deployment critique."},{"cited_title":"Harnessing Neuron Stability to Improve DNN Verification","cited_arxiv_id":"2401.14412","evidence_quote":"Provides VERA, the verbalized anomaly-detection method whose 86.55% AUC on DAD and UCF-Crime anchors the medium-scale reasoning end of the trade-off."},{"cited_title":"Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages","cited_arxiv_id":"2402.12204","evidence_quote":"Provides LA V AD, the training-free lightweight method whose 85.0% AUC and deployment-ready label anchor the efficiency end of the spectrum."},{"cited_title":"Dad: A dashcam accident dataset,","cited_arxiv_id":null,"evidence_quote":"DAD is one of the two primary crash datasets used for training and evaluation; its size figures appear in the comparison tables and in the conclusion."},{"cited_title":"Bulk-Boundary Correspondence in the Quantum Hall Effect","cited_arxiv_id":"1801.03759","evidence_quote":"UCF-Crime is the anomaly benchmark used to compare VERA, Holmes-VAD, LA V AD, and Video-LLaMA, and its road-accident subset underpins several reported AUC numbers."},{"cited_title":"ScVLM: Enhancing Vision-Language Model for Safety-Critical Event Understanding","cited_arxiv_id":"2410.00982","evidence_quote":"ScVLM and SHRP 2 NDS provide the hybrid supervised-contrastive method and naturalistic driving dataset used to argue for moving beyond binary crash detection."}],"review_version":1}