{"id":"cc4d8542-6dc7-41a7-8aa4-036f5cfa19ca","arxiv_id":"2507.21649","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper organizes VAD methods into a five-dimension framework spanning task objective, modality, input, architecture, and optimization, with emphasis on MLLM/LLM-era work.","lead":"This paper surveys video anomaly detection (VAD) from traditional deep networks to LLM/MLLM-based methods and proposes a five-part framework to organize all approaches. It claims to be the first comprehensive VAD survey focused on large language models, which is useful for newcomers and researchers tracking the field's shift to semantic understanding.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed five-dimension 'unified' framework is not applied to OSVAD/OVVAD; Sections X and XI explicitly decline to categorize these methods, so the exhaustiveness claim in Section IV-A is internally contradicted.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the framework is claimed to be unified and exhaustive, but the paper itself declines to apply it to OSVAD and OVVAD. This is the most direct threat to the central contribution because it is an internal contradiction rather than a scope or novelty disagreement. Section IV-A claims compatibility with all VAD tasks; Sections X and XI explicitly say those methods are described without categorizing them in the proposed framework. The concern is load-bearing: a reader relying on the paper's framing would believe every VAD method can be compared in the five dimensions, but the paper's own organization shows otherwise. The paper still has real strengths—recent MLLM/LLM coverage, useful performance tables, and a plausible partition of closed-set paradigms—so rejection is not warranted. The appropriate remedy is a revised scope statement or a completed mapping of OSVAD/OVVAD methods onto the five dimensions. Since the reader already issued CONDITIONAL and my read does not move that verdict, UNCHANGED is the honest recommendation. The concrete test, classifying Section X and XI methods through Fig. 5, would settle whether the framework is salvageable as stated.","tokens_in":47683,"tokens_out":3490,"duration_ms":40751,"concrete_test":"Independently take the decision tree in Fig. 5 and classify every method listed in Sections X and XI into the five dimensions, recording whether each dimension can be filled from an existing leaf of the tree without an 'other' placeholder. If any method, such as few-shot VAD under Task Objective or 'semantic openness' under Task Modality, requires a new branch, the exhaustiveness claim fails. Complements: require a supplementary table mapping all surveyed methods in Sections V–XI onto the five dimensions; if OSVAD and OVVAD rows stay blank or use 'not applicable,' the unified-framework claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the five-dimension framework (Task Objective, Task Modality, Video Input, Model Architecture, Model Optimization; Section IV-A, Fig. 5) is 'compatible with all existing types of VAD tasks' and establishes a unified analytical system for both DNN-based and VLM/LLM-based methods. But the paper explicitly exempts two of the seven task types it enumerates in Section III-B. Section X states that OSVAD methods are described 'rather than categorizing them according to the framework proposed earlier,' and Section XI repeats this for OVVAD. These are not fringe tasks: OSVAD and OVVAD are among the categories listed in the Introduction and Section III-B, and the Conclusion claims the survey covers open-set and open-vocabulary paradigms. If the framework cannot accommodate methods the paper itself identifies as VAD, the 'compatible with all existing types' claim fails as an internal-consistency matter, not merely a scope matter. A weaker claim—for example, that the framework covers DNN and LLM/MLLM methods sharing closed-set grounding or understanding objectives—would be supported by the content, but that is not what Section IV-A claims. The paper also does not provide a per-method mapping onto the five dimensions for the methods it does cover, so the framework's operational value as a 'common analytical language' is asserted rather than demonstrated. The contradiction is specific, located in the paper's own text, and directly touches the strongest claimed contribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews video anomaly detection (VAD) from traditional DNN-based methods to modern LLM/MLLM-based approaches. It proposes a 'unified analytical framework' with five dimensions—Task Objective, Task Modality, Video Input, Model Architecture, and Model Optimization (Section IV-A, Fig. 5)—claimed to be compatible with all existing types of VAD tasks. The paper enumerates seven VAD task types under different supervision settings (Section III-B), surveys datasets and metrics, and then reviews methods for each task type, with dedicated emphasis on training-free and instruction-tuned large-model paradigms, comparative performance tables, challenges, and future research directions.","tokens_in":47963,"tokens_out":3810,"duration_ms":46943,"significance":"If the five-dimension framework were consistently instantiated, it could provide the field with a common analytical language for positioning and comparing both classical and LLM-based VAD methods. The survey's coverage of recent MLLM/LLM VAD works is timely and useful, and the inclusion of performance tables and a comparison of prior surveys (Table I) adds practical value. However, the central claim of framework compatibility is currently undermined by internal contradictions and a lack of demonstrated application, so the contribution is promising but not yet fully realized.","major_comments":[{"comment":"The central claim in Section IV-A that the five-dimension framework is 'compatible with all existing types of VAD tasks' (Fig. 5) is directly contradicted by Sections X and XI. Both sections explicitly state that the few existing OSVAD and OVVAD methods are described 'rather than categorizing them according to the framework proposed earlier.' Since OSVAD and OVVAD are enumerated as two of the seven task types in Section III-B and are included in the Conclusion's coverage statement, the framework as presented is not exhaustive. Please either apply the five-dimension analysis to these sections, for example by adding a table that maps representative OSVAD/OVVAD methods onto the five dimensions, or revise the claim to state that the framework covers closed-set and semantically grounded DNN/LLM methods, leaving open-set and open-vocabulary paradigms as boundary cases.","section":"§IV-A vs §X and §XI"},{"comment":"The framework is asserted but not operationalized. Section IV-A claims that the framework establishes a 'unified analytical system,' but the subsequent method sections do not provide an explicit per-method mapping onto the five dimensions. For example, Section V includes subsections for Video Input, Model Architecture, and Model Optimization, but the 'Task Objective' and 'Task Modality' dimensions are not systematically coded for the surveyed methods, and no worked example or summary table places representative methods into the five-dimensional space. Without such a demonstration, the framework's practical value as a common analytical language remains unsubstantiated; please add a mapping for at least a representative subset of methods.","section":"§IV-A and §§V-IX"},{"comment":"The quantitative trend claims in Fig. 1 (increasing proportion of MLLM/LLM-based publications) and Fig. 2 (performance saturation of DNN methods) are not accompanied by a description of the search strategy, databases, inclusion criteria, or the selection of representative methods and benchmarks. The text mentions only that papers were collected from Google Scholar (Section I), which is insufficient for verification. Please specify the methodology behind these figures or soften the claims to clearly indicate that they are illustrative rather than systematic measurements.","section":"§I, Figs. 1-2"}],"minor_comments":[{"comment":"There is a grammatical error: 'we proposes' should be 'we propose.'","section":"§IV-A"},{"comment":"The spacing between 'V' and 'AD' is inconsistent throughout the manuscript (e.g., 'V AD' vs. 'VAD'); please unify the notation.","section":"§III-C and elsewhere"},{"comment":"The method name 'Nomral-Pose V AD' appears to be a typo for 'Normal-Pose VAD'.","section":"Table III"},{"comment":"The metric definitions contain formatting artifacts such as 'B Regiongt' and 'B Trackgt'; these should be cleaned up.","section":"§III-C"},{"comment":"The text refers to 'AnomalyRule' while the cited method is 'AnomalyRuler'; please ensure consistent naming.","section":"§VIII-D"},{"comment":"The claim of being the 'first comprehensive survey' of MLLM/LLM-based VAD should be reconciled with the prior work denoted as [24], which is listed in Table I as covering LLMs/VLMs in VAD; the distinction in scope and depth should be stated explicitly.","section":"Abstract and Introduction"}],"recommendation":"major_revision","confidential_remarks":"The internal contradiction between the claimed unified framework and the explicit exemption of OSVAD/OVVAD is real and load-bearing; the authors can address it either by extending the framework to those sections or by carefully restating the claim's scope. The survey itself is a useful compilation, so the manuscript is worth a major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it. The paper is a survey of video anomaly detection with an emphasis on LLM/MLLM methods. The useful part is the organized walk through semi-supervised, weakly supervised, unsupervised, training-free, instruction-tuning, open-set, and open-vocabulary VAD, with performance tables and representative method descriptions. As a newcomer's roadmap to recent large-model VAD literature it is genuinely handy.\n\nWhat is actually new is modest: a five-dimension taxonomy (task objective, task modality, video input, model architecture, model optimization) that tries to unify DNN and LLM/MLLM paradigms. That taxonomy is reasonable for closed-set-style methods and is a fine way to structure a survey. The survey does a good job describing training-free and instruction-tuned approaches, and the paradigm figures are clear.\n\nThe soft spots are in the claims, not the descriptions. The paper says the framework is 'compatible with all existing types of VAD tasks,' but Sections X and XI explicitly say OSVAD and OVVAD methods are described 'rather than categorizing them according to the framework proposed earlier.' That is a direct internal contradiction. The fix is easy: weaken the claim to cover methods that share closed-set grounding or understanding objectives, or actually map OSVAD/OVVAD onto the five dimensions. Also, the 'first comprehensive survey dedicated to MLLM/LLM-based approaches' is not true: Ding et al. [24] is cited in their own Table I as covering LLMs/VLMs in VAD. The temporal trend statistics in Figs 1 and 2 have no methodology or data release, so they are not reproducible. Minor: 'we proposes' and a few typos; the per-method mapping onto the five dimensions is asserted rather than demonstrated.\n\nNone of this is fatal to the survey's value as a broad informative overview. The framework is a taxonomy, not a new algorithm, and the authors do not hide that. If the overclaims are trimmed and the OSVAD/OVVAD contradiction resolved, this is a solid survey that a serious referee should engage with. I would send it to review, and I would probably cite it for coverage of training-free and instruction-tuned VAD methods.","headline":"A useful survey with real overclaims: the five-dimension framework is a reasonable organizational tool, but it explicitly does not cover two of the seven task types it lists, and the 'first comprehensive survey' claim is undercut by a cited prior survey.","tokens_in":48498,"tokens_out":1763,"would_cite":true,"duration_ms":20622,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that all video anomaly detection (VAD) methods, from classic DNN pipelines to MLLM/LLM systems, can be described and compared within one five-dimension framework built from task objective, task modality, video input…","keywords":["video anomaly detection","survey","unified classification framework","large language models","multimodal large language models","training-free VAD","instruction tuning","open-vocabulary anomaly detection"],"falsifier":"Check whether the open-set and open-vocabulary methods reviewed in Sections X and XI can each be assigned a complete profile under the five dimensions in Fig. 5; the compatibility claim is falsified if any method's defining property, such as zero-shot querying by arbitrary text, has no corresponding node in the framework.","tokens_in":3404,"feed_emoji":"🎥","tokens_out":1685,"duration_ms":97641,"temperature":0.7,"pith_summary":"This paper argues that video anomaly detection has entered a large-model era and that the field needs a common analytical language to describe both old and new methods. It proposes a five-dimension classification framework—task objective, task modality, video input, model architecture, and model optimization—intended to cover both traditional DNN-based VAD and newer VLM/LLM/MLLM-based systems. The paper claims that large-model methods shift anomaly detection from learning classification boundaries in visual feature space to reasoning in semantic space, which improves generalization and interpretability. It then uses that framework to review MLLM/LLM-driven VAD, including training-free and instruction-tuned paradigms, and to compare methods on standard benchmarks. If the framework holds, researchers gain a shared vocabulary for positioning any VAD method and for seeing which parts of the pipeline large models have transformed.","feed_headline":"Five axes now unify DNN and LLM video anomaly detection","feed_subtitle":"The survey's five-dimension framework gives every VAD method, old or new, one shared place to be compared.","key_machinery":"The load-bearing object is the five-dimension framework tree (Fig. 5), which decomposes any VAD method into Task Objective, Task Modality, Video Input, Model Architecture, and Model Optimization. Task Objective splits into Video Anomaly Grounding (locating when and where anomalies occur) and Video Anomaly Understanding (classifying, describing, and explaining them), while the other dimensions record what signals the method consumes, what network carries it, and how it is optimized. The companion mechanism is the distinction between learning classification boundaries in visual feature space (traditional DNN methods) and in semantic space (VLM/LLM methods, which use pretrained knowledge and prompt-based interaction). That distinction is what lets the survey place methods such as LAVAD, SUVAD, VERA, and instruction-tuned systems as new paradigm branches while still putting them on the same tree.","core_discovery":"The paper's central claim is that the apparent rupture between DNN-era and LLM-era VAD can be captured in one compatible classification system organized by five dimensions: Task Objective (anomaly grounding versus anomaly understanding, with understanding expanded to classification, question-answering, and causal analysis), Task Modality, Video Input, Model Architecture, and Model Optimization. Within this system, traditional methods are characterized by mapping annotation semantics into the visual feature space, while VLM/LLM methods construct classification boundaries directly in semantic space using pretrained knowledge and prompt interaction. The survey claims that this semantic-space shift is the driving force behind changes in data annotation, input modalities, model architecture, and task objectives, and that the resulting new categories—training-free VAD, instruction fine-tuning VAD, open-vocabulary VAD—can be compared with semi-supervised, weakly supervised, and unsupervised VAD under the same dimensions. It presents per-paradigm performance tables on benchmarks such as Ped2, Avenue, ShanghaiTech, UCF-Crime, and XD-Violence to support the compatibility claim.","pith_inferences":["An implication the paper leaves implicit is that the five dimensions could be applied to neighboring video-understanding tasks, such as temporal action localization or video question-answering, where the same task-objective and task-modality split appears.","The paper describes open-set and open-vocabulary VAD outside its five-dimension taxonomy, so a testable retrofit would be to formally map those methods onto the framework; success would confirm the unified claim, while failure would imply a sixth dimension is needed for open semantics.","If the semantic-space thesis is correct, performance on classic closed-set benchmarks will become a weaker signal, and zero-shot or open-vocabulary benchmarks will be the settings that actually discriminate between methods.","The framework could support a quantitative transformation map: annotate each surveyed method with its five-dimension profile and measure which dimensions large models actually change, turning a qualitative survey into a structured dataset."],"forward_implications":["Every VAD method can be positioned and compared along five shared dimensions, so future papers can state their contribution by specifying which dimension or node they change.","The shift to semantic-space classification means anomaly detection systems can offer explanations, answer questions, and adapt to new scenes without retraining, functions that traditional DNN methods did not provide.","Training-free VAD and instruction-tuned VAD become first-class paradigms, so benchmarks and baselines should expand to include zero-shot and text-guided settings rather than only frame-level AUC.","The survey's performance tables indicate that conventional DNN methods are saturating on simpler datasets, while MLLM/LLM methods are opening new task objectives such as causal analysis and anomaly question-answering.","Future progress will concentrate on the bottlenecks the paper lists: multimodal dataset scale, hallucination suppression, computational efficiency, generalization to unseen scenarios, and intrinsic anomaly reasoning in large models."],"supporting_citations":[{"why":"Supplies the reconstruction-based semi-supervised baseline whose paradigm the framework's Model Architecture and Model Optimization branches classify.","marker":"[7]"},{"why":"Introduces the weakly supervised MIL paradigm that anchors the weakly supervised VAD branch of the framework.","marker":"[9]"},{"why":"Defines the causal-understanding task objective and a learned frame selector, load-bearing for the anomaly-understanding and instruction-tuning taxonomy.","marker":"[10]"},{"why":"Provides an open-world instruction-tuning method whose knowledge-graph, pseudo-anomaly, and edge-deployment components populate the framework's optimization branches.","marker":"[11]"},{"why":"Supplies an instruction-tuning dataset and memory-bank mechanism used to illustrate semantic-space optimization in VAD.","marker":"[13]"},{"why":"Prior survey whose taxonomy of supervision settings is the baseline the five-dimension framework claims to unify and extend.","marker":"[22]"},{"why":"Shows a CLIP-based weakly supervised method constructing classification boundaries in semantic space, central to the paper's VLM paradigm analysis.","marker":"[33]"},{"why":"Defines training-free VAD, the key new paradigm category built on LLM/MLLM priors.","marker":"[34]"},{"why":"Refines training-free VAD with coarse-to-fine analysis and hallucination smoothing, serving as another framework exemplar.","marker":"[38]"},{"why":"First open-vocabulary VAD framework; its presence tests whether the five dimensions can absorb methods the survey handles outside its taxonomy.","marker":"[244]"}],"fun_headline_variants":["Five axes unify DNN and MLLM video anomaly detection","One framework spans DNN to MLLM for VAD","Five dimensions bridge traditional and LLM video anomaly detection","Unified VAD framework from DNN to MLLM","Survey: five-axis classification for all video anomaly detection"],"cache_read_input_tokens":50560,"weakest_assumption_plain":"The load-bearing premise is that every VAD method, including open-set and open-vocabulary ones, can be fully described by the five dimensions, so the framework is genuinely unified rather than a taxonomy for only the DNN and LLM methods the survey chooses to categorize.","fun_headline_variants_meta":{"raw":{"variants":["Five axes unify DNN and MLLM video anomaly detection","One framework spans DNN to MLLM for VAD","Five dimensions bridge traditional and LLM video anomaly detection","Unified VAD framework from DNN to MLLM","Survey: five-axis classification for all video anomaly detection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000231,"raw_usage":{"total_tokens":1542,"prompt_tokens":1060,"completion_tokens":482,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":676,"completion_tokens_details":{"reasoning_tokens":399}},"tokens_in":676,"tokens_out":482,"duration_ms":5178,"temperature":1.0,"reasoning_tokens":399,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T12:30:08.469399+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check whether the open-set and open-vocabulary methods reviewed in Sections X and XI can each be assigned a complete profile under the five dimensions in Fig. 5; the compatibility claim is falsified if any method's defining property, such as zero-shot querying by arbitrary text, has no corresponding node in the framework.","supporting_citations":[{"cited_title":"Adaptive anomaly detection network for unseen scene without fine-tuning,","cited_arxiv_id":null,"evidence_quote":"First open-vocabulary VAD framework; its presence tests whether the five dimensions can absorb methods the survey handles outside its taxonomy."}],"review_version":1}