Pith. sign in

REVIEW 5 cited by

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.19877 v1 pith:TRXQETWH submitted 2025-05-26 cs.CV

classification cs.CV
keywords anomalyreasoningvad-r1videomllmsproposeanomaliescapability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in reasoning capability of Multimodal Large Language Models (MLLMs) demonstrate its effectiveness in tackling complex visual tasks. However, existing MLLM-based Video Anomaly Detection (VAD) methods remain limited to shallow anomaly descriptions without deep reasoning. In this paper, we propose a new task named Video Anomaly Reasoning (VAR), which aims to enable deep analysis and understanding of anomalies in the video by requiring MLLMs to think explicitly before answering. To this end, we propose Vad-R1, an end-to-end MLLM-based framework for VAR. Specifically, we design a Perception-to-Cognition Chain-of-Thought (P2C-CoT) that simulates the human process of recognizing anomalies, guiding the MLLM to reason anomaly step-by-step. Based on the structured P2C-CoT, we construct Vad-Reasoning, a dedicated dataset for VAR. Furthermore, we propose an improved reinforcement learning algorithm AVA-GRPO, which explicitly incentivizes the anomaly reasoning capability of MLLMs through a self-verification mechanism with limited annotations. Experimental results demonstrate that Vad-R1 achieves superior performance, outperforming both open-source and proprietary models on VAD and VAR tasks. Codes and datasets will be released at https://github.com/wbfwonderful/Vad-R1.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    MAVEN pipeline generates multi-scale spatio-temporal event descriptions from videos using agentic adaptation and refinement, then produces training data that lets a fine-tuned 8B model outperform Gemini baselines on p...

  2. ESOM: Efficiently Understanding Streaming Video Anomalies with Open-world Dynamic Definitions

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    ESOM is a training-free streaming model for open-world video anomaly detection with dynamic definitions that achieves real-time single-GPU efficiency and state-of-the-art results on a new benchmark.

  3. O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An object-centric, training-free agentic pipeline that tracks object state changes and reasons over them with a vision-language model achieves strong video-level AUROC on Phys-AD, LiquidAD, and IPAD, while producing i...

  4. MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

    cs.CV 2026-05 conditional novelty 6.0 of 10

    An agentic three-stage video annotation pipeline with an MSTED intermediate and top-down domain adaptation produces CoT training data that lifts Cosmos-Reason2 past Gemini on traffic event reasoning.

  5. DAMS:Dual-Branch Adaptive Multiscale Spatiotemporal Framework for Video Anomaly Detection

    cs.CV 2025-07 conditional novelty 4.0 of 10

    DAMS, a dual-branch architecture fusing adaptive temporal pyramids, CBAM attention, and CLIP pseudo-labels, reports 94.67 AUC on UCF-Crime and 84.00 AP on XD-Violence for weakly supervised video anomaly detection.

Pith tools