REVIEW 5 cited by
DVD: A Comprehensive Dataset for Advancing Violence Detection in Real-World Scenarios
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Violence Detection (VD) has become an increasingly vital area of research. Existing automated VD efforts are hindered by the limited availability of diverse, well-annotated databases. Existing databases suffer from coarse video-level annotations, limited scale and diversity, and lack of metadata, restricting the generalization of models. To address these challenges, we introduce DVD, a large-scale (500 videos, 2.7M frames), frame-level annotated VD database with diverse environments, varying lighting conditions, multiple camera sources, complex social interactions, and rich metadata. DVD is designed to capture the complexities of real-world violent events.
Forward citations
Cited by 5 Pith papers
-
Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework
A multi-agent iterative-questioning framework plus a 605-video benchmark for detecting developmentally inappropriate risks in AI-generated children's videos.
-
EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding
A new egocentric safety benchmark shows current video-language models can describe scenes well but fail at multi-step causal reasoning about blind spots and covert actions.
-
Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation
Distance-aware soft prompts over a 3×3 emotion grid with CLIP text prototypes and audio-visual GRU fusion achieve CCC_mean 0.5361 on Aff-Wild2, beating only the paper's self-defined baselines.
-
Team RAS in 9th ABAW Competition: Multimodal Compound Expression Recognition Approach
A six-modality zero-shot pipeline with CLIP, Qwen-VL, WavLM, Mamba, and new fusion/aggregation modules reports F1 scores of 46.95 (AffWild2), 49.02 (AFEW), and 34.85 (C-EXPR-DB) without target-domain fine-tuning.
-
TAGF: Time-aware Gated Fusion for Multimodal Valence-Arousal Estimation
TAGF adds a BiLSTM-based gate that reweights recursive cross-attention outputs for valence-arousal prediction, with results slightly below several existing methods on Aff-Wild2.
Discussion (0). Sign in to comment.