Pith. sign in

REVIEW 2 cited by

Learning Weakly Supervised Audio-Visual Violence Detection in Hyperbolic Space

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.18797 v3 pith:QNSPD6LG submitted 2023-05-30 cs.CV

classification cs.CV
keywords spacehyperbolicframeworkaudio-visualdetectioneffectivelyfeaturefusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, the task of weakly supervised audio-visual violence detection has gained considerable attention. The goal of this task is to identify violent segments within multimodal data based on video-level labels. Despite advances in this field, traditional Euclidean neural networks, which have been used in prior research, encounter difficulties in capturing highly discriminative representations due to limitations of the feature space. To overcome this, we propose HyperVD, a novel framework that learns snippet embeddings in hyperbolic space to improve model discrimination. Our framework comprises a detour fusion module for multimodal fusion, effectively alleviating modality inconsistency between audio and visual signals. Additionally, we contribute two branches of fully hyperbolic graph convolutional networks that excavate feature similarities and temporal relationships among snippets in hyperbolic space. By learning snippet representations in this space, the framework effectively learns semantic discrepancies between violent and normal events. Extensive experiments on the XD-Violence benchmark demonstrate that our method outperforms state-of-the-art methods by a sizable margin.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HyPCV-Former: Hyperbolic Spatio-Temporal Transformer for 3D Point Cloud Video Anomaly Detection

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    HyPCV-Former embeds point cloud video features in Lorentzian hyperbolic space and uses hyperbolic attention to improve video anomaly detection on two benchmarks.

  2. The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM

    cs.CV 2025-07 conditional novelty 4.0 of 10

    The paper organizes VAD methods into a five-dimension framework spanning task objective, modality, input, architecture, and optimization, with emphasis on MLLM/LLM-era work.

Pith tools