Visual features modulate audio Mamba temporal operators layer-wise with a learned routing gate and Adaptive AV-InfoNCE, claiming SOTA violence detection on NTU-CCTV and DVD.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.CV 2years
2026 2verdicts
CONDITIONAL 2roles
method 1polarities
use method 1representative citing papers
Trajectory-aware local temporal modeling before spatial aggregation plus CTC viseme supervision and EMA consistency lifts DVS-Lip word accuracy to 77.49%.
citing papers explorer
-
AViS-Mamba: Adaptive Visual Steering of Audio State-Space Dynamics for Violence Detection
Visual features modulate audio Mamba temporal operators layer-wise with a learned routing gate and Adaptive AV-InfoNCE, claiming SOTA violence detection on NTU-CCTV and DVD.
-
TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading
Trajectory-aware local temporal modeling before spatial aggregation plus CTC viseme supervision and EMA consistency lifts DVS-Lip word accuracy to 77.49%.