REVIEW 4 cited by
Hear Me Out: Fusional Approaches for Audio Augmented Temporal Action Localization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
State of the art architectures for untrimmed video Temporal Action Localization (TAL) have only considered RGB and Flow modalities, leaving the information-rich audio modality totally unexploited. Audio fusion has been explored for the related but arguably easier problem of trimmed (clip-level) action recognition. However, TAL poses a unique set of challenges. In this paper, we propose simple but effective fusion-based approaches for TAL. To the best of our knowledge, our work is the first to jointly consider audio and video modalities for supervised TAL. We experimentally show that our schemes consistently improve performance for state of the art video-only TAL approaches. Specifically, they help achieve new state of the art performance on large-scale benchmark datasets - ActivityNet-1.3 (54.34 mAP@0.5) and THUMOS14 (57.18 mAP@0.5). Our experiments include ablations involving multiple fusion schemes, modality combinations and TAL architectures. Our code, models and associated data are available at https://github.com/skelemoa/tal-hmo.
Forward citations
Cited by 4 Pith papers
-
UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization
Skip-scanning Mamba with unified audio-visual sequences reaches 63.4% AP@0.95 on LAV-DF and 63.58% mAP on AV-Deepfake1M by regularizing toward low/mid-frequency forgery cues.
-
EVAS: Efficient Multimodal Temporal Forgery Localization via Audio-Visual Synergy and Steered Boundary Calibration
EVAS localizes sparse multimodal forgeries via multi-stage audio-visual synergy and decoupled boundary-aware refinement, reporting SOTA AP and AR on LAV-DF, AV-Deepfake1M, and TVIL.
-
Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization
UniCaCLF trains temporal instant features to separate real from forged moments relative to each sample's global context, achieving state-of-the-art temporal forgery localization on five public datasets.
-
DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding
DEL is a new audio-visual transformer framework that reports state-of-the-art temporal action localization on UnAV-100, THUMOS14, ActivityNet 1.3, and EPIC-Kitchens-100.
Discussion (0). Continue with ORCID to comment.