Pith. sign in

REVIEW 4 cited by

Hear Me Out: Fusional Approaches for Audio Augmented Temporal Action Localization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.14118 v4 pith:5CBGFORV submitted 2021-06-27 cs.CV cs.MM

classification cs.CVcs.MM
keywords audioactionapproachesstatearchitecturesfusionlocalizationmodalities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

State of the art architectures for untrimmed video Temporal Action Localization (TAL) have only considered RGB and Flow modalities, leaving the information-rich audio modality totally unexploited. Audio fusion has been explored for the related but arguably easier problem of trimmed (clip-level) action recognition. However, TAL poses a unique set of challenges. In this paper, we propose simple but effective fusion-based approaches for TAL. To the best of our knowledge, our work is the first to jointly consider audio and video modalities for supervised TAL. We experimentally show that our schemes consistently improve performance for state of the art video-only TAL approaches. Specifically, they help achieve new state of the art performance on large-scale benchmark datasets - ActivityNet-1.3 (54.34 mAP@0.5) and THUMOS14 (57.18 mAP@0.5). Our experiments include ablations involving multiple fusion schemes, modality combinations and TAL architectures. Our code, models and associated data are available at https://github.com/skelemoa/tal-hmo.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Skip-scanning Mamba with unified audio-visual sequences reaches 63.4% AP@0.95 on LAV-DF and 63.58% mAP on AV-Deepfake1M by regularizing toward low/mid-frequency forgery cues.

  2. EVAS: Efficient Multimodal Temporal Forgery Localization via Audio-Visual Synergy and Steered Boundary Calibration

    cs.CV 2026-07 conditional novelty 6.0 of 10

    EVAS localizes sparse multimodal forgeries via multi-stage audio-visual synergy and decoupled boundary-aware refinement, reporting SOTA AP and AR on LAV-DF, AV-Deepfake1M, and TVIL.

  3. Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization

    cs.CV 2025-06 conditional novelty 6.0 of 10

    UniCaCLF trains temporal instant features to separate real from forged moments relative to each sample's global context, achieving state-of-the-art temporal forgery localization on five public datasets.

  4. DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    DEL is a new audio-visual transformer framework that reports state-of-the-art temporal action localization on UnAV-100, THUMOS14, ActivityNet 1.3, and EPIC-Kitchens-100.

Pith tools