CLIP-AE uses CLIP features, audio, cross-attention fusion, and two self-supervised losses to improve unsupervised temporal action localization, with new top scores on THUMOS14 and ActivityNet v1.2.
Datasets THUMOS14consists of 200 validation and 213 test videos across 20 action classes, averaging 15 action segments per video
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization
CLIP-AE uses CLIP features, audio, cross-attention fusion, and two self-supervised losses to improve unsupervised temporal action localization, with new top scores on THUMOS14 and ActivityNet v1.2.