TeMTG fuses CLIP/CLAP text embeddings with audio and visual features and applies K-hop graph attention to achieve state of the art segment-level event parsing on the LLP dataset.
ColeaF: A Contrastive-Collaborative Learning Framework for Weakly Supervised Audio- Visual Video Parsing
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.MM 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing
TeMTG fuses CLIP/CLAP text embeddings with audio and visual features and applies K-hop graph attention to achieve state of the art segment-level event parsing on the LLP dataset.