TeMTG fuses CLIP/CLAP text embeddings with audio and visual features and applies K-hop graph attention to achieve state of the art segment-level event parsing on the LLP dataset.
Anchor-aware Deep Metric Learning for Audio-visual Retrieval
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.MM 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing
TeMTG fuses CLIP/CLAP text embeddings with audio and visual features and applies K-hop graph attention to achieve state of the art segment-level event parsing on the LLP dataset.