A single-stage contrastive alignment of audio, visual, and text using both audio and visual captions improves audio-to-visual retrieval recall@10 from 0.27 to 0.52 on AVCaps.
Multimodal learn- ing with deep boltzmann machines,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
A single-stage contrastive alignment of audio, visual, and text using both audio and visual captions improves audio-to-visual retrieval recall@10 from 0.27 to 0.52 on AVCaps.