Cascaded co-attention between CLAP audio features and RoBERTa text features improves text-to-audio retrieval mAP by about 16 percent on Clotho and 15 percent on AudioCaps relative to the authors' earlier GPTtar method.
Dynamic modality interaction modeling for image-text retrieval,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Language-based Audio Retrieval with Co-Attention Networks
Cascaded co-attention between CLAP audio features and RoBERTa text features improves text-to-audio retrieval mAP by about 16 percent on Clotho and 15 percent on AudioCaps relative to the authors' earlier GPTtar method.