A four-modality contrastive model with token-level alignment improves text, video, and newly introduced audio-to-motion retrieval on HumanML3D and KIT-ML.
Cross-modal quantization for co-speech gesture generation,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space
A four-modality contrastive model with token-level alignment improves text, video, and newly introduced audio-to-motion retrieval on HumanML3D and KIT-ML.