MoCLIP fine-tunes CLIP's text encoder on motion-text pairs using contrastive learning and a distillation loss, and swapping it into MoMask and BAMM improves R-Precision by about 1 to 2 percent while FID stays roughly the same.
Text2action: Generative adversarial synthesis from language to action, 2017
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MoCLIP: Motion-Aware Fine-Tuning and Distillation of CLIP for Human Motion Generation
MoCLIP fine-tunes CLIP's text encoder on motion-text pairs using contrastive learning and a distillation loss, and swapping it into MoMask and BAMM improves R-Precision by about 1 to 2 percent while FID stays roughly the same.