CoM sequentially prompts a VLM over video, force/audio, and hand pose, yielding roughly threefold better extraction of task plans and control parameters, with real robots succeeding in 73% of trials.
Activitynet: A large-scale video benchmark for human activity understanding,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models
CoM sequentially prompts a VLM over video, force/audio, and hand pose, yielding roughly threefold better extraction of task plans and control parameters, with real robots succeeding in 73% of trials.