A benchmark and framework for generating key-step video clips of human skills from one initial image and a skill description.
V AST: A vision-audio-subtitle-text omni-modality foundation model and dataset
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Learning Human Skill Generators at Key-Step Levels
A benchmark and framework for generating key-step video clips of human skills from one initial image and a skill description.