Pith. sign in

GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We address the task of generating temporally consistent and physically plausible images of actions and object state transformations. Given an input image and a text prompt describing the targeted transformation, our generated images preserve the environment and transform objects in the initial image. Our contributions are threefold. First, we leverage a large body of instructional videos and automatically mine a dataset of triplets of consecutive frames corresponding to initial object states, actions, and resulting object transformations. Second, equipped with this data, we develop and train a conditioned diffusion model dubbed GenHowTo. Third, we evaluate GenHowTo on a variety of objects and actions and show superior performance compared to existing methods. In particular, we introduce a quantitative evaluation where GenHowTo achieves 88% and 74% on seen and unseen interaction categories, respectively, outperforming prior work by a large margin.

citation-role summary

baseline 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

baseline 1

polarities

baseline 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • Learning Human Skill Generators at Key-Step Levels cs.CV · 2025-02-12 · conditional · none · ref 41 · internal anchor

    A benchmark and framework for generating key-step video clips of human skills from one initial image and a skill description.