Pith. sign in

REVIEW 1 cited by

GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.07322 v2 pith:YN7ONYIE submitted 2023-12-12 cs.CV

classification cs.CV
keywords actionsgenhowtoobjecttransformationsimageimagesinitialinstructional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We address the task of generating temporally consistent and physically plausible images of actions and object state transformations. Given an input image and a text prompt describing the targeted transformation, our generated images preserve the environment and transform objects in the initial image. Our contributions are threefold. First, we leverage a large body of instructional videos and automatically mine a dataset of triplets of consecutive frames corresponding to initial object states, actions, and resulting object transformations. Second, equipped with this data, we develop and train a conditioned diffusion model dubbed GenHowTo. Third, we evaluate GenHowTo on a variety of objects and actions and show superior performance compared to existing methods. In particular, we introduce a quantitative evaluation where GenHowTo achieves 88% and 74% on seen and unseen interaction categories, respectively, outperforming prior work by a large margin.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Human Skill Generators at Key-Step Levels

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A benchmark and framework for generating key-step video clips of human skills from one initial image and a skill description.

Pith tools