Pith. sign in

REVIEW 3 cited by

LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03849 v2 pith:UPBHCMUE submitted 2023-12-06 cs.CV

classification cs.CV
keywords actionegocentricimageframevisualgenerationinstructionlego
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generating instructional images of human daily actions from an egocentric viewpoint serves as a key step towards efficient skill transfer. In this paper, we introduce a novel problem -- egocentric action frame generation. The goal is to synthesize an image depicting an action in the user's context (i.e., action frame) by conditioning on a user prompt and an input egocentric image. Notably, existing egocentric action datasets lack the detailed annotations that describe the execution of actions. Additionally, existing diffusion-based image manipulation models are sub-optimal in controlling the state transition of an action in egocentric image pixel space because of the domain gap. To this end, we propose to Learn EGOcentric (LEGO) action frame generation via visual instruction tuning. First, we introduce a prompt enhancement scheme to generate enriched action descriptions from a visual large language model (VLLM) by visual instruction tuning. Then we propose a novel method to leverage image and text embeddings from the VLLM as additional conditioning to improve the performance of a diffusion model. We validate our model on two egocentric datasets -- Ego4D and Epic-Kitchens. Our experiments show substantial improvement over prior image manipulation models in both quantitative and qualitative evaluation. We also conduct detailed ablation studies and analysis to provide insights in our method. More details of the dataset and code are available on the website (https://bolinlai.github.io/Lego_EgoActGen/).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generative Physical AI in Vision: A Survey

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A structured review that categorizes physics-aware generative models in vision into explicit-simulation and implicit-learning families and proposes six integration paradigms.

  2. HANDI: Hand-Centric Text-and-Image Conditioned Video Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    HANDI generates hand-centric videos from an image and text prompt via automatic motion-area localization and a hand refinement loss.

  3. ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A video diffusion model generates scene-conditioned, step-by-step visual instructions from an input image and text prompts, trained on a new 0.6M-sequence dataset.

Pith tools