Pith. sign in

REVIEW 4 cited by

EditWorld: Simulating World Dynamics for Instruction-Following Image Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.14785 v1 pith:ZCYWYTPZ submitted 2024-05-23 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords editingimageworldeditworlddatasetdynamicsinstructionsexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models have significantly improved the performance of image editing. Existing methods realize various approaches to achieve high-quality image editing, including but not limited to text control, dragging operation, and mask-and-inpainting. Among these, instruction-based editing stands out for its convenience and effectiveness in following human instructions across diverse scenarios. However, it still focuses on simple editing operations like adding, replacing, or deleting, and falls short of understanding aspects of world dynamics that convey the realistic dynamic nature in the physical world. Therefore, this work, EditWorld, introduces a new editing task, namely world-instructed image editing, which defines and categorizes the instructions grounded by various world scenarios. We curate a new image editing dataset with world instructions using a set of large pretrained models (e.g., GPT-3.5, Video-LLava and SDXL). To enable sufficient simulation of world dynamics for image editing, our EditWorld trains model in the curated dataset, and improves instruction-following ability with designed post-edit strategy. Extensive experiments demonstrate our method significantly outperforms existing editing methods in this new task. Our dataset and code will be available at https://github.com/YangLing0818/EditWorld

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Autoregressive Images Watermarking through Lexical Biasing: An Approach Resistant to Regeneration Attack

    cs.CR 2025-06 conditional novelty 6.0 of 10

    LBW embeds watermarks into autoregressive image token maps by biasing token sampling toward a secret green list and detects them with a z-test on green-token counts.

  2. OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data

    cs.CV 2025-05 conditional novelty 6.0 of 10

    OmniConsistency is a style-agnostic consistency module for Flux that preserves structure and details during stylization with arbitrary LoRAs, reaching GPT-4o-level content consistency.

  3. ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A released 6.4 million pair dataset and 613 sample benchmark for instruction-guided image editing of non-rigid motions, plus a Flux.1-dev based baseline that outperforms open-source methods on the new benchmark.

  4. CausalDiffTab: Mixed-Type Causal-Aware Diffusion for Tabular Data Generation

    cs.CL 2025-06 reject novelty 4.0 of 10

    CausalDiffTab adds a causal graph learned from the data as a regularizer to a mixed-type diffusion tabular generator and claims gains on seven datasets, but the evaluation contains a metric direction error.

Pith tools