Pith. sign in

REVIEW 35 cited by

Guiding Instruction-based Image Editing via Multimodal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.17102 v2 pith:Q34QCN3K submitted 2023-09-29 cs.CV

classification cs.CV
keywords editingimageinstructionsinstruction-basedmgieexpressivehumanlanguage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Instruction-based image editing improves the controllability and flexibility of image manipulation via natural commands without elaborate descriptions or regional masks. However, human instructions are sometimes too brief for current methods to capture and follow. Multimodal large language models (MLLMs) show promising capabilities in cross-modal understanding and visual-aware response generation via LMs. We investigate how MLLMs facilitate edit instructions and present MLLM-Guided Image Editing (MGIE). MGIE learns to derive expressive instructions and provides explicit guidance. The editing model jointly captures this visual imagination and performs manipulation through end-to-end training. We evaluate various aspects of Photoshop-style modification, global photo optimization, and local editing. Extensive experimental results demonstrate that expressive instructions are crucial to instruction-based image editing, and our MGIE can lead to a notable improvement in automatic metrics and human evaluation while maintaining competitive inference efficiency.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 35 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. B-repLer: Language-guided Editing of CAD Models

    cs.GR 2025-08 conditional novelty 7.0 of 10

    B-repLer fine-tunes a multimodal LLM to locate edits and trains a transformer to modify a B-rep latent code, achieving 53.4% exact-match success on synthetic text-guided editing.

  2. Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

    cs.CV 2025-04 conditional novelty 7.0 of 10

    Diffusion timestep tokens give large language models a recursive visual language that improves unified multimodal comprehension and generation relative to spatial patch tokens.

  3. Instruction-based Image Manipulation by Watching How Things Move

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A diffusion editing model, InstructMove, is trained on video frame pairs annotated by MLLMs using spatial conditioning, enabling non-rigid edits and viewpoint changes.

  4. SAFIRE: Segment Any Forged Image Region

    cs.CV 2024-12 conditional novelty 7.0 of 10

    SAFIRE uses point prompting and feature clustering to partition forged images into multiple source regions, and reports state-of-the-art results on both binary forgery localization and a new multi-source partitioning task.

  5. Making Image Editing Easier via Adaptive Task Reformulation with Agentic Executions

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    An MLLM agent that profiles, routes, and reformulates image-editing queries into better-conditioned multi-step operations consistently improves existing editors on hard cases.

  6. Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing

    cs.CV 2025-09 conditional novelty 6.0 of 10

    VARIN uses a Location-aware Argmax Inversion pseudo-inverse of Gumbel-max sampling to extract editable discrete noises, enabling training-free prompt-guided editing for visual autoregressive models.

  7. After the Party: Navigating the Mapping From Color to Ambient Lighting

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A new paired dataset and Retinex-based network, RLN2, for restoring images captured under multiple colored light sources to ambient-normalized versions.

  8. Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing

    cs.CV 2025-07 conditional novelty 6.0 of 10

    X-Planner, an MLLM-based planner, decomposes complex image-editing instructions into localized sub-edits with masks and boxes, improving editing quality on standard and new complex benchmarks.

  9. Balancing Preservation and Modification: A Region and Semantic Aware Metric for Instruction-Based Image Editing

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A region and semantic aware metric for instruction-based image editing, built from LLM parsing plus detection, segmentation, and CLIP directional similarity, reports the highest human alignment among compared metrics.

  10. ORIDa: Object-centric Real-world Image Composition Dataset

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ORIDa is a public real-world dataset of 200 objects in 30,000+ images with multiple positions per scene, designed for object compositing training and evaluation.

  11. Image Editing As Programs with Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    IEAP decomposes complex editing instructions into atomic operations executed sequentially on a diffusion transformer, and reports state-of-the-art results on MagicBrush and AnyEdit.

  12. VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VideoREPA adds a token-relation distillation loss that aligns a text-to-video diffusion model's internal features with VideoMAEv2, boosting physical commonsense scores on VideoPhy and VideoPhy2.

  13. Does Feasibility Matter? Understanding the Impact of Feasibility on Synthetic Training Data

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Feasibility of synthetic images has little effect on fine-tuned CLIP accuracy; the edited attribute (background, color, or texture) matters more than whether the attribute is realistic.

  14. InstructAttribute: Fine-grained Object Attributes editing with Instruction

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new attention-manipulation scheme (SPAA) is used to generate 1.1M synthetic color/material edit pairs, and an InstructPix2Pix-style model trained on them outperforms prior instruction editors in the authors' benchmarks.

  15. $\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A new benchmark shows that as image editing instructions get more complex, models increasingly fail to preserve identity and quality, with models trained on synthetic data producing more artificial-looking results.

  16. Seeing World Dynamics in a Nutshell

    cs.CV 2025-02 conditional novelty 6.0 of 10

    NutWorld is a feed-forward model that represents a monocular video as structured dynamic 3D Gaussians in a canonical orthographic space, trained with depth and flow priors.

  17. OmniEraser: Remove Objects and Their Effects in Images with Paired Video-Frame Data

    cs.CV 2025-01 conditional novelty 6.0 of 10

    OmniEraser removes objects along with their shadows and reflections by conditioning a FLUX diffusion model on separate object and background latents, trained on a 134,281-sample video-derived dataset.

  18. Text2Relight: Creative Portrait Relighting with Text Guidance

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Text2Relight learns to re-light portrait photos from text prompts using a synthetic dataset generated by a three-stage pipeline.

  19. BrushEdit: All-In-One Image Inpainting and Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    BrushEdit couples a multimodal language model and an object detector with a single arbitrary-mask inpainting model to turn free-form text instructions into interactive, multi-turn image edits.

  20. HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    HumanEdit provides 5,751 human-annotated, high-resolution image editing pairs with masks and a six-type instruction taxonomy, plus baseline benchmark results.

  21. InsightEdit: Towards Better Instruction Following for Image Editing

    cs.CV 2024-11 conditional novelty 6.0 of 10

    InsightEdit uses multimodal language model features in a two-stream adapter to improve complex instruction following and background consistency in image editing.

  22. AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A large automatically collected image editing dataset with 25 editing types and a task-aware diffusion model trained on it achieve new state-of-the-art results on two standard image editing benchmarks.

  23. ColorEdit: Training-free Image-Guided Color editing with diffusion model

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A training-free method changes an object's color by aligning cross-attention value matrices with a reference image during early denoising steps, plus a new COLORBENCH benchmark.

  24. Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models

    cs.CV 2026-07 conditional novelty 5.0 of 10

    ATLAS adds a Think–Plan–Paint loop with shared positional tokens to unified MLLMs, plus RL-based layout alignment, achieving large reported gains over prior layout-based unified models on compositional image generatio...

  25. Instant Preference Alignment for Text-to-Image Diffusion Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    An MLLM-driven, training-free pipeline extracts preference keywords from a reference image and modulates diffusion cross-attention at global and regional levels for instant, multi-round preference-aligned image generation.

  26. ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A released 6.4 million pair dataset and 613 sample benchmark for instruction-guided image editing of non-rigid motions, plus a Flux.1-dev based baseline that outperforms open-source methods on the new benchmark.

  27. ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback

    cs.AI 2025-05 conditional novelty 5.0 of 10

    A ComfyUI-based multi-agent system with semantic workflow modules and tree-based local-feedback planning reports near-perfect pass rates on ComfyBench and competitive scores on GenEval and Reason-Edit.

  28. X-Edit: Detecting and Localizing Edits in Images Altered by Text-Guided Diffusion Models

    cs.CV 2025-05 conditional novelty 5.0 of 10

    X-Edit uses Stable Diffusion inversion features with a U-Net and attention to predict edited-region masks, and contributes a paired 167,026-image dataset for the task.

  29. Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

    cs.CV 2025-05 reject novelty 5.0 of 10

    Selftok encodes images as diffusion-time-indexed discrete tokens, enabling a pure autoregressive VLM and visual RL with strong GenEval and DPG scores, though its claim that spatial tokens cannot support RL is not proven.

  30. Preliminary Explorations with GPT-4o(mni) Native Image Generation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A qualitative exploration showing GPT-4o image generation excels at stylization, editing, and personalization but struggles with spatial reasoning, knowledge-based accuracy, and temporal prediction.

  31. MagicQuill: An Intelligent Interactive Image Editing System

    cs.CV 2024-11 conditional novelty 5.0 of 10

    MagicQuill combines brush-based edge and color control with an MLLM that guesses user intent, enabling fast interactive image edits without typing prompts.

  32. MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection

    cs.CV 2025-05 reject novelty 4.0 of 10

    MIND-Edit combines instruction rewriting with MLLM-derived visual embeddings to guide diffusion-based image editing, but the reported numbers only partly support the claim of state-of-the-art performance.

  33. R-Genie: Reasoning-Guided Generative Image Editing

    cs.CV 2025-05 conditional novelty 4.0 of 10

    R-Genie couples a multimodal LLM with a discrete diffusion model to perform image edits that require commonsense reasoning, and introduces a 1,070-triple benchmark called REditBench.

  34. Hands-off Image Editing: Language-guided Editing without any Task-specific Labeling, Masking or even Training

    cs.CL 2025-02 conditional novelty 4.0 of 10

    An instruction-guided image editor that needs no training, labels, or masks: an LLM writes before/after captions and their embedding difference guides Stable Diffusion.

  35. Mapping the Mind of an Instruction-based Image Editing using SMILE

    cs.AI 2024-12 reject novelty 4.0 of 10

    SMILE applies LIME-style prompt perturbation with image-embedding distances to create word-level heatmaps for instruction-based image editing models.

Pith tools