Pith. sign in

REVIEW 13 cited by

ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.03107 v2 pith:CFSF4IVJ submitted 2025-06-03 cs.CV

ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions

classification cs.CV
keywords editingimagenon-rigidbytemorphmotionsbytemorph-6mcomprehensivedataset
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Editing images with instructions to reflect non-rigid motions, camera viewpoint shifts, object deformations, human articulations, and complex interactions, poses a challenging yet underexplored problem in computer vision. Existing approaches and datasets predominantly focus on static scenes or rigid transformations, limiting their capacity to handle expressive edits involving dynamic motion. To address this gap, we introduce ByteMorph, a comprehensive framework for instruction-based image editing with an emphasis on non-rigid motions. ByteMorph comprises a large-scale dataset, ByteMorph-6M, and a strong baseline model built upon the Diffusion Transformer (DiT), named ByteMorpher. ByteMorph-6M includes over 6 million high-resolution image editing pairs for training, along with a carefully curated evaluation benchmark ByteMorph-Bench. Both capture a wide variety of non-rigid motion types across diverse environments, human figures, and object categories. The dataset is constructed using motion-guided data generation, layered compositing techniques, and automated captioning to ensure diversity, realism, and semantic coherence. We further conduct a comprehensive evaluation of recent instruction-based image editing methods from both academic and commercial domains.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences

    cs.CV 2026-05 unverdicted novelty 7.0

    CV-Arena is a new 12K-pair benchmark for instruction-guided real-image editing with 16 task types, CogRetriever curation, and Active Elo mixed human-AI evaluation that finds gaps in 21 models and presents CV-Agent.

  2. WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing

    cs.CV 2026-03 conditional novelty 7.0

    WeEdit trains a glyph-guided, RL-optimized image editor on a 330K-pair synthetic multilingual dataset and reports open-source SOTA on its own bilingual and multilingual text-editing benchmarks.

  3. Taming I2V models for Image HOI Editing: A Cognitive Benchmark and Agentic Self-Correcting Framework

    cs.CV 2026-06 unverdicted novelty 6.0

    Introduces HOI-Edit benchmark with HOI-Eval metric and SCPE self-correcting framework leveraging I2V models for competitive HOI image editing performance.

  4. Taming I2V models for Image HOI Editing: A Cognitive Benchmark and Agentic Self-Correcting Framework

    cs.CV 2026-06 unverdicted novelty 6.0

    HOI-Edit benchmarks human–object interaction image editing with a VLM-based metric, and SCPE steers I2V models via agentic self-correcting prompts to produce competitive HOI edits.

  5. SIGMA: Semantic-Difference Instruction-Grounding Mask Annotator for Text-Driven Image Manipulation Localization

    cs.CV 2026-05 unverdicted novelty 6.0

    SIGMA generates accurate IML masks via semantic feature differencing and instruction-guided cross-modal refinement, yielding a 1.1M training set that boosts six detectors by 18.34% F1 on five datasets.

  6. DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing

    cs.CV 2026-04 unverdicted novelty 6.0

    DDA-Thinker decouples planning from generation and applies dual-atomic RL with checklist-based rewards to boost reasoning in image editing, yielding competitive results on RISE-Bench and KRIS-Bench.

  7. LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing

    cs.CV 2026-04 unverdicted novelty 6.0

    LIVE achieves state-of-the-art instruction-based video editing by jointly training on image and video data with a frame-wise token noise strategy to bridge domain gaps and a new benchmark of over 60 tasks.

  8. EditCaption: Human-Refined SFT and HAE-DPO for Image Editing Instruction Synthesis

    cs.CV 2026-04 unverdicted novelty 6.0

    EditCaption reduces critical errors in automated image editing instructions from 47.75% to 23% via SFT and DPO, yielding fine-tuned models that match or exceed closed-source VLMs on Eval-400 and ByteMorph-Bench.

  9. EditCaption: Human-Refined SFT and HAE-DPO for Image Editing Instruction Synthesis

    cs.CV 2026-04 conditional novelty 6.0

    A 235B VLM trained with human-refined SFT and hardness-adaptive error-aware DPO cuts critical instruction errors from ~48% to ~18% and beats Gemini-3-Pro on three editing-instruction benchmarks.

  10. PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing

    cs.CV 2026-04 unverdicted novelty 6.0

    PhyEdit improves physical accuracy in image object manipulation by using explicit geometric simulation as 3D-aware guidance combined with joint 2D-3D supervision.

  11. Towards Robust Sequential Decomposition for Complex Image Editing

    cs.CV 2026-05 unverdicted novelty 5.0

    Sequential decomposition trained on synthetic editing tasks improves robustness for complex image instructions and transfers to real images via co-training.

  12. Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications

    cs.CV 2026-07 conditional novelty 4.0

    A survey of instruction-based image editing plus a new 21-task benchmark, CDD-IIE, on which ten open models are scored by human experts.

  13. Towards Robust Sequential Decomposition for Complex Image Editing

    cs.CV 2026-05 unverdicted novelty 4.0

    Develops a synthetic data pipeline for training sequential decomposition in generative image editing, showing robust gains with complexity and sim-to-real transfer via co-training.