Pith. sign in

REVIEW 9 cited by

Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.02826 v4 pith:FIJZQA5I submitted 2025-04-03 cs.CV

classification cs.CV
keywords editingvisualmodelsrisebenchreasoningappearancechallengesconsistency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but they still face challenges in General Visual Editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats. To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE). RISEBench focuses on four key reasoning categories: Temporal, Causal, Spatial, and Logical Reasoning. We curate high-quality test cases for each category and propose an robust evaluation framework that assesses Instruction Reasoning, Appearance Consistency, and Visual Plausibility with both human judges and the LMM-as-a-judge approach. We conducted experiments evaluating nine prominent visual editing models, comprising both open-source and proprietary models. The evaluation results demonstrate that current models face significant challenges in reasoning-based editing tasks. Even the most powerful model evaluated, GPT-4o-Image, achieves an accuracy of merely 28.8%. RISEBench effectively highlights the limitations of contemporary editing models, provides valuable insights, and indicates potential future directions for the field of reasoning-aware visual editing. Our code and data have been released at https://github.com/PhoenixZ810/RISEBench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do Image Editing Models Understand Lighting?

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    New 3DLP benchmark with real-world 1K HDR pairs shows state-of-the-art image editing models vary in physical lighting consistency, with best models close to reality but error-prone in low-light regions.

  2. WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing

    cs.CV 2026-03 conditional novelty 7.0 of 10

    WeEdit trains a glyph-guided, RL-optimized image editor on a 330K-pair synthetic multilingual dataset and reports open-source SOTA on its own bilingual and multilingual text-editing benchmarks.

  3. R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A 3,068-prompt benchmark with per-instance Q&A scoring shows that current text-to-image models, including reasoning-enhanced ones, handle reasoning-driven prompts poorly, with mathematical reasoning near zero.

  4. OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A new benchmark and evaluation framework for instruction-based video editing that covers spatial, temporal, audio, reference, and reasoning edits, with an accuracy-aware penalty to prevent inflated scores for incorrect edits.

  5. Image-Space Rule Discovery

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A new benchmark shows frontier image-editing models can sometimes follow instructions rendered inside an image and answer on the same canvas, with the best model, Nano Banana Pro, reaching 48.7% strict proxy accuracy ...

  6. GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GOBench measures how well multimodal AI models generate and understand geometric optics, finding that even top models make frequent physical errors.

  7. OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A 57-task benchmark for multimodal image generation that uses automated visual parsers and an LLM judge to show GPT-4o-Native leads current models.

  8. KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new benchmark, KRIS-Bench, evaluates image editing models on knowledge-grounded reasoning across factual, conceptual, and procedural tasks, and finds large performance gaps in current models.

  9. R-Genie: Reasoning-Guided Generative Image Editing

    cs.CV 2025-05 conditional novelty 4.0 of 10

    R-Genie couples a multimodal LLM with a discrete diffusion model to perform image edits that require commonsense reasoning, and introduces a 1,070-triple benchmark called REditBench.

Pith tools