REVIEW 9 cited by
Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but they still face challenges in General Visual Editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats. To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE). RISEBench focuses on four key reasoning categories: Temporal, Causal, Spatial, and Logical Reasoning. We curate high-quality test cases for each category and propose an robust evaluation framework that assesses Instruction Reasoning, Appearance Consistency, and Visual Plausibility with both human judges and the LMM-as-a-judge approach. We conducted experiments evaluating nine prominent visual editing models, comprising both open-source and proprietary models. The evaluation results demonstrate that current models face significant challenges in reasoning-based editing tasks. Even the most powerful model evaluated, GPT-4o-Image, achieves an accuracy of merely 28.8%. RISEBench effectively highlights the limitations of contemporary editing models, provides valuable insights, and indicates potential future directions for the field of reasoning-aware visual editing. Our code and data have been released at https://github.com/PhoenixZ810/RISEBench.
Forward citations
Cited by 9 Pith papers
-
Do Image Editing Models Understand Lighting?
New 3DLP benchmark with real-world 1K HDR pairs shows state-of-the-art image editing models vary in physical lighting consistency, with best models close to reality but error-prone in low-light regions.
-
WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
WeEdit trains a glyph-guided, RL-optimized image editor on a 330K-pair synthetic multilingual dataset and reports open-source SOTA on its own bilingual and multilingual text-editing benchmarks.
-
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
A 3,068-prompt benchmark with per-instance Q&A scoring shows that current text-to-image models, including reasoning-enhanced ones, handle reasoning-driven prompts poorly, with mathematical reasoning near zero.
-
OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing
A new benchmark and evaluation framework for instruction-based video editing that covers spatial, temporal, audio, reference, and reasoning edits, with an accuracy-aware penalty to prevent inflated scores for incorrect edits.
-
Image-Space Rule Discovery
A new benchmark shows frontier image-editing models can sometimes follow instructions rendered inside an image and answer on the same canvas, with the best model, Nano Banana Pro, reaching 48.7% strict proxy accuracy ...
-
GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
GOBench measures how well multimodal AI models generate and understand geometric optics, finding that even top models make frequent physical errors.
-
OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
A 57-task benchmark for multimodal image generation that uses automated visual parsers and an LLM judge to show GPT-4o-Native leads current models.
-
KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
A new benchmark, KRIS-Bench, evaluates image editing models on knowledge-grounded reasoning across factual, conceptual, and procedural tasks, and finds large performance gaps in current models.
-
R-Genie: Reasoning-Guided Generative Image Editing
R-Genie couples a multimodal LLM with a discrete diffusion model to perform image edits that require commonsense reasoning, and introduces a 1,070-triple benchmark called REditBench.
Discussion (0). Continue with ORCID to comment.