Pith. sign in

REVIEW 12 cited by

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.16707 v1 pith:QEABN7JE submitted 2025-05-22 cs.CV

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models

classification cs.CV
keywords editingmodelsreasoningimageknowledgekris-benchtasksbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models produce visually plausible outputs, their capacity for knowledge-based reasoning editing tasks remains under-explored. In this paper, we introduce KRIS-Bench (Knowledge-based Reasoning in Image-editing Systems Benchmark), a diagnostic benchmark designed to assess models through a cognitively informed lens. Drawing from educational theory, KRIS-Bench categorizes editing tasks across three foundational knowledge types: Factual, Conceptual, and Procedural. Based on this taxonomy, we design 22 representative tasks spanning 7 reasoning dimensions and release 1,267 high-quality annotated editing instances. To support fine-grained evaluation, we propose a comprehensive protocol that incorporates a novel Knowledge Plausibility metric, enhanced by knowledge hints and calibrated through human studies. Empirical results on 10 state-of-the-art models reveal significant gaps in reasoning performance, highlighting the need for knowledge-centric benchmarks to advance the development of intelligent image editing systems.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Do Image Editing Models Understand Lighting?

    cs.CV 2026-06 conditional novelty 7.0

    A 1,000-pair real-world HDR benchmark with two new affine-invariant error scores shows the best image-editing models reproduce the relative structure of real light transport but degrade in dim regions, and that VLMs f...

  2. Do Image Editing Models Understand Lighting?

    cs.CV 2026-06 unverdicted novelty 7.0

    New 3DLP benchmark with real-world 1K HDR pairs shows state-of-the-art image editing models vary in physical lighting consistency, with best models close to reality but error-prone in low-light regions.

  3. PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing

    cs.CV 2026-06 unverdicted novelty 7.0

    PhyEditBench is a new benchmark with real-world and synthetic instances that reveals limitations in current image editing models' physics reasoning and proposes a video-generation-based baseline called PhyWorld.

  4. PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing

    cs.CV 2026-06 unverdicted novelty 7.0

    PhyEditBench is a new benchmark for physics-aware image editing with real and synthetic instances plus a training-free PhyWorld baseline that uses test-time scaling to outperform SOTA models.

  5. EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement

    cs.CV 2026-05 unverdicted novelty 7.0

    EditRefiner uses a perception-reasoning-action-evaluation agent loop and the EditFHF-15K human feedback dataset to refine text-guided image edits more accurately than prior methods.

  6. CAMEO: A Conditional and Quality-Aware Multi-Agent Image Editing Orchestrator

    cs.CV 2026-04 unverdicted novelty 7.0

    CAMEO uses coordinated agents for planning, prompting, generation, and quality feedback to achieve higher structural reliability in conditional image editing than single-step models.

  7. DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model

    cs.CV 2026-02 unverdicted novelty 7.0

    DLEBench is the first benchmark for small-scale object editing in instruction-based image editing models, using 1889 samples, seven instruction types, and a dual-mode evaluation protocol to reveal performance gaps in ...

  8. DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing

    cs.CV 2026-04 unverdicted novelty 6.0

    DDA-Thinker decouples planning from generation and applies dual-atomic RL with checklist-based rewards to boost reasoning in image editing, yielding competitive results on RISE-Bench and KRIS-Bench.

  9. Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing

    cs.HC 2026-04 conditional novelty 6.0

    SOTA diffusion image editors score poorly on implicit physical, environmental, cultural, causal, and referential constraints, which a lightweight reasoning-guided post-edit can partially fix.

  10. VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On

    cs.CV 2026-03 conditional novelty 6.0

    VTEdit-Bench and VTEdit-QA show top universal multi-reference editors match specialized VTON models on standard tasks and transfer more stably to harder multi-person/multi-cloth settings, yet still fail under complex ...

  11. CAMEO: A Conditional and Quality-Aware Multi-Agent Image Editing Orchestrator

    cs.CV 2026-04 unverdicted novelty 5.0

    A closed-loop multi-agent image editor (CAMEO) reports ~20% higher average win rates than strong one-shot editors on anomaly insertion and pose switching.

  12. Emerging Properties in Unified Multimodal Pretraining

    cs.CV 2025-05 unverdicted novelty 5.0

    BAGEL is a unified decoder-only model that develops emerging complex multimodal reasoning abilities after pretraining on large-scale interleaved data and outperforms prior open-source unified models.