{"paper":{"title":"Making Image Editing Easier via Adaptive Task Reformulation with Agentic Executions","license":"http://creativecommons.org/licenses/by/4.0/","headline":"Reformulating vague image editing instructions into adaptive operation sequences with an MLLM agent lifts performance without changing the model.","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Bo Zhao, Haiyang Sun, Huan Yang, Kairui Guo, Kun Gai, Pengshan Wang, Runnan Du, Wei Ji, Yixin Cao","submitted_at":"2026-04-17T10:17:22Z","abstract_excerpt":"Instruction guided image editing has advanced substantially with recent generative models, yet it still fails to produce reliable results across many seemingly simple cases. We observe that a large portion of these failures stem not from insufficient model capacity, but from poorly formulated editing tasks, such as those involving small targets, implicit spatial relations, or under-specified instructions. In this work, we frame image editing failures as a task formulation problem and propose an adaptive task reformulation framework that improves editing performance without modifying the underl"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Experiments on multiple benchmarks, including ImgEdit, PICA, and RePlan, across diverse editing backbones such as Qwen Image Edit and Nano Banana, show consistent improvements, with especially large gains on challenging cases.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"A large portion of these failures stem not from insufficient model capacity, but from poorly formulated editing tasks, such as those involving small targets, implicit spatial relations, or under-specified instructions.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"An MLLM agent reformulates image editing tasks into executable operation sequences to improve reliability on challenging cases across existing generative backbones.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Reformulating vague image editing instructions into adaptive operation sequences with an MLLM agent lifts performance without changing the model.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"22d824ed9697d93d334ef94d1a6f3028dd4a9928dda1cda9da20cd577d915bb7"},"source":{"id":"2604.15917","kind":"arxiv","version":2},"verdict":{"id":"e3ca54e1-34b5-494a-9aba-39735ab0eeb9","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-10T08:19:07.102990Z","strongest_claim":"Experiments on multiple benchmarks, including ImgEdit, PICA, and RePlan, across diverse editing backbones such as Qwen Image Edit and Nano Banana, show consistent improvements, with especially large gains on challenging cases.","one_line_summary":"An MLLM agent reformulates image editing tasks into executable operation sequences to improve reliability on challenging cases across existing generative backbones.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"A large portion of these failures stem not from insufficient model capacity, but from poorly formulated editing tasks, such as those involving small targets, implicit spatial relations, or under-specified instructions.","pith_extraction_headline":"Reformulating vague image editing instructions into adaptive operation sequences with an MLLM agent lifts performance without changing the model."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2604.15917/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}