{"id":"fc0d2736-fa10-4051-808d-61944d24cc93","arxiv_id":"2608.02841","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A client-side landmark mask and feathered composite confined image-editing API outputs to the requested facial region in a 15-face pilot, at the cost of little on-target change, while no tested editor moved identity embeddings closer to postoperative photographs.","lead":"This pilot study tests whether a cosmetic-surgery preview can be kept inside the requested facial region when the editing model is a commercial black-box API. A client-side mask-and-composite step confined off-target pixel changes across all six editors tested, but the gain is partly built into the method and the sample is small.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claim that compositing preserves target-region change rests on one sample per cell; stochastic variation could turn 'localize' into 'suppress'.","rationale":"The reader's weakest assumption—single generation per cell with unknown stochastic variation—is the same concern I find most load-bearing, and I focus it on the one non-tautological quantitative claim: that compositing preserves target-region pixel change. Off-mask preservation is guaranteed by construction, so the localization improvement is not evidence about editor behavior. What would make the method practically useful is that the target region still receives a comparable edit; that claim is supported only by median Delta E_tgt 4.5 versus 4.0 over at most 8 faces per cell, each sampled once. Two models ignore seeds, so repeatability is not assured. The paper is unusually honest about this limitation and explicitly calls itself a pilot, so this concern does not warrant rejection; it reinforces the CONDITIONAL verdict. I therefore set verdict_should_be to UNCHANGED and agree with the reader's assessment.","tokens_in":13034,"tokens_out":5371,"duration_ms":50074,"concrete_test":"Pick two representative editors (e.g., GPT Image 2 and Nano Banana Pro) and re-run the full facelift/rhinoplasty face set with K=10 independent generations per face-model-control cell, recording per-cell median Delta E_tgt for prompt-only and composite, plus per-generation loc and Delta E_off. Compute a paired, face-clustered bootstrap interval for the median within-cell ratio (composite Delta E_tgt / prompt-only Delta E_tgt). If the interval excludes 0.8-1.25, or if composite median Delta E_tgt is more than 20% lower than prompt-only in a majority of cells, the 'about as much target change' claim fails and the method should be described as suppression rather than localization. Use the same prompts, masks, and frozen pipeline to isolate stochastic variation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two parts: (a) client-side compositing enforces off-mask preservation, which is true by construction and acknowledged in Sec. IV-B, and (b) it does so 'without materially reducing on-target pixel change', with editor differences and a negative post-op identity shift as supporting evidence. Part (b) is the load-bearing empirical content. It rests on one output per model-face-control cell (Sec. V: 'One output is generated per cell'), with no within-model variance estimate, and Sec. VII states that two of the six configurations ignore the supplied seed. The headline comparison is median Delta E_tgt 4.5 for composite versus 4.0 for prompt-only over 12 frontal faces (Sec. VI-A); with 8 or fewer faces per cell and a single draw, an atypical generation could flip or widen that comparison. If repeated sampling shows composite Delta E_tgt is substantially lower than prompt-only, the method is not localizing edits but suppressing them, and the previews would understate the requested change. The localization gain of 0.446 is not the vulnerable part—it is forced by copying off-mask pixels—so the only non-tautological quantitative support for usefulness is exactly the part threatened by single-sample stochasticity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether region-confined editing of black-box commercial image-editing APIs can be achieved entirely on the client side, for cosmetic-surgery previews. It proposes a control ladder: prompt-only, a masked composite that aligns the model's output back onto the input and blends it through a landmark-derived feathered mask, and native masked inpainting where supported. On a pilot benchmark of six commercial editing configurations and one inpainting model, the paper reports that masked compositing increases a CIELAB localization ratio by a median of 0.446 on 12 frontal faces while leaving target-region pixel change roughly unchanged (median ΔE_tgt 4.5 vs 4.0), that the tested inpainting model did not beat the composite, and that no editor moved ArcFace embeddings closer to postoperative photographs. The authors explicitly scope the study as a pilot, disclaim clinical validity, and note that each cell contains one generated output and that two models ignore the supplied seed.","tokens_in":13196,"tokens_out":6333,"duration_ms":54577,"significance":"If the descriptive claims hold, the paper makes a useful systems-level observation: for black-box editing APIs, off-target pixel change can be suppressed entirely on the client, at trivial cost and without model internals. The control ladder is a simple, reproducible baseline, and the paper is transparent about its limitations, provides code and a canonical score table, and avoids overclaiming clinical relevance. The main non-tautological empirical claim—that compositing preserves on-target pixel change—is not, however, supported by the single-sample design. The localization gain itself is a mathematical consequence of the compositing construction, so the paper's significance rests on the target-region and editor-difference results, which are exactly the parts threatened by generation stochasticity. The paper is an honest pilot, but its abstract and conclusion currently state the target-preservation result more strongly than the evidence allows.","major_comments":[{"comment":"The claim in Sec. VI-A that masked compositing retains target-region pixel change (median ΔE_tgt 4.5 vs 4.0) is not supported by the experimental design, because Sec. V states that one output is generated per cell and Sec. VII confirms that two models ignore the supplied seed. The face-clustered bootstrap interval for the localization gain does not account for generation stochasticity, so a single atypical draw could flip or widen the 0.5-unit difference. The authors should either add repeated generations (e.g., 5–10 per cell) and report within-model variance, or explicitly downgrade this claim to a preliminary observation that requires further testing.","section":"Sec. V and Sec. VI-A"},{"comment":"The localization ratio loc is forced toward 1 for the masked composite because compositing copies off-mask pixels from the input by construction, making ΔE_off approximately zero. The paper acknowledges this in Sec. IV-B and Sec. VII, but the abstract and conclusion present the 0.446 gain as a primary result. Since the gain is a mathematical consequence of the construction rather than an empirical property of the editors, the non-tautological content resides in the target-region and editor-difference results, exactly the parts threatened by single-sample stochasticity. The contribution should be reframed so the localization gain is reported as a consistency check of the compositing implementation, not as a headline improvement.","section":"Sec. IV-B, Eq. (1)"},{"comment":"The comparison between Qwen-Image-Edit's masked inpainting and client-side compositing rests on a single model, a single output per cell, and shows the inpainter producing near-copies on rhinoplasty (ΔE_tgt 1.9, identity 0.965). This near-copy behavior suggests the model may not have executed the requested edit, so the conclusion that 'client-side compositing produced stronger measured keep-zone preservation than provider-side masking' is confounded by possible instruction-following failure. The authors should either add repeated generations, test additional inpainting models, or report qualitative evidence that the requested anatomical change was actually made, to allow a fair comparison.","section":"Sec. VI-C"}],"minor_comments":[{"comment":"The abstract states that the masked composite improved localization 'while changing the requested region about as much,' but the single-sample design and the tautological nature of the localization metric (Sec. IV-B) make this statement stronger than the evidence supports; the abstract should carry the same caveats as Sec. VII.","section":"Abstract"},{"comment":"The definition of the keep zone for rhinoplasty deserves clarification: the text says the keep zone is 'the upper-face region minus the target,' but the nose region is defined by a hull that likely extends below the lower-eyelid line, so it is not obvious that subtracting the target from the upper-face zone yields a meaningful disjoint keep zone. Please spell out how the keep zone is computed for the nose procedure.","section":"Sec. III-B"},{"comment":"The column header 'Lat.' is ambiguous; it presumably denotes latency in seconds, but the abbreviation is unusual. Spell it out or use 'Latency (s)'.","section":"Table II"},{"comment":"The description of the face set is difficult to follow: 'The benchmark matrix uses 8 facelift faces and 8 rhinoplasty faces (the 4 frontal cross-views, which the regional metric can score, plus 4 profiles...)' – it would help to state explicitly that the main analysis contains 8 facelift and 7 rhinoplasty faces after one extreme-profile face failed detection.","section":"Sec. V"},{"comment":"Reporting 'only 15% of outputs score above their face's baseline' as a percentage is unhelpful without the denominator; report the count (e.g., 29 of 194) alongside the percentage.","section":"Sec. VI-D"},{"comment":"The conclusion's phrase 'uniformly across the six editors tested' refers to off-target preservation, which is a by-product of the compositing construction; consider rephrasing to avoid implying that the editors themselves behaved uniformly.","section":"Sec. IX"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-scoped pilot with exemplary transparency about its own limitations. The main blocker is that the non-tautological empirical claims rest on one generation per cell; the authors' own Sec. VII acknowledges this, but the abstract and conclusion do not carry the same caution. The localization-gain result is essentially a verification of the compositing construction and should be reframed. The editor-difference observations (e.g., FLUX.2's low identity) are potentially interesting but equally underpowered. I would be willing to revisit a revised version that either adds repeated generations or downgrades the strength of the claims to match the pilot design."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — here's my read on 2608.02841. It's a small, honest pilot. The thing to know: the headline localization gain is real but almost tautological, because the composite copies off-mask pixels from the input, so the measured keep zone is preserved by construction. The paper says this itself, but it matters because the only non-tautological evidence that the method is useful is the claim that it doesn't suppress the intended edit. That claim rests on one generated output per cell, with two of six models ignoring the supplied seed. A single atypical draw could flip the median target-region comparison (Delta E_tgt 4.0 vs 4.5). So the load-bearing empirical part is thinner than the abstract's confident tone suggests.\n\nWhat's actually new: a controlled comparison of client-side region confinement across six commercial black-box editing APIs for cosmetic-surgery previews. That specific setup isn't in the cited literature — Envisage hosts its own inpainter, and the general benchmarks don't vary control. The paper does this cleanly: a three-rung control ladder, face-clustered bootstrap for the primary contrast, explicit accounting of rejected and failed cells, and a limitations section that names the real threats (single generation, no surgeon ratings, unreleased source images). The postoperative identity reference with per-face input baseline is a nice touch, and the negative median shift is reported honestly even though it doesn't help the method.\n\nSoft spots, in proportion. The single-sample issue is the main one. The stress-test note is right: repeated generations could show that compositing suppresses on-target change, which would turn 'localize' into 'suppress.' The paper's own numbers hint at a small suppression (4.0 vs 4.5) but with no variance estimate. The face set is 15 main-analysis faces, demographically uncharacterized, and the regional metric abstains on profiles — precisely where rhinoplasty previews matter. No surgeon ratings, so clinical plausibility is untouched. Reproducibility is partial: analysis scripts and score tables are public, but source photographs and the full generation pipeline are not. These are all acknowledged limitations; I don't see a hidden error. The citation pattern looks fair, with Envisage and the leakage benchmarks properly distinguished.\n\nWho benefits: engineers building client-side guardrails on hosted editing APIs, and benchmark designers who want a careful example of where a simple post-hoc blend works and where the evidence stops. It deserves a serious referee — the question is genuine and the execution is transparent — but the right outcome is a conditional revise, not acceptance as is. The revisions that matter: repeated generations per cell (even 3–5), release of source images or a documented subset, and a whole-image or pose-robust localization measure so profiles aren't dropped.\n\nWould I cite it? Probably, as the first cautious benchmark of client-side confinement on black-box APIs, with the caveat that the on-target preservation claim is unverified. Take it to reading group as a case study in honest pilot design.","headline":"A small, candid pilot whose localization gain is near-tautological; the real empirical claim—that compositing doesn't suppress the intended edit—rests on one sample per cell and needs replication before the benchmark is trusted.","tokens_in":13790,"tokens_out":3006,"would_cite":true,"duration_ms":28275,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Client-side mask-and-paste confines AI face edits to the intended region without any access to the model.","keywords":["client-side image editing","masked compositing","black-box editing APIs","region-confined editing","cosmetic surgery previews","identity preservation","CIELAB localization metric","instruction-guided image editing"],"falsifier":"Re-run the benchmark with repeated generations per cell on the same 12 frontal faces, for example 10 outputs per model and control, and recompute the paired localization gain; the central claim weakens if the face-clustered interval no longer excludes zero or if any editor's composited outputs show off-mask $\\Delta E$ comparable to its prompt-only outputs.","tokens_in":12757,"feed_emoji":"✂️","tokens_out":9877,"duration_ms":78570,"temperature":0.7,"pith_summary":"Commercial image-editing APIs change far more of a face than a prompt names, and existing methods that pin an edit to one region need the model's internals, which these services do not expose. This paper asks how much region control can be achieved from the client side alone and reports that a simple post-processing step delivers it: cut the edited region out of the API response, align it to the original photo, and paste it back through a feathered mask derived from facial landmarks. Across six commercial editing configurations and one mask-based inpainting service, the masked composite raised the measured localization ratio over paired prompt-only outputs by a median of 0.446 on 12 frontal faces while leaving target-region pixel change nearly unchanged (median $\\Delta E_{\\mathrm{tgt}}$ 4.5 versus 4.0). The same benchmark found that no tested editor moved outputs closer to the face's own postoperative photograph in face-recognition embedding space, and the one tested inpainting model did not beat the client-side composite. The study is explicitly a control diagnostic with one generation per cell and no clinical ratings, but its core point is that model access is not needed to enforce region confinement.","feed_headline":"Mask-and-paste confines AI face edits to the target region","feed_subtitle":"A client-side composite beat the tested model's own inpainting at keeping edits inside the requested zone.","key_machinery":"The load-bearing mechanism is the masked composite, a three-step client-side wrapper. First, a region mask is built from a 478-point face-mesh landmark set: the face splits at the lower-eyelid line into an upper-face keep zone and a jaw--neck target zone, and the nose region is a dilated perimeter hull. Then the edited response is aligned back onto the original photograph with a similarity warp estimated from stable landmarks. Finally it is blended as $\\mathrm{out} = x(1-w) + e\\,w$, where $x$ is the input, $e$ is the aligned edit, and $w$ is a feathered per-pixel weight; compositing is skipped when estimated yaw exceeds 0.35 to avoid ghost seams on profile views. The diagnostic carrying the quantitative claim is the localization ratio $\\mathrm{loc} = \\Delta E_{\\mathrm{tgt}} / (\\Delta E_{\\mathrm{tgt}} + \\Delta E_{\\mathrm{off}})$, a CIELAB pixel-change fraction saying how much measured change lands inside the requested region rather than in the protected upper-face zone.","core_discovery":"On the paper's own terms, the discovery is that text instructions alone cannot confine a black-box editor to the requested facial region, while a client-side masked composite can. For the six primary editing configurations, composited outputs had a median localization ratio of 0.985 (range 0.968--0.998) against 0.538 (range 0.347--0.680) for the same prompts without compositing; the within-face, within-model median gain was 0.446 with a face-clustered bootstrap interval of 0.421--0.457. Median off-target pixel change $\\Delta E_{\\mathrm{off}}$ fell from 3.50 to 0.07, while median target-region change $\\Delta E_{\\mathrm{tgt}}$ stayed at roughly the same level (4.5 versus 4.0), meaning the composite removed leakage without erasing the requested edit. The tested mask-based inpainting model did not surpass this post-hoc control, and against each face's input-to-postoperative baseline all editors showed a negative median identity change of $-0.029$, so none moved closer to the postoperative photograph in embedding space. The paper treats these as system-level measurements of control, not clinical evidence of surgical accuracy.","pith_inferences":["A natural extension is repeated sampling: because the paper generates one output per cell and two editors ignore the supplied seed, running many generations per cell could convert the observed 0.446 median gain into a distribution and reveal editors where compositing matters less.","The near-perfect localization of the composite is partly by construction, since off-mask pixels are copied verbatim; the more informative follow-up would measure whether inside-mask edits remain anatomically plausible, for example by comparing landmark geometry before and after.","The same wrapper should transfer to other black-box editing tasks beyond faces, such as product or document edits, whenever a mask, an alignment, and a feathered blend can be derived from the image content.","The negative identity movement relative to the postoperative baseline may reflect session differences in pose, expression, and lighting rather than anatomical failure; a pose-normalized geometry comparison would separate identity-embedding drift from true outcome drift."],"forward_implications":["Any editor reachable through a public endpoint can be region-confined without provider cooperation, at the cost of one alignment and one feathering operation per edit.","Prompt guardrails alone are not sufficient for surgical previews: prompt-only outputs still changed the protected keep zone, so localization has to be enforced in pixel space.","Native mask-based inpainting is not automatically better than client-side compositing; the one inpainting model tested produced a near-copy on rhinoplasty and lower localization than the composite on the same task.","Choosing an editor for preview work becomes an edit-strength-versus-identity trade-off rather than a pure quality ranking, since editors span a wide range of target-region change and identity retention.","On the embedding-level outcome check, current editors moved outputs slightly away from the postoperative photograph rather than toward it, so these APIs do not yet act as outcome predictors under the protocol tested."],"supporting_citations":[{"why":"establishes the clinical motivation: over-idealized previews create expectations surgery cannot meet, so off-target changes matter.","marker":"[3]"},{"why":"supplies the pixel-space final blend that the masked composite adapts to a post-hoc, internals-free setting.","marker":"[4]"},{"why":"represents the mask-guided diffusion editing methods that require model internals, defining why the black-box setting is different.","marker":"[5]"},{"why":"the closest prior system for composited rhinoplasty visualization; its mask-decomposed evaluation frames why full-face identity scores are confounded after compositing.","marker":"[7]"},{"why":"prior benchmark showing commercial editors over-edit irrelevant regions, motivating the off-target measurement and the choice of editors.","marker":"[8]"},{"why":"supplies the face-recognition embedding whose cosine similarity measures identity preservation.","marker":"[32]"},{"why":"supplies the 478-point face-mesh landmarks from which the target and keep-zone masks are derived.","marker":"[33]"},{"why":"supplies the profile-view face detection and landmarks used as fallback for nose masks when the frontal mesh fails.","marker":"[35]"},{"why":"defines the color-difference formula on which the CIELAB localization ratio is based.","marker":"[36]"}],"fun_headline_variants":["Client-side mask confines AI face edits, no model access needed","Mask-and-composite beats inpainting for localized face edits","No model internals needed: client mask localizes cosmetic edits","Client-side control localizes face edits, beats inpainting","Localize edit with mask, not model: client-side wins"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that one generated output per model-and-photo combination fairly represents how that editor behaves; if the single output is atypical, the ranking of editors, the composite-versus-inpainting comparison, and the claim that target-region change stays level could all shift.","fun_headline_variants_meta":{"raw":{"variants":["Client-side mask confines AI face edits, no model access needed","Mask-and-composite beats inpainting for localized face edits","No model internals needed: client mask localizes cosmetic edits","Client-side control localizes face edits, beats inpainting","Localize edit with mask, not model: client-side wins"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000224,"raw_usage":{"total_tokens":1564,"prompt_tokens":1149,"completion_tokens":415,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":765,"completion_tokens_details":{"reasoning_tokens":329}},"tokens_in":765,"tokens_out":415,"duration_ms":4181,"temperature":1.0,"reasoning_tokens":329,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:57:55.466407+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the benchmark with repeated generations per cell on the same 12 frontal faces, for example 10 outputs per model and control, and recompute the paired localization gain; the central claim weakens if the face-clustered interval no longer excludes zero or if any editor's composited outputs show off-mask $\\Delta E$ comparable to its prompt-only outputs.","supporting_citations":[{"cited_title":"Realistic expectations: To morph or not to morph?","cited_arxiv_id":null,"evidence_quote":"establishes the clinical motivation: over-idealized previews create expectations surgery cannot meet, so off-target changes matter."},{"cited_title":"Envisage: Diffusion-based rhinoplasty goal visualization with mask-decomposed evaluation,","cited_arxiv_id":null,"evidence_quote":"the closest prior system for composited rhinoplasty visualization; its mask-decomposed evaluation frames why full-face identity scores are confounded after compositing."},{"cited_title":"GIE-Bench: Towards grounded evaluation for text-guided image editing,","cited_arxiv_id":null,"evidence_quote":"prior benchmark showing commercial editors over-edit irrelevant regions, motivating the off-target measurement and the choice of editors."},{"cited_title":"MediaPipe: A framework for building perception pipelines,","cited_arxiv_id":null,"evidence_quote":"supplies the 478-point face-mesh landmarks from which the target and keep-zone masks are derived."},{"cited_title":"The CIEDE2000 color-difference formula: Implementation notes, supplementary test data, and mathemat- ical observations,","cited_arxiv_id":null,"evidence_quote":"defines the color-difference formula on which the CIELAB localization ratio is based."}],"review_version":2}