{"id":"a9bfafe8-4089-40f1-bf38-b678fa5c68be","arxiv_id":"2607.10301","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"ChartSync formalizes visuo-logical cascading chart editing and finds only two frontier proprietary models show strong text-to-geometry synchronization; open-source models largely fail.","lead":"Most image editors only rewrite chart labels and fail to resize bars, slices, or other geometry when the numbers change. ChartSync is a public 870-example benchmark that isolates this cascading text-to-geometry skill and shows only two frontier proprietary models currently handle it well.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged judge-family risk.","rationale":"The reader's weakest_assumption correctly isolates the only load-bearing soft spot: reliance on a single proprietary VLM family for VLCS/TESR/BFS. Everything else that would have to be true for the strongest claim is well supported—deterministic code-rendered GT, expert QA retention 93%, human–metric correlation, public release, and consistent multi-metric evidence of the TESR–VLCS gap. I do not find a deeper technical flaw (e.g., circular construction of GT, mis-specified VLCS rubric, or confounded chart-category sampling) that would overturn the ranking. Therefore the appropriate verdict remains CONDITIONAL with high confidence; no adjustment is warranted.","tokens_in":19634,"tokens_out":467,"duration_ms":4135,"concrete_test":"Re-score the full 235 VLCE predictions with an independent non-Gemini VLM judge (e.g., GPT-5.x or Claude) under the same Fig. 12 rubric; if Nano Banana Pro / GPT-Image-2 VLCS remain ≥70 and the open-source cluster remains ≤20, the ranking claim holds; a reordering or >15-point absolute shift would weaken it.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is an empirical ranking gap on the 235 VLCE instances (TESR vs VLCS drops for most models; only Nano Banana Pro and GPT-Image-2 show strong VLCS). That ranking rests on Gemini-3.1-Pro VLCS scores. The paper already reports the right mitigations: deterministic programmatic GT, expert unanimous QA (Table 4), OCR/SSIM objective tier, and blind human calibration on 200 samples with VLCS Pearson r=0.892 and ICC 0.87 (Table 3). Limitations explicitly flag single-family judge bias. No stronger internal inconsistency appears: the code-mediated baseline, category/task breakdowns (Fig. 5), and qualitative cases (Figs. 13–16) all point the same direction. Residual risks (judge family, proprietary data exposure) are real but already correctly treated as CONDITIONAL rather than REJECT.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper formalizes Visuo-Logical Cascading Editing (VLCE) for statistical charts—edits where a textual value change must trigger synchronized geometric deformation—and introduces ChartSync, an expert-validated benchmark of 870 image–instruction–GT triplets across 9 chart types and 4 task types, of which 235 are geometry-coupled VLCE instances. Ground truth is produced by a three-stage programmatic pipeline (instruction synthesis, code-driven re-rendering from ChartMimic sources, unanimous three-expert QA retaining 870/935 candidates). Evaluation uses a two-tier protocol (OCR F1, SSIM plus VLM-as-a-Judge TESR/VLCS/BFS) on 14 image editors and one chart-to-code pipeline. Main empirical claim: most open-source and several proprietary models show large TESR–VLCS drops (e.g., Qwen-Image-Edit-2511: 61.81 vs 13.83), while only Nano Banana Pro and GPT-Image-2 exhibit strong emerging text-to-geometry synchronization (VLCS 83.71 and 74.47), with residual failures in semantic isolation and background corruption. Failure modes are distilled into three hierarchical meta-abilities; data and code are released.","tokens_in":19923,"tokens_out":1490,"duration_ms":18926,"significance":"If the reported ranking holds, ChartSync is a useful diagnostic benchmark that isolates value-to-geometry synchronization—an ability prior chart-editing suites (ChartEdit, ChartM3, ChartE³, ChartEditVista) do not score as a first-class target (Table 1). Strengths that raise the contribution above a routine leaderboard paper include: deterministic code-rendered GT with expert unanimous QA (Appendix A, Table 4), public release of dataset and construction code, a calibrated two-tier metric suite with blind human correlation (ICC 0.87; VLCS Pearson r=0.892 on 200 samples, Table 3), and a code-mediated baseline that shows reconstruction loss is not a free solution. The TESR–VLCS gap and the qualitative cases (Figs. 13–16) give a concrete target for multimodal editing architectures.","major_comments":[{"comment":"§4.2 and Table 2: The central ranking claim—that only Nano Banana Pro and GPT-Image-2 show strong VLCE—rests primarily on Gemini-judged VLCS over the 235 VLCE instances. Table 3 reports strong overall human correlation (VLCS r=0.892, n=200), but the paper does not break out calibration statistics restricted to the VLCE subset or report inter-judge agreement for VLCS specifically. Given that Limitations already flags single-family judge risk, a modest expansion (e.g., second-family judge or human VLCS on a stratified VLCE subsample with reported agreement) is needed so the TESR–VLCS gap ranking is not over-dependent on one evaluator family.","section":"§4.2, Table 2, Table 3"},{"comment":"§5.2 and Fig. 6: The three hierarchical meta-abilities (foundation perception, synchronization/reasoning, high-fidelity generation) are presented as general architectural guidance, yet the error taxonomy is derived only from Qwen-Image-Edit-2511 vs Nano Banana Pro. Limitations item (4) acknowledges this, but the main text still generalizes from a two-model multi-label breakdown. Either restrict the meta-ability claims to “representative failure modes of the top open-source and top proprietary systems,” or add a lighter taxonomy pass over a few additional models so the hierarchy is not load-bearing on an n=2 comparison.","section":"§5.2, Figure 6"},{"comment":"§3.4.2 / Overall Score definition: Overall Score is the equal average of OCR F1, SSIM, TESR, VLCS, and BFS, while VLCS is defined only on the 235 VLCE samples and is not assigned on text-only items. The paper notes this is a “metric-level summary,” but Table 2 rankings and the abstract’s capability narrative mix full-set and VLCE-only signals. Please report a VLCE-subset Overall (or primary VLCS/TESR/BFS on the 235) alongside the full-set summary so the cascading-reasoning claim is not diluted or inflated by the 635 text-only triplets.","section":"§3.4.2, Table 2"}],"minor_comments":[{"comment":"Model identifiers such as “Nano Banana Pro” (footnote 2: gemini-3-pro-image) and “GPT-Image-2” should be stated once with canonical API names in the main text of §4.1 for reproducibility; the footnotes alone are easy to miss.","section":"§4.1, Table 2"},{"comment":"Figure 5 compares only two models across categories/tasks; a compact appendix table of VLCS by chart family for all models (or at least all proprietary + top open-source) would make the “fluctuations on bar/pie/multidiff” claim checkable.","section":"§5.1, Figure 5"},{"comment":"Eq. (1) writes ΔG = S(ΔV) without specifying the type of S; a one-sentence note that S is the deterministic renderer-induced map (not a learned operator) would avoid misreading S as a model component.","section":"§3.1, Eq. (1)"},{"comment":"Appendix D judge prompt omits 0.75 for VLCS by design; a short pointer in §3.4.2 to that design choice (already in the appendix) would help readers who only skim the main metrics section.","section":"§3.4.2, Appendix D"},{"comment":"Table 1 “Value-to-Geometry Sync. / Sync. Metric” columns correctly distinguish ChartSync; ensure ChartE³ and ChartEditVista citations remain accurate if those arXiv versions evolve before camera-ready.","section":"Table 1"}],"recommendation":"minor_revision","confidential_remarks":"Fit for a solid CV/ML venue as a diagnostic benchmark paper. The empirical gap is well supported; the main risk is over-claiming generality of meta-abilities and judge-dependent ranking without a second evaluator family. I would not reject on that basis—the mitigations (programmatic GT, expert QA, human correlation, public code) are above average for this genre. No novelty or citation-pattern concerns stood out. Recommend minor revision rather than major: the three major points are fixable without new model training or a full re-benchmark."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline first: ChartSync isolates a concrete, under-measured capability—when you change a data label, the coupled geometry has to move with it—and shows that almost every current instruction editor still fails that cascade. That gap is real and useful to have measured cleanly.\n\nWhat is actually new is the focus, not the general idea of chart editing. Prior work already covers code-mediated edits, text fidelity, and layout. ChartSync’s contribution is deterministic ground truth from source-code edits plus re-render, a 235-instance VLCE subset that forces text-to-geometry synchronization, and a two-tier protocol (OCR/SSIM plus calibrated VLM scores, especially VLCS). Construction is careful: programmatic pipeline, unanimous three-expert QA (870/935 retained), public code and data. The main table is readable: large TESR–VLCS drops for open-source and several proprietary models; Nano Banana Pro and GPT-Image-2 are the only ones with strong VLCS; residual errors are semantic isolation and background bleed. The code-mediated baseline is a fair control and correctly shows reconstruction loss. Human calibration (ICC 0.87, VLCS r≈0.89) is better than most VLM-judge papers bother with.\n\nSoft spots are real but proportionate. The judge is still a single Gemini-family model; they flag it and calibrate, so treat rankings as directionally solid rather than absolute. Error taxonomy is only fully worked on two models. Dataset scale is modest. Proprietary training exposure cannot be audited. None of these overturn the central empirical claim.\n\nMath is light (definitional ΔG = S(ΔV)); data and citations look solid for a systems benchmark. This is for people building or evaluating multimodal editors on structured documents. I would bring it to a reading group if we are working on chart or document editing; I would cite the benchmark and the VLCS gap if I am writing in that lane. Send it to peer review—it deserves referee time, with normal pressure on judge robustness and broader failure analysis.","headline":"Clean diagnostic for a real failure mode: most editors still treat chart numbers as free text, not as geometry drivers; only two frontier systems show emerging VLCE.","tokens_in":20536,"tokens_out":515,"would_cite":true,"duration_ms":15193,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Most image editors can change chart labels but fail to update the linked bars, slices, and geometry that the labels describe.","keywords":["chart editing","visuo-logical cascading editing","VLCE","image editing benchmark","text-to-geometry synchronization","vision-language models","structured document understanding"],"falsifier":"Re-score the same 235 VLCE predictions with a different-family human panel or multi-model judge ensemble; if the large TESR-to-VLCS gaps reverse or the two frontier models lose their lead, the central capability claim fails.","tokens_in":20550,"feed_emoji":"📊","tokens_out":915,"duration_ms":15790,"temperature":0.7,"pith_summary":"Instruction-based image editors work on natural photos, yet statistical charts demand a harder skill: when a number changes, the drawn geometry must change with it. The authors formalize that requirement as Visuo-Logical Cascading Editing (VLCE) and release ChartSync, an expert-checked set of 870 original-instruction-ground-truth triplets spanning nine chart types, including 235 cases that force text-to-geometry synchronization. Ground truth is produced by editing the original plotting code and re-rendering, so the correct geometry is deterministic rather than hand-drawn. A two-tier scoreboard pairs ordinary image metrics with a vision-language judge that separately rates textual success, geometric consistency, and background preservation. Across fourteen editors and one chart-to-code pipeline, most models still treat the task as local text replacement; only two frontier proprietary systems show clear cascading ability, and even they leave residual isolation and background errors. The gap isolates three meta-abilities future multimodal systems will need: grounding text to marks, reasoning over data-driven shape changes, and editing without bleeding into the rest of the figure.","feed_headline":"Chart editors change labels but not the bars they name","feed_subtitle":"A new 870-triplet benchmark shows only two frontier models can keep geometry in sync with data edits.","key_machinery":"Visuo-Logical Cascading Editing (VLCE) plus the ChartSync VLCE subset and VLCS score: a textual value change must induce the matching geometric deformation (bar height, pie angle, etc.) while non-target regions stay intact, measured on 235 geometry-coupled instances with deterministic code-rendered ground truth.","core_discovery":"On ChartSync, reliable text-to-geometry cascading is rare. Textual edit success often remains moderate to high while Visuo-Logical Consistency collapses for open-source and many proprietary editors (for example one strong open model drops from about 62 textual success to under 14 geometric consistency). Only two frontier proprietary models reach strong VLCS scores (roughly 75 and 84); residual failures for those systems cluster in semantic isolation and background corruption. A code-mediated reconstruct-edit-render path can rewrite labels yet still loses geometric and layout fidelity when the original source is unavailable.","pith_inferences":["Training regimes that only reward literal text replacement will keep producing the observed TESR–VLCS cliff until geometric coupling is an explicit objective.","The same cascading test could be ported to other rigid diagrams (timelines, flowcharts, scientific schematics) where labels control shape.","Residual background bleeding even in strong models suggests that mask-free diffusion editors still lack reliable region isolation for dense symbolic graphics."],"forward_implications":["Pixel-space chart editors must be tested on value-to-geometry coupling, not only OCR or global SSIM.","Code-only pipelines are not a free fix when source plots are unavailable, because reconstruction loses layout and geometry.","Future architectures need three stacked meta-abilities: perception grounding, data-driven geometric reasoning, and artifact-free isolation.","Public release of the 870 triplets and rendering pipeline gives a fixed diagnostic for measuring progress on structured document editing."],"fun_headline_variants":["Only two models keep bars in sync when charts change","Chart edits break geometry for almost every model","ChartSync shows cascading bar-label updates rarely work","Text edits succeed while chart geometry collapses","Most editors fail visuo-logical chart synchronization"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The ranking of cascading skill rests on a single proprietary vision-language judge for geometric consistency and background fidelity, even though that judge was checked against a few hundred human ratings.","fun_headline_variants_meta":{"raw":{"variants":["Only two models keep bars in sync when charts change","Chart edits break geometry for almost every model","ChartSync shows cascading bar-label updates rarely work","Text edits succeed while chart geometry collapses","Most editors fail visuo-logical chart synchronization"]},"model":"grok-4.5","effort":"low","cost_usd":0.003196,"raw_usage":{"total_tokens":1084,"prompt_tokens":827,"num_sources_used":0,"completion_tokens":53,"cost_in_usd_ticks":31960000,"prompt_tokens_details":{"text_tokens":827,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":204,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":827,"tokens_out":53,"duration_ms":3384,"temperature":1.0,"reasoning_tokens":204,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T12:47:01.291807+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-score the same 235 VLCE predictions with a different-family human panel or multi-model judge ensemble; if the large TESR-to-VLCS gaps reverse or the two frontier models lose their lead, the central capability claim fails.","supporting_citations":[],"review_version":1}