{"id":"aafd6189-f15d-4d24-8be5-3a502f32fdc5","arxiv_id":"2608.07570","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"The paper introduces the first large-scale crop-composition-explanation benchmark for aesthetic cropping and shows that an SFT+GRPO pipeline trained on it beats prior explainable-cropping methods on IoU and explanation quality.","lead":"COMEX is a new benchmark of 33,161 image-crop-explanation triples that ties aesthetic cropping decisions to explicit composition categories, together with a two-stage SFT+GRPO training framework. The authors report that adding composition supervision improves both crop accuracy and the quality of generated cropping explanations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"COMEX's IO-reversal ground truth is unvalidated and may be solvable via real-vs-generated boundary cues, undermining the composition-intermediate claim.","rationale":"The reader's weakest_assumption names the IO-reversal premise in Section 3.1, and I agree that this is the most load-bearing point, though I would sharpen it: the label-transfer half of the concern is actually valid by construction, because the crop box contains exactly the original pixels, so the inherited PICD label does apply to the crop. The vulnerable premise is instead that the original photograph remains the ideal crop after outpainting, and that the task cannot be solved by detecting the real/synthetic boundary. The paper's Limitations section acknowledges synthetic outpainting and annotation noise, but that acknowledgment does not address the stronger confound that the ground-truth box is defined as the original image rectangle. The reported composition benefit for crop quality is also numerically small (+0.0045 IoU in Table 3) and lacks uncertainty quantification, which makes the boundary-detection confound more consequential. I keep the reader's CONDITIONAL verdict unchanged because a targeted control experiment—training a boundary classifier on COMEX—could settle the concern, and because the FCDB transfer experiments provide partial real-world evidence that the learned representation is not purely a synthetic-context detector. The most appropriate path is therefore to condition acceptance on releasing the benchmark and running this control, exactly as the conditional verdict implies.","tokens_in":15947,"tokens_out":7996,"duration_ms":85413,"concrete_test":"On a held-out subset of COMEX expanded images, train a small patch-level classifier to distinguish original-PICD pixels from FLUX-outpainted pixels (binary mask supervision), then take the tightest bounding box of the predicted original-region mask and compute mIoU on the COMEX test set against the ground-truth crop boxes. If this composition-free, boundary-based baseline reaches mIoU within 10% of the reported 0.7765, the benchmark's crop task is confounded by synthetic-context detection and the central composition-intermediate claim is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—composition as an explicit intermediate improves cropping quality and explanation faithfulness—depends on COMEX, whose ground-truth crop boxes are fixed by the IO-reversal construction in Section 3.1: each original PICD image is 'treated as the ideal crop target' of a larger FLUX-outpainted canvas, and the composition label is inherited by definition. This conflates two distinct tasks. The actual task learned is 'locate the original photograph inside the synthesized canvas,' not 'choose the aesthetically best crop.' Nothing in the pipeline verifies that the original photo remains the best crop after outpainting; new salient context can make a crop including part of the outpainted region more aesthetic. Moreover, because only the context is synthetic, low-level real-vs-FLUX boundary artifacts provide a learnable shortcut for box prediction that has nothing to do with composition. Since the no-composition vs. with-composition ablation (Table 3), the SFT/GRPO comparisons, and all COMEX metrics share this ground truth, the small IoU gain (+0.0045) and large explanation win rates could reflect recovery of the original region or style rather than compositional reasoning. FCDB transfer is real-world evidence, but it does not isolate the composition-intermediate benefit and still leaves COMEX's construct validity unresolved.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reformulates explainable aesthetic image cropping as a structured crop-composition-explanation problem, introduces the COMEX benchmark constructed by outpainting PICD images and treating the original photographs as ground-truth crops, and proposes a two-stage SFT+GRPO training framework with rewards for box accuracy, composition consistency, and explanation quality. Experiments on COMEX and FCDB report that the proposed framework outperforms prior cropping methods and general-purpose MLLMs, with composition supervision improving explanation preference and GRPO improving all metrics.","tokens_in":16228,"tokens_out":6525,"duration_ms":63034,"significance":"If the COMEX ground truth is valid, the paper would make a useful contribution: it is the first large-scale benchmark to jointly provide crop boxes, composition categories, and composition-grounded explanations, and the two-stage SFT+GRPO recipe is practical and shows consistent gains over multiple backbones. The FCDB transfer results and the human preference study are positive evidence. However, the benchmark construction rests on an unvalidated IO-reversal assumption, and the explanation evaluation is entangled with the Seed model family, so the central claims about composition being a faithful intermediate layer are not yet established at the level of confidence the paper claims.","major_comments":[{"comment":"The IO-reversal premise is unvalidated. The paper asserts that the original PICD image, preserved pixel-for-pixel, is the ideal crop of the outpainted canvas, but no experiment checks whether humans or established aesthetic criteria prefer that region over alternative crops of the expanded image. Because only the surrounding context is synthesized, the model can learn to solve the task by detecting real-vs-FLUX boundary artifacts, and the statement in §3.1 that the domain gap is confined to the context does not rule out such low-level shortcut cues. Since Tables 1, 3, and 4 all use this ground truth, the in-domain IoU gains and the composition-intermediate ablation inherit this concern. Please add a human preference validation on a sample of COMEX images and a diagnostic that removes or blurs boundary cues (for example, applying a common JPEG or color transform to the full canvas) to test whether box prediction collapses.","section":"§3.1, Figure 2"},{"comment":"Explanation references and explanation judges come from the same model family: references were generated with Seed-1.8, and the judges in Table 3 are Seed2.0Pro, Seed1.8, and Seed1.6. In addition, the with-composition condition explicitly receives the composition category that the reference explanations were instructed to mention. METEOR and the MLLM win rates may therefore reward stylistic mimicry of Seed-generated text and the presence of the category token rather than faithfulness to the cropping rationale. The human study, with 100 images and 81 participants, is too small to resolve this concern, and no inter-annotator agreement or confidence intervals are reported. Please evaluate with judge models from outside the Seed family and with a human protocol that defines faithfulness criteria and blinds the composition condition.","section":"§3.2(4), Table 3"},{"comment":"All quantitative comparisons are point estimates without variance, significance tests, or confidence intervals. Several load-bearing differences are small, such as IoU 0.7484 vs. 0.7529 in Table 3, IoU 0.7778 vs. 0.7765 in Table 5, and Comp-ACC 0.7448 vs. 0.7411 in Table 4. With single-seed evaluation and no error bars, the claims of consistent improvements are not substantiated. Please report bootstrap confidence intervals or multiple-seed standard deviations for the main metrics, especially for the ablation comparisons that drive the central claims.","section":"Tables 1–5"},{"comment":"The zero-shot transfer to FCDB is important evidence for generalization, but the paper does not check for image overlap or near-duplicates between the FCDB test set and the PICD-derived training data. PICD sources include Flickr-derived aesthetic datasets, and FCDB is also Flickr-derived, so near-duplicate leakage could inflate the transfer IoU reported in Table 2 (0.6251 for Stage I+II zero-shot). Please quantify source overlap and remove near-duplicate images before reporting transfer results.","section":"§5.2.2, Table 2"}],"minor_comments":[{"comment":"The reward weights (w_bde, lambda_s, alpha_g, alpha_s, alpha_f) are fixed on validation performance, but no sensitivity analysis is reported; a small table varying these weights would strengthen the claim that the method is robust to hyperparameter choice.","section":"§4.3, Eqs. (6)–(8)"},{"comment":"The paper reports that source filtering retains 44,201 images and that quality control removes approximately 25% of samples, but exact counts after each pipeline stage would improve reproducibility and clarify how the final 33,161 samples are reached.","section":"§3.2"},{"comment":"The pairwise preference protocol is underspecified: please state whether ties were allowed, how the 100 test images were sampled, how the 71 general users and 10 experts were recruited, whether evaluators were blind to condition, and report inter-annotator agreement.","section":"§5.3.1, Table 3"},{"comment":"The text refers to FLUX.1 [dev], while reference [22] cites FLUX.1 Kontext; please clarify which model version was used and include the exact outpainting prompt and sampling settings.","section":"References [22] and §3.2"},{"comment":"Several details essential for reproducibility, such as the full prompt template and the definitions of R_bde, R_exp, R_comp, and R_fmt, are deferred to the supplementary material; please ensure these are fully available at publication time.","section":"§4.2, §4.3"}],"recommendation":"major_revision","confidential_remarks":"The paper has a potentially useful benchmark and a clean training pipeline, but the central claims currently rest on unvalidated IO-reversal ground truth and evaluation metrics that are partly self-referential. The requested validation experiments—human preference checks on COMEX crops, boundary-artifact diagnostics, external explanation judges, and statistical significance reporting—are feasible within a revision. If the authors cannot provide them, I would not support acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid empirical paper for the aesthetic-cropping community. The crop-composition-explanation formulation is genuinely new, COMEX is the first large benchmark with crop-level composition annotations, and the two-stage SFT+GRPO recipe is clean and gives consistent gains across three backbones. The experiments are extensive: 15 MLLM baselines, FCDB transfer, ablations, and a human preference study. I would send this to a serious referee.\n\nThe main soft spot is the IO-reversal construction. Treating the original PICD image as the ideal crop of a FLUX-outpainted canvas turns the task into 'locate the original photograph inside the expanded image.' That is a well-posed supervised problem, but it is not obviously aesthetic cropping. The model could exploit boundary artifacts or layout priors instead of compositional reasoning. The manual quality control and the pixel-preserving design weaken that worry, and the zero-shot transfer to FCDB (0.6251 IoU) suggests the learned representation is not just about synthetic boundaries. Still, the composition-grounded claim is not isolated: the composition on/off ablation only runs on COMEX, and the explanation judges are Seed-family models that also produced the reference texts, so there is a circular flavor. Human preference (70%/76%) is encouraging but based on only 100 images.\n\nLesser issues: metrics are point estimates without variance or significance tests; the reward design has five validation-tuned weights; code and data are not released. The limitations section is honest about the synthetic data and proprietary annotations.\n\nThe framework's core empirical result—GRPO after SFT beats continued SFT—survives independent of the benchmark: it holds when training directly on FCDB and sets a new SOTA there. That is a real finding. The benchmark itself needs more validation: a real-image subset with human-annotated crops, or at least a non-Seed judge and significance testing, would address the main concern. I would accept this for peer review with a request for major revision, primarily to strengthen the benchmark's construct validity.","headline":"A solid empirical paper with a genuinely new benchmark and a clean two-stage training recipe, but the synthetic IO-reversal ground truth raises a real construct-validity question that the paper only partially answers.","tokens_in":16718,"tokens_out":3926,"would_cite":false,"duration_ms":37591,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By inserting an explicit composition category between the crop box and the explanation, this paper claims that models both crop better and explain their crops more faithfully, with a 2B-parameter model surpassing far larger rivals.","keywords":["aesthetic image cropping","image composition","vision-language model","GRPO","reinforcement learning","benchmark construction","image outpainting","explainable cropping"],"falsifier":"Ask expert photographers to independently choose the best crop of a sample of COMEX's expanded images without being told where the original photo lies; if their chosen crops frequently disagree with the stored ground-truth box, the IO-reversal supervision is not capturing what humans consider the ideal crop. A second direct check: re-annotate a random subset of COMEX crops with fresh composition labels by independent experts and measure agreement with the inherited PICD labels - low agreement would mean the composition supervision itself is unreliable and the explanation-grounding claims rest on noisy labels.","tokens_in":15770,"feed_emoji":"📐","tokens_out":6980,"duration_ms":52883,"temperature":0.7,"pith_summary":"The paper argues that aesthetic image cropping should not stop at predicting a crop box; it should also say why that crop is good, and that \"why\" should be grounded in photographic composition. It reformulates the task as a structured crop-composition-explanation problem and builds COMEX, a 33,161-sample benchmark of expanded image, crop box, composition category, and composition-grounded explanation quadruples. The central claim is that inserting composition as an explicit intermediate layer makes explanations more specific and more faithful to the cropping rationale while also improving the crop itself. A two-stage SFT+GRPO pipeline on a 2B-parameter vision-language backbone reaches the best results on both COMEX and the real-photo FCDB benchmark, suggesting that a small model plus task-aligned rewards can beat scale for explainable crop selection.","feed_headline":"Composition grounding lifts AI crop quality and explanation fidelity","feed_subtitle":"With composition labels and RL rewards, a 2B vision model beats far larger rivals at where and why to crop.","key_machinery":"The load-bearing object is the structured triplet $y=(b,c,e)$ - crop box, composition category, explanation - generated as one text sequence, with the composition category acting as the bridge that ties geometry to language. Under it sits the IO-reversal construction pipeline: the original photograph is preserved pixel-for-pixel inside an outpainted canvas and treated as the ideal crop, so composition labels transfer from image to crop box without manual re-annotation. Training runs in two stages: supervised fine-tuning teaches the output format, and Group Relative Policy Optimization (GRPO), a value-model-free reinforcement learning update that normalizes rewards within a sampled group, optimizes the total reward $R_{all} = 0.5 R_{box} + 0.175 R_{sem} + 0.325 R_{fmt}$, where $R_{box}$ combines IoU with boundary displacement error, $R_{sem}$ mixes explanation similarity with composition-category agreement, and $R_{fmt}$ enforces valid syntax. The ablation that carries the argument shows that with composition information, IoU rises from 0.7484 to 0.7529 and explanation win rates jump from roughly a third to two-thirds or more across all evaluator groups.","core_discovery":"The central claim is that composition is the missing intermediate layer in explainable aesthetic cropping: instead of generating an explanation after the fact from a predicted box, the model should first decide where to crop, name the composition category that justifies the placement (rule of thirds, centered single shape, horizontal arranged shapes, and so on), and then produce an explanation grounded in that category. To make this trainable, the authors construct COMEX by taking 49,123 expert-labeled composition images from PICD, filtering them, outpainting each one into a larger canvas with FLUX so the original photograph becomes the ground-truth crop, and using Seed-1.8 to write explanations conditioned on the crop box and composition category. They then train a vision-language model in two stages: structured supervised fine-tuning to establish the output protocol, followed by GRPO with rewards for box IoU and boundary accuracy, explanation-composition consistency, and output format. Their Qwen3-VL-2B model reaches 0.7765 mIoU, 0.7448 composition accuracy, and 0.5194 METEOR on COMEX and 0.7225 IoU on FCDB, and the ablations show that removing composition supervision cuts explanation preference sharply while composition-aware explanations win 66-76% of pairwise comparisons.","pith_inferences":["A risk the paper leaves implicit: because the ground-truth crop is always the original photo embedded in the outpainted canvas, the model may partly learn to find the real-photo region rather than a general aesthetic skill; the FCDB transfer results argue against this but do not fully rule it out, and a test on real photos with human-composed, crop-grounded explanations would settle it.","The 24-category composition taxonomy constrains what explanations can say; a finer-grained or hierarchical composition vocabulary could yield more specific explanations, and the same SFT+GRPO pipeline should transfer to it directly.","The win-rate gains suggest a cheap testable extension: prompt generic vision-language models with the composition category as an extra input at inference time and measure how much explanation quality improves without any retraining.","If the claims hold, the recipe - structured intermediate labels plus RL rewards aligned to those labels - should generalize to other perception tasks with human-interpretable intermediate concepts, such as design layout assessment or document formatting."],"forward_implications":["Composition supervision transfers across domains: a model trained only on the synthetic COMEX reaches 0.6251 IoU on real FCDB photos zero-shot, and fine-tuning on FCDB pushes IoU to 0.7225, ahead of prior methods.","GRPO is a genuine second stage: five epochs of reinforcement learning improve IoU from 0.7532 to 0.7765 and METEOR from 0.5129 to 0.5194, whereas five more epochs of SFT barely move any metric.","The semantic reward is what keeps explanations honest: removing $R_{sem}$ costs 2.25 points of composition accuracy and 1.34 points of METEOR while gaining only 0.13% IoU.","Small backbones suffice: with composition supervision, even a 0.8B-parameter model reaches 0.6927 mIoU after SFT, and the 2B model matches or beats prior dedicated cropping models while additionally producing composition labels and explanations."],"supporting_citations":[{"why":"PICD supplies all source images and expert composition labels that COMEX inherits through the IO-reversal procedure.","marker":"[60]"},{"why":"FLUX.1 provides the outpainting model used to build the larger-context canvas around each original photograph.","marker":"[22]"},{"why":"Seed-1.8 generates the composition-grounded explanations for each crop under composition-category constraints.","marker":"[7]"},{"why":"DeepSeek-R1 introduces GRPO, the value-model-free reinforcement learning update used as Stage II.","marker":"[13]"},{"why":"InstructCrop defines the prior crop-and-explain setting and the evaluation protocol that COMEX extends and outperforms.","marker":"[41]"},{"why":"Venus is the strongest prior explainable-cropping baseline that the proposed framework surpasses on COMEX metrics.","marker":"[12]"},{"why":"FCDB provides the real-world benchmark used to test cross-dataset transfer of the learned cropping ability.","marker":"[9]"},{"why":"CACNet is the strongest classic cropping baseline used for comparison on both COMEX and FCDB.","marker":"[15]"}],"fun_headline_variants":["Composition-first crops make AI explain itself better","A 2B model beats larger rivals by grounding crops in composition","COMEX benchmark ties crop location to composition reasoning","Why composition is the missing step in explainable cropping"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pipeline assumes that the original photograph, inside its artificially outpainted canvas, is the ideal crop for that expanded image, and that the composition label inherited from the source dataset remains correct at the crop-box level - if outpainting changes what a good crop is, or the label does not transfer to the crop, every downstream metric inherits the error.","fun_headline_variants_meta":{"raw":{"variants":["Composition-first crops make AI explain itself better","A 2B model beats larger rivals by grounding crops in composition","COMEX benchmark ties crop location to composition reasoning","Why composition is the missing step in explainable cropping"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00057,"raw_usage":{"total_tokens":2736,"prompt_tokens":1027,"completion_tokens":1709,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":643,"completion_tokens_details":{"reasoning_tokens":1645}},"tokens_in":643,"tokens_out":1709,"duration_ms":10520,"temperature":1.0,"reasoning_tokens":1645,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T00:33:07.338561+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask expert photographers to independently choose the best crop of a sample of COMEX's expanded images without being told where the original photo lies; if their chosen crops frequently disagree with the stored ground-truth box, the IO-reversal supervision is not capturing what humans consider the ideal crop. A second direct check: re-annotate a random subset of COMEX crops with fresh composition labels by independent experts and measure agreement with the inherited PICD labels - low agreement would mean the composition supervision itself is unreliable and the explanation-grounding claims rest on noisy labels.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"PICD supplies all source images and expert composition labels that COMEX inherits through the IO-reversal procedure."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Seed-1.8 generates the composition-grounded explanations for each crop under composition-category constraints."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"InstructCrop defines the prior crop-and-explain setting and the evaluation protocol that COMEX extends and outperforms."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Venus is the strongest prior explainable-cropping baseline that the proposed framework surpasses on COMEX metrics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"FCDB provides the real-world benchmark used to test cross-dataset transfer of the learned cropping ability."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"CACNet is the strongest classic cropping baseline used for comparison on both COMEX and FCDB."}],"review_version":1}