{"id":"318edf9d-6e5e-45f9-9e65-7c656400ae54","arxiv_id":"2605.30257","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":2,"one_line_summary":"Stable-Layers applies Flow-GRPO with LoRA and a two-stage VLM scoring pipeline to improve layer decomposition without paired supervision, yielding stronger separation and lower reconstruction error on Crello.","lead":"Stable-Layers fine-tunes a pretrained image layer decomposition model using reinforcement learning where a vision-language model provides the reward signal instead of paired human labels. A smart generalist might read it to understand how AI feedback loops can reduce the need for expensive annotated datasets in computer vision tasks.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"VLM reward reliability unverified against human or independent metrics","rationale":"The reader's weakest assumption correctly isolates the single untested link required for the RL improvement to be credible. Full-text details on training or additional metrics do not remove the need for external validation of the VLM signal itself.","tokens_in":1674,"tokens_out":266,"duration_ms":16128,"concrete_test":"Sample 100 Crello images, generate 4 decompositions each from base and fine-tuned models, collect blinded human ratings on separation/artifact criteria, and compute Spearman rank correlation between the VLM two-stage scores and the human ratings; correlation <0.5 indicates the reward is misaligned.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim of improved layer separation and lower reconstruction error rests on Flow-GRPO successfully optimizing from the two-stage VLM scores. The per-sample + grid-calibration pipeline is described as solving score compression, yet the paper supplies no correlation analysis between those VLM scores and either human expert ratings or an objective proxy (e.g., layer-wise IoU against Crello ground-truth masks) on the same samples. Without that link, any measured drop in reconstruction error could be incidental rather than caused by a trustworthy reward signal.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes Stable-Layers, a reinforcement learning framework for fine-tuning a pretrained image layer decomposition model (Qwen-Image-Layered) using only VLM feedback via Flow-GRPO with LoRA. It introduces a two-stage VLM scoring pipeline (per-sample scoring on five criteria followed by grid-based side-by-side calibration) to address score compression and low variance in rewards. The method is claimed to produce better layer decompositions with stronger separation, fewer artifacts, and lower reconstruction error on the Crello dataset compared to the base model.","tokens_in":1785,"tokens_out":379,"duration_ms":23680,"significance":"If the VLM-based reward signal proves reliable, this approach could enable effective fine-tuning of decomposition models without requiring paired supervision data, which is often scarce. However, the current manuscript provides no quantitative results or validation of the reward, limiting the ability to assess its significance.","major_comments":[{"comment":"Abstract: the abstract states that Stable-Layers produces decompositions with stronger layer separation, fewer blank or artifact-heavy layers, and lower per-layer reconstruction error on the Crello dataset, but supplies no quantitative results, error bars, ablation details, or dataset statistics to support these claims.","section":"Abstract"},{"comment":"Abstract: the central claim relies on the two-stage VLM evaluation pipeline supplying a reliable reward signal, but no correlation analysis is provided between the VLM scores and human expert ratings or objective proxies such as layer-wise IoU against Crello ground-truth masks.","section":"Abstract"}],"minor_comments":[{"comment":"The description of the five edit-centric criteria could be expanded for reproducibility.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for highlighting the need for stronger quantitative support in the abstract and validation of the VLM-based reward. We will revise the manuscript accordingly to address these points while preserving the core contributions of the two-stage scoring pipeline and Flow-GRPO fine-tuning approach.","responses":[{"response":"We agree the abstract should be more self-contained with quantitative backing. In the revision we will insert concrete metrics drawn from the experimental section (e.g., mean per-layer reconstruction error reduction, layer-separation scores, and Crello dataset statistics) together with error bars from repeated runs and a brief reference to the ablation studies already present in the main text.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the abstract states that Stable-Layers produces decompositions with stronger layer separation, fewer blank or artifact-heavy layers, and lower per-layer reconstruction error on the Crello dataset, but supplies no quantitative results, error bars, ablation details, or dataset statistics to support these claims."},{"response":"We accept that explicit validation of the reward signal strengthens the central claim. The revised manuscript will include a new subsection reporting layer-wise IoU correlations against Crello ground-truth masks and, where feasible, a small-scale human rating study. If resource constraints limit the human study, we will clearly state this as a limitation while still providing the IoU analysis.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central claim relies on the two-stage VLM evaluation pipeline supplying a reliable reward signal, but no correlation analysis is provided between the VLM scores and human expert ratings or objective proxies such as layer-wise IoU against Crello ground-truth masks."}],"tokens_in":1301,"tokens_out":376,"duration_ms":16904,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The new piece is the two-stage scoring pipeline: first rate each decomposition independently on five edit-centric criteria, then re-score the whole group side-by-side in a grid so the VLM can spread its scores. That directly tackles the compression problem that usually kills within-group variance for GRPO. They apply it with Flow-GRPO plus LoRA on top of Qwen-Image-Layered and claim cleaner layers and lower reconstruction error on Crello. The approach is a narrow but concrete engineering step for anyone trying to use VLMs as reward models without paired data.\n\nThe execution looks honest on its own terms. They identify the variance issue, propose a fix, and run the RL loop. No obvious circularity in the reward definition.\n\nThe main weakness is that nothing in the abstract or stress-test note shows the VLM scores actually predict better decompositions. There are no correlation numbers against human raters, no layer-wise IoU against Crello masks, no ablation on the grid step, and no error bars or dataset sizes. Without that link, the reported gains could come from anything. The five criteria and grid parameters are also free knobs that might have been tuned to the test set.\n\nThis is for researchers already working on image layer decomposition or VLM-based RL in vision. A reader who needs a practical recipe for reward calibration might get something out of the method section. It is not a broad advance, but the core idea is clear enough that a serious referee could check whether the missing validation holds up in the full paper.","headline":"The two-stage VLM calibration gives a workable fix for reward variance in Flow-GRPO, but the paper shows no evidence that those scores track actual layer quality.","tokens_in":2282,"tokens_out":392,"would_cite":false,"duration_ms":14219,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Stable-Layers fine-tunes layer decomposition models with VLM-scored reinforcement learning and no paired supervision.","keywords":["layer decomposition","reinforcement learning","vision-language models","fine-tuning","image editing","GRPO","Crello dataset"],"falsifier":"Running the identical training loop but omitting the grid-based calibration step and checking whether within-group score variance collapses to the point that policy updates become negligible, or measuring whether the reported gains on layer separation and reconstruction error vanish on a fresh held-out image set.","tokens_in":2588,"feed_emoji":"🖼️","tokens_out":674,"duration_ms":23876,"temperature":0.7,"pith_summary":"The paper presents a reinforcement learning approach that starts from a pretrained layer decomposition model and improves it solely through feedback from a vision-language model. It solves the problem of narrow score ranges by using a two-stage pipeline: first scoring each candidate decomposition on five edit-centric criteria, then calibrating all candidates together in a grid view to restore variance. With this reward signal the method runs Flow-GRPO plus LoRA adaptation, sampling multiple decompositions per image and updating the policy from group-relative advantages. The result is decompositions that separate layers more cleanly, contain fewer blank or artifact-laden outputs, and show lower per-layer reconstruction error on the Crello dataset.","feed_headline":"VLM feedback alone fine-tunes layer decomposition models","feed_subtitle":"Two-stage scoring restores variance for GRPO updates, producing cleaner separated layers and lower errors on Crello.","key_machinery":"The two-stage VLM evaluation pipeline that pairs structured per-sample scoring across five edit-centric criteria with a grid-based calibration step in which the VLM re-scores all candidates side-by-side.","core_discovery":"Starting from Qwen-Image-Layered, Stable-Layers applies Flow-GRPO with LoRA adaptation, sampling multiple candidate decompositions per image, scoring them with a VLM through a two-stage pipeline of per-sample criteria scoring followed by grid-based side-by-side calibration, and optimising the policy from group-relative advantages, yielding stronger layer separation, fewer blank or artifact-heavy layers, and lower per-layer reconstruction error on Crello compared with the base model.","pith_inferences":["The same calibrated VLM reward construction could be tested on other generative tasks where human preference data is expensive but visual quality is easy to judge.","If the calibration step proves essential, future VLM-as-reward pipelines may need an explicit group comparison stage rather than isolated scoring.","The method could be re-run with different base VLMs to test whether the quality of the reward model itself limits further gains."],"forward_implications":["Decompositions exhibit stronger layer separation than the base model.","The number of blank or artifact-heavy layers decreases.","Per-layer reconstruction error drops on the Crello dataset.","Fine-tuning becomes possible without any paired ground-truth decompositions."],"fun_headline_variants":["VLM RL fine-tunes layer models","Stable-Layers uses VLM-scored GRPO","VLM two-stage scoring supports GRPO","Layer decomposition via VLM reinforcement learning"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The two-stage VLM evaluation pipeline supplies a sufficiently reliable and high-variance reward signal that enables effective policy improvement via Flow-GRPO.","fun_headline_variants_meta":{"raw":{"variants":["VLM RL fine-tunes layer models","Stable-Layers uses VLM-scored GRPO","VLM two-stage scoring supports GRPO","Layer decomposition via VLM reinforcement learning"]},"model":"grok-4.3","cost_usd":0.007415,"raw_usage":{"total_tokens":3394,"prompt_tokens":641,"num_sources_used":0,"completion_tokens":54,"cost_in_usd_ticks":74149500,"prompt_tokens_details":{"text_tokens":641,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2699,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":641,"tokens_out":54,"duration_ms":25634,"temperature":1.0,"reasoning_tokens":2699,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T08:23:57.520892+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the identical training loop but omitting the grid-based calibration step and checking whether within-group score variance collapses to the point that policy updates become negligible, or measuring whether the reported gains on layer separation and reconstruction error vanish on a fresh held-out image set.","supporting_citations":[],"review_version":1}