{"id":"25737d74-390e-4766-9fbe-54ba156bcedd","arxiv_id":"2505.11178","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"CompAlign is a new benchmark and evaluation pipeline for compositional text-to-image generation that also demonstrates moderate accuracy gains in smaller diffusion models after aligning them with its own evaluator.","lead":"This paper introduces CompAlign, a set of 900 complex text-to-image prompts that require placing up to six objects in 3D spatial layouts with specific colors and textures. It also presents CompQuest, an automatic evaluator, and a fine-tuning method that uses the evaluator's feedback to improve smaller open-source image models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No human validation of the gpt-4o-mini judge, which both creates preference labels and measures the reported gains, so the ACA improvements may reflect reward hacking rather than compositional ability; Appendix D's naturalness failures show the metric can be gamed.","rationale":"The reader's weakest assumption correctly identifies the gpt-4o-mini judge as the load-bearing premise. My stress-test sharpens this into a circularity problem: the same judge provides the training reward and the evaluation metric, so any judge-specific bias inflates apparent gains. Appendix D's naturalness analysis is direct, self-admitted evidence that ACA can be gamed by collage-like layouts, which is exactly the kind of shortcut an aligned model might exploit. The paper gives no human validation, no inter-annotator agreement, and no alternative metric, so the central claim is not yet established. This supports the reader's CONDITIONAL verdict: the method is plausible, but the evidence is insufficient without an independent check of the judge. I therefore keep the verdict unchanged rather than moving to accept or reject. The concrete test—human annotation of atomic questions plus an independent naturalness check—would settle whether the reported improvements reflect real compositional fidelity or reward hacking.","tokens_in":13247,"tokens_out":2722,"duration_ms":31053,"concrete_test":"Sample 100 generated images stratified across base, CompAlign-aligned, and CompBench++-aligned SD2 outputs, plus a few from closed-source models. Have three human annotators answer the same atomic questions used in CompQuest. Compute Cohen's kappa between gpt-4o-mini and the majority human label. Then recompute ACA with human labels and compare the aligned-vs-base gain. Also run the Appendix D naturalness judge on the aligned SD2 images; if the gain shrinks or reverses under human labels, or if naturalness drops while ACA rises, the concern lands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on CompQuest's atomic gpt-4o-mini feedback being an accurate measure of compositional correctness (Section 3.2). This judge is used twice: to construct per-image win/lose preferences (Eq. 2, Section 4.1) and to compute the reported ACA outcomes (Section 5.2). Because the training reward and the evaluation metric are identical, any systematic bias in the judge—e.g., overlooking wrong spatial layout, accepting separately cropped objects as a single scene, or ignoring attribute mismatches—is baked into both sides. The paper provides no human agreement analysis, no calibration against an independent metric, and no error bars. Appendix D directly demonstrates that high-ACA generations can be unnatural collages, and its naturalness scores (Table 8) show that even top models fail naturalness. This is evidence that ACA can be 'hacked' by layout tricks, and the same vulnerability applies to the alignment signal: an SD2 model optimized on ACA may learn to produce images that satisfy the judge's binary questions while remaining poor compositions. Without independent validation, the reported improvements (SD1.5: 43.31 to 46.94; SD2: 46.28 to 53.08, Table 2) are not trustworthy evidence for the method's effectiveness.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CompAlign, a 900-prompt benchmark for compositional text-to-image generation, emphasizing 3+ subjects, numerical/3D-spatial relations, and attribute bindings, together with CompQuest, an evaluation framework that decomposes prompts into atomic binary questions answered by an MLLM (gpt-4o-mini) and aggregates them into an ACA score. The authors evaluate 9 T2I models on CompAlign, finding that closed-source models outperform open-source ones and that performance degrades with spatial complexity. They then use ACA-based per-image win/lose signals (Eq. 2) to fine-tune SD1.5 and SD2 via a diffusion-alignment objective, reporting ACA improvements (SD1.5: 43.31 to 46.94; SD2: 46.28 to 53.08) that outperform the best CompBench++ checkpoint. The paper also includes an appendix case study showing that high-ACA images can be unnatural collages, and naturalness scores for five strong models.","tokens_in":13487,"tokens_out":2175,"duration_ms":22569,"significance":"If the ACA metric and the alignment results are valid, CompAlign would be a useful, more demanding benchmark for compositional generation, and the proposed feedback-driven alignment would offer a scalable, interpretable way to improve diffusion models. The benchmark construction is careful: 900 prompts balanced across five 3D-spatial configurations and multiple generation categories, with prompt templates that combine attributes and spatial layouts. The interpretability of atomic decomposition is a genuine strength, as is the reproducible, automated nature of the pipeline. However, the paper's central contribution rests on the accuracy of the gpt-4o-mini judge, which is used both as the training reward and as the evaluation metric. The paper provides no human validation, no agreement analysis, and its own Appendix D demonstrates that the metric can be gamed by models that produce collaged, unnatural scenes. Until the judge is validated against human judgments or an independent metric, the reported ACA numbers, including the claimed improvements, are not fully trustworthy evidence of compositional ability.","major_comments":[{"comment":"The same MLLM (gpt-4o-mini) is used to construct the per-image win/lose preference labels in Eq. (2) and to compute the reported ACA evaluation scores in Table 2. The paper provides no human agreement analysis, no calibration against an independent metric, and no error bars. Appendix D shows that models with near-perfect ACA can produce unnatural, collaged scenes, which is direct evidence that the judge can be gamed. Consequently, the reported improvements (e.g., SD2 from 46.28 to 53.08) may partly measure how well the fine-tuned model satisfies this particular judge rather than genuine compositional ability. The authors should validate the atomic binary feedback against human judgments on a sampled subset of CompAlign and, ideally, evaluate the aligned models with an independent metric (e.g., human preference or a second, different judge). This validation is load-bearing because it affects both the benchmark's evaluation claim and the alignment method's effectiveness claim.","section":"§3.2, §4.1, §5.2 (Eq. 2 and Table 2)"},{"comment":"The claim that CompAlign 'effectively and consistently improves the performance for both base diffusion models' is contradicted by the table itself. For SD1.5, object_color drops from 34.78 to 31.00 and 2rows×2sub drops from 35.00 to 31.67; for SD2, object_color_bathroom and object_color_kitchen show no improvement (both 57.14 and 60.00, respectively). The paper should either soften the consistency claim or explain why these subcategories are not expected to improve. In addition, the evaluation of different subcategories is based on small samples (50 entries for people_only and bathroom/kitchen categories), so the reported percentage differences may not be statistically meaningful. The authors should provide confidence intervals or significance tests.","section":"§5.2, Table 2"},{"comment":"The ACA metric in Eq. (1) aggregates independent yes/no judgments about each entity's presence, attribute, and position. This formulation cannot capture compositional coherence, such as whether multiple entities appear together in a single unified scene rather than as separately positioned objects or collaged pieces. Appendix D demonstrates exactly this failure mode: models can achieve high ACA by placing entities in separate scenes while violating naturalness. The authors should either extend the metric to include relational or scene-level judgments, or explicitly state that ACA measures per-entity accuracy only and does not claim to measure overall image-prompt alignment. As written, the abstract and Section 3.3 describe ACA as a measure of 'alignment between generated images and compositional prompts,' which overstates what the metric captures.","section":"§3.3, Eq. (1), Appendix D"}],"minor_comments":[{"comment":"The threshold tau in Eq. (2) is introduced as adjustable, but the paper never reports a sensitivity analysis or states how the specific value 0.5 was chosen. Adding a short analysis of how tau affects the win/lose balance and downstream ACA would make the method more reproducible.","section":"§4.1"},{"comment":"The comparison baselines are limited to CompBench++ checkpoints provided by Huang et al. Several relevant alignment methods for diffusion models (e.g., D3PO, DPO-based diffusion alignment, ImageReward, DreamSync) are discussed in Section 6.2 but not compared experimentally. Including at least one additional baseline would strengthen the claim that the proposed approach outperforms previous methods.","section":"§5.2 and Table 2"},{"comment":"The naturalness analysis is a valuable addition, but it relies solely on gpt-4o judgments without human validation. Since the main message of Appendix D is that automated metrics can be fooled, it would be more convincing to include human naturalness ratings or show at least a few human examples.","section":"Appendix D"},{"comment":"There are several typos and formatting issues: 'leverageing' in Section 3, 'we proposeCompQuest' in Section 3, 'compositionally' in Section 5.2 and Appendix D, and missing spaces around references in Sections 5.2 and 6.1. These should be corrected.","section":"Throughout"},{"comment":"The paper states that the 90-10 train-test split is 'balanced across sub-categories,' but does not describe the split procedure. For example, does each 3D-spatial configuration receive exactly 162 training and 18 test prompts? Adding the split details would facilitate reproduction and ensure no data leakage through prompt templates.","section":"§2.2, Table 4"}],"recommendation":"major_revision","confidential_remarks":"The central concern is circularity: the training reward and the evaluation metric are generated by the same unvalidated MLLM judge. The reader's stress-test note about this is well-founded and is, in my reading, the single most important issue. If the authors add human validation of the judge, report error bars, and temper the consistency claim, the paper would come much closer to being acceptable. The benchmark itself is a useful resource, but the reported alignment gains should be taken as preliminary until the metric is grounded."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the benchmark is real, the evaluation framework is clean, and the alignment results are plausible but not yet trustworthy. The judge is the load-bearing piece, and it is doing double duty as both the training reward and the reported metric.\n\nWhat is actually new: CompAlign extends T2I-CompBench++ in a meaningful way—900 prompts with up to six subjects, five 3D spatial configurations (rows/columns), and combined numerical, spatial, and attribute bindings. That is a harder and more structured test than the 2-entity prompts used in most prior benchmarks. CompQuest's atomic decomposition into binary questions is a good idea: it gives interpretable, per-entity feedback rather than a single opaque score. The alignment framework is a sensible adaptation of Li et al.'s per-image binary preference optimization to compositional generation, with a threshold tau that controls strictness. The authors also deserve credit for being upfront about the naturalness limitation in Appendix D—they show that high-ACA images can look like collages, and they report low naturalness scores even for strong models.\n\nWhere the soft spots are: the central issue is that the gpt-4o-mini judge is never validated against human judgments. The same binary answers produce the preference labels (Eq. 2, Sec. 4.1) and the ACA evaluation scores (Sec. 3.3). If the judge misses wrong spatial layouts or accepts separately cropped objects as a single scene, that bias is baked into both training and evaluation. Appendix D is not just a side note—it is direct evidence that ACA can be gamed, and the same gaming likely applies to the alignment signal. So the improvements in Table 2 (SD1.5: 43.31 to 46.94; SD2: 46.28 to 53.08) should be read as \"the model got better at satisfying this particular judge,\" not necessarily as improved compositional ability. The experimental scope is also narrow: only two older UNet models, no error bars or significance tests, and baselines are limited to Huang et al.'s checkpoints rather than other modern alignment methods like D3PO or ImageReward. Some subcategories actually drop (SD1.5 object_color goes from 34.78 to 31.00), so the gains are not as consistent as the abstract suggests. Data and code are promised but not released, which limits reproducibility.\n\nWho this is for: people working on compositional T2I evaluation or alignment. The benchmark is a useful resource once released, and the atomic decomposition framework could be adopted independently of the alignment recipe. The paper deserves serious peer review, but the referees should ask for a human agreement study on the judge, an independent evaluator for the alignment results, and a broader baseline set.","headline":"A genuinely harder compositional T2I benchmark with an interpretable atomic evaluation framework, but the alignment gains rest on an unvalidated MLLM judge that is used as both training reward and evaluation metric.","tokens_in":14011,"tokens_out":2191,"would_cite":false,"duration_ms":23725,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Decomposing complex prompts into atomic yes/no questions answered by a multimodal LLM yields an interpretable compositional-accuracy metric, and turning that metric into win/lose preference labels improves open-weight diffusion models on…","keywords":["compositional text-to-image generation","benchmark","3D-spatial relationships","attribute binding","multimodal large language model feedback","diffusion model alignment","preference optimization","compositional accuracy"],"falsifier":"Take a random sample of CompAlign prompts, have human annotators independently answer the same atomic questions, and compute agreement with gpt-4o-mini's answers; if agreement falls below roughly 85 percent, both the ACA metric and the preference labels derived from it lose their grounding. A cheaper probe is to test whether the judge credits an object as 'on the left' when that object also appears on the right side of the image.","tokens_in":13042,"feed_emoji":"🖼️","tokens_out":4150,"duration_ms":37331,"temperature":0.7,"pith_summary":"Compositional text-to-image generation, where a prompt specifies multiple objects, colors, textures, and 3D arrangements, still trips up strong models. This paper introduces CompAlign, 900 prompts that combine numeracy, attribute binding, and escalating 3D-spatial layouts, and CompQuest, an evaluator that breaks each prompt into atomic yes/no sub-questions answered by a multimodal LLM and aggregates them into a compositional accuracy score. It then argues that these scores, thresholded into per-image win/lose labels, can fine-tune diffusion models more effectively than previous preference pipelines. If that holds, the benchmark and the alignment recipe together give open-weight models a concrete route to close part of the gap with commercial systems.","feed_headline":"Atomic visual feedback lifts diffusion-model compositional accuracy","feed_subtitle":"A 900-prompt benchmark with 3D layouts shows open models can close part of the gap using win/lose preference signals.","key_machinery":"The load-bearing object is the atomic sub-question. Each CompAlign prompt is decomposed into n sub-questions, one per entity, each checking presence, attribute, and 3D position (for example, \"Is there a green eraser on the left of the second row?\"). A multimodal LLM returns a binary yes or no for each, and the Aggregated Compositional Accuracy is the fraction of yes answers. That same binary signal is thresholded to a win or lose label (win if ACA is at least 0.5) and fed into a KL-regularized binary preference objective, Equations 2 and 3, that updates the diffusion model's sampling policy.","core_discovery":"On the paper's own terms, the central discovery is that a benchmark built around 3+ subjects, five escalating 3D layouts, and natural attribute bindings separates current text-to-image models sharply, with closed-source systems leading and U-Net diffusion models falling below 50% average accuracy, and that the gap is actionable: training SD1.5 and SD2 with binary preference labels derived from CompQuest's atomic feedback raises average compositional accuracy from 43.31 to 46.94 and from 46.28 to 53.08 respectively, beating checkpoints trained through the previous benchmark's pipeline. A secondary claim is that CompQuest's decompose-then-verify design is more interpretable and robust than direct scalar scoring by a multimodal LLM or VQA-based counting.","pith_inferences":["Since the judge's binary labels are the only training signal, the same recipe could be repurposed to optimize other decomposable prompt properties, such as style consistency or object counts, simply by rewriting the atomic questions.","If the judge's yes/no answers are miscalibrated for rare objects or attributes, the preference signal would amplify that bias; checking whether alignment gains survive when the judge is swapped or human-validated would separate benchmark difficulty from judge artifacts.","The paper's own naturalness analysis shows that collage-like scenes can score high ACA, so folding a naturalness sub-question into the atomic set would make the metric harder to game and could change which models rank best."],"forward_implications":["Model performance degrades consistently as 3D-spatial layouts grow from one row of two subjects to two rows of three, so composition difficulty is monotone in layout complexity.","Closed-source commercial models outperform open-weight models by a wide margin on CompAlign, with the best open-weight model at 83.92% and the best closed-source model at 93.51%.","Fine-tuning with CompAlign feedback improves both SD1.5 and SD2 overall, with the largest gains on two-row spatial configurations such as +11.11% in the 2-row-by-1-subject setting, while checkpoints from the previous benchmark often underperform the base model.","The alignment pipeline is data-agnostic and can be augmented with generations from stronger models such as SD3.5 and DALL-E 3 to bootstrap weaker diffusion models.","Because CompQuest yields a per-image score rather than pairwise comparisons, the same pipeline can scale to larger training sets without human preference annotation."],"supporting_citations":[{"why":"Provides the prior compositional benchmark, its attribute vocabulary, and the comparison checkpoints that CompAlign extends and beats.","marker":"[13]"},{"why":"Supplies the per-image binary feedback alignment framework and the objective that Equation 3 directly builds on.","marker":"[16]"},{"why":"Introduces VQA-based feedback for aligning diffusion models, the approach CompQuest's atomic MLLM feedback is designed to improve upon.","marker":"[22]"},{"why":"Provides the prospect-theoretic utility optimization objective that grounds the win/lose preference training.","marker":"[6]"},{"why":"Establishes question-answering based faithfulness evaluation for text-to-image, which CompQuest adapts into atomic compositional checks.","marker":"[11]"},{"why":"SD3 is both one of the nine evaluated models and a source of augmented alignment data.","marker":"[5]"},{"why":"DALL-E 3 is a leading closed-source model used in evaluation and as a generator of augmented training data.","marker":"[18]"}],"fun_headline_variants":["Atomic feedback lifts T2I compositional accuracy on 3D spatial prompts","New benchmark grades T2I on 3D relationships; feedback improves diffusion","Diffusion models gain with fine-grained feedback on 3D composition","CompQuest feedback sharpens diffusion models on 3D-spatial prompts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire argument depends on gpt-4o-mini's binary answers actually tracking whether each entity, its attribute, and its position are correct in the generated image; the paper offers no human validation of this judge, and its own naturalness analysis shows that high scores can be earned by scenes that look like cut-and-paste collages.","fun_headline_variants_meta":{"raw":{"variants":["Atomic feedback lifts T2I compositional accuracy on 3D spatial prompts","New benchmark grades T2I on 3D relationships; feedback improves diffusion","Diffusion models gain with fine-grained feedback on 3D composition","CompQuest feedback sharpens diffusion models on 3D-spatial prompts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001237,"raw_usage":{"total_tokens":5104,"prompt_tokens":999,"completion_tokens":4105,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":615,"completion_tokens_details":{"reasoning_tokens":4026}},"tokens_in":615,"tokens_out":4105,"duration_ms":28585,"temperature":1.0,"reasoning_tokens":4026,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:55:34.785486+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of CompAlign prompts, have human annotators independently answer the same atomic questions, and compute agreement with gpt-4o-mini's answers; if agreement falls below roughly 85 percent, both the ACA metric and the preference labels derived from it lose their grounding. A cheaper probe is to test whether the judge credits an object as 'on the left' when that object also appears on the right side of the image.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the per-image binary feedback alignment framework and the objective that Equation 3 directly builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes question-answering based faithfulness evaluation for text-to-image, which CompQuest adapts into atomic compositional checks."},{"cited_title":"Dall ·e 3 system card, Oct 2023","cited_arxiv_id":null,"evidence_quote":"DALL-E 3 is a leading closed-source model used in evaluation and as a generator of augmented training data."}],"review_version":1}