{"id":"913f790a-6e15-4e5c-a7c4-6703dc66be6f","arxiv_id":"2411.10183","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A text-to-image evaluation metric that combines ChatGPT-generated yes/no questions answered by a VQA model with a no-reference image quality score, with adjustable weights.","lead":"This paper proposes an automatic scoring method for text-to-image generation: ChatGPT turns the prompt into simple yes/no questions, a visual question answering model looks at the generated image and answers them, and the share of correct answers becomes the alignment score. It adds a no-reference image quality score and lets users weight the two components, but the evidence shown is only a few qualitative examples.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim depends on unvalidated ChatGPT-question/VQA yes/no channel; no human agreement or calibration shown for S_tia, so 'superior metric' claim lacks support.","rationale":"The paper proposes a plausible pipeline, but the central claim is strong: the proposed metric is 'the superior metric' for simultaneously assessing fine-grained text-image alignment and image quality with adjustable ratios. For that claim to hold, the TIA score must actually track alignment. The most load-bearing link is the unsupervised question-generation plus zero-shot VQA channel, because every subsequent comparison inherits its errors. The paper does not validate that channel on generated images: there is no human correlation, no calibration of BEIT-3, no ablation of the question-generation prompt, and no comparison with established VQA-based metrics. The qualitative figures show only a few examples, so they cannot establish superiority. Since the reader's weakest assumption identifies the same unvalidated VQA reliability, I agree with the reader's verdict. The missing evidence is not a matter of consensus; it is an absence of empirical support for a causal chain on which the whole metric depends. Therefore the reader's REJECT verdict remains appropriate.","tokens_in":5450,"tokens_out":3519,"duration_ms":37424,"concrete_test":"Build a benchmark from 100 COCO and 100 CUB prompts; for each prompt generate 3 images with a fixed text-to-image model. Have ChatGPT produce questions exactly as in Sec. II-A, then ask 3 human annotators to give the yes/no ground truth per question per image. Compute (1) BEIT-3 per-question accuracy relative to human labels and (2) Spearman correlation between S_tia and mean human alignment ratings. Also compare with TIFA and VQAScore on the same data. If per-question accuracy is below 90% or the Spearman correlation is below 0.7 with a nontrivial p-value, the proposed S_tia cannot support the superiority claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the proposed score is a superior metric requires S_tia to be a faithful measure of text-image alignment. That in turn requires two unvalidated links. First, the ChatGPT protocol in Sec. II-A must produce question sets whose yes/no answers are sufficient to detect every entity, attribute, and relation in the prompt; the paper reports no check of question coverage or ambiguity, and the instruction that all questions be answerable yes can produce trivial questions that are uninformative. Second, BEIT-3 as used in Sec. II-B must answer those questions accurately and without bias on synthetic generated images; the paper offers no calibration, no per-question accuracy, and no comparison with existing VQA-based metrics such as TIFA, DSG-1, or VQAScore. The only evidence is a few hand-selected qualitative rank comparisons (Figs. 2 and 3), with no human agreement data, error bars, or sample size. The heuristics for question count and the default weights W1=W2=0.5 are stated to be 'based on experimental results' but those results are not shown. If the VQA yes/no channel is noisy or biased, S_tia does not measure alignment, and the claimed superiority over ImageReward, CLIPScore, and MANIQA is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an automatic evaluation metric for text-to-image generation and text-guided image manipulation. The method uses ChatGPT to generate yes/no questions from the input prompt, uses the BEIT-3 VQA model to answer those questions on the generated image, and computes a Text-Image Alignment score as the fraction of 'yes' answers. This score is combined with a no-reference image quality score from MANIQA via a weighted sum, with user-adjustable weights. The authors claim the resulting metric is superior to CLIPScore, ImageReward, and MANIQA because it simultaneously evaluates fine-grained text-image alignment and image quality, and because the relative weighting can be adjusted. The experimental section presents qualitative comparisons in Figures 2 and 3 but provides no quantitative aggregation, no error bars, no human correlation, and no comparison with existing VQA-based metrics.","tokens_in":5724,"tokens_out":2651,"duration_ms":27516,"significance":"If the claims were properly validated, the proposed metric could be a practically useful tool, since it combines a fine-grained question-answering approach to text-image alignment with a no-reference quality score and offers an adjustable trade-off. The use of ChatGPT for automatic question generation is a reasonable way to avoid manual question templates, and the components (BEIT-3, MANIQA) are established models. However, the manuscript currently provides no quantitative evidence that the metric agrees with human judgments or outperforms existing metrics on any statistically meaningful sample. The central claim of superiority is therefore unsupported by the presented experiments.","major_comments":[{"comment":"The question-generation protocol is not validated. The paper states that ChatGPT generates simple yes/no questions but provides no check that the generated questions cover all entities, attributes, and relations in the prompt, and no analysis of how often questions are trivial or ambiguous. The rule for the number of questions (one for 1-2 words, plus one per additional 6 words) is said to be 'based on experimental results' but those results are not shown. Since the TIA score is simply the fraction of 'yes' answers, any bias or incompleteness in the question set directly undermines the validity of the metric.","section":"Section II-A"},{"comment":"The accuracy of the BEIT-3 VQA model on synthetic generated images is never established. No calibration, per-question accuracy, or comparison against human annotations is reported. If the VQA model systematically answers 'yes' or 'no' incorrectly on generated images, the resulting Stia does not measure text-image alignment. This is a load-bearing gap because the claimed 'superior metric' relies entirely on the faithfulness of the VQA channel.","section":"Section II-B"},{"comment":"The experimental evaluation is purely qualitative. Figures 2 and 3 show a handful of hand-picked examples, with no sample size, no error bars, no aggregate statistics, and no correlation with human ratings. The claim that the proposed method is superior to CLIPScore and ImageReward is not supported by any quantitative comparison. The paper also does not compare against existing VQA-based metrics such as TIFA, DSG-1, or VQAScore, which are directly related to the proposed approach.","section":"Section III"},{"comment":"Equation (1) combines Stia and Siqa with weights W1 and W2, but the paper does not discuss the numerical ranges or distributions of the two component scores. If Siqa and Stia are on different scales, the default W1=W2=0.5 may not actually give equal importance, and the adjustable weighting may not behave as claimed. A sensitivity analysis of the weights is also missing, even though the authors state the defaults are 'based on experimental results' without showing those results.","section":"Section II-D"},{"comment":"The abstract claims the metric assesses alignment 'for every individual object,' but the method does not identify or isolate objects. The score is a single aggregate fraction of 'yes' answers across all questions; there is no per-object breakdown or object-level evaluation. This overstates what the metric actually computes and should be corrected or substantiated with object-level results.","section":"Abstract and Section II"}],"minor_comments":[{"comment":"The capitalization of 'chatGPT' is inconsistent; the standard form is 'ChatGPT' when referring to the model.","section":"Throughout"},{"comment":"The phrase 'the proposed network' in Section II introduction is unclear, since no network is proposed; consider replacing with 'the proposed method.'","section":"Section II-C"},{"comment":"The figures are difficult to evaluate in grayscale because they rely on bold text and subtle score differences; adding numbered annotations or a table of scores would improve readability.","section":"Figures 2 and 3"},{"comment":"The paper says experiments use CUB and COCO datasets but does not state how many prompts or images were used, how text modifications were chosen, or whether the main results are reproducible from the described protocol.","section":"Section III"},{"comment":"The paper does not cite or discuss existing VQA-based evaluation metrics for text-to-image generation, such as TIFA, DSG-1, or VQAScore; these should be included in the related work and experimental comparison.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The core idea is not unreasonable, but the paper is far from a publishable state. The revision needs a full quantitative experimental section: human correlation, calibration of the VQA channel, comparison with TIFA/DSG-1/VQAScore, and statistical significance. If the authors cannot provide such evidence, the claim of superiority should be withdrawn. The current manuscript is more of a position statement than a validated method."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Zhou, quick read on arXiv:2411.10183. The short version: this is a re-implementation of TIFA (LLM-generated yes/no questions scored by a VQA model) with an NR-IQA term added, and neither TIFA nor VQAScore is cited. The adjustable weighted sum is the only new knob, and that doesn't change the underlying mechanism.\n\nWhat the paper does well is show, in a few qualitative examples, that CLIPScore can rank a mismatched text-image pair higher than a matched one, while the proposed score ranks correctly. That's a real failure mode of CLIPScore, and the proposed method does avoid it in those examples. Also, separating the TIA and IQA components and letting the user weight them is a sensible thing to want in a benchmark.\n\nThe soft spots are large. The empirical section is a handful of hand-picked rows in Figures 2 and 3. There is no human correlation, no error bars, no sample size, and no aggregate comparison with any VQA-based metric. The two most relevant baselines, TIFA and VQAScore, are not mentioned at all. The question-count heuristic and the default weights W1=W2=0.5 are said to be 'based on experimental results,' but those results are not shown. The instruction that all ChatGPT questions be answerable 'yes' is inherited from TIFA and brings the same ambiguity problem: a question like 'Is there a bird?' is guaranteed yes if the image has any bird, which doesn't verify the attribute or relation mentioned in the prompt. BEIT-3 is adopted without calibration on generated images, so the whole Stia channel is unvalidated.\n\nI agree with the reader's skeptical take. The central claim of superiority over ImageReward rests on a few images and no statistical evidence. That said, the failure mode they identify in CLIPScore is real, and the joint TIA+IQA score could be a practical addition to the toolbox if it were validated. The right move is to go back and compare against TIFA and VQAScore, add human judgments, and report the fitting of question counts and weights.\n\nFor a serious editor, I don't think this deserves referee time in its current form. The core method is prior art, the validation is absent, and the citation gap is damning. I'd desk reject with a clear message about existing VQA-based metrics. If the authors rework it into a proper comparative study, it could become a workshop-level note.","headline":"A re-implementation of TIFA with an added weighted IQA term, missing both TIFA and VQAScore citations and any real validation.","tokens_in":6205,"tokens_out":2513,"would_cite":false,"duration_ms":23951,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a VQA-based metric, using ChatGPT-generated yes/no questions answered by BEIT-3 plus a MANIQA quality score, evaluates text-to-image alignment more finely than CLIPScore and allows adjustable weighting of alignment…","keywords":["visual question answering","text-to-image generation","evaluation metrics","text-image alignment","no-reference image quality assessment","ChatGPT question generation","BEIT-3","MANIQA"],"falsifier":"Take a controlled set of generated images in which an object named in the prompt is definitely absent, ask BEIT-3 the corresponding yes/no question, and compare its answer with the known ground truth; if the VQA model answers 'yes' on a substantial share of absent-object questions, the TIA score is not a faithful alignment measure. A human-ranking study on the same images, comparing proposed TIA scores with human judgments of prompt fidelity, would also directly test the paper's central claim.","tokens_in":5290,"feed_emoji":"🖼️","tokens_out":5770,"duration_ms":52086,"temperature":0.7,"pith_summary":"The paper proposes a new automatic evaluation metric for text-to-image generation and text-guided image manipulation. Instead of comparing global image and text features as CLIPScore does, it has ChatGPT turn the input prompt into a list of simple yes/no questions about individual details, feeds each question together with the generated image to the BEIT-3 visual question answering model, and scores text-image alignment as the fraction of questions answered 'yes'. A no-reference image-quality score from MANIQA is then combined with the alignment score through user-adjustable weights, so the same metric can cover prompt fidelity and image quality. The paper reports that this approach tracks fine-grained alignment more reliably than CLIPScore and matches ImageReward while adding separate, reweightable quality and alignment scores.","feed_headline":"New metric asks yes/no questions to score text-to-image fidelity","feed_subtitle":"It pairs ChatGPT-generated questions with BEIT-3 answers and an image-quality score, beating CLIPScore on fine-grained alignment.","key_machinery":"The carrying mechanism is a two-stage scoring pipeline. An LLM (ChatGPT) generates simple yes/no questions, about seven words each, from the input prompt, with question count scaled to prompt length; the VQA model BEIT-3 answers each question about the generated image; and the Text-Image Alignment score is the proportion of 'yes' answers. The no-reference quality assessor MANIQA produces the image-quality score. The final score is the weighted sum $S_{\\mathrm{out}} = S_{\\mathrm{tia}} W_1 + S_{\\mathrm{iqa}} W_2$, with default weights $W_1 = W_2 = 0.5$.","core_discovery":"The paper's central claim is that a question-answering pipeline can measure whether each detail in a prompt survives into a generated image, something global embedding-similarity metrics miss. In the reported comparisons, CLIPScore sometimes assigns higher scores to image-text pairs whose text actually contradicts the image, whereas the proposed Text-Image Alignment score decreases as mismatched words are introduced, matching the intended ranking. The authors further claim that combining this TIA score with MANIQA's no-reference image-quality score gives a metric that simultaneously tracks alignment and quality, and that unlike ImageReward, it keeps the two components separate so users can reweight or view them independently.","pith_inferences":["Beyond the paper's own claims, the all-yes question design could make the score sensitive to how ChatGPT phrases questions; rephrasing or asking the same question in multiple ways would test robustness.","A natural extension is calibrating the VQA model on generated images, since BEIT-3 was not verified here on synthetic images; confidence-weighted answers could reduce systematic yes/no bias.","The pipeline could be used with other LLMs or applied to text-guided editing by restricting questions to edited regions, but these are extensions the paper does not demonstrate.","An independent human-ranking study on a larger prompt set would tell whether the TIA score agrees with human judgment beyond the few examples shown."],"forward_implications":["The metric can rank generated images by whether each object and detail in the prompt is actually present, not just by overall similarity in embedding space.","Because alignment and quality scores are computed separately, users can inspect either aspect alone or combine them with chosen weights.","The method applies to both text-to-image generation and text-guided image manipulation, since the same question-answering logic works on any text-image pair.","If the reported comparisons hold, the metric offers a practical alternative to CLIPScore for fine-grained prompt fidelity and to ImageReward when separate quality reporting is needed."],"supporting_citations":[{"why":"ChatGPT generates the yes/no questions from the prompt; the whole scoring pipeline starts from these questions.","marker":"[1]"},{"why":"BEIT-3 is the VQA model that answers the questions and produces the alignment signal.","marker":"[14]"},{"why":"MANIQA supplies the no-reference image-quality score combined into the final metric.","marker":"[12]"},{"why":"CLIPScore is the primary baseline for text-image alignment comparisons.","marker":"[10]"},{"why":"ImageReward is the state-of-the-art baseline it compares against for combined alignment and quality.","marker":"[13]"},{"why":"COCO provides one of the two datasets and prompt annotations used in the experiments.","marker":"[6]"},{"why":"CUB provides the other dataset and prompt annotations used in the experiments.","marker":"[7]"}],"fun_headline_variants":["VQA queries per object reveal skipped details in text-to-image","Per-object question answers beat CLIPScore for prompt fidelity","Ask ChatGPT for quiz, then test images on each answer","Alignment and quality scores stay separate for easy weighting","Contradictory images drop this metric while CLIPScore stays high"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The metric assumes that BEIT-3 answers ChatGPT-generated yes/no questions about synthetic generated images accurately, and that the fraction of 'yes' answers is a faithful measure of how well the prompt's details are reflected; the paper does not calibrate or verify this accuracy on generated images.","fun_headline_variants_meta":{"raw":{"variants":["VQA queries per object reveal skipped details in text-to-image","Per-object question answers beat CLIPScore for prompt fidelity","Ask ChatGPT for quiz, then test images on each answer","Alignment and quality scores stay separate for easy weighting","Contradictory images drop this metric while CLIPScore stays high"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000235,"raw_usage":{"total_tokens":1458,"prompt_tokens":860,"completion_tokens":598,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":476,"completion_tokens_details":{"reasoning_tokens":525}},"tokens_in":476,"tokens_out":598,"duration_ms":7113,"temperature":1.0,"reasoning_tokens":525,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:51:44.023531+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a controlled set of generated images in which an object named in the prompt is definitely absent, ask BEIT-3 the corresponding yes/no question, and compare its answer with the known ground truth; if the VQA model answers 'yes' on a substantial share of absent-object questions, the TIA score is not a faithful alignment measure. A human-ranking study on the same images, comparing proposed TIA scores with human judgments of prompt fidelity, would also directly test the paper's central claim.","supporting_citations":[{"cited_title":"Maniqa: Multi-dimension attention network for no-reference image quality assessment,","cited_arxiv_id":null,"evidence_quote":"MANIQA supplies the no-reference image-quality score combined into the final metric."}],"review_version":1}