{"id":"4ca9c24f-d2ec-4f37-83e4-8371793b1c6e","arxiv_id":"2412.15677","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"AIGI-VC is a 2,500-image benchmark for judging AI-generated ads on clarity and emotional impact, and current IQA metrics and open LMMs mostly fail at it.","lead":"This paper introduces AIGI-VC, a dataset of 2,500 AI-generated advertisement images with human preference ratings on information clarity and emotional interaction. It shows that most current quality metrics and open-source vision-language models perform poorly on judging ad effectiveness, while GPT-4o performs best.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground-truth preference labels are validated only at M=4 (≈80–100% coverage of exhaustive within-prompt pairs), yet the actual 2,000-pair protocol samples only ≈40% of the 5,000 possible pairs; if MAP estimates at that sparsity deviate from exhaustive human preferences, every benchmark comparison…","rationale":"The central claim is that AIGI-VC provides valid human-preference ground truth for information clarity and emotional interaction, and that current models perform poorly against it. All model comparisons—especially the headline result that GPT-4o significantly outperforms competitors—are evaluated against preference probabilities derived from MAP estimation over a subset of pairwise comparisons. If those probabilities do not reproduce exhaustive human preferences, the dataset's core contribution and the benchmark conclusions both collapse. The reader's weakest_assumption identifies this same dependence, and I agree with it. My stress-test sharpens the concern: the pilot validation is run at M=4, which corresponds to roughly 80–100% coverage of the 500 exhaustive pairs in the 250-image subset, while the actual annotation protocol samples only 2,000 of the 5,000 full-dataset pairs, i.e., 40% coverage. The paper gives no evidence that MAP estimation at 40% coverage recovers true preferences; it only shows recovery at M=4. This is a concrete, testable gap in the argument. Secondary issues—the partial circularity of GPT-4o-generated golden descriptions for the interpretation/reasoning benchmark, the absence of significance tests, and unclear dataset release—also weaken peripheral claims, but the label-reliability issue is more load-bearing because it underpins every quantitative result in the paper. The concern is addressable: a direct validation at the actual sampling density, or release of per-comparison data, would settle it. I therefore keep the reader's CONDITIONAL verdict unchanged rather than escalating to rejection, since the central dataset contribution may still be sound if the 40%-coverage validation passes.","tokens_in":15245,"tokens_out":7652,"duration_ms":69037,"concrete_test":"Re-run the pilot on the 250-image/50-prompt subset using the exact sampling density of the full protocol: retain 2,000/5,000 = 40% of the 500 same-prompt pairs (i.e., 200 pairs) with the same random-pairing procedure, estimate Q via Eq. (1), and compute rank correlation and choice accuracy against the exhaustive 500-pair ground truth, separately for all pairs and for pairs with preference probability outside [0.3,0.7]. Report both the M=4 result and the 40%-coverage result; if accuracy at 40% drops materially below the M=4 value, the 'same preferences' claim is unsupported. Also report how many of the full 5,000 pairs receive at least one label under the 2,000-comparison protocol.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The dataset's ground truth—and therefore every model ranking in Tables 2–5—depends on Eq. (1) MAP estimates from a partial sample of within-prompt pairwise comparisons. There are 5 images per prompt (2,500 images / 500 prompts), so exhaustive comparison requires 10 pairs per prompt, i.e., 5,000 pairs. The validation uses 250 images / 50 prompts, and the described M-round matching makes M=4 correspond to roughly 80–100% of the 500 exhaustive pairs (400–500 comparisons); Fig. 3 reports accuracy at M=4. The final annotation labels 2,000 pairs from the full 5,000-pair set, i.e., only 40% coverage, closer to M≈2 in the pilot. The statement that labeling 2,000 pairs reduces exhaustive comparisons by 60% 'while producing the same preferences' extrapolates from an 80–100%-density validation to a 40%-density protocol without any reported check. If the 2,000 labels are comparisons rather than distinct pairs, many of the 5,000 possible pairs receive zero or one vote, making MAP estimates especially noisy near the [0.3,0.7] ambiguity region; yet the paper evaluates both the full 'Dall' set and the strong-preference subset, so the Dall numbers rely on exactly these unvalidated labels. This is not a minor implementation detail: if 40%-density MAP preferences diverge from exhaustive human judgment, the central dataset contribution and the claim that GPT-4o significantly outperforms other models lose their ground-truth basis.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces AIGI-VC, a dataset of 2,500 AI-generated advertisement images produced from 500 prompts by five text-to-image models, spanning 14 ad topics and 8 emotion categories. Human annotators provided pairwise preferences along two dimensions, information clarity and emotional interaction; the paper then applies a Thurstone-based MAP model (Eq. 1) to estimate preference probabilities from a sampled subset of the exhaustive pairs. The dataset also includes fine-grained natural-language descriptions generated by GPT-4o and verified by human experts, intended to explain the reasons behind human preferences. The paper benchmarks 14 baselines (IQA metrics and large multimodal models, including GPT-4o) on three tasks: preference prediction, interpretation of human choices, and reasoning about image pairs. The main empirical claims are that existing IQA methods and open-source LMMs perform poorly on this task, while GPT-4o significantly outperforms all competitors, and that the dataset is the first to study communicability of AIGIs in visual communication.","tokens_in":15531,"tokens_out":5941,"duration_ms":46693,"significance":"If the ground-truth preferences are reliable, AIGI-VC fills a genuine gap: existing AIGI datasets focus on general quality or aesthetics, whereas this one targets information clarity and emotional interaction in an advertising context. The dataset construction covers diverse topics and emotions, and the authors provide code, three challenge subsets (human-object interaction, fantastical ads, positive/negative emotions), and a broad set of baseline models. These are strengths. However, the central empirical claims depend on two load-bearing assumptions: that MAP estimates from a 40%-coverage pair sample reproduce exhaustive human preferences, and that the GPT-assisted interpretation/reasoning evaluation is not biased toward GPT-4o. Both need stronger support before the dataset and the model rankings can be fully trusted.","major_comments":[{"comment":"The validation of MAP-estimated preferences is performed at M=4 rounds, which corresponds to roughly 80-100% of the exhaustive within-prompt pairs for the pilot (50 prompts with 5 images each imply 10 pairs per prompt, i.e., about 400-500 comparisons at M=4), yet the actual annotation protocol samples only 2,000 of the 5,000 possible pairs, i.e., 40% coverage, which is closer to M=2 density. The claim that labeling 2,000 pairs reduces exhaustive comparisons by 60% 'while producing the same preferences' is therefore an extrapolation from a denser regime. Please report the MAP accuracy at the actual sampling density (e.g., the M=2 point in Fig. 3, or a direct simulation on the pilot data that samples exactly 4 pairs per prompt and compares against exhaustive preferences). Without this, the preference probabilities used in all benchmark tables are not demonstrated to be reliable.","section":"Dataset Construction / Human Preference Annotation (Fig. 3)"},{"comment":"Model comparisons are reported only as point estimates with no confidence intervals or significance tests. For example, the claim that GPT-4o 'significantly outperforms all other competing models' rests on accuracy differences such as 0.7928 vs. 0.7518 (Table 3, IC Dall), with no measure of variability. Because the ground-truth labels are themselves MAP estimates from a sparse sample, label uncertainty propagates into the model rankings. Please provide bootstrap confidence intervals or paired significance tests (e.g., McNemar's test for accuracy) for α, ρ, and κ, so that 'significantly outperforms' is statistically supported.","section":"Experimental Settings / Tables 2-5"},{"comment":"The interpretation and reasoning benchmarks score LMM outputs against golden descriptions that were initially generated by GPT-4o and then verified by human experts, and the evaluation itself is described as 'GPT-assisted' following Q-Bench, without identifying the judge model. This design risks favoring GPT-4o in two ways: (i) the golden descriptions may retain GPT-4o's style and priorities even after human verification, and (ii) if the judge is also GPT-4o, the scoring is not independent. Please identify the judge model and either use a different model as the judge or provide a human-evaluated subset to demonstrate that the rankings in Tables 6-7 are not an artifact of self-scoring.","section":"Performance on Interpretation and Reasoning / Fine-grained Descriptions (Tables 6-7)"}],"minor_comments":[{"comment":"The PSA topics are listed as 'environment protection, animal rights, social welfare, safety, healthcare, and self-esteem,' which is six topics, while the dataset is described as having 14 topics and Figure 1 includes 'Bullying & violence' as the seventh PSA topic; please correct the enumeration.","section":"Data Collection"},{"comment":"The AIGI-VC row contains the typo 'Commnunication' for 'Communication'; the table header has the same typo.","section":"Table 1"},{"comment":"Equation (2) defines ρ as PLCC between ground-truth and predicted preference probabilities, but the conditioning notation P(X,Y)|Z is not defined clearly; please clarify that it denotes the preference probability distribution over pairs given the reference text or emotion.","section":"Experimental Settings / Eq. (2)"},{"comment":"The text states that 'HPSv2 achieves higher λ and ρ values,' but no λ is defined in the paper; presumably α is meant.","section":"Table 4 discussion"},{"comment":"The paper says 'we employ 14 objective metrics' but then lists LMMs (e.g., GPT-4o) which are not typically called objective metrics; consider renaming to '14 baseline models.'","section":"Experimental Settings"},{"comment":"Figure 2 does not label the columns with the generative model names; adding column labels would improve readability. Also, Stable Diffusion 2.0 and Dreamlike Photoreal 2.0 are both cited to Rombach et al. 2022, which is the original latent diffusion paper; more specific references would be appropriate.","section":"Figure 2 and model references"},{"comment":"The abbreviations 'Dall' and 'Dsub' are used without definition in the table captions; please define them when first used in the main text.","section":"Tables 2-3"},{"comment":"The conclusion does not mention any limitations of the dataset or the evaluation; a short limitations paragraph would help readers interpret the claims.","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The paper's main risk is the validation gap between the pilot MAP validation and the actual annotation protocol. If the authors can validate the 40%-coverage protocol directly (or at least report the M=2 accuracy), and if they address the potential self-scoring issue in the language evaluation, the dataset could be a valuable contribution. The statistical support for the model comparisons also needs to be added. These are fixable within revision, so major_revision seems appropriate rather than reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The dataset is a real contribution. It fills a gap: no prior benchmark looks at AIGIs specifically for visual communication in advertising, with information clarity and emotional interaction as separate dimensions. The two-dimensional coarse-grained preferences plus fine-grained descriptions are well-motivated, and the empirical finding that open-source LMMs are inconsistent while GPT-4o leads is plausible and likely robust. I would use this dataset in my own work if I were evaluating AIGI quality.\n\nThe soft spots are real but addressable. The biggest one is the validation of the 60% reduction claim. The pilot used 250 images from 50 prompts, where exhaustive comparison requires 10 pairs per prompt (500 total). At M=4, the pilot is effectively validating on roughly the full set of within-prompt pairs. The actual annotation samples 2,000 pairs from the 5,000 possible pairs — only 40% coverage, closer to M=2 in the pilot. No validation is reported at that density. The claim that the subsampled MAP estimates \"produce the same preferences\" is therefore not backed by the evidence. This matters because the ground-truth preferences used for every model ranking in Tables 2–5 depend on those MAP estimates. In the near-ambiguous [0.3, 0.7] region, 40% coverage could give noisy estimates. This is not a fatal flaw in the dataset's existence, but it is a load-bearing gap in the reliability argument. The authors should either validate at 2,000-pair density or release the full exhaustive labels.\n\nThere is also no significance testing or confidence intervals anywhere. \"GPT-4o significantly outperforms\" is stated without statistical support. The magnitude of the gap makes the conclusion likely, but a paired test across prompts would be easy and should be added.\n\nOn the interpretation/reasoning benchmark: the golden descriptions are initially generated by GPT-4o and then human-verified. Scoring GPT-4o against those descriptions is partially circular — the model is being rewarded for matching its own style of critique. Human verification reduces the problem but doesn't eliminate it. The authors should discuss this or regenerate descriptions with a different model.\n\nMinor: the dataset availability is not stated explicitly (only a code link), and Table 1 has a typo. Neither affects the contribution.\n\nOverall, this paper deserves a serious referee. The dataset is new, the annotations are thoughtful, and the failure mode of current LMMs is worth reporting. But the validation density issue and the circularity need to be fixed before the benchmark is relied on. I would send it to review with a request for revision.","headline":"A genuinely useful new benchmark for AI-generated image communicability, but the ground-truth preference labels are validated at a different sampling density than the one used, and the interpretation benchmark has an avoidable circularity.","tokens_in":16079,"tokens_out":2569,"would_cite":true,"duration_ms":23012,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces AIGI-VC, a 2,500-image benchmark arguing that current image-quality models cannot judge whether AI-generated ads communicate clearly and emotionally.","keywords":["AI-generated image quality assessment","visual communication","advertising","human preference dataset","information clarity","emotional interaction","large multimodal models","benchmark"],"falsifier":"Re-run a random sample of the 500 prompts with exhaustive pairwise comparisons among the five generated images using new raters, then compare the full ranking to the paper's MAP-estimated preferences; if accuracy drops well below the level reported in the pilot, the ground-truth labels are not stable enough to support the benchmark conclusions.","tokens_in":15013,"feed_emoji":"🖼️","tokens_out":8818,"duration_ms":70649,"temperature":0.7,"pith_summary":"The paper is trying to establish that the quality of an AI-generated advertisement cannot be judged by how realistic or aesthetically pleasing it is: what matters is whether it clearly delivers the intended message and triggers the intended emotion. To make this measurable, the paper introduces AIGI-VC, a 2,500-image database of AI-generated ads spanning 14 topics and 8 emotions, with human pairwise preferences on those two dimensions and written explanations of the preferences. On this database, the paper reports that existing image-quality metrics and open-source multimodal models rank far below human judgment, while the proprietary multimodal model GPT-4o comes closest across preference prediction, interpretation, and reasoning. If the benchmark is reliable, it gives automated image evaluation a concrete new target: communicative effectiveness rather than low-level fidelity.","feed_headline":"AI-generated ads defeat today's image-quality metrics","feed_subtitle":"A 2,500-image advertising benchmark with human labels for clarity and emotion exposes a wide gap in automated judges.","key_machinery":"The annotation protocol carries the argument. Raters compare pairs of images generated from the same prompt and choose which one conveys the text more clearly or evokes the intended emotion more strongly; a Thurstone Case V model with maximum-a-posteriori estimation turns the sparse pairwise votes into a global preference score, allowing 2,000 labeled pairs to stand in for the full set of comparisons. For fine-grained labels, the lower-ranked image is explained against a pseudo-reference: a large multimodal model drafts reasons using prescribed visual cues, and human experts verify and supplement those drafts into golden descriptions. The evaluation layer then measures whether automated models can predict, explain, and reason about these human preferences using accuracy, correlation, order-consistency, and GPT-assisted completeness, preciseness, and relevance scores.","core_discovery":"In the paper's own telling, the core discovery is that communicability is a measurable quality axis that current IQA tools miss. AIGI-VC contains 2,500 images generated from 500 advertisement prompts by five text-to-image models, covering 14 ad topics and 8 emotion categories, and it annotates each image pair for information clarity and emotional interaction separately. Benchmarking 14 metrics and 7 LMMs, the paper finds that CLIP-based preference metrics reach only about 0.75 accuracy on clarity and 0.69 on emotion, that open-source LMMs are close to chance or inconsistent when the order of the two images is flipped, and that GPT-4o reaches roughly 0.79 to 0.88 accuracy with much higher consistency. The paper also reports a strong correlation between the two dimensions (SRCC 0.9371, PLCC 0.9360), which it reads as evidence that clarity and emotional impact are closely linked. The conclusion is that automated evaluators are not yet effective for the communicability of AIGIs in visual communication.","pith_inferences":["A practical extension the authors do not draw: merging the two axes into one composite communicability score would likely preserve most ranking information given the 0.94 correlation, and could halve future annotation costs.","The same pairwise-protocol design could be reused for news illustrations, educational graphics, or social-media creatives, where clarity and emotion also determine whether an image works.","If a future open-source model matches GPT-4o's accuracy and consistency on AIGI-VC, that would show the current gap is a modeling problem rather than a missing benchmark.","The order-consistency failure of open-source LMMs suggests that any deployable ad-quality judge should be required to pass a counterbalanced presentation test."],"forward_implications":["Ad teams can use AIGI-VC to filter generated images for clarity and emotional impact before publication, since the benchmark provides human preference ground truth for both axes.","Future quality metrics for AI-generated content should be tested on preference prediction and on interpretation and reasoning, because the paper shows these abilities are not the same.","The high correlation between clarity and emotion preferences implies that a single ranking score captures much of what humans care about in ads, although the two dimensions remain separately labeled.","Open-source multimodal models need order-consistency improvements before they can serve as dependable automatic judges of visual communication."],"supporting_citations":[{"why":"Supplies Thurstone's Case V paired-comparison model used to estimate global preference scores from sparse human choices.","marker":"(Tsukida, Gupta et al. 2011)"},{"why":"Establishes pairwise-preference image quality assessment and the sampling practice the annotation protocol follows.","marker":"(Prashnani et al. 2018)"},{"why":"Suggests MAP estimation from a subset of comparisons and the two-alternative forced-choice prompting used in the experiments.","marker":"(Zhu et al. 2024a)"},{"why":"Provides Pick-a-pic, a prior preference dataset, and PickScore, one of the strongest benchmarked CLIP-based baselines.","marker":"(Kirstain et al. 2024)"},{"why":"Provides HPD v2 and HPSv2, a human-preference dataset and metric the comparison must beat.","marker":"(Wu et al. 2023)"},{"why":"Provides ImageReward, a CLIP-based AIGI preference metric and expert-comparison dataset used as a baseline.","marker":"(Xu et al. 2024)"},{"why":"Describes the proprietary multimodal model used for description generation and reported as the strongest judge on the benchmark.","marker":"(Achiam et al. 2023)"},{"why":"Defines the completeness, preciseness, and relevance criteria used to grade interpretation and reasoning against golden descriptions.","marker":"(Wu et al. 2024a)"},{"why":"Supplies the eight emotion categories used to define intended emotional interaction in the advertisement prompts.","marker":"(Mikels et al. 2005)"}],"fun_headline_variants":["AI ad quality benchmark stumps today's image metrics","New benchmark: AI ads hide quality from automated judges","Clarity and emotion in AI ads outwit current metrics","For AI ads, automated quality checks miss the point"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's ground-truth labels are only as trustworthy as the assumption that preference scores estimated from a sampled subset of pairwise comparisons match the preferences people would give in exhaustive comparison, and that assumption was tested on only 250 images.","fun_headline_variants_meta":{"raw":{"variants":["AI ad quality benchmark stumps today's image metrics","New benchmark: AI ads hide quality from automated judges","Clarity and emotion in AI ads outwit current metrics","For AI ads, automated quality checks miss the point"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000274,"raw_usage":{"total_tokens":1633,"prompt_tokens":929,"completion_tokens":704,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":640}},"tokens_in":545,"tokens_out":704,"duration_ms":6033,"temperature":1.0,"reasoning_tokens":640,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:11:14.430948+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run a random sample of the 500 prompts with exhaustive pairwise comparisons among the five generated images using new raters, then compare the full ranking to the paper's MAP-estimated preferences; if accuracy drops well below the level reported in the pilot, the ground-truth labels are not stable enough to support the benchmark conclusions.","supporting_citations":[{"cited_title":"R.; et al","cited_arxiv_id":null,"evidence_quote":"Supplies Thurstone's Case V paired-comparison model used to estimate global preference scores from sparse human choices."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes pairwise-preference image quality assessment and the sampling practice the annotation protocol follows."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides Pick-a-pic, a prior preference dataset, and PickScore, one of the strongest benchmarked CLIP-based baselines."},{"cited_title":"A.; Fredrickson, B","cited_arxiv_id":null,"evidence_quote":"Supplies the eight emotion categories used to define intended emotional interaction in the advertisement prompts."}],"review_version":1}