{"id":"700500fc-0360-43ee-9015-fd089bb86f23","arxiv_id":"2608.11452","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new human-annotated benchmark and rubric-conditioned evaluator show that text-to-image models can depict a poem's surface imagery but consistently fail to evoke its implicit emotion.","lead":"The paper introduces TangPoetryBench, a benchmark of 1,280 images that tests how well four AI image generators illustrate 320 classical Chinese Tang poems, rated by humans on ten dimensions. It also trains an open evaluator, PoemAutoEvaluator, that reproduces human quality judgments and can rank a poem's illustrations as well as a strong proprietary AI judge.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PAE and the per-image quality scores rest on un-modeled single-rater noise for 82% of images; the reported MAE and parity with Claude may be measured against unstable ground truth.","rationale":"The paper has real strengths: a two-phase annotation protocol, adjudicated text labels, an explicit rubric, bootstrap stability for model-level comparisons, and detailed implementation reporting. These make the per-model diagnostic partly credible, because per-model means average over 320 images and uncorrelated rater noise tends to cancel. The evaluator layer is not similarly protected: PAE is trained and evaluated per image on single-rater labels, so rater noise directly affects the MAE, the within-0.05 fraction, and the ranking tau used to claim parity with Claude. This is exactly the reader's weakest assumption, and I agree with it. The Song Ci generalization, which rests on 20 images from one generator, is also thin, but it is a secondary generalization claim, whereas the rater-noise issue affects the principal quantitative results of the evaluator and the benchmark's scalable-annotation promise. Both issues are addressable with additional data and analysis, so the reader's CONDITIONAL verdict remains appropriate; no stronger verdict change is needed.","tokens_in":15736,"tokens_out":8011,"duration_ms":77950,"concrete_test":"Randomly select a representative set of held-out images (e.g., 50 poems x 4 models = 200 images) and have at least three fresh, qualified annotators rate each image with the same rubric and quality-control filters. Compute the multi-rater mean per dimension and per image, then re-evaluate PAE, Claude, and the open baseline against these consensus scores, reporting MAE, within-0.05 fraction, Kendall tau, and bootstrap 95% confidence intervals. Also compute the single-rater-versus-consensus noise floor as the mean absolute difference between one rater and the multi-rater mean. If PAE's MAE against consensus rises substantially above 0.129, or if its tau advantage over Claude disappears when compared on consensus rather than single-rater labels, the parity and scalability claims need to be weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 and Appendix D state that 1,527 ratings from 191 annotators cover 1,280 images, with only 230 images multi-rated. For the remaining 82% of images, the Phase-2 quality scores come from a single annotator, and those single-rater scores are used both as PAE's supervised training targets and as the human reference in Table 4. The paper reports only within-one-level agreement on the 230 multi-rated images: 73.1% for overall impression and 75.3% for emotional resonance. For 3- to 5-point rubrics this is a lenient statistic; it does not quantify how far a single rater typically deviates from the unobserved multi-rater mean, and no ICC, weighted kappa, or rater-noise floor is reported. The bootstrap in Section 3 resamples images while holding the single rating fixed, so it captures sampling variation across poems but not annotator variation. If single-rater noise is comparable to or larger than the claimed MAE of 0.129, then PAE's parity with Claude, its per-dimension diagnostics, and the claim that the benchmark can scale without fresh human annotation are not established. The paper itself acknowledges that residual subjectivity on emotion sets a natural ceiling, but it never converts that ceiling into an uncertainty estimate for the headline numbers.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces TangPoetryBench, a benchmark of 1,280 poem-to-image generations (320 Tang poems × 4 T2I models) annotated along ten dimensions, together with a diagnostic analysis of current T2I models and a rubric-conditioned evaluator, PoemAutoEvaluator (PAE). The benchmark uses a two-phase protocol: a poem-recognition phase and a multi-dimensional scoring phase, with 1,527 ratings from 191 annotators, of which 230 images are multi-rated. The diagnostic analysis reports a depict-to-evoke gap: models master visual surface dimensions but score lowest on emotional resonance and overall impression, with the paper attributing the emotion gap to a 'no readable face' failure mode. PAE fine-tunes Qwen3-VL-8B with LoRA and a GRPO stage on the human per-dimension scores, and is evaluated on held-out poems, reaching a per-poem ranking Kendall τ of 0.431, close to zero-shot Claude (0.444), while standard metrics such as CLIPScore, BLIPScore, and VQAScore show weak or no correlation with human ranking. The paper also reports generalization to an unseen generator (Kolors) and to a second poetic tradition (Song Ci, 20 images).","tokens_in":15839,"tokens_out":7773,"duration_ms":66538,"significance":"If the findings hold, TangPoetryBench would be a valuable resource for evaluating culturally grounded and affectively complex text-to-image generation, a domain that existing compositional benchmarks do not address. The manuscript is unusually transparent about its annotation process: it provides a detailed rubric, describes QC filters, adjudicates text-integrity labels, reports inter-annotator agreement, and includes bootstrap stability checks for the model ranking. The release of the benchmark, annotations, and an open evaluator is a practical contribution. The diagnostic claims—particularly the depict-to-evoke gap and the face-avoidance bottleneck—are actionable for model development. However, the strength of the paper's quantitative claims is currently limited because the per-image human scores that serve as PAE's training targets and as the evaluation ground truth are single-rater for 82% of images, no rater-noise floor is reported, and the cross-tradition generalization rests on a very small sample. These issues are fixable with additional analysis and reporting, so the central contributions are defensible but require revision.","major_comments":[{"comment":"The per-image quality scores used as PAE's supervised training targets and as the human reference in Table 4 are single-rater for 82% of images (1,050 of 1,280), yet the manuscript reports only within-one-level inter-annotator agreement on the 230 multi-rated images. This statistic is lenient: on the 3–4 level rubrics, a one-level difference can be as large as 0.33, which is much larger than PAE's reported MAE of 0.129. To interpret the headline numbers, the paper must report a rater-noise floor on the multi-rated subset—for example, the mean absolute deviation of single ratings from the multi-rater mean, an ICC, or a weighted kappa—and then compare PAE's and Claude's MAE against that floor. Without this, the claims of 'reproduces human per-dimension judgment' and 'parity with Claude' are not established, because the apparent agreement may be dominated by the noise of the unstable ground truth.","section":"§3, Appendix D, Table 5"},{"comment":"The generalization to Song Ci is based on only 20 images, each produced by a single generator, and the appendix does not state how many annotators rated these images or whether the same PAE checkpoint was used zero-shot or after any rubric-specific calibration. The main text claims that extending PAE to a new tradition 'required only a tradition-specific rubric, not retraining,' but the experimental detail needed to verify a zero-shot or minimally adapted transfer is missing. Please report the number of raters per Song Ci image, the exact protocol, and whether any calibration labels or additional fine-tuning were used; with 20 images, the current evidence is too thin to support the generalizable rubric-conditioning claim.","section":"§5, Appendix A"},{"comment":"The paper's central diagnostic finding—that the emotion bottleneck is caused by models avoiding a readable human face—is supported by a 'no readable face' category that does not appear in the rubric of Table 1 or Appendix C. No annotation protocol, definition, or inter-annotator reliability is given for this category, so the reader cannot tell whether the 15% figure and the per-model rates are reproducible expert judgments or informal post-hoc coding. Please specify who assigned this label, under what instructions, and with what agreement; this is load-bearing because the paper's main actionable guidance about expressive human faces rests on it.","section":"§4 ('What All Four Models Share')"},{"comment":"The parity claim between PAE (τ=0.431) and Claude (τ=0.444) is based on a difference of 0.013, but no confidence intervals or significance tests are reported for any of the Kendall τ values. Given the small number of held-out poems (64 poems, 256 images), the difference between the two models could easily be within sampling noise, and the claim that PAE 'reaches parity' or 'matches' Claude is not yet supported. Please provide bootstrap confidence intervals for τ and, ideally, a paired test across poems, for PAE, the open baseline, and Claude.","section":"Table 4, §5 ('Results')"}],"minor_comments":[{"comment":"The table is garbled in the current text: 'Seedream368 44' and 'MJ 27173' should be formatted as separate columns with clear counts (e.g., Seedream: 36 leakage, 8 fake, 44 total; MJ: 27 leakage, 46 fake, 73 total).","section":"Table 3"},{"comment":"The paper lists the QC filters (response-time filtering, pattern detection, accuracy floor, consistency checks) but does not report how many annotators or ratings were rejected by these filters; adding these counts would help the reader assess the quality of the retained annotations.","section":"§3 ('Annotation and adjudication')"},{"comment":"The term 'Quality' for the unweighted mean over applicable dimensions is easy to confuse with the 'Overall Impression' dimension; consider renaming the aggregate to 'Mean Dimension Score' or 'Aggregate Quality' for clarity.","section":"§3 ('Two reporting metrics')"},{"comment":"The bootstrap resampling over images holds each image's rating fixed and therefore does not capture annotator variability; this should be stated explicitly in the text so the reliability claim is not over-interpreted.","section":"§3 ('Reliability of the model comparison')"},{"comment":"The caption should explicitly state that PAE is fine-tuned on the benchmark's training poems while the open baseline and Claude are zero-shot; this framing clarifies the comparison and prevents an unfair 'fine-tuned vs. zero-shot' misreading.","section":"§5, Table 4 caption"},{"comment":"The abstract and introduction state that the benchmark, annotations, and evaluator are released, but no URL or repository is provided in the manuscript; please add an availability statement.","section":"§1, Data release"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid contribution to an underexplored evaluation area, and the authors are unusually transparent about their annotation procedure. However, the three load-bearing issues—missing rater-noise floor, thin Song Ci generalization evidence, and the undocumented face-avoidance category—need to be addressed before the paper can be accepted. The first is the most important: without a noise floor, the central quantitative claims about PAE are uninterpretable, even though the qualitative pattern of the benchmark (depict-to-evoke gap) is likely robust to single-rater noise because it is computed over means of 320 images per model. I recommend major revision rather than rejection, as all issues can be resolved with additional analysis and reporting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real contribution. TangPoetryBench is the first multi-dimensional, human-grounded benchmark for poetry-to-image generation that I know of, and the annotation protocol is more careful than most T2I evaluation work: two-phase design, adjudicated text labels, a Phase-1 accuracy floor, and bootstrap stability checks. The central diagnostic finding, that strong models handle the visual surface but fall short on implicit emotion, is internally consistent with the tables. The 'depict-to-evoke gap' is a useful name for something real, and the face-avoidance observation is plausible even if the coding protocol is under-described. PAE is a sensible extension of VLM-as-judge, and the comparison against scalar metrics and a zero-shot baseline is fair. The paper deserves a serious referee.\n\nThe soft spots are real but addressable. The stress-test note is right: about 82% of images have a single rating, and those single-rater scores serve as both PAE's training targets and the human reference in Table 4. The paper reports only within-one-level agreement on the 230 multi-rated images, which is a lenient statistic; that does not bound the deviation of a single rater from the unobserved multi-rater mean. The bootstrap resamples images but not raters, so it does not capture annotator noise. Given the claimed MAE of 0.129, a rater-noise floor of that size would undermine the parity claim with Claude. I don't think this sinks the paper, because the ranking claims (Kendall tau) average over 320 images per model and are more robust, but the MAE and 'parity' language overstates precision. The Song Ci generalization also rests on 20 images from one generator; the appendix reports per-image agreement, not ranking, so it is suggestive, not a validation. The face-avoidance diagnosis is stated numerically without a described coding protocol for 'no readable face'. And the code/data release is promised but not linked, which makes verification impossible right now.\n\nThe citation pattern looks fine; the related work is accurate and the paper does not oversell its novelty. I disagree with any framing of the single-rater issue as fatal; the central logic holds. But the authors should report ICC or weighted kappa, release the artifacts, and soften the parity claim until rater noise is quantified. If they do that, this is a solid benchmark paper worth citing.","headline":"A genuinely useful benchmark and a clever rubric-conditioned evaluator, but the headline parity claim needs stronger ground-truth uncertainty analysis before I'd trust it.","tokens_in":16482,"tokens_out":2051,"would_cite":true,"duration_ms":18554,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new benchmark shows text-to-image models render a poem's scene but miss its emotion, and an open evaluator can grade the gap.","keywords":["poetry-to-image generation","Tang poetry","text-to-image evaluation","benchmark","emotional resonance","rubric-conditioned evaluator","depict-to-evoke gap","multimodal LLM judge"],"falsifier":"Re-annotate a random sample of the 1,050 single-rated images with at least three raters and recompute the per-image quality scores; if PAE's per-dimension mean absolute error (0.129) and per-poem ranking τ (0.431) degrade substantially against this multi-rater ground truth, the parity-with-Claude claim is an artifact of annotation noise.","tokens_in":15392,"feed_emoji":"🖼️","tokens_out":7340,"duration_ms":54868,"temperature":0.7,"pith_summary":"The paper tries to establish that poetry-to-image generation can be judged along multiple human dimensions, and that when judged this way, current models master a poem's literal scene but fail to convey its implicit emotion. It argues that standard text-image alignment metrics cannot detect this failure, often inverting the human ranking, and that a new open evaluator conditioned on a written rubric can replicate human per-dimension judgments closely enough to rank images as well as a strong proprietary judge. If true, the field gains both a diagnostic target—emotional resonance, which is bottlenecked by the models' avoidance of readable human faces—and a way to scale that evaluation to new images and traditions without fresh human annotation.","feed_headline":"AI paints a poem's scene but misses its sorrow","feed_subtitle":"A new 1,280-image benchmark and open evaluator prove the gap—and measure it.","key_machinery":"The load-bearing object is the benchmark's two-stage annotation protocol: Phase 1 recognizability (a viewer identifies the correct poem from the image alone among four candidates) and Phase 2 per-dimension scoring on ten dimensions from safety and technical quality to core imagery and emotional resonance. The diagnostic that carries the argument is the depict-to-evoke gap, the consistent drop between near-ceiling surface dimensions (scene, cultural coherence, style) and meaning-bearing ones (core imagery, emotion). The mechanism attributed to this gap is face avoidance: models place figures in the right scene but omit a readable expression, and the rate of this failure tracks each model's emotion score. The evaluator argument is carried by rubric conditioning: PAE takes an image, its poem, and a written rubric defining score anchors, and outputs per-dimension scores after supervised fine-tuning plus a reinforcement stage on a Qwen3-VL-8B backbone.","core_discovery":"On its own terms, the paper's central discovery is that current text-to-image models share a 'depict-to-evoke gap': they score near the ceiling on the visual surface of a Tang poem—scene, cultural coherence, artistic style—but drop sharply on core imagery and more sharply on emotional resonance, the dimension that correlates most with holistic human judgment. The paper attributes this to a concrete rendering failure: models avoid drawing a readable human face, which is the natural carrier of feeling, and the frequency of this avoidance tracks each model's emotion score. The paper further claims that this gap is invisible to standard scalar metrics, which reward literal correspondence and rank the best and worst models almost identically, and that an open, rubric-conditioned evaluator (PAE) reproduces human per-dimension judgments, reaches a per-poem ranking τ of 0.431 against 0.444 for a strong proprietary judge, and transfers to an unseen generator and to the Song Ci tradition.","pith_inferences":["The face-avoidance finding suggests an intervention the paper does not run: taking a high-scoring image and inserting a legible expressive face (or measuring face detectability) should shift emotional-resonance scores; this would test whether the bottleneck is causal rather than correlational.","If the depict-to-evoke gap is a real property of current models, it likely reflects training on literal caption-image pairs, so a similar gap should appear in other implicit-meaning generation tasks (proverbs, abstract art, film stills) and could be tested with the same rubric-conditioned recipe.","Because PAE matches the proprietary judge on ranking but not on absolute per-image error, its practical role is comparative evaluation—ordering models and versions—rather than certifying a single image's quality.","All reported scores come from one fixed prompt per model, so the benchmark measures each generator at a single operating point; varying prompt specificity could reorder the models, a scoping limit the paper acknowledges and a natural next measurement."],"forward_implications":["Standard alignment metrics (CLIPScore, BLIPScore, VQAScore) cannot rank poetry illustrations and should be retired for this task in favor of per-dimension rubric scoring.","Improving a model's ability to render expressive human faces should directly raise emotional resonance, the dimension that most strongly drives human overall-impression ratings.","PAE's rubric conditioning means the benchmark can be extended to new dimensions, new generators, and new poetic traditions by supplying a rubric, not by retraining the evaluator.","For model builders, the data say effort is better spent on core imagery and emotion than on already near-ceiling surface dimensions such as style and safety.","Because poem difficulty tracks abstraction and allusion rather than length, the hardest poems are the natural target for knowledge-grounded or retrieval-augmented generation."],"supporting_citations":[{"why":"Supplies CLIPScore, the baseline scalar metric shown to be near-constant across good and bad poetry illustrations.","marker":"Hessel et al. 2021"},{"why":"Supplies the BLIP matching score, which correlates near zero with human quality in this task.","marker":"Li et al. 2022"},{"why":"Supplies VQAScore, the strongest scalar baseline, which still cannot separate the best model from the worst.","marker":"Lin et al. 2024"},{"why":"Provides Qwen3-VL-8B, the open backbone that PAE fine-tunes.","marker":"Qwen Team 2025"},{"why":"Supplies LoRA, the parameter-efficient fine-tuning method used to train PAE.","marker":"Hu et al. 2022"},{"why":"Supplies GRPO, the reinforcement stage that lifts ranking performance when initialized from supervised fine-tuning.","marker":"Shao et al. 2024"},{"why":"Provides EvalMuse-40K, a prior benchmark reporting annotator consistency used to contextualize the paper's agreement numbers.","marker":"Han et al. 2026"},{"why":"Surveys T2I evaluation papers and finds that none report inter-annotator agreement, motivating the benchmark's reporting.","marker":"Otani et al. 2023"},{"why":"Supplies Kolors, the unseen generator used in PAE's out-of-domain generalization test.","marker":"Kuaishou Kolors Team 2024"},{"why":"Supplies InternVL3.5 backbones used to show PAE's training recipe transfers across base models.","marker":"Wang et al. 2025"}],"fun_headline_variants":["AI depicts Tang poems but misses their emotion","Poetry-to-image AI flunks emotional resonance","AI renders poems but not their feelings","1,280 images show AI misses poetry's emotion","Benchmark reveals AI's poetry-to-image emotion gap"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quality scores that serve as PAE's training targets and as ground truth in the comparisons rest mostly on a single annotator per image: only 230 of the 1,280 images were multi-rated, and no correction for rater noise is applied to the rest.","fun_headline_variants_meta":{"raw":{"variants":["AI depicts Tang poems but misses their emotion","Poetry-to-image AI flunks emotional resonance","AI renders poems but not their feelings","1,280 images show AI misses poetry's emotion","Benchmark reveals AI's poetry-to-image emotion gap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001099,"raw_usage":{"total_tokens":4614,"prompt_tokens":1000,"completion_tokens":3614,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":616,"completion_tokens_details":{"reasoning_tokens":3544}},"tokens_in":616,"tokens_out":3614,"duration_ms":24994,"temperature":1.0,"reasoning_tokens":3544,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:12:35.421196+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a random sample of the 1,050 single-rated images with at least three raters and recompute the per-image quality scores; if PAE's per-dimension mean absolute error (0.129) and per-poem ranking τ (0.431) degrade substantially against this multi-rater ground truth, the parity-with-Claude claim is an artifact of annotation noise.","supporting_citations":[],"review_version":1}