A new human-annotated benchmark and rubric-conditioned evaluator show that text-to-image models can depict a poem's surface imagery but consistently fail to evoke its implicit emotion.
InAdvances in Neural Information Processing Systems (NeurIPS)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TangPoetryBench: A Multi-Dimensional Benchmark and Rubric-Conditioned Evaluator for Poetry-to-Image Generation
A new human-annotated benchmark and rubric-conditioned evaluator show that text-to-image models can depict a poem's surface imagery but consistently fail to evoke its implicit emotion.