Pith. sign in

REVIEW 4 cited by

Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.04251 v3 pith:6V55NJ7N submitted 2024-04-05 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords metricspromptfaithfulnessimageserroradvancescoherencemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With advances in the quality of text-to-image (T2I) models has come interest in benchmarking their prompt faithfulness -- the semantic coherence of generated images to the prompts they were conditioned on. A variety of T2I faithfulness metrics have been proposed, leveraging advances in cross-modal embeddings and vision-language models (VLMs). However, these metrics are not rigorously compared and benchmarked, instead presented with correlation to human Likert scores over a set of easy-to-discriminate images against seemingly weak baselines. We introduce T2IScoreScore, a curated set of semantic error graphs containing a prompt and a set of increasingly erroneous images. These allow us to rigorously judge whether a given prompt faithfulness metric can correctly order images with respect to their objective error count and significantly discriminate between different error nodes, using meta-metric scores derived from established statistical tests. Surprisingly, we find that the state-of-the-art VLM-based metrics (e.g., TIFA, DSG, LLMScore, VIEScore) we tested fail to significantly outperform simple (and supposedly worse) feature-based metrics like CLIPScore, particularly on a hard subset of naturally-occurring T2I model errors. TS2 will enable the development of better T2I prompt faithfulness metrics through more rigorous comparison of their conformity to expected orderings and separations under objective criteria.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can Vision Language Models Understand Mimed Actions?

    cs.CL 2025-06 conditional novelty 7.0 of 10

    Vision-language models identify real actions with context far better than they identify mimed actions performed by 3D avatars, while humans are equally accurate on both.

  2. What makes a good metric? Evaluating automatic metrics for text-to-image consistency

    cs.CL 2024-12 conditional novelty 6.0 of 10

    None of the four tested text-to-image consistency metrics satisfies all proposed validity criteria, and the VQA-based metrics appear to rely largely on text priors such as yes-bias.

  3. EvalGIM: A Library for Evaluating Generative Image Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    EvalGIM packages text-to-image evaluation into a single extensible library with four 'Evaluation Exercises', two of which introduce new analysis methods for ranking robustness and balanced prompt-style comparisons.

  4. ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A human-scored benchmark shows that even GPT-4o averages below 4/5 correctness and all tested models struggle with scientific diagram prompts that combine spatial, numeric, and attribute requirements.

Pith tools