REVIEW 7 cited by
T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
read the original abstract
We propose T2I-ReasonBench, a benchmark evaluating reasoning capabilities of text-to-image (T2I) models. It consists of four dimensions: Idiom Interpretation, Textual Image Design, Entity-Reasoning and Scientific-Reasoning. We propose a two-stage evaluation protocol to assess the reasoning accuracy and image quality. We benchmark various T2I generation models, and provide comprehensive analysis on their performances.
Forward citations
Cited by 7 Pith papers
-
Do Image Editing Models Understand Lighting?
New 3DLP benchmark with real-world 1K HDR pairs shows state-of-the-art image editing models vary in physical lighting consistency, with best models close to reality but error-prone in low-light regions.
-
Do Image Editing Models Understand Lighting?
A 1,000-pair real-world HDR benchmark with two new affine-invariant error scores shows the best image-editing models reproduce the relative structure of real light transport but degrade in dim regions, and that VLMs f...
-
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage
Vision-language models exhibit literal superiority bias on noun compounds, with photorealistic visuals linked to poorer idiomatic grounding via new DIVA benchmark and Δ metric.
-
SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation
Introduces SciIR-82k dataset and SciIR-Bench for scientific image reasoning generation organized by Peirce's semiotic triad, with fine-tuning raising model score from 35% to 43%.
-
Faithful, Enriched, and Precise: Benchmarking Natural-Science Illustration Generation by T2I models
Introduces FEPBench benchmark to evaluate T2I models on instruction faithfulness, reasoning enrichment, and semantic precision for natural-science illustrations using atom set annotations.
-
Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation
Qwen-Image-Bench introduces a hierarchical creator-centric benchmark with 1000 prompts, 23 sub-capabilities, and a Q-Judger model that scores images on 56 verifiable facets to distinguish T2I models on fidelity and cr...
-
SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
A unified multimodal model can improve its own text-to-image generation by using its understanding module as a rewarder in a global-plus-local reward-weighted training loop.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.