Pith. sign in

REVIEW 16 cited by

T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.06350 v3 pith:QFLFEWDW submitted 2023-07-12 cs.CV

classification cs.CV
keywords benchmarkcompositionalmodelsrelationshipst2i-compbenchtext-to-imageenhancedmetrics
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Despite the impressive advances in text-to-image models, they often struggle to effectively compose complex scenes with multiple objects, displaying various attributes and relationships. To address this challenge, we present T2I-CompBench++, an enhanced benchmark for compositional text-to-image generation. T2I-CompBench++ comprises 8,000 compositional text prompts categorized into four primary groups: attribute binding, object relationships, generative numeracy, and complex compositions. These are further divided into eight sub-categories, including newly introduced ones like 3D-spatial relationships and numeracy. In addition to the benchmark, we propose enhanced evaluation metrics designed to assess these diverse compositional challenges. These include a detection-based metric tailored for evaluating 3D-spatial relationships and numeracy, and an analysis leveraging Multimodal Large Language Models (MLLMs), i.e. GPT-4V, ShareGPT4v as evaluation metrics. Our experiments benchmark 11 text-to-image models, including state-of-the-art models, such as FLUX.1, SD3, DALLE-3, Pixart-${\alpha}$, and SD-XL on T2I-CompBench++. We also conduct comprehensive evaluations to validate the effectiveness of our metrics and explore the potential and limitations of MLLMs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Using a synthetic alphabet video testbed, balanced data mixing and caption precision are shown to dominate T2V model quality, while CFG and fine-tuning only partially compensate for corrupted captions.

  2. Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    XTC-Bench reveals that strong performance on generation or understanding tasks in unified multimodal models does not guarantee cross-task semantic consistency, which instead depends on how tightly coupled the learning...

  3. Simile Understanding in Text-to-Image Models: An Evaluation Framework

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A YOLO-based evaluation framework shows that text-to-image models consistently render the literal vehicle of a simile, a failure that CLIPScore and PickScore miss.

  4. Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Poplar-9K is a curated human-centric image dataset generated through a reproducible attribute sampling, diffusion rendering, and vision-language inspection pipeline, with 9,401 accepted pairs from 11,765 candidates.

  5. The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    Safety-aligned T2I diffusion models exhibit semantic collapse in text embeddings causing TIFA drops; SAGE regularization restores structured utility while retaining safety.

  6. Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    CLVR couples verified logical planning with pixel diffusion, uses proxy reinforcement learning on distilled histories, and merges weights to cut inference to 4 NFEs while outperforming open-source T2I models on comple...

  7. Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    CLVR framework adds closed-loop visual verification, proxy prompt reinforcement learning, and delta-space weight merge to improve complex text-to-image generation over single-step or unverified multi-step baselines.

  8. LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    LLaSO releases a 3.8B speech-language model, 25.5M training instances, and an evaluation benchmark, claiming a normalized score of 0.72.

  9. Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A curated GPT-4o synthetic image dataset improves open-source generation models on instruction-following, surreal scenes, and multi-reference synthesis, plus two new benchmarks to measure those skills.

  10. Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

    cs.CL 2024-11 conditional novelty 6.0 of 10

    MoT decouples non-embedding parameters by modality in transformers to match dense multi-modal performance with roughly one-third to one-half the FLOPs.

  11. Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

    cs.CV 2024-03 conditional novelty 6.0 of 10

    Biased noise sampling for rectified flows combined with a bidirectional text-image transformer architecture yields state-of-the-art high-resolution text-to-image results that scale predictably with model size.

  12. MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding

    cs.CV 2026-08 conditional novelty 5.0 of 10

    MultiCompose combines embedding regularization, cross-attention suppression, and mask-guided denoising to compose independently personalized subjects into one image while keeping each subject's attributes exclusive.

  13. Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models

    cs.CV 2026-07 conditional novelty 5.0 of 10

    ATLAS adds a Think–Plan–Paint loop with shared positional tokens to unified MLLMs, plus RL-based layout alignment, achieving large reported gains over prior layout-based unified models on compositional image generatio...

  14. Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    SKD-CAG erases adversarial text triggers from diffusion models by distilling the model's own clean outputs through cross-attention guidance, claiming 100% and 93% removal for pixel and style backdoors.

  15. TextPixs: Glyph-Conditioned Diffusion with Character-Aware Attention and OCR-Guided Supervision

    cs.CV 2025-07 reject novelty 5.0 of 10

    The GCDA framework claims state-of-the-art text rendering in diffusion images via dual-stream encoding, attention segregation, and OCR supervision, but the paper lacks verifiable artifacts and contains internal incons...

  16. Towards Evaluating Robustness of Prompt Adherence in Text to Image Models

    cs.CV 2025-07 conditional novelty 4.0 of 10

    New benchmark results show that Stable Diffusion 3.x and Janus Pro models struggle to place simple geometric shapes in the correct image quadrant, with best F1 scores around 0.41 to 0.5.

Pith tools