Pith. sign in

REVIEW 6 cited by

ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.14339 v1 pith:IZ2XYB5X submitted 2024-08-26 cs.CV

classification cs.CV
keywords conceptmixconceptsmodelspromptstextcompositionalgenerationimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Compositionality is a critical capability in Text-to-Image (T2I) models, as it reflects their ability to understand and combine multiple concepts from text descriptions. Existing evaluations of compositional capability rely heavily on human-designed text prompts or fixed templates, limiting their diversity and complexity, and yielding low discriminative power. We propose ConceptMix, a scalable, controllable, and customizable benchmark which automatically evaluates compositional generation ability of T2I models. This is done in two stages. First, ConceptMix generates the text prompts: concretely, using categories of visual concepts (e.g., objects, colors, shapes, spatial relationships), it randomly samples an object and k-tuples of visual concepts, then uses GPT4-o to generate text prompts for image generation based on these sampled concepts. Second, ConceptMix evaluates the images generated in response to these prompts: concretely, it checks how many of the k concepts actually appeared in the image by generating one question per visual concept and using a strong VLM to answer them. Through administering ConceptMix to a diverse set of T2I models (proprietary as well as open ones) using increasing values of k, we show that our ConceptMix has higher discrimination power than earlier benchmarks. Specifically, ConceptMix reveals that the performance of several models, especially open models, drops dramatically with increased k. Importantly, it also provides insight into the lack of prompt diversity in widely-used training datasets. Additionally, we conduct extensive human studies to validate the design of ConceptMix and compare our automatic grading with human judgement. We hope it will guide future T2I model development.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A 3,068-prompt benchmark with per-instance Q&A scoring shows that current text-to-image models, including reasoning-enhanced ones, handle reasoning-driven prompts poorly, with mathematical reasoning near zero.

  2. How Do Diffusion Classifiers Decide? A Bias-Centric Evaluation

    cs.CV 2026-07 accept novelty 6.5 of 10

    Diffusion classifiers show lower attribute-misbinding CAB than OpenCLIP but larger size-order gaps and background-driven accuracy drops, traced to pixel-aggregated reconstruction error and cross-attention routing.

  3. ConceptMix++: Leveling the Playing Field in Text-to-Image Benchmarking via Iterative Prompt Optimization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Iteratively optimized prompts improve compositional text-to-image scores by up to 20% across three models, and the improved prompts transfer across models.

  4. OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A 57-task benchmark for multimodal image generation that uses automated visual parsers and an LLM judge to show GPT-4o-Native leads current models.

  5. Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ABP evaluates and improves how well text-to-image models render implicit real-world knowledge.

  6. Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Current image-text alignment metrics, including CLIPScore and DSGScore, produce unstable model rankings under random seeds and are highly sensitive to tiny image perturbations.

Pith tools