Pith. sign in

REVIEW 23 cited by

GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.11513 v1 pith:QBNO4ZP6 submitted 2023-10-17 cs.CV cs.LG

GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

classification cs.CV cs.LG
keywords modelsgenevaltext-to-imageevaluateframeworkobjectalignmentautomated
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recent breakthroughs in diffusion models, multimodal pretraining, and efficient finetuning have led to an explosion of text-to-image generative models. Given human evaluation is expensive and difficult to scale, automated methods are critical for evaluating the increasingly large number of new models. However, most current automated evaluation metrics like FID or CLIPScore only offer a holistic measure of image quality or image-text alignment, and are unsuited for fine-grained or instance-level analysis. In this paper, we introduce GenEval, an object-focused framework to evaluate compositional image properties such as object co-occurrence, position, count, and color. We show that current object detection models can be leveraged to evaluate text-to-image models on a variety of generation tasks with strong human agreement, and that other discriminative vision models can be linked to this pipeline to further verify properties like object color. We then evaluate several open-source text-to-image models and analyze their relative generative capabilities on our benchmark. We find that recent models demonstrate significant improvement on these tasks, though they are still lacking in complex capabilities such as spatial relations and attribute binding. Finally, we demonstrate how GenEval might be used to help discover existing failure modes, in order to inform development of the next generation of text-to-image models. Our code to run the GenEval framework is publicly available at https://github.com/djghosh13/geneval.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance

    cs.LG 2026-04 unverdicted novelty 8.0

    FMRG reformulates guidance as deterministic optimal control, deriving a single-trajectory method using the flow map that matches or exceeds baselines on reward-guided generation and inverse problems with 3 NFEs at tex...

  2. RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution

    cs.CV 2026-05 conditional novelty 7.0

    RankE co-evolves AR policy and decoder via alternating ranking optimization, improving both FID and CLIP scores on LlamaGen-XL and Janus-Pro where policy-only RL degrades FID.

  3. How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance

    cs.LG 2026-04 unverdicted novelty 7.0

    FMRG is a training-free, single-trajectory guidance method for flow models derived from optimal control that achieves strong reward alignment with only 3 NFEs.

  4. Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models

    cs.CV 2026-04 unverdicted novelty 7.0

    XTC-Bench reveals that strong performance on generation or understanding tasks in unified multimodal models does not guarantee cross-task semantic consistency, which instead depends on how tightly coupled the learning...

  5. Oracle Noise: Faster Semantic Spherical Alignment for Interpretable Latent Optimization

    cs.CV 2026-04 unverdicted novelty 7.0

    Oracle Noise optimizes diffusion model noise on a Riemannian hypersphere guided by key prompt words to preserve the Gaussian prior, eliminate norm inflation, and achieve faster semantic alignment than Euclidean methods.

  6. Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning

    cs.CV 2026-04 unverdicted novelty 7.0

    Process-driven image generation decomposes text-to-image synthesis into interleaved cycles of textual planning, visual drafting, textual reflection, and visual refinement with dense consistency supervision.

  7. Reflective Flow Sampling Enhancement

    cs.CV 2026-03 unverdicted novelty 7.0

    RF-Sampling enhances flow matching models by implicitly performing gradient ascent on text-image alignment scores via linear textual combinations and flow inversion.

  8. Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching

    cs.LG 2025-09 conditional novelty 7.0

    Derives exact guidance transition rates for discrete flow matching models that require only one model evaluation per sampling step and unify prior approximation-based methods.

  9. RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

    cs.AI 2026-07 conditional novelty 6.0

    A unified transformer model generates language-and-image planning sequences for embodied tasks, and shows real-robot manipulation without large-scale action pretraining.

  10. Improving Sample Diversity in Autoregressive Text-to-Image Generation via Cluster Truncation

    cs.CV 2026-07 conditional novelty 6.0

    p-less cluster decoding, which truncates and samples over K-means clusters of visual tokens rather than individual tokens, yields higher per-prompt sample diversity than default or dynamic-temperature baselines on mos...

  11. The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models

    cs.CV 2026-07 unverdicted novelty 6.0

    Safety-aligned T2I diffusion models exhibit semantic collapse in text embeddings causing TIFA drops; SAGE regularization restores structured utility while retaining safety.

  12. NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning

    cs.LG 2026-06 conditional novelty 6.0

    NormGuard, a hinge penalty on excess velocity norm during RL post-training of flow models, improves perceptual quality and realism without sacrificing reward.

  13. NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning

    cs.LG 2026-06 unverdicted novelty 6.0

    NormGuard adds a training-time hinge penalty on velocity norm inflation in flow-matching RL to improve MLLM-judged image quality and forensic realism while preserving reward across multiple setups.

  14. NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning

    cs.LG 2026-06 conditional novelty 6.0

    A hinge regularizer that penalizes velocity-norm growth beyond the reference model improves perceptual quality and realism in RL-finetuned image flow models without sacrificing reward.

  15. Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

    cs.CV 2026-06 unverdicted novelty 6.0

    Qwen-Image-Agent bridges the context gap in text-to-image models via Context-Aware Planning and Context Grounding that integrate plan, reason, search, memory and feedback, achieving SOTA on IA-Bench and related benchmarks.

  16. Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

    cs.CV 2026-06 unverdicted novelty 6.0

    Qwen-Image-Agent is a unified agent framework that progressively builds sufficient generation context for T2I models via Context-Aware Planning and Context Grounding, achieving SOTA on IA-Bench, Mindbench, and WISE-Verified.

  17. Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning

    cs.CV 2026-05 unverdicted novelty 6.0

    CLVR couples verified logical planning with pixel diffusion, uses proxy reinforcement learning on distilled histories, and merges weights to cut inference to 4 NFEs while outperforming open-source T2I models on comple...

  18. Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning

    cs.CV 2026-05 unverdicted novelty 6.0

    CLVR framework adds closed-loop visual verification, proxy prompt reinforcement learning, and delta-space weight merge to improve complex text-to-image generation over single-step or unverified multi-step baselines.

  19. How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance

    cs.LG 2026-04 unverdicted novelty 6.0

    FMRG is a training-free single-trajectory guidance framework for flow-based models that matches or exceeds baselines on reward-guided tasks and inverse problems using as few as 3 NFEs.

  20. Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

    cs.CV 2024-03 conditional novelty 6.0

    Biased noise sampling for rectified flows combined with a bidirectional text-image transformer architecture yields state-of-the-art high-resolution text-to-image results that scale predictably with model size.

  21. VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment

    cs.CV 2026-03 conditional novelty 5.5

    A VGGT-based pointwise reprojection reward with geometry-aware sampling improves video geometric consistency via SFT/DPO and causal test-time search.

  22. Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding

    cs.CV 2026-04 unverdicted novelty 5.0

    UniRect-CoT is a training-free rectification chain-of-thought framework that treats diffusion denoising as visual reasoning and uses the model's inherent understanding to align and correct intermediate generation results.

  23. Optimizing Few-Step Generation with Adaptive Matching Distillation

    cs.CV 2026-02 conditional novelty 5.0

    Adaptive Matching Distillation uses reward-model scores to reweight teacher and fake-teacher gradients, improving few-step diffusion distillation on image and video benchmarks.