Pith. sign in

REVIEW 5 cited by

Consistency-diversity-realism Pareto fronts of conditional image generative models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.10429 v1 pith:YP5EVY3C submitted 2024-06-14 cs.CV cs.AI

classification cs.CVcs.AI
keywords modelsconsistencydiversityparetoworldconsistency-diversity-realismfrontsgenerative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Building world models that accurately and comprehensively represent the real world is the utmost aspiration for conditional image generative models as it would enable their use as world simulators. For these models to be successful world models, they should not only excel at image quality and prompt-image consistency but also ensure high representation diversity. However, current research in generative models mostly focuses on creative applications that are predominantly concerned with human preferences of image quality and aesthetics. We note that generative models have inference time mechanisms - or knobs - that allow the control of generation consistency, quality, and diversity. In this paper, we use state-of-the-art text-to-image and image-and-text-to-image models and their knobs to draw consistency-diversity-realism Pareto fronts that provide a holistic view on consistency-diversity-realism multi-objective. Our experiments suggest that realism and consistency can both be improved simultaneously; however there exists a clear tradeoff between realism/consistency and diversity. By looking at Pareto optimal points, we note that earlier models are better at representation diversity and worse in consistency/realism, and more recent models excel in consistency/realism while decreasing significantly the representation diversity. By computing Pareto fronts on a geodiverse dataset, we find that the first version of latent diffusion models tends to perform better than more recent models in all axes of evaluation, and there exist pronounced consistency-diversity-realism disparities between geographical regions. Overall, our analysis clearly shows that there is no best model and the choice of model should be determined by the downstream application. With this analysis, we invite the research community to consider Pareto fronts as an analytical tool to measure progress towards world models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation

    cs.CV 2026-02 conditional novelty 6.0 of 10

    GASS enhances fixed-prompt diversity in T2I models by expanding CLIP embedding spread along the text direction and a computed orthogonal background direction.

  2. Scaling Group Inference for Diverse and High-Quality Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Groups of generated images become more diverse while staying high-quality when K outputs are chosen from M candidates via a quadratic integer program with progressive pruning.

  3. Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A per-category best-of-K RL reward, multi-axis max@K, shifts SD3.5-M perceived-appearance distributions toward uniform coverage (Fairness Score +0.23 to +0.36) without quality loss.

  4. EMAG: Self-Rectifying Diffusion Sampling with Exponential Moving Average Guidance

    cs.CV 2025-12 conditional novelty 5.0 of 10

    EMAG replaces selected attention maps with their exponential moving average during diffusion sampling, reporting +0.46 HPS over CFG on SD3 and composing with APG/CADS.

  5. Breaking the Lock-in: Diversifying Text-to-Image Generation via Representation Modulation

    cs.CV 2026-06 unverdicted novelty 4.0 of 10

    Early DC component convergence in text-to-image Transformer features causes output homogeneity; selective early attenuation via DAVE improves diversity without retraining or extra cost.

Pith tools