REVIEW 5 cited by
Consistency-diversity-realism Pareto fronts of conditional image generative models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Building world models that accurately and comprehensively represent the real world is the utmost aspiration for conditional image generative models as it would enable their use as world simulators. For these models to be successful world models, they should not only excel at image quality and prompt-image consistency but also ensure high representation diversity. However, current research in generative models mostly focuses on creative applications that are predominantly concerned with human preferences of image quality and aesthetics. We note that generative models have inference time mechanisms - or knobs - that allow the control of generation consistency, quality, and diversity. In this paper, we use state-of-the-art text-to-image and image-and-text-to-image models and their knobs to draw consistency-diversity-realism Pareto fronts that provide a holistic view on consistency-diversity-realism multi-objective. Our experiments suggest that realism and consistency can both be improved simultaneously; however there exists a clear tradeoff between realism/consistency and diversity. By looking at Pareto optimal points, we note that earlier models are better at representation diversity and worse in consistency/realism, and more recent models excel in consistency/realism while decreasing significantly the representation diversity. By computing Pareto fronts on a geodiverse dataset, we find that the first version of latent diffusion models tends to perform better than more recent models in all axes of evaluation, and there exist pronounced consistency-diversity-realism disparities between geographical regions. Overall, our analysis clearly shows that there is no best model and the choice of model should be determined by the downstream application. With this analysis, we invite the research community to consider Pareto fronts as an analytical tool to measure progress towards world models.
Forward citations
Cited by 5 Pith papers
-
GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation
GASS enhances fixed-prompt diversity in T2I models by expanding CLIP embedding spread along the text direction and a computed orthogonal background direction.
-
Scaling Group Inference for Diverse and High-Quality Generation
Groups of generated images become more diverse while staying high-quality when K outputs are chosen from M candidates via a quadratic integer program with progressive pruning.
-
Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation
A per-category best-of-K RL reward, multi-axis max@K, shifts SD3.5-M perceived-appearance distributions toward uniform coverage (Fairness Score +0.23 to +0.36) without quality loss.
-
EMAG: Self-Rectifying Diffusion Sampling with Exponential Moving Average Guidance
EMAG replaces selected attention maps with their exponential moving average during diffusion sampling, reporting +0.46 HPS over CFG on SD3 and composing with APG/CADS.
-
Breaking the Lock-in: Diversifying Text-to-Image Generation via Representation Modulation
Early DC component convergence in text-to-image Transformer features causes output homogeneity; selective early attenuation via DAVE improves diversity without retraining or extra cost.
Discussion (0). Sign in to comment.