REVIEW 7 cited by
Rethinking FID: Towards a Better Evaluation Metric for Image Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
As with many machine learning problems, the progress of image generation methods hinges on good evaluation metrics. One of the most popular is the Frechet Inception Distance (FID). FID estimates the distance between a distribution of Inception-v3 features of real images, and those of images generated by the algorithm. We highlight important drawbacks of FID: Inception's poor representation of the rich and varied content generated by modern text-to-image models, incorrect normality assumptions, and poor sample complexity. We call for a reevaluation of FID's use as the primary quality metric for generated images. We empirically demonstrate that FID contradicts human raters, it does not reflect gradual improvement of iterative text-to-image models, it does not capture distortion levels, and that it produces inconsistent results when varying the sample size. We also propose an alternative new metric, CMMD, based on richer CLIP embeddings and the maximum mean discrepancy distance with the Gaussian RBF kernel. It is an unbiased estimator that does not make any assumptions on the probability distribution of the embeddings and is sample efficient. Through extensive experiments and analysis, we demonstrate that FID-based evaluations of text-to-image models may be unreliable, and that CMMD offers a more robust and reliable assessment of image quality.
Forward citations
Cited by 7 Pith papers
-
Drop-In Perceptual Optimization for 3D Gaussian Splatting
WD-R, a lightly regularized Wasserstein Distortion loss, is preferred by humans 2.3× over the original 3DGS loss and 1.5× over Perceptual-GS while matching or reducing splat count and generalizing to other 3DGS framew...
-
MorphUNet: Alpha-Controlled Biometric Transport for Diffusion-Based Face Morphing Attacks
Parent-separated dual cross-attention inside a diffusion U-Net yields stronger, more balanced two-identity face morphs than StableMorph, MIPGAN-II, and MorDIFF on FEI and FRLL.
-
ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling
A single-step IMLE generator with per-stage supervision and a robust loss reports FID 2.56 on ImageNet-256 by filtering ~5% of samples at test time.
-
LeDiFlow: Learned Distribution-guided Flow Matching to Accelerate Image Generation
A learned distribution prior, predicted by a VAE-style decoder, lets flow-matching image models generate in fewer ODE steps than with a Gaussian prior.
-
Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation
SKD-CAG erases adversarial text triggers from diffusion models by distilling the model's own clean outputs through cross-attention guidance, claiming 100% and 93% removal for pixel and style backdoors.
-
Evaluation of Human Visual Privacy Protection: A Three-Dimensional Framework and Benchmark Dataset
The HR-VISPR dataset plus a privacy, utility, and practicality framework ranks 11 visual anonymization methods, with privacy scores based on classifier accuracy loss rather than measured human perception.
-
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models
A probabilistic overlap measure for object positions yields a human-aligned spatial relationship metric and a training-free generation guidance method for text-to-image models.
Discussion (0). Sign in to comment.