Pith. sign in

REVIEW 13 cited by

Rethinking FID: Towards a Better Evaluation Metric for Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.09603 v2 pith:UEFV3TYN submitted 2023-11-30 cs.CV

classification cs.CV
keywords distancegeneratedimageimagesmetricmodelssampletext-to-image
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As with many machine learning problems, the progress of image generation methods hinges on good evaluation metrics. One of the most popular is the Frechet Inception Distance (FID). FID estimates the distance between a distribution of Inception-v3 features of real images, and those of images generated by the algorithm. We highlight important drawbacks of FID: Inception's poor representation of the rich and varied content generated by modern text-to-image models, incorrect normality assumptions, and poor sample complexity. We call for a reevaluation of FID's use as the primary quality metric for generated images. We empirically demonstrate that FID contradicts human raters, it does not reflect gradual improvement of iterative text-to-image models, it does not capture distortion levels, and that it produces inconsistent results when varying the sample size. We also propose an alternative new metric, CMMD, based on richer CLIP embeddings and the maximum mean discrepancy distance with the Gaussian RBF kernel. It is an unbiased estimator that does not make any assumptions on the probability distribution of the embeddings and is sample efficient. Through extensive experiments and analysis, we demonstrate that FID-based evaluations of text-to-image models may be unreliable, and that CMMD offers a more robust and reliable assessment of image quality.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. JSCGC: Joint Source-Channel-Generation Coding for Wireless Generative Communications

    cs.IT 2026-06 unverdicted novelty 7.0 of 10

    JSCGC replaces the conventional decoder with a generative model for controlled sampling, reformulating communication as mutual information maximization under perceptual constraints and showing improved semantic and di...

  2. Drop-In Perceptual Optimization for 3D Gaussian Splatting

    cs.CV 2026-03 accept novelty 6.5 of 10

    WD-R, a lightly regularized Wasserstein Distortion loss, is preferred by humans 2.3× over the original 3DGS loss and 1.5× over Perceptual-GS while matching or reducing splat count and generalizing to other 3DGS framew...

  3. MorphUNet: Alpha-Controlled Biometric Transport for Diffusion-Based Face Morphing Attacks

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Parent-separated dual cross-attention inside a diffusion U-Net yields stronger, more balanced two-identity face morphs than StableMorph, MIPGAN-II, and MorDIFF on FEI and FRLL.

  4. ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A single-step IMLE generator with per-stage supervision and a robust loss reports FID 2.56 on ImageNet-256 by filtering ~5% of samples at test time.

  5. Flow Matching in Feature Space for Stochastic World Modeling

    cs.CV 2026-06 conditional novelty 6.0 of 10

    FlowWM trains a flow-matching model in frozen DINOv3 feature space, using a one-step projection for temporal and task-driven losses, improving stochastic future prediction on a Waymo-derived benchmark.

  6. Flow Matching in Feature Space for Stochastic World Modeling

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    FlowWM applies flow matching directly in pretrained feature space with a one-step projection mechanism, improving perception accuracy, mode coverage, and horizon robustness on synthetic and real-world benchmarks.

  7. Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping

    cs.CV 2026-05 conditional novelty 6.0 of 10

    Mixing real UAV imagery with 2101 AI-generated image-mask pairs improves semantic segmentation F1 scores for fine-grained forest species by over 15 percentage points overall and up to 30 points for rare classes.

  8. Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping

    cs.CV 2026-05 conditional novelty 6.0 of 10

    Prompt-generated image-mask pairs, mixed with real UAV imagery at a 40:60 ratio, lift forest-regeneration segmentation by >15 F1 points over supervised baselines and sharply improve rare-species F1.

  9. Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics

    cs.CY 2026-04 unverdicted novelty 6.0 of 10

    Community members from the UK blind community, Kerala, and Tamil Nadu helped define what counts as culturally appropriate depictions of artifacts, and the authors tested whether those definitions can be turned into re...

  10. Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model

    cs.CV 2025-03 unverdicted novelty 6.0 of 10

    Seedream 2.0 is a native Chinese-English bilingual diffusion model that integrates a self-developed LLM text encoder, Glyph-Aligned ByT5, and Scaled ROPE to reach claimed state-of-the-art results in prompt following, ...

  11. Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics

    cs.CY 2026-04 unverdicted novelty 5.0 of 10

    Case studies with blind UK residents and people from Kerala and Tamil Nadu demonstrate that community input at the systematization stage produces culturally grounded definitions of appropriateness for text-to-image mo...

  12. Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    SKD-CAG erases adversarial text triggers from diffusion models by distilling the model's own clean outputs through cross-attention guidance, claiming 100% and 93% removal for pixel and style backdoors.

  13. Evaluation of Human Visual Privacy Protection: A Three-Dimensional Framework and Benchmark Dataset

    cs.CV 2025-07 conditional novelty 5.0 of 10

    The HR-VISPR dataset plus a privacy, utility, and practicality framework ranks 11 visual anonymization methods, with privacy scores based on classifier accuracy loss rather than measured human perception.

Pith tools