Pith. sign in

REVIEW 3 cited by

Effectively Unbiased FID and Inception Score and where to find them

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.07023 v3 pith:3QGDMNT6 submitted 2019-11-16 cs.CV cs.LG

classification cs.CVcs.LG
keywords scorefinitemodelscorescomputedeffectivelyinceptionnumber
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

This paper shows that two commonly used evaluation metrics for generative models, the Fr\'echet Inception Distance (FID) and the Inception Score (IS), are biased -- the expected value of the score computed for a finite sample set is not the true value of the score. Worse, the paper shows that the bias term depends on the particular model being evaluated, so model A may get a better score than model B simply because model A's bias term is smaller. This effect cannot be fixed by evaluating at a fixed number of samples. This means all comparisons using FID or IS as currently computed are unreliable. We then show how to extrapolate the score to obtain an effectively bias-free estimate of scores computed with an infinite number of samples, which we term $\overline{\textrm{FID}}_\infty$ and $\overline{\textrm{IS}}_\infty$. In turn, this effectively bias-free estimate requires good estimates of scores with a finite number of samples. We show that using Quasi-Monte Carlo integration notably improves estimates of FID and IS for finite sample sets. Our extrapolated scores are simple, drop-in replacements for the finite sample scores. Additionally, we show that using low discrepancy sequence in GAN training offers small improvements in the resulting generator.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Decomposable Probe for Few-Step Diffusion Models: Prompt, Latent, and Score Selectivity across Backbone Families and Distillation Paradigms

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A three-layer perturbation probe shows latent selectivity is a near-binary rectified-flow fingerprint that survives ADD distillation, while score selectivity tracks distillation objective across 23 T2I models.

  2. Augmented Conditioning Is Enough For Effective Training Image Generation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Conditioning a pretrained diffusion model on augmented real images and class labels produces synthetic training data that improves downstream classification accuracy without any generator fine-tuning.

  3. HistoFID- Calibrating Frechet-distance evaluation across pathology foundation models

    eess.IV 2026-07 conditional novelty 5.0 of 10

    Raw Fréchet distances in pathology vary ~30-fold across encoders; dividing by each encoder's own within-cohort floor cuts cross-encoder variation by ~89% within and ~58% across cohorts.

Pith tools