Pith. sign in

REVIEW 7 cited by

The Role of ImageNet Classes in Fr\'echet Inception Distance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.06026 v3 pith:6CI2O6CI submitted 2022-03-11 cs.CV cs.AIcs.LGcs.NEstat.ML

classification cs.CVcs.AIcs.LGcs.NEstat.ML
keywords imagenetaccidentalclassificationsdistanceechetgeneratedhumanimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Fr\'echet Inception Distance (FID) is the primary metric for ranking models in data-driven generative modeling. While remarkably successful, the metric is known to sometimes disagree with human judgement. We investigate a root cause of these discrepancies, and visualize what FID "looks at" in generated images. We show that the feature space that FID is (typically) computed in is so close to the ImageNet classifications that aligning the histograms of Top-$N$ classifications between sets of generated and real images can reduce FID substantially -- without actually improving the quality of results. Thus, we conclude that FID is prone to intentional or accidental distortions. As a practical example of an accidental distortion, we discuss a case where an ImageNet pre-trained FastGAN achieves a FID comparable to StyleGAN2, while being worse in terms of human evaluation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    H-Adapter uses a region-specific loss to induce disentangled cross-attention from which source-aligned hair masks are derived to guide diffusion inpainting, achieving strong results on pose-different hairstyle transfer.

  2. Scalable GANs with Transformers

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A transformer-only GAN trained in VAE latent space with multi-level noise supervision and width-scaled learning rates achieves FID 2.96 on ImageNet-256 in 40 epochs.

  3. Breast Cancer Detection in Thermographic Images via Diffusion-Based Augmentation and Nonlinear Feature Fusion

    cs.CV 2025-09 reject novelty 6.0 of 10

    Diffusion-based augmentation plus nonlinear feature fusion yields a reported 98.0% accuracy on a small breast thermogram dataset, but the evaluation protocol is not fully disclosed.

  4. Revisiting Diffusion Models: From Generative Pre-training to One-Step Generation

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Fine-tuning a pretrained diffusion model with a GAN objective and most weights frozen yields a one-step generator that matches or beats prior distillation methods on several datasets.

  5. FreeScene: Mixed Graph Diffusion for 3D Scene Synthesis from Free Prompts

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FreeScene parses free-form text and image prompts into scene graphs via a VLM-based Graph Designer, then generates 3D indoor layouts with a mixed graph diffusion transformer, reporting improved quality and controllabi...

  6. First-Place Solution to NeurIPS 2024 Invisible Watermark Removal Challenge

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A competition-winning pipeline removes 95.7% of StegaStamp and TreeRing watermarks on the NeurIPS 2024 benchmark by combining VAE fine-tuning, diffusion purification, and translation tricks.

  7. Learning from Limited and Imperfect Data

    cs.LG 2025-07 unverdicted novelty 3.0 of 10

    A doctoral thesis compiling nine peer-reviewed papers on long-tailed image generation, long-tailed recognition, semi-supervised learning, and domain adaptation.

Pith tools