REVIEW 7 cited by
The Role of ImageNet Classes in Fr\'echet Inception Distance
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Fr\'echet Inception Distance (FID) is the primary metric for ranking models in data-driven generative modeling. While remarkably successful, the metric is known to sometimes disagree with human judgement. We investigate a root cause of these discrepancies, and visualize what FID "looks at" in generated images. We show that the feature space that FID is (typically) computed in is so close to the ImageNet classifications that aligning the histograms of Top-$N$ classifications between sets of generated and real images can reduce FID substantially -- without actually improving the quality of results. Thus, we conclude that FID is prone to intentional or accidental distortions. As a practical example of an accidental distortion, we discuss a case where an ImageNet pre-trained FastGAN achieves a FID comparable to StyleGAN2, while being worse in terms of human evaluation.
Forward citations
Cited by 7 Pith papers
-
H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks
H-Adapter uses a region-specific loss to induce disentangled cross-attention from which source-aligned hair masks are derived to guide diffusion inpainting, achieving strong results on pose-different hairstyle transfer.
-
Scalable GANs with Transformers
A transformer-only GAN trained in VAE latent space with multi-level noise supervision and width-scaled learning rates achieves FID 2.96 on ImageNet-256 in 40 epochs.
-
Breast Cancer Detection in Thermographic Images via Diffusion-Based Augmentation and Nonlinear Feature Fusion
Diffusion-based augmentation plus nonlinear feature fusion yields a reported 98.0% accuracy on a small breast thermogram dataset, but the evaluation protocol is not fully disclosed.
-
Revisiting Diffusion Models: From Generative Pre-training to One-Step Generation
Fine-tuning a pretrained diffusion model with a GAN objective and most weights frozen yields a one-step generator that matches or beats prior distillation methods on several datasets.
-
FreeScene: Mixed Graph Diffusion for 3D Scene Synthesis from Free Prompts
FreeScene parses free-form text and image prompts into scene graphs via a VLM-based Graph Designer, then generates 3D indoor layouts with a mixed graph diffusion transformer, reporting improved quality and controllabi...
-
First-Place Solution to NeurIPS 2024 Invisible Watermark Removal Challenge
A competition-winning pipeline removes 95.7% of StegaStamp and TreeRing watermarks on the NeurIPS 2024 benchmark by combining VAE fine-tuning, diffusion purification, and translation tricks.
-
Learning from Limited and Imperfect Data
A doctoral thesis compiling nine peer-reviewed papers on long-tailed image generation, long-tailed recognition, semi-supervised learning, and domain adaptation.
Discussion (0). Sign in to comment.