REVIEW 7 cited by
Analyzing and Improving the Image Quality of StyleGAN
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The style-based GAN architecture (StyleGAN) yields state-of-the-art results in data-driven unconditional generative image modeling. We expose and analyze several of its characteristic artifacts, and propose changes in both model architecture and training methods to address them. In particular, we redesign the generator normalization, revisit progressive growing, and regularize the generator to encourage good conditioning in the mapping from latent codes to images. In addition to improving image quality, this path length regularizer yields the additional benefit that the generator becomes significantly easier to invert. This makes it possible to reliably attribute a generated image to a particular network. We furthermore visualize how well the generator utilizes its output resolution, and identify a capacity problem, motivating us to train larger models for additional quality improvements. Overall, our improved model redefines the state of the art in unconditional image modeling, both in terms of existing distribution quality metrics as well as perceived image quality.
Forward citations
Cited by 7 Pith papers
-
crossMoDA Challenge: Evolution of Cross-Modality Domain Adaptation Techniques for Vestibular Schwannoma and Cochlea Segmentation from 2021 to 2023
The 2022 and 2023 crossMoDA challenge results show that training on heterogeneous multi-institutional data reduces segmentation outliers on homogeneous test sets, while cochlea Dice declined in 2023.
-
Proto-LeakNet: Towards Signal-Leak Aware Attribution in Synthetic Human Face Imagery
Proto-LeakNet reaches 98.13% closed-set Macro AUC but only about 57% AUROC for open-set separation, directly contradicting the abstract's claim of strong generalization.
-
From Coarse to Fine: Learnable Discrete Wavelet Transforms for Efficient 3D Gaussian Splatting
AutoOpti3DGS uses learnable discrete wavelet transforms on input images to train 3DGS from coarse to fine, reducing peak Gaussian counts by about 18 to 23 percent with modest quality trade-offs.
-
Masked Conditioning for Deep Generative Models
Masking conditions during training with varying sparsity schedules lets small VAEs and latent diffusion models generate engineering designs from partially specified inputs.
-
Beyond Sliders: Mastering the Art of Diffusion-based Image Manipulation
Beyond Sliders augments Concept Sliders with perceptual, adversarial, and an undefined triplet loss, claiming better real-world edits, but the evidence is weak and the derivation is not valid.
-
Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward
A survey paper reviews multimodal generative AI and autoregressive LLMs for text-driven human motion generation, with comparative tables of models, datasets, and metrics.
-
Tackling fake images in cybersecurity -- Interpretation of a StyleGAN and lifting its black-box
A StyleGAN trained on CelebA is pruned and its latent dimensions tweaked, showing many weights are redundant and some dimensions correlate with features, but results are anecdotal and lack external validation.
Discussion (0). Sign in to comment.