Pith. sign in

REVIEW 2 cited by

Designing an Encoder for StyleGAN Image Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.02766 v1 pith:J76XVN2W submitted 2021-02-04 cs.CV

classification cs.CV
keywords editingimagelatentstyleganimagesrealspaceallows
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, there has been a surge of diverse methods for performing image editing by employing pre-trained unconditional generators. Applying these methods on real images, however, remains a challenge, as it necessarily requires the inversion of the images into their latent space. To successfully invert a real image, one needs to find a latent code that reconstructs the input image accurately, and more importantly, allows for its meaningful manipulation. In this paper, we carefully study the latent space of StyleGAN, the state-of-the-art unconditional generator. We identify and analyze the existence of a distortion-editability tradeoff and a distortion-perception tradeoff within the StyleGAN latent space. We then suggest two principles for designing encoders in a manner that allows one to control the proximity of the inversions to regions that StyleGAN was originally trained on. We present an encoder based on our two principles that is specifically designed for facilitating editing on real images by balancing these tradeoffs. By evaluating its performance qualitatively and quantitatively on numerous challenging domains, including cars and horses, we show that our inversion method, followed by common editing techniques, achieves superior real-image editing quality, with only a small reconstruction accuracy drop.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TokenVerse: Versatile Multi-concept Personalization in Token Modulation Space

    cs.CV 2025-01 conditional novelty 7.0 of 10

    TokenVerse personalizes multiple visual concepts, including non-object concepts like pose and lighting, by learning per-text-token modulation offsets in a pretrained text-to-image Diffusion Transformer.

  2. DFCon: Attention-Driven Supervised Contrastive Learning for Robust Deepfake Detection

    cs.CV 2025-01 conditional novelty 3.0 of 10

    An ensemble of three pretrained vision transformers trained with supervised contrastive loss and majority voting reports 95.83% validation accuracy on the DFWild-Cup 2025 deepfake detection dataset.

Pith tools