Pith. sign in

REVIEW 4 cited by

StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.17249 v1 pith:QU5FZ23F submitted 2021-03-31 cs.CV cs.CLcs.GRcs.LG

classification cs.CVcs.CLcs.GRcs.LG
keywords manipulationlatentstyleganimageimagesinputtexttext-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Inspired by the ability of StyleGAN to generate highly realistic images in a variety of domains, much recent work has focused on understanding how to use the latent spaces of StyleGAN to manipulate generated and real images. However, discovering semantically meaningful latent manipulations typically involves painstaking human examination of the many degrees of freedom, or an annotated collection of images for each desired manipulation. In this work, we explore leveraging the power of recently introduced Contrastive Language-Image Pre-training (CLIP) models in order to develop a text-based interface for StyleGAN image manipulation that does not require such manual effort. We first introduce an optimization scheme that utilizes a CLIP-based loss to modify an input latent vector in response to a user-provided text prompt. Next, we describe a latent mapper that infers a text-guided latent manipulation step for a given input image, allowing faster and more stable text-based manipulation. Finally, we present a method for mapping a text prompts to input-agnostic directions in StyleGAN's style space, enabling interactive text-driven image manipulation. Extensive results and comparisons demonstrate the effectiveness of our approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TokenVerse: Versatile Multi-concept Personalization in Token Modulation Space

    cs.CV 2025-01 conditional novelty 7.0 of 10

    TokenVerse personalizes multiple visual concepts, including non-object concepts like pose and lighting, by learning per-text-token modulation offsets in a pretrained text-to-image Diffusion Transformer.

  2. GCA-3D: Towards Generalized and Consistent Domain Adaptation of 3D Generators

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GCA-3D adapts 3D generators to text or one-shot image domains without dataset synthesis, using depth-aware score distillation and hierarchical spatial consistency losses.

  3. DFCon: Attention-Driven Supervised Contrastive Learning for Robust Deepfake Detection

    cs.CV 2025-01 conditional novelty 3.0 of 10

    An ensemble of three pretrained vision transformers trained with supervised contrastive loss and majority voting reports 95.83% validation accuracy on the DFWild-Cup 2025 deepfake detection dataset.

  4. A Decade of Deep Learning: A Survey on The Magnificent Seven

    cs.LG 2024-12 reject novelty 2.0 of 10

    This is a review of seven influential deep learning models that does not present new research results and suffers from methodological and integrity issues.

Pith tools