REVIEW 4 cited by
StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Inspired by the ability of StyleGAN to generate highly realistic images in a variety of domains, much recent work has focused on understanding how to use the latent spaces of StyleGAN to manipulate generated and real images. However, discovering semantically meaningful latent manipulations typically involves painstaking human examination of the many degrees of freedom, or an annotated collection of images for each desired manipulation. In this work, we explore leveraging the power of recently introduced Contrastive Language-Image Pre-training (CLIP) models in order to develop a text-based interface for StyleGAN image manipulation that does not require such manual effort. We first introduce an optimization scheme that utilizes a CLIP-based loss to modify an input latent vector in response to a user-provided text prompt. Next, we describe a latent mapper that infers a text-guided latent manipulation step for a given input image, allowing faster and more stable text-based manipulation. Finally, we present a method for mapping a text prompts to input-agnostic directions in StyleGAN's style space, enabling interactive text-driven image manipulation. Extensive results and comparisons demonstrate the effectiveness of our approaches.
Forward citations
Cited by 4 Pith papers
-
TokenVerse: Versatile Multi-concept Personalization in Token Modulation Space
TokenVerse personalizes multiple visual concepts, including non-object concepts like pose and lighting, by learning per-text-token modulation offsets in a pretrained text-to-image Diffusion Transformer.
-
GCA-3D: Towards Generalized and Consistent Domain Adaptation of 3D Generators
GCA-3D adapts 3D generators to text or one-shot image domains without dataset synthesis, using depth-aware score distillation and hierarchical spatial consistency losses.
-
DFCon: Attention-Driven Supervised Contrastive Learning for Robust Deepfake Detection
An ensemble of three pretrained vision transformers trained with supervised contrastive loss and majority voting reports 95.83% validation accuracy on the DFWild-Cup 2025 deepfake detection dataset.
-
A Decade of Deep Learning: A Survey on The Magnificent Seven
This is a review of seven influential deep learning models that does not present new research results and suffers from methodological and integrity issues.
Discussion (0). Continue with ORCID to comment.