Pith. sign in

REVIEW 1 cited by

Conditional Image Generation and Manipulation for User-Specified Content

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.04909 v1 pith:GYK44SG4 submitted 2020-05-11 cs.CV

classification cs.CV
keywords imageconditioningcontentfacialgenerationimagesmanipulationpipeline
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, Generative Adversarial Networks (GANs) have improved steadily towards generating increasingly impressive real-world images. It is useful to steer the image generation process for purposes such as content creation. This can be done by conditioning the model on additional information. However, when conditioning on additional information, there still exists a large set of images that agree with a particular conditioning. This makes it unlikely that the generated image is exactly as envisioned by a user, which is problematic for practical content creation scenarios such as generating facial composites or stock photos. To solve this problem, we propose a single pipeline for text-to-image generation and manipulation. In the first part of our pipeline we introduce textStyleGAN, a model that is conditioned on text. In the second part of our pipeline we make use of the pre-trained weights of textStyleGAN to perform semantic facial image manipulation. The approach works by finding semantic directions in latent space. We show that this method can be used to manipulate facial images for a wide range of attributes. Finally, we introduce the CelebTD-HQ dataset, an extension to CelebA-HQ, consisting of faces and corresponding textual descriptions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Disentangling 3D from Large Vision-Language Models for Controlled Portrait Generation

    cs.CV 2025-06 conditional novelty 8.0 of 10

    CLIPortrait disentangles camera and geometry information from CLIP embeddings via 2D canonicalization, then prevents distribution collapse with a Jacobian regularizer, enabling text-guided 3D portrait generation from ...

Pith tools