Pith. sign in

REVIEW 1 cited by

LatentKeypointGAN: Controlling Images via Latent Keypoints

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.15812 v5 pith:FNVUXWJQ submitted 2021-03-29 cs.CV

classification cs.CV
keywords imageskeypointsimagelatentkeypointganappearancecontrolembeddingsgenerated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative adversarial networks (GANs) have attained photo-realistic quality in image generation. However, how to best control the image content remains an open challenge. We introduce LatentKeypointGAN, a two-stage GAN which is trained end-to-end on the classical GAN objective with internal conditioning on a set of space keypoints. These keypoints have associated appearance embeddings that respectively control the position and style of the generated objects and their parts. A major difficulty that we address with suitable network architectures and training schemes is disentangling the image into spatial and appearance factors without domain knowledge and supervision signals. We demonstrate that LatentKeypointGAN provides an interpretable latent space that can be used to re-arrange the generated images by re-positioning and exchanging keypoint embeddings, such as generating portraits by combining the eyes, nose, and mouth from different images. In addition, the explicit generation of keypoints and matching images enables a new, GAN-based method for unsupervised keypoint detection.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single-image 3D keypoint estimator trained without 3D labels, using multi-view diffusion-generated views and features as self-supervision.

Pith tools