Pith. sign in

REVIEW 1 cited by

Annotated Hands for Generative Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.15075 v1 pith:HJ4VDDGA submitted 2024-01-26 cs.CV cs.AIcs.GR

classification cs.CVcs.AIcs.GR
keywords generativehandsimagesmodelshandadditionalannotationsapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative models such as GANs and diffusion models have demonstrated impressive image generation capabilities. Despite these successes, these systems are surprisingly poor at creating images with hands. We propose a novel training framework for generative models that substantially improves the ability of such systems to create hand images. Our approach is to augment the training images with three additional channels that provide annotations to hands in the image. These annotations provide additional structure that coax the generative model to produce higher quality hand images. We demonstrate this approach on two different generative models: a generative adversarial network and a diffusion model. We demonstrate our method both on a new synthetic dataset of hand images and also on real photographs that contain hands. We measure the improved quality of the generated hands through higher confidence in finger joint identification using an off-the-shelf hand detector.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HANDI: Hand-Centric Text-and-Image Conditioned Video Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    HANDI generates hand-centric videos from an image and text prompt via automatic motion-area localization and a hand refinement loss.

Pith tools