Pith. sign in

REVIEW 2 cited by

Joker: Conditional 3D Head Synthesis with Extreme Facial Expressions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.16395 v1 pith:AUSDH5LB submitted 2024-10-21 cs.CV cs.GR

classification cs.CVcs.GR
keywords extremeexpressionsmethodpriorarticulationconditionaldistillationexpression
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Joker, a new method for the conditional synthesis of 3D human heads with extreme expressions. Given a single reference image of a person, we synthesize a volumetric human head with the reference identity and a new expression. We offer control over the expression via a 3D morphable model (3DMM) and textual inputs. This multi-modal conditioning signal is essential since 3DMMs alone fail to define subtle emotional changes and extreme expressions, including those involving the mouth cavity and tongue articulation. Our method is built upon a 2D diffusion-based prior that generalizes well to out-of-domain samples, such as sculptures, heavy makeup, and paintings while achieving high levels of expressiveness. To improve view consistency, we propose a new 3D distillation technique that converts predictions of our 2D prior into a neural radiance field (NeRF). Both the 2D prior and our distillation technique produce state-of-the-art results, which are confirmed by our extensive evaluations. Also, to the best of our knowledge, our method is the first to achieve view-consistent extreme tongue articulation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions

    cs.CV 2026-07 conditional novelty 6.5 of 10

    EmoteGPT regresses FLAME 3DMM expression parameters from explicit or implicit text using an MLLM with a dedicated <Expr> token, trained on the new Txt2Emote dataset plus image data, outperforming prior text-to-3D face...

  2. CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    CAP4D combines a morphable multi-view diffusion model with 3D Gaussian splatting to build animatable 4D head avatars from 1 to 100 reference images, claiming state-of-the-art results.

Pith tools