Pith. sign in

REVIEW 1 cited by

DiffusionAct: Controllable Diffusion Autoencoder for One-shot Face Reenactment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.17217 v2 pith:ELGXZC27 submitted 2024-03-25 cs.CV cs.AI

classification cs.CVcs.AI
keywords reenactmentdiffusionfacefacialimagesposeappearanceautoencoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video-driven neural face reenactment aims to synthesize realistic facial images that successfully preserve the identity and appearance of a source face, while transferring the target head pose and facial expressions. Existing GAN-based methods suffer from either distortions and visual artifacts or poor reconstruction quality, i.e., the background and several important appearance details, such as hair style/color, glasses and accessories, are not faithfully reconstructed. Recent advances in Diffusion Probabilistic Models (DPMs) enable the generation of high-quality realistic images. To this end, in this paper we present DiffusionAct, a novel method that leverages the photo-realistic image generation of diffusion models to perform neural face reenactment. Specifically, we propose to control the semantic space of a Diffusion Autoencoder (DiffAE), in order to edit the facial pose of the input images, defined as the head pose orientation and the facial expressions. Our method allows one-shot, self, and cross-subject reenactment, without requiring subject-specific fine-tuning. We compare against state-of-the-art GAN-, StyleGAN2-, and diffusion-based methods, showing better or on-par reenactment performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Navigating Large-Pose Challenge for High-Fidelity Face Reenactment with Video Diffusion Model

    cs.CV 2025-07 conditional novelty 5.0 of 10

    FRVD warps a source face toward driving poses with implicit keypoints, then repairs lost details inside Stable Video Diffusion's latent space, reporting gains over seven baselines on large-pose reenactment benchmarks.

Pith tools