Pith. sign in

REVIEW 3 cited by

Audio-driven Talking Face Video Generation with Learning-based Personalized Head Pose

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.10137 v2 pith:XG3CWBB2 submitted 2020-02-24 cs.CV cs.GR

classification cs.CVcs.GR
keywords headtalkingfacevideoframespersonalizedposeperson
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this problem by proposing a deep neural network model that takes an audio signal A of a source person and a very short video V of a target person as input, and outputs a synthesized high-quality talking face video with personalized head pose (making use of the visual information in V), expression and lip synchronization (by considering both A and V). The most challenging issue in our work is that natural poses often cause in-plane and out-of-plane head rotations, which makes synthesized talking face video far from realistic. To address this challenge, we reconstruct 3D face animation and re-render it into synthesized frames. To fine tune these frames into realistic ones with smooth background transition, we propose a novel memory-augmented GAN module. By first training a general mapping based on a publicly available dataset and fine-tuning the mapping using the input short video of target person, we develop an effective strategy that only requires a small number of frames (about 300 frames) to learn personalized talking behavior including head pose. Extensive experiments and two user studies show that our method can generate high-quality (i.e., personalized head movements, expressions and good lip synchronization) talking face videos, which are naturally looking with more distinguishing head movement effects than the state-of-the-art methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync

    cs.CV 2025-07 conditional novelty 6.0 of 10

    JOLT3D jointly trains a 3DMM reconstruction network with a talking head generator, then uses FACS mouth blendshapes from a diffusion model to lip-sync videos while preserving the original chin contour.

  2. Few-Shot Identity Adaptation for 3D Talking Heads via Global Gaussian Field

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A shared global Gaussian field plus identity embeddings lets a 3D talking head model adapt to new speakers with a few seconds of footage while improving quality over prior per-identity models.

  3. Robust Deepfake Detection for Electronic Know Your Customer Systems Using Registered Images

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A video deepfake detector for eKYC that combines temporal identity-vector differences with differences against a registered photo, and shows robustness to image degradation.

Pith tools