Pith. sign in

REVIEW 2 cited by

StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.10922 v2 pith:BTNUEL24 submitted 2022-08-23 cs.CV cs.LGeess.ASeess.IV

classification cs.CVcs.LGeess.ASeess.IV
keywords headtalkingvideoaudio-drivenimagelatentstyletalkeraccurately
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose StyleTalker, a novel audio-driven talking head generation model that can synthesize a video of a talking person from a single reference image with accurately audio-synced lip shapes, realistic head poses, and eye blinks. Specifically, by leveraging a pretrained image generator and an image encoder, we estimate the latent codes of the talking head video that faithfully reflects the given audio. This is made possible with several newly devised components: 1) A contrastive lip-sync discriminator for accurate lip synchronization, 2) A conditional sequential variational autoencoder that learns the latent motion space disentangled from the lip movements, such that we can independently manipulate the motions and lip movements while preserving the identity. 3) An auto-regressive prior augmented with normalizing flow to learn a complex audio-to-motion multi-modal latent space. Equipped with these components, StyleTalker can generate talking head videos not only in a motion-controllable way when another motion source video is given but also in a completely audio-driven manner by inferring realistic motions from the input audio. Through extensive experiments and user studies, we show that our model is able to synthesize talking head videos with impressive perceptual quality which are accurately lip-synced with the input audios, largely outperforming state-of-the-art baselines.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos

    cs.CV 2025-08 reject novelty 4.0 of 10

    The paper previews a claimed 2M-clip multimodal benchmark for whole-body talking avatar video generation, with standard metrics and an initial evaluation of eight open-source models.

  2. JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1

    cs.CV 2025-07 reject novelty 4.0 of 10

    A paper announcing a large-scale whole-body talking avatar benchmark and evaluation protocol, but with insufficient details to verify the dataset or the joint audio-video evaluation.

Pith tools