Pith. sign in

REVIEW 2 cited by

EDTalk: Efficient Disentanglement for Emotional Talking Head Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.01647 v1 pith:JQ4KL6EL submitted 2024-04-02 cs.CV

classification cs.CV
keywords headedtalkspacetalkingbasesefficientfacialinput
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Achieving disentangled control over multiple facial motions and accommodating diverse input modalities greatly enhances the application and entertainment of the talking head generation. This necessitates a deep exploration of the decoupling space for facial features, ensuring that they a) operate independently without mutual interference and b) can be preserved to share with different modal input, both aspects often neglected in existing methods. To address this gap, this paper proposes a novel Efficient Disentanglement framework for Talking head generation (EDTalk). Our framework enables individual manipulation of mouth shape, head pose, and emotional expression, conditioned on video or audio inputs. Specifically, we employ three lightweight modules to decompose the facial dynamics into three distinct latent spaces representing mouth, pose, and expression, respectively. Each space is characterized by a set of learnable bases whose linear combinations define specific motions. To ensure independence and accelerate training, we enforce orthogonality among bases and devise an efficient training strategy to allocate motion responsibilities to each space without relying on external knowledge. The learned bases are then stored in corresponding banks, enabling shared visual priors with audio input. Furthermore, considering the properties of each space, we propose an Audio-to-Motion module for audio-driven talking head synthesis. Experiments are conducted to demonstrate the effectiveness of EDTalk. We recommend watching the project website: https://tanshuai0219.github.io/EDTalk/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion Space

    cs.CV 2024-11 conditional novelty 5.0 of 10

    LES-Talker defines emotions as 41-dimensional vectors over facial action units and uses them to edit talking-head videos with continuous emotion levels and per-muscle control.

  2. A Review of Human Emotion Synthesis Based on Generative Technology

    cs.LG 2024-12 conditional novelty 3.0 of 10

    A systematic review that taxonomizes roughly 230 papers on generative-model-based emotion synthesis across faces, speech, and text, and catalogs datasets, metrics, and future directions.

Pith tools