Pith. sign in

REVIEW 2 cited by

DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits Animation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.03786 v2 pith:V6IQDGHY submitted 2023-01-10 cs.CV

classification cs.CV
keywords difftalktalkinggeneralizedheadaudio-drivensynthesisaudiodiffusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Talking head synthesis is a promising approach for the video production industry. Recently, a lot of effort has been devoted in this research area to improve the generation quality or enhance the model generalization. However, there are few works able to address both issues simultaneously, which is essential for practical applications. To this end, in this paper, we turn attention to the emerging powerful Latent Diffusion Models, and model the Talking head generation as an audio-driven temporally coherent denoising process (DiffTalk). More specifically, instead of employing audio signals as the single driving factor, we investigate the control mechanism of the talking face, and incorporate reference face images and landmarks as conditions for personality-aware generalized synthesis. In this way, the proposed DiffTalk is capable of producing high-quality talking head videos in synchronization with the source audio, and more importantly, it can be naturally generalized across different identities without any further fine-tuning. Additionally, our DiffTalk can be gracefully tailored for higher-resolution synthesis with negligible extra computational cost. Extensive experiments show that the proposed DiffTalk efficiently synthesizes high-fidelity audio-driven talking head videos for generalized novel identities. For more video results, please refer to \url{https://sstzal.github.io/DiffTalk/}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SemTalk generates co-speech body motion by separating rhythm-based base gestures from semantically important sparse gestures and blending them with a learned frame-level semantic score.

  2. ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal Guidance

    cs.CV 2024-11 conditional novelty 6.0 of 10

    ConsistentAvatar aligns a Fourier high-frequency detail map through a diffusion model and uses it, with normals and emotion text, to condition talking-head avatar generation, reducing temporal and expression inconsistency.

Pith tools