REVIEW 5 cited by
DREAM-Talk: Diffusion-based Realistic Emotional Audio-driven Method for Single Image Talking Face Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The generation of emotional talking faces from a single portrait image remains a significant challenge. The simultaneous achievement of expressive emotional talking and accurate lip-sync is particularly difficult, as expressiveness is often compromised for the accuracy of lip-sync. As widely adopted by many prior works, the LSTM network often fails to capture the subtleties and variations of emotional expressions. To address these challenges, we introduce DREAM-Talk, a two-stage diffusion-based audio-driven framework, tailored for generating diverse expressions and accurate lip-sync concurrently. In the first stage, we propose EmoDiff, a novel diffusion module that generates diverse highly dynamic emotional expressions and head poses in accordance with the audio and the referenced emotion style. Given the strong correlation between lip motion and audio, we then refine the dynamics with enhanced lip-sync accuracy using audio features and emotion style. To this end, we deploy a video-to-video rendering module to transfer the expressions and lip motions from our proxy 3D avatar to an arbitrary portrait. Both quantitatively and qualitatively, DREAM-Talk outperforms state-of-the-art methods in terms of expressiveness, lip-sync accuracy and perceptual quality.
Forward citations
Cited by 5 Pith papers
-
Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
A text-only talking face generator maps text to visemes, hallucinates pseudo-audio, and renders lip-synced video; the paper's headline metrics are partially contradicted by its own tables.
-
MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation
MEMO introduces memory-guided linear attention and emotion-aware multi-modal attention for audio-driven talking video generation, reporting state-of-the-art quality on self-collected test sets.
-
INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations
A unified two-stage model uses dual-track audio and learnable memory banks to generate expressive head motions for an agent that freely switches between speaking and listening.
-
LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion Space
LES-Talker defines emotions as 41-dimensional vectors over facial action units and uses them to edit talking-head videos with continuous emotion levels and per-muscle control.
-
A Review of Human Emotion Synthesis Based on Generative Technology
A systematic review that taxonomizes roughly 230 papers on generative-model-based emotion synthesis across faces, speech, and text, and catalogs datasets, metrics, and future directions.
Discussion (0). Continue with ORCID to comment.