Pith. sign in

REVIEW 8 cited by

EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.17485 v3 pith:SAID3FNS submitted 2024-02-27 cs.CV

classification cs.CV
keywords facialvideosexpressiveexpressivenessportraitrealismstylesvideo
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we tackle the challenge of enhancing the realism and expressiveness in talking head video generation by focusing on the dynamic and nuanced relationship between audio cues and facial movements. We identify the limitations of traditional techniques that often fail to capture the full spectrum of human expressions and the uniqueness of individual facial styles. To address these issues, we propose EMO, a novel framework that utilizes a direct audio-to-video synthesis approach, bypassing the need for intermediate 3D models or facial landmarks. Our method ensures seamless frame transitions and consistent identity preservation throughout the video, resulting in highly expressive and lifelike animations. Experimental results demonsrate that EMO is able to produce not only convincing speaking videos but also singing videos in various styles, significantly outperforming existing state-of-the-art methodologies in terms of expressiveness and realism.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    SyncBreaker jointly attacks image and audio streams with Multi-Interval Sampling and Cross-Attention Fooling to degrade speech-driven talking head generation more than single-modality baselines.

  2. Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MF-Talk, a mask-free and identity-reference-free three-stage pipeline, improves visual quality and identity preservation in talking-face generation while remaining competitive on lip-sync.

  3. EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis

    cs.CV 2025-08 conditional novelty 5.0 of 10

    EDTalk++ disentangles talking-head video into four orthogonal motion banks (mouth, pose, eyes, expression) and drives them from either video or audio inputs.

  4. LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion

    cs.CV 2025-07 conditional novelty 5.0 of 10

    LiON-LoRA adds a learned scaling token to video-diffusion LoRA adapters, enabling linear and independent control of camera trajectory and object motion strength.

  5. MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    MoDiT, a diffusion transformer conditioned on 3DMM coefficients and Wav2Lip references, produces talking-head videos with improved same-identity lip sync and more natural blinks in its reported benchmarks.

  6. FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases

    cs.CV 2025-07 conditional novelty 5.0 of 10

    FixTalk adds two modules to a real-time GAN talking-head model, decoupling identity from motion to stop identity leakage while using a memory to recover details and reduce artifacts.

  7. SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SyncTalk++ synthesizes speech-driven talking-head videos via 3D Gaussian Splatting and reports state-of-the-art synchronization and quality at up to 101 FPS.

  8. Multi-View Face and Gesture Animation with Dynamic Gaussians

    cs.CV 2026-08 conditional novelty 4.0 of 10

    Combining separate face and hand models with a parametric body and Gaussian splatting enables multi-view-consistent upper-body avatars that can be re-animated with new expressions and gestures.

Pith tools