Pith. sign in

REVIEW 3 cited by

MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.01808 v2 pith:MLLC2DYZ submitted 2025-01-03 cs.CV

classification cs.CV
keywords emotionsemotionalexpressionscontrolaudiocompoundemotionmoee
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The generation of talking avatars has achieved significant advancements in precise audio synchronization. However, crafting lifelike talking head videos requires capturing a broad spectrum of emotions and subtle facial expressions. Current methods face fundamental challenges: a) the absence of frameworks for modeling single basic emotional expressions, which restricts the generation of complex emotions such as compound emotions; b) the lack of comprehensive datasets rich in human emotional expressions, which limits the potential of models. To address these challenges, we propose the following innovations: 1) the Mixture of Emotion Experts (MoEE) model, which decouples six fundamental emotions to enable the precise synthesis of both singular and compound emotional states; 2) the DH-FaceEmoVid-150 dataset, specifically curated to include six prevalent human emotional expressions as well as four types of compound emotions, thereby expanding the training potential of emotion-driven models. Furthermore, to enhance the flexibility of emotion control, we propose an emotion-to-latents module that leverages multimodal inputs, aligning diverse control signals-such as audio, text, and labels-to ensure more varied control inputs as well as the ability to control emotions using audio alone. Through extensive quantitative and qualitative evaluations, we demonstrate that the MoEE framework, in conjunction with the DH-FaceEmoVid-150 dataset, excels in generating complex emotional expressions and nuanced facial details, setting a new benchmark in the field. These datasets will be publicly released.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A hierarchical motion autoencoder with a conditional diffusion decoder reconstructs 16-frame videos from latents as small as 0.07% of the input size while maintaining competitive PSNR and perceptual scores.

  2. Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A two-stage text-guidance framework, Think-Before-Draw, uses chain-of-thought prompting to convert emotion labels into facial muscle descriptions and then progressively guides a diffusion model from coarse emotion to ...

  3. ChronoTailor: Harnessing Attention Guidance for Fine-Grained Video Virtual Try-On

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ChronoTailor combines region-aware attention guidance, temporal feature fusion, and multi-scale garment-pose alignment to produce state-of-the-art video virtual try-on results, and contributes the StyleDress dataset.

Pith tools