Pith. sign in

REVIEW 5 cited by

MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.04448 v1 pith:7ROMWIEI submitted 2024-12-05 cs.CV

classification cs.CV
keywords talkingattentionaudioconsistencydiffusionidentitymemomemory-guided
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency, and producing natural, audio-aligned expressions in generated talking videos remain significant challenges. To address these challenges, we propose Memory-guided EMOtion-aware diffusion (MEMO), an end-to-end audio-driven portrait animation approach to generate identity-consistent and expressive talking videos. Our approach is built around two key modules: (1) a memory-guided temporal module, which enhances long-term identity consistency and motion smoothness by developing memory states to store information from a longer past context to guide temporal modeling via linear attention; and (2) an emotion-aware audio module, which replaces traditional cross attention with multi-modal attention to enhance audio-video interaction, while detecting emotions from audio to refine facial expressions via emotion adaptive layer norm. Extensive quantitative and qualitative results demonstrate that MEMO generates more realistic talking videos across diverse image and audio types, outperforming state-of-the-art methods in overall quality, audio-lip synchronization, identity consistency, and expression-emotion alignment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A diffusion-based talking-face generator uses 3D blendshape coefficients to continuously control the emotion intensity of generated facial expressions.

  2. LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    LeapTalk distills a multi-step diffusion teacher into a one-step Brownian-bridge student and reports stable streaming talking-head generation at up to 200 FPS.

  3. ARIG: Autoregressive Interactive Head Generation for Real-time Conversations

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ARIG introduces a real-time, frame-wise autoregressive head generation framework with diffusion-based continuous motion prediction, improving interactive realism over clip-wise methods.

  4. Matrix-Game: Interactive World Foundation Model

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A 17B-parameter diffusion model generates controllable, physically consistent Minecraft video from a reference image and user actions, beating Oasis and MineWorld on a new benchmark.

  5. SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

    cs.CV 2025-06 conditional novelty 5.0 of 10

    An audio-conditioned video diffusion transformer that animates portraits from image, video, text, and audio inputs with a sliding-window fusion for long videos.

Pith tools