Pith. sign in

REVIEW 61 cited by

Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08801 v2 pith:IQXLMXAD submitted 2024-06-13 cs.CV

classification cs.CV
keywords visualaudio-drivenhierarchicalimagesynthesisalignmentanimationapproach
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The field of portrait image animation, driven by speech audio input, has experienced significant advancements in the generation of realistic and dynamic portraits. This research delves into the complexities of synchronizing facial movements and creating visually appealing, temporally consistent animations within the framework of diffusion-based methodologies. Moving away from traditional paradigms that rely on parametric models for intermediate facial representations, our innovative approach embraces the end-to-end diffusion paradigm and introduces a hierarchical audio-driven visual synthesis module to enhance the precision of alignment between audio inputs and visual outputs, encompassing lip, expression, and pose motion. Our proposed network architecture seamlessly integrates diffusion-based generative models, a UNet-based denoiser, temporal alignment techniques, and a reference network. The proposed hierarchical audio-driven visual synthesis offers adaptive control over expression and pose diversity, enabling more effective personalization tailored to different identities. Through a comprehensive evaluation that incorporates both qualitative and quantitative analyses, our approach demonstrates obvious enhancements in image and video quality, lip synchronization precision, and motion diversity. Further visualization and access to the source code can be found at: https://fudan-generative-vision.github.io/hallo.

Discussion (0). Continue with ORCID to comment.

Forward citations

Showing 60 of 61 Pith papers that cite this

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. See all 61 Pith citations

  1. Every Image Listens, Every Image Dances: Music-Driven Image Animation

    cs.CV 2025-01 conditional novelty 7.0 of 10

    MuseDance animates a reference image into a music-synchronized dance video conditioned only on the audio track and a text description, and contributes a new 2,904-video dataset.

  2. Avatar-Forever: Decoupled Parallel Training for High-Quality Real-Time Infinite Avatars

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Decoupled parallel training of few-step distillation and rollout-based long-horizon adaptation, together with chunk-wise history feature caching, enables real-time 768x512 infinite audio-driven avatars at 27.2 FPS.

  3. EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot

    cs.CL 2026-08 conditional novelty 6.0 of 10

    EmpaAva is an open-source, LLM-orchestrated 3D avatar chatbot that perceives user affect from speech and video, plans empathetic replies, and delivers them with synchronized emotional speech and facial motion.

  4. Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A diffusion-based talking-face generator uses 3D blendshape coefficients to continuously control the emotion intensity of generated facial expressions.

  5. ViDS: Video Diffusion Shader using 3D Face Tracking

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Fine-tuning a video diffusion model on dense 3DMM normal maps from Pixel3DMM lets a single portrait photo be animated with a driving video's expressions and pose, surpassing landmark- and latent-based portrait animati...

  6. SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    SyncBreaker jointly attacks image and audio streams with Multi-Interval Sampling and Cross-Attention Fooling to degrade speech-driven talking head generation more than single-modality baselines.

  7. STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

    cs.CV 2025-12 conditional novelty 6.0 of 10

    STARCaster is a 2D spatio-temporal video diffusion model that unifies identity-conditioned audio-driven portrait animation and novel-view synthesis without explicit 3D reconstruction.

  8. OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    OmniHuman-1.5 combines MLLM-based planning with a multimodal diffusion transformer and pseudo-last-frame identity conditioning to generate context-aware avatar videos from audio and a reference image.

  9. TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis

    cs.CV 2025-08 conditional novelty 6.0 of 10

    TalkVid is a 1,244-hour, 7,729-speaker, 15-language talking-head video dataset with a stratified evaluation benchmark, and models trained on it generalize better across demographics.

  10. Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering

    cs.CV 2025-08 reject novelty 6.0 of 10

    A text-only talking face generator maps text to visemes, hallucinates pseudo-audio, and renders lip-synced video; the paper's headline metrics are partially contradicted by its own tables.

  11. X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention

    cs.CV 2025-07 conditional novelty 6.0 of 10

    X-NeMo trains a 1D identity-agnostic motion descriptor end-to-end with a diffusion model, enabling zero-shot portrait animation with improved identity and expression fidelity.

  12. DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single-DiT portrait animation framework with style and emotion branches plus parallel audio-style cross-attention claims faster, controllable talking-head generation without a Reference Net.

  13. JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync

    cs.CV 2025-07 conditional novelty 6.0 of 10

    JOLT3D jointly trains a 3DMM reconstruction network with a talking head generator, then uses FACS mouth blendshapes from a diffusion model to lip-sync videos while preserving the original chin contour.

  14. MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MagicAnime is a 400k-clip multimodal cartoon dataset with hierarchical annotations and benchmarks for image-to-video, pose-driven, face reenactment, and audio-driven animation generation.

  15. Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A two-stage text-guidance framework, Think-Before-Draw, uses chain-of-thought prompting to convert emotion labels into facial muscle descriptions and then progressively guides a diffusion model from coarse emotion to ...

  16. ARIG: Autoregressive Interactive Head Generation for Real-time Conversations

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ARIG introduces a real-time, frame-wise autoregressive head generation framework with diffusion-based continuous motion prediction, improving interactive realism over clip-wise methods.

  17. Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Silencer adds a nearly invisible disturbance to portraits that makes LDM-based talking-head models keep the mouth silent, and it survives several image-purification countermeasures.

  18. Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MultiTalk is the first framework to generate multi-person conversational videos from multi-stream audio, using Label Rotary Position Embedding to bind each voice to the correct person.

  19. Exploring Timeline Control for Facial Motion Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A diffusion model generates natural facial motions from user-specified multi-track timelines, using TICC-based frame-level action interval annotation for training and evaluation.

  20. HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

    cs.CV 2025-05 conditional novelty 6.0 of 10

    HunyuanVideo-Avatar is an audio-driven video generator that enables emotion-controllable and multi-character animation by injecting character images, routing audio via face masks, and transferring emotion from referen...

  21. MAVOS-DD: Multilingual Audio-Video Open-Set Deepfake Detection Benchmark

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MAVOS-DD, a multilingual audio-video deepfake dataset with controlled open-set splits for unseen languages and generators, shows that state-of-the-art detectors degrade substantially outside their training distribution.

  22. A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    PAHA improves audio-driven avatar video generation by re-weighting training loss toward hands and face and by adding audio-video consistency classifiers during inference.

  23. Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A motion-prior diffusion model with archived-frame memory improves identity, lip-sync, and head-motion consistency in long talking-face videos.

  24. OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A mixed-condition training recipe, combining text, audio, and pose signals, lets a single diffusion transformer animate photos into realistic talking or singing videos at unprecedented data scale.

  25. EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    EMO2 generates talking-head videos by first predicting hand poses from audio and then using those hand signals to drive a video diffusion model that synthesizes face and upper-body motion.

  26. MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A mixture-of-experts model and a new 150-hour dataset improve emotion control in audio-driven talking head videos.

  27. UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting Control

    cs.CV 2024-12 conditional novelty 6.0 of 10

    UniAvatar integrates FLAME-based 3D motion rendering and SH-based illumination rendering into a diffusion talking-head model, enabling separate or combined control of motion and lighting in generated videos.

  28. FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FADA distills a diffusion-based talking avatar model into a 6-step student that mimics multi-condition classifier-free guidance with learnable tokens, achieving 4.17 to 12.5 times NFE speedup with comparable quality.

  29. Real-time One-Step Diffusion-based Expressive Portrait Videos Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    OSA-LCM distills a portrait video diffusion model into a single-step generator that matches the quality of a 20-step teacher on FID/FVD, enabling near real-time talking-head generation.

  30. LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LatentSync achieves state-of-the-art lip sync with an end-to-end latent diffusion model, a redesigned SyncNet supervisor, and a temporal representation alignment loss.

  31. MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MEMO introduces memory-guided linear attention and emotion-aware multi-modal attention for audio-driven talking video generation, reporting state-of-the-art quality on self-collected test sets.

  32. INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A unified two-stage model uses dual-track audio and learnable memory banks to generate expressive head motions for an agent that freely switches between speaking and listening.

  33. SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SINGER attaches wavelet-based multi-scale spectral and self-adaptive filter modules to the frozen Hallo diffusion backbone and reports better singing-video generation than seven baselines on two datasets.

  34. Playable Game Generation

    cs.AI 2024-12 conditional novelty 6.0 of 10

    An autoregressive latent diffusion system, PlayGen, generates real-time playable Super Mario Bros and Doom sessions on an RTX 2060, with accuracy of game mechanics measured by action-recognition metrics.

  35. Sonic: Shifting Focus to Global Audio Perception in Portrait Animation

    cs.MM 2024-11 conditional novelty 6.0 of 10

    Sonic generates audio-driven portrait videos from a single image using only global audio cues, with a new time-aware shift fusion that improves long-video stability and lip sync.

  36. EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion

    cs.CV 2024-11 conditional novelty 6.0 of 10

    EmotiveTalk generates talking-head videos by decoupling audio into lip and expression latents and conditioning a video diffusion model on the separate signals.

  37. TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Anchor-guided fixed-capacity multimodal memory plus stage-parallel few-step denoising stabilizes long-form causal audio-video digital humans at real-time speed.

  38. FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian Head

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A one-shot model for animatable 3D/4D Gaussian head reconstruction that adds attention regularization, decoupled reconstruction-animation training, and autoregressive visibility-gated fusion, reporting consistent metr...

  39. InfinityHuman: Towards Long-Term Audio-Driven Human

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A coarse-to-fine audio-driven animation framework that uses pose-guided refinement and hand-specific reward learning to generate long, identity-stable talking videos.

  40. InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Sparse-frame dubbing with adjacent-chunk keyframe sampling lets a streaming audio-video model produce full-body motion synchronized to new audio while preserving identity and camera motion.

  41. EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis

    cs.CV 2025-08 conditional novelty 5.0 of 10

    EDTalk++ disentangles talking-head video into four orthogonal motion banks (mouth, pose, eyes, expression) and drives them from either video or audio inputs.

  42. StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A diffusion-based avatar generator that uses a timestep-aware audio adapter, audio-adaptive guidance, and weighted sliding-window fusion to produce audio-synced talking-head videos several minutes long with reduced id...

  43. Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation

    cs.SD 2025-07 conditional novelty 5.0 of 10

    A lightweight VAE adapter bridges a frozen text-to-speech model and a frozen talking-head model to generate face-consistent speech and animation from a single image and text.

  44. MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    MoDiT, a diffusion transformer conditioned on 3DMM coefficients and Wav2Lip references, produces talking-head videos with improved same-identity lip sync and more natural blinks in its reported benchmarks.

  45. MoDA: Multi-modal Diffusion Architecture for Talking Head Generation

    cs.GR 2025-07 conditional novelty 5.0 of 10

    MoDA uses flow matching in a compact face-motion space with a progressively fused multi-modal transformer to generate expressive, lip-synced talking-head videos from a single image and audio.

  46. FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases

    cs.CV 2025-07 conditional novelty 5.0 of 10

    FixTalk adds two modules to a real-time GAN talking-head model, decoupling identity from motion to stop identity leakage while using a memory to recover details and reduce artifacts.

  47. MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MirrorMe adapts the LTX video diffusion transformer to generate real-time, high-fidelity audio-driven halfbody animations with identity preservation and hand pose control.

  48. GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    GGTalker combines large-scale audio-to-expression and expression-to-texture priors with rapid per-identity fine-tuning to create high-quality 3D talking heads from a short video.

  49. OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    An audio-driven avatar generator that injects Wav2Vec2 audio features as additive latents into multiple DiT layers of a LoRA-fine-tuned Wan2.1 model, improving lip-sync and enabling prompt-controlled full-body animation.

  50. SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SyncTalk++ synthesizes speech-driven talking-head videos via 3D Gaussian Splatting and reports state-of-the-art synchronization and quality at up to 101 FPS.

  51. Speaking images. A novel framework for the automated self-description of artworks

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A four-stage open-source AI pipeline turns a digitized artwork into a short video where a depicted person animates and narrates the scene.

  52. FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A single framework can edit predefined facial attributes in audio-synchronized talking head videos while preserving identity and lip-sync quality.

  53. DATA: Multi-Disentanglement based Contrastive Learning for Open-World Semi-Supervised Deepfake Attribution

    cs.CV 2025-05 conditional novelty 5.0 of 10

    DATA combines learned orthonormal 'deepfake bases' and an augmented memory clustering module to improve open-world semi-supervised deepfake attribution accuracy.

  54. Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation

    cs.CV 2025-04 conditional novelty 5.0 of 10

    DICE-Talk improves emotional talking-head generation by combining an audio-visual Gaussian emotion prior, a vector-quantized emotion bank, and an auxiliary emotion classifier in a diffusion model.

  55. Joint Learning of Depth and Appearance for Portrait Image Animation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A single diffusion model jointly generates portrait RGB images and aligned depth maps, and its fine-tuned variants can estimate depth, edit from depth, relight, and produce audio-driven talking heads with depth.

  56. Identity-Preserving Video Dubbing Using Motion Warping

    cs.CV 2025-01 reject novelty 5.0 of 10

    IPTalker dubs videos by aligning audio with reference mouth images, warping them to match the target lip shape, and inpainting the result.

  57. Human Motion Video Generation: A Survey

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A comprehensive survey with a five-phase pipeline model for human motion video generation, covering over 200 papers and adding a new benchmark comparison of nine pose-guided methods.

  58. LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Using consistency distillation, INT8 quantization, and pipeline parallelism, the LLIA system generates portrait video from audio at 78 FPS, with 140 ms initial latency.

  59. SyncAnimation: A Real-Time End-to-End Framework for Audio-Driven Human Pose and Talking Head Animation

    cs.CV 2025-01 conditional novelty 4.0 of 10

    SyncAnimation introduces a NeRF-based system that generates audio-synchronized upper-body and head animations with facial expressions in real time.

  60. JoyGen: Audio-Driven 3D Depth-Aware Talking-Face Video Editing

    cs.CV 2025-01 conditional novelty 4.0 of 10

    A two-stage talking-face video editing method that conditions a single-step latent-space UNet on audio-derived 3D mouth depth maps, plus a new 130-hour Chinese dataset, claims state-of-the-art lip sync and visual quality.

See all 61 Pith citations

Pith tools