REVIEW 36 cited by
Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
With the introduction of diffusion-based video generation techniques, audio-conditioned human video generation has recently achieved significant breakthroughs in both the naturalness of motion and the synthesis of portrait details. Due to the limited control of audio signals in driving human motion, existing methods often add auxiliary spatial signals to stabilize movements, which may compromise the naturalness and freedom of motion. In this paper, we propose an end-to-end audio-only conditioned video diffusion model named Loopy. Specifically, we designed an inter- and intra-clip temporal module and an audio-to-latents module, enabling the model to leverage long-term motion information from the data to learn natural motion patterns and improving audio-portrait movement correlation. This method removes the need for manually specified spatial motion templates used in existing methods to constrain motion during inference. Extensive experiments show that Loopy outperforms recent audio-driven portrait diffusion models, delivering more lifelike and high-quality results across various scenarios.
Forward citations
Cited by 36 Pith papers
-
SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation
SyncBreaker jointly attacks image and audio streams with Multi-Interval Sampling and Cross-Attention Fooling to degrade speech-driven talking head generation more than single-modality baselines.
-
Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation
A causal diffusion-forcing model generates interactive head-avatar motion with 500ms motion-generation latency and learns expressive reactions via DPO with synthetic negative samples.
-
STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
STARCaster is a 2D spatio-temporal video diffusion model that unifies identity-conditioned audio-driven portrait animation and novel-view synthesis without explicit 3D reconstruction.
-
Multi-human Interactive Talking Dataset
The paper contributes a 12-hour multi-person conversational video dataset with pose and speaking annotations, plus a baseline model for generating full-body talking videos of two to four people.
-
ARIG: Autoregressive Interactive Head Generation for Real-time Conversations
ARIG introduces a real-time, frame-wise autoregressive head generation framework with diffusion-based continuous motion prediction, improving interactive realism over clip-wise methods.
-
FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation
FlowMo reduces temporal artifacts in video generation by guiding the denoising process to lower the maximum patch-wise variance of consecutive-frame differences in the latent space.
-
Semantics-Aware Human Motion Generation from Audio Instructions
An end-to-end audio-conditioned model generates 3D human motion from spoken instructions, with performance close to text-conditioned baselines on synthetic speech datasets.
-
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
MultiTalk is the first framework to generate multi-person conversational videos from multi-stream audio, using Label Rotary Position Embedding to bind each voice to the correct person.
-
Exploring Timeline Control for Facial Motion Generation
A diffusion model generates natural facial motions from user-specified multi-track timelines, using TICC-based frame-level action interval annotation for training and evaluation.
-
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
HunyuanVideo-Avatar is an audio-driven video generator that enables emotion-controllable and multi-character animation by injecting character images, routing audio via face masks, and transferring emotion from referen...
-
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
HunyuanCustom reports the strongest face and subject identity consistency among compared open and commercial video-customization models, on a self-selected 200-video benchmark, while also supporting audio- and video-c...
-
Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
A motion-prior diffusion model with archived-frame memory improves identity, lip-sync, and head-motion consistency in long talking-face videos.
-
OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models
A mixed-condition training recipe, combining text, audio, and pose signals, lets a single diffusion transformer animate photos into realistic talking or singing videos at unprecedented data scale.
-
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
EMO2 generates talking-head videos by first predicting hand poses from audio and then using those hand signals to drive a video diffusion model that synthesizes face and upper-body motion.
-
MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation
A mixture-of-experts model and a new 150-hour dataset improve emotion control in audio-driven talking head videos.
-
FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation
FADA distills a diffusion-based talking avatar model into a 6-step student that mimics multi-condition classifier-free guidance with learnable tokens, achieving 4.17 to 12.5 times NFE speedup with comparable quality.
-
MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation
MEMO introduces memory-guided linear attention and emotion-aware multi-modal attention for audio-driven talking video generation, reporting state-of-the-art quality on self-collected test sets.
-
Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
Hallo3 builds a portrait animation system on CogVideoX, using an identity reference network, audio cross-attention, and motion-frame extrapolation to generate dynamic talking videos.
-
Sonic: Shifting Focus to Global Audio Perception in Portrait Animation
Sonic generates audio-driven portrait videos from a single image using only global audio cues, with a new time-aware shift fusion that improves long-video stability and lip sync.
-
InfinityHuman: Towards Long-Term Audio-Driven Human
A coarse-to-fine audio-driven animation framework that uses pose-guided refinement and hand-specific reward learning to generate long, identity-stable talking videos.
-
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
A new autoregressive video-generation framework for interactive digital humans, with a 64x compression autoencoder and a diffusion renderer, claims real-time multimodal control but is only demonstrated for audio input.
-
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
Sparse-frame dubbing with adjacent-chunk keyframe sampling lets a streaming audio-video model produce full-body motion synchronized to new audio while preserving identity and camera motion.
-
EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis
EDTalk++ disentangles talking-head video into four orthogonal motion banks (mouth, pose, eyes, expression) and drives them from either video or audio inputs.
-
StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation
A diffusion-based avatar generator that uses a timestep-aware audio adapter, audio-adaptive guidance, and weighted sliding-window fusion to produce audio-synced talking-head videos several minutes long with reduced id...
-
FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases
FixTalk adds two modules to a real-time GAN talking-head model, decoupling identity from motion to stop identity leakage while using a memory to recover details and reduce artifacts.
-
MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation
MirrorMe adapts the LTX video diffusion transformer to generate real-time, high-fidelity audio-driven halfbody animations with identity preservation and hand pose control.
-
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
An audio-driven avatar generator that injects Wav2Vec2 audio features as additive latents into multiple DiT layers of a LoRA-fine-tuned Wan2.1 model, improving lip-sync and enabling prompt-controlled full-body animation.
-
SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers
An audio-conditioned video diffusion transformer that animates portraits from image, video, text, and audio inputs with a sliding-window fusion for long videos.
-
VRAG: Learning World Models for Interactive Video Generation
VRAG improves long-horizon interactive video generation by conditioning autoregressive diffusion on retrieved historical frames and explicit global state, outperforming long-context baselines on the tested Minecraft a...
-
Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation
DICE-Talk improves emotional talking-head generation by combining an audio-visual Gaussian emotion prior, a vector-quantized emotion bank, and an auxiliary emotion classifier in a diffusion model.
-
Joint Learning of Depth and Appearance for Portrait Image Animation
A single diffusion model jointly generates portrait RGB images and aligned depth maps, and its fine-tuned variants can estimate depth, edit from depth, relight, and produce audio-driven talking heads with depth.
-
Multi-View Face and Gesture Animation with Dynamic Gaussians
Combining separate face and hand models with a parametric body and Gaussian splatting enables multi-view-consistent upper-body avatars that can be re-animated with new expressions and gestures.
-
Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos
The paper previews a claimed 2M-clip multimodal benchmark for whole-body talking avatar video generation, with standard metrics and an initial evaluation of eight open-source models.
-
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1
A paper announcing a large-scale whole-body talking avatar benchmark and evaluation protocol, but with insufficient details to verify the dataset or the joint audio-video evaluation.
-
LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models
Using consistency distillation, INT8 quantization, and pipeline parallelism, the LLIA system generates portrait video from audio at 78 FPS, with 140 ms initial latency.
-
JoyGen: Audio-Driven 3D Depth-Aware Talking-Face Video Editing
A two-stage talking-face video editing method that conditions a single-step latent-space UNet on audio-derived 3D mouth depth maps, plus a new 130-hour Chinese dataset, claims state-of-the-art lip sync and visual quality.
Discussion (0). Continue with ORCID to comment.