Pith. sign in

REVIEW 34 cited by

LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.03168 v2 pith:U7UULCU6 submitted 2024-07-03 cs.CV

LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

classification cs.CV
keywords animationcontrollabilityframeworkgenerationliveportraitportraitbettercomputational
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Portrait Animation aims to synthesize a lifelike video from a single source image, using it as an appearance reference, with motion (i.e., facial expressions and head pose) derived from a driving video, audio, text, or generation. Instead of following mainstream diffusion-based methods, we explore and extend the potential of the implicit-keypoint-based framework, which effectively balances computational efficiency and controllability. Building upon this, we develop a video-driven portrait animation framework named LivePortrait with a focus on better generalization, controllability, and efficiency for practical usage. To enhance the generation quality and generalization ability, we scale up the training data to about 69 million high-quality frames, adopt a mixed image-video training strategy, upgrade the network architecture, and design better motion transformation and optimization objectives. Additionally, we discover that compact implicit keypoints can effectively represent a kind of blendshapes and meticulously propose a stitching and two retargeting modules, which utilize a small MLP with negligible computational overhead, to enhance the controllability. Experimental results demonstrate the efficacy of our framework even compared to diffusion-based methods. The generation speed remarkably reaches 12.8ms on an RTX 4090 GPU with PyTorch. The inference code and models are available at https://github.com/KwaiVGI/LivePortrait

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 34 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Chehre: An Emoji-Prompted Video Dataset for Perceptually Diverse Facial Expression Recognition

    cs.CV 2026-06 unverdicted novelty 7.0

    Chehre introduces a new emoji-prompted video dataset with multi-annotator labels to benchmark models on dominant and distributional facial expression recognition tasks.

  2. Loki: Representation over Architecture for Diffusion-Based Portrait Animation

    cs.CV 2026-05 unverdicted novelty 7.0

    Loki replaces RGB conditioning stacks with identity-orthogonal parametric face encodings rasterized for diffusion, achieving efficient cross-ID portrait animation without cross-ID training data.

  3. Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling

    cs.CV 2026-04 unverdicted novelty 7.0

    Talker-T2AV achieves better lip-sync accuracy, video quality, and audio quality than dual-branch baselines by separating high-level shared autoregressive modeling from modality-specific low-level diffusion refinement ...

  4. AvatarPointillist: AutoRegressive 4D Gaussian Avatarization

    cs.CV 2026-04 unverdicted novelty 7.0

    AvatarPointillist autoregressively generates adaptive 3D point clouds via Transformer for photorealistic 4D Gaussian avatars from one image, jointly predicting animation bindings and using a conditioned Gaussian decoder.

  5. UIKA: Fast Universal Head Avatar from Pose-Free Images

    cs.CV 2026-01 conditional novelty 7.0

    UIKA is a feed-forward animatable Gaussian head model using UV-guided correspondence estimation and learnable UV tokens with dual-level attention, trained on large-scale synthetic data to handle pose-free inputs.

  6. ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body

    cs.CV 2025-12 unverdicted novelty 7.0

    ViBES introduces a speech-language-behavior model using modality-specific transformer experts that jointly generates dialogue and 3D body actions, showing gains over separate co-speech and text-to-motion baselines on ...

  7. Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization

    cs.CV 2025-12 unverdicted novelty 7.0

    Omni-Attribute is a new open-vocabulary image attribute encoder trained on semantically linked pairs with dual objectives to produce disentangled representations for personalization and compositional generation.

  8. Unmasking Puppeteers: Leveraging Biometric Leakage to Expose Impersonation in AI-Based Videoconferencing

    cs.CV 2025-10 unverdicted novelty 7.0

    A pose-conditioned large-margin contrastive encoder isolates persistent biometric identity cues from transmitted latents in talking-head videoconferencing to flag impersonation attacks via cosine similarity without in...

  9. FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling

    cs.CV 2025-09 unverdicted novelty 7.0

    Phoneme-guided autoregressive framework for talking-head animation that reduces inter-frame flicker via causal keyframe generation and timestamp-aware interpolation, outperforming diffusion baselines on FVD and a new ...

  10. Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer

    cs.CV 2025-09 conditional novelty 7.0

    Durian introduces a dual-reference diffusion model trained via self-reconstruction on video frames to enable cross-identity attribute transfer in portrait animations, supporting multi-attribute composition and interpolation.

  11. TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment

    cs.CV 2026-07 conditional novelty 6.0

    A two-stage pipeline (Gaussian reenactment plus geometry-anchored masked diffusion) transfers cross-identity tongue dynamics, roughly doubling tongue-specific metrics over prior reenactment baselines.

  12. ViDS: Video Diffusion Shader using 3D Face Tracking

    cs.CV 2026-07 conditional novelty 6.0

    Fine-tuning a video diffusion model on dense 3DMM normal maps from Pixel3DMM lets a single portrait photo be animated with a driving video's expressions and pose, surpassing landmark- and latent-based portrait animati...

  13. Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation

    cs.CV 2026-06 unverdicted novelty 6.0

    InterTalk is a motion-based real-time framework for flexible multi-round multi-person conversational talking face generation using motion feedback, iterative strategies, and facial component disentanglement, supported...

  14. Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs

    cs.CV 2026-06 unverdicted novelty 6.0

    A causal VAE with variable reference guidance and a Rectified Flow Transformer enables real-time streamable high-quality talking portrait video generation from audio and images.

  15. Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation

    cs.GR 2026-05 unverdicted novelty 6.0

    Reformulates evaluation of audio-driven talking head generation as a sequence alignment problem using Soft DTW, showing improved robustness and consistency across 20 methods and seven datasets.

  16. CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning

    cs.CV 2026-05 unverdicted novelty 6.0

    CogPortrait uses MLLM-based hierarchical planning to convert high-level labels into eye keypoints and a conditioned DiT model to produce portrait animations with improved eye-region accuracy on the new EMH benchmark.

  17. Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    TT-SAC is a parameter-free inference framework that uses a generator-encoder feedback loop to adapt conditioning representations and stabilize identity and motion in audio-driven talking-head videos.

  18. TOPOS: High-Fidelity and Efficient Industry-Grade 3D Head Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    TOPOS creates high-fidelity 3D heads with fixed industry topology from single images via a specialized VAE with Perceiver Resampler and a rectified flow transformer.

  19. The Alpha Blending Hypothesis: Compositing Shortcut in Deepfake Detection

    cs.CV 2026-05 unverdicted novelty 6.0

    Deepfake detectors act as alpha blending searchers; training solely on self-blended real images yields top cross-dataset generalization on 15 datasets without using synthetic deepfakes.

  20. SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    SocialDirector uses spatiotemporal actor masking and directional reweighting on cross-attention maps to reduce actor-action mismatches and improve target-directed interactions in generated multi-person videos.

  21. MeshLAM: Feed-Forward One-Shot Animatable Textured Mesh Avatar Reconstruction

    cs.CV 2026-04 unverdicted novelty 6.0

    MeshLAM reconstructs high-fidelity animatable textured mesh head avatars from a single image via a feed-forward dual shape-texture architecture with iterative GRU decoding and reprojection-based guidance.

  22. ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception

    cs.CV 2026-04 unverdicted novelty 6.0

    ARGen generates high-fidelity dynamic facial expression videos using affective semantic injection and adaptive reinforcement diffusion to improve emotion recognition models facing data scarcity and long-tail distributions.

  23. AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors

    cs.CV 2026-03 conditional novelty 6.0

    Identity-finetuned video diffusion plus RF-Inversion can supply multi-view body supervision that lets 3D Gaussian avatars be completed and animated from heavily occluded monocular video.

  24. JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching

    cs.CV 2025-06 unverdicted novelty 6.0

    JAM-Flow introduces a unified flow-matching model with a Multi-Modal Diffusion Transformer that jointly synthesizes facial motion and speech from text, audio, or motion inputs.

  25. Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars

    cs.CV 2026-07 conditional novelty 5.0

    A single-image 3DGS head avatar with internalized motion encoding and three region-specialized Gaussian branches runs real-time end-to-end and matches or beats recent baselines on reenactment metrics.

  26. TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation

    cs.CV 2026-07 conditional novelty 5.0

    Anchor-guided fixed-capacity multimodal memory plus stage-parallel few-step denoising stabilizes long-form causal audio-video digital humans at real-time speed.

  27. Real-Time Human-Centric World Modeling for Upper-Body Human-Object Interaction

    cs.CV 2026-07 conditional novelty 5.0

    A distilled video model jointly controls multi-scale upper-body motion latents and two discrete contact states to generate real-time human–object interaction at 25 FPS.

  28. Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation

    cs.CV 2026-06 unverdicted novelty 5.0

    Two-stage pipeline with region-aware attention and Mamba-enhanced diffusion achieves SOTA accuracy, naturalness and temporal coherence on audio-driven portrait animation benchmarks using a new 380-hour dataset.

  29. SteerFace: Debiasing Synthetic Face Generation via Adaptive Residue Perturbation

    cs.CV 2026-05 unverdicted novelty 5.0

    SteerFace perturbs identity embeddings toward random orthogonal directions on the hypersphere with an adaptive strategy to mitigate visual tendency in synthetic faces and improve downstream recognition performance.

  30. EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation

    cs.CV 2026-05 unverdicted novelty 5.0

    EasyVFX decouples VFX generation via frequency-aware Mixture-of-Experts and test-time training to achieve realistic effects with limited resources.

  31. PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment

    cs.CV 2026-04 unverdicted novelty 5.0

    PortraitDirector uses hierarchical disentanglement of spatial physical motions and semantic emotions to deliver controllable, high-fidelity real-time facial reenactment at 20 FPS.

  32. JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation

    cs.CV 2024-11 unverdicted novelty 5.0

    JoyVASA decouples static 3D facial representations from identity-independent dynamic motion sequences generated by a diffusion transformer to produce audio-driven animations for humans and animals.

  33. OpenWorldLib: A Unified Codebase and Definition of Advanced World Models

    cs.CV 2026-04 unverdicted novelty 4.0

    OpenWorldLib offers a standardized codebase and definition for world models that combine perception, interaction, and memory to understand and predict the world.

  34. OpenWorldLib: A Unified Codebase and Definition of Advanced World Models

    cs.CV 2026-04 conditional novelty 4.0

    OpenWorldLib defines world models as perception-centered systems with interaction and long-term memory, and provides a modular inference codebase unifying interactive video, 3D, reasoning, and VLA tasks.