Pith. sign in

REVIEW 9 cited by

Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.01900 v3 pith:YTROFYQH submitted 2024-06-04 cs.CV

classification cs.CV
keywords portraitanimationexpressionfollow-your-emojireferencelandmarkmodelmotion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Follow-Your-Emoji, a diffusion-based framework for portrait animation, which animates a reference portrait with target landmark sequences. The main challenge of portrait animation is to preserve the identity of the reference portrait and transfer the target expression to this portrait while maintaining temporal consistency and fidelity. To address these challenges, Follow-Your-Emoji equipped the powerful Stable Diffusion model with two well-designed technologies. Specifically, we first adopt a new explicit motion signal, namely expression-aware landmark, to guide the animation process. We discover this landmark can not only ensure the accurate motion alignment between the reference portrait and target motion during inference but also increase the ability to portray exaggerated expressions (i.e., large pupil movements) and avoid identity leakage. Then, we propose a facial fine-grained loss to improve the model's ability of subtle expression perception and reference portrait appearance reconstruction by using both expression and facial masks. Accordingly, our method demonstrates significant performance in controlling the expression of freestyle portraits, including real humans, cartoons, sculptures, and even animals. By leveraging a simple and effective progressive generation strategy, we extend our model to stable long-term animation, thus increasing its potential application value. To address the lack of a benchmark for this field, we introduce EmojiBench, a comprehensive benchmark comprising diverse portrait images, driving videos, and landmarks. We show extensive evaluations on EmojiBench to verify the superiority of Follow-Your-Emoji.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Instant Expressive Gaussian Head Avatars at Over 100 FPS

    cs.CV 2025-12 conditional novelty 7.0 of 10

    A single-photo avatar encoder with per-Gaussian feature-space deformation animates faces at 107 FPS with expression quality competitive with diffusion models.

  2. Conversational Human Audio-visual Talking Dialogue Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    CHAT generates mutually responsive dyadic audio-visual dialogue clips from a single text prompt and yields a 50k synthetic pre-training set that improves facial reaction models on REACT 2024.

  3. X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention

    cs.CV 2025-07 conditional novelty 6.0 of 10

    X-NeMo trains a 1D identity-agnostic motion descriptor end-to-end with a diffusion model, enabling zero-shot portrait animation with improved identity and expression fidelity.

  4. OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A benchmark and five-million-clip dataset for evaluating and training subject-to-video generation models, with three new metrics for subject consistency, naturalness, and text alignment.

  5. Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A motion-prior diffusion model with archived-frame memory improves identity, lip-sync, and head-motion consistency in long talking-face videos.

  6. FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases

    cs.CV 2025-07 conditional novelty 5.0 of 10

    FixTalk adds two modules to a real-time GAN talking-head model, decoupling identity from motion to stop identity leakage while using a memory to recover details and reduce artifacts.

  7. FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers

    cs.CV 2025-06 conditional novelty 5.0 of 10

    FullDiT2 accelerates FullDiT-style in-context conditioning for video by dynamic token selection and selective context caching, cutting per-step time by 2-3x with minimal quality loss.

  8. Multi-View Face and Gesture Animation with Dynamic Gaussians

    cs.CV 2026-08 conditional novelty 4.0 of 10

    Combining separate face and hand models with a parametric body and Gaussian splatting enables multi-view-consistent upper-body avatars that can be re-animated with new expressions and gestures.

  9. LIA-X: Interpretable Latent Portrait Animator

    cs.CV 2025-08 conditional novelty 4.0 of 10

    Adding an L1 sparsity penalty to the motion dictionary of the LIA portrait animator produces disentangled, human-interpretable motion vectors that support controllable image and video editing and scale to roughly one ...

Pith tools