Pith. sign in

REVIEW 30 cited by

ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.15275 v3 pith:YSKK5QE4 submitted 2024-04-23 cs.CV

ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

classification cs.CV
keywords generationvideoid-animatorhumanidentityfacialmodelstraining
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Generating high-fidelity human video with specified identities has attracted significant attention in the content generation community. However, existing techniques struggle to strike a balance between training efficiency and identity preservation, either requiring tedious case-by-case fine-tuning or usually missing identity details in the video generation process. In this study, we present \textbf{ID-Animator}, a zero-shot human-video generation approach that can perform personalized video generation given a single reference facial image without further training. ID-Animator inherits existing diffusion-based video generation backbones with a face adapter to encode the ID-relevant embeddings from learnable facial latent queries. To facilitate the extraction of identity information in video generation, we introduce an ID-oriented dataset construction pipeline that incorporates unified human attributes and action captioning techniques from a constructed facial image pool. Based on this pipeline, a random reference training strategy is further devised to precisely capture the ID-relevant embeddings with an ID-preserving loss, thus improving the fidelity and generalization capacity of our model for ID-specific video generation. Extensive experiments demonstrate the superiority of ID-Animator to generate personalized human videos over previous models. Moreover, our method is highly compatible with popular pre-trained T2V models like animatediff and various community backbone models, showing high extendability in real-world applications for video generation where identity preservation is highly desired. Our codes and checkpoints are released at https://github.com/ID-Animator/ID-Animator.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 30 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A Comprehensive Ecosystem for Open-Domain Customized Video Generation

    cs.CV 2026-06 unverdicted novelty 7.0

    Introduces PexelsCustom-1M dataset, CustoMDiT parameter-efficient model, and OpenCustom benchmark for open-domain customized video generation.

  2. EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation

    cs.CV 2026-05 conditional novelty 7.0

    EntityBench is a new benchmark with detailed per-shot entity schedules from real media, and the EntityMem baseline using persistent per-entity memory achieves the highest character fidelity with Cohen's d of +2.33.

  3. MoZoo:Unleashing Video Diffusion power in animal fur and muscle simulation

    cs.GR 2026-04 unverdicted novelty 7.0

    MoZoo generates high-fidelity animal videos with fur and muscle dynamics from coarse meshes by extending video diffusion with role-aware RoPE and asymmetric decoupled attention, trained on a new synthetic-to-real dataset.

  4. ID-V2V: Identity-Preserving Video Restylization

    cs.CV 2026-07 conditional novelty 6.0

    ID-V2V restyles video by conditioning a diffusion model on edited keyframes, depth, relit faces, and face normals, so scene edits propagate while facial identity and performance are preserved.

  5. GroupVideo: Multi-Identity Customized Text-to-Video Generation

    cs.CV 2026-07 conditional novelty 6.0

    GroupVideo generates multi-person videos from reference photos plus text, using multimodal identity alignment and ID localization to keep each person's identity consistent.

  6. HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

    cs.CV 2026-07 conditional novelty 6.0

    HOMIE unifies inter- and intra-subject video personalization by injecting MLLM-derived relational features into DiT self-attention (GMG) and tagging tokens with modality/reference embeddings (MRE), reporting SOTA on a...

  7. Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization

    cs.CV 2026-07 conditional novelty 6.0

    TC-UAP learns a shared multi-frame adversarial perturbation that protects videos of the same identity from both fine-tuning-based and reference-based video customization, remaining effective on unseen clips and under ...

  8. Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment

    cs.CV 2026-07 conditional novelty 6.0

    Aura combines VLM meta-queries, T5-teacher alignment, subject-aware RoPE shifts, memory tokens, and a large AIGC-curated dataset to claim SOTA multi-element subject-to-video generation under OpenS2V-Eval Total score.

  9. GeoEdit: Geometry-Aware Object Editing via Dual-Branch Denoising

    cs.CV 2026-06 unverdicted novelty 6.0

    GeoEdit introduces a Lift-Manipulate-Render-Denoise pipeline with dual-branch denoising and variance-homogeneous injection for 3D-consistent object editing in single photos.

  10. Customizing Video Portraits via Identity-ActionDecoupling

    cs.CV 2026-06 unverdicted novelty 6.0

    Proposes IaD framework with Identity Decoupling Loss and Text Alignment Loss for richer, identity-consistent IPT2V without subject-specific fine-tuning.

  11. ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation

    cs.CV 2026-06 unverdicted novelty 6.0

    ARGUS converts MLLM-selected identity evidence into a synchronized 3x3 mosaic injected as negative-time memory in a diffusion model, plus supporting training techniques, to achieve SOTA subject preservation on human v...

  12. FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization

    cs.CV 2026-05 unverdicted novelty 6.0

    FashionChameleon achieves interactive multi-garment video customization at 23.8 FPS via in-context teacher models, streaming distillation, and training-free KV cache rescheduling while using only single-garment data.

  13. FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization

    cs.CV 2026-05 unverdicted novelty 6.0

    FashionChameleon achieves interactive multi-garment video customization in real time by training a teacher model with in-context learning on single-garment pairs, applying streaming distillation, and using training-fr...

  14. FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    FaithfulFaces introduces a pose-faithful identity aligner with a shared dictionary and invariance constraint to maintain facial identity in text-to-video generation under large pose changes and occlusions.

  15. CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration

    cs.MM 2026-04 unverdicted novelty 6.0

    CineAGI is a multi-agent LLM framework that generates multi-scene movies with improved character consistency, narrative coherence, and audio-visual alignment.

  16. MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation

    cs.CV 2026-04 unverdicted novelty 6.0

    MMControl adds multi-modal controls for identity, timbre, pose, and layout to unified audio-video diffusion models via dual-stream injection and adjustable guidance scaling.

  17. MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation

    cs.CV 2026-04 conditional novelty 6.0

    Translation function vectors extracted from a single English→X direction transfer across unseen target languages in three multilingual LLMs, extending language-agnosticity findings to task-level representations.

  18. OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

    cs.SD 2026-02 conditional novelty 6.0

    A zero-shot model that generates a video of a reference face speaking user-chosen text with a reference voice timbre.

  19. HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

    cs.CV 2025-09 conditional novelty 6.0

    HuMo uses a two-stage training scheme and a face-focus trick to generate human videos that follow text, keep a reference person's identity, and sync speech to audio, beating several single-task systems on benchmarks.

  20. Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute

    cs.CV 2025-04 unverdicted novelty 6.0

    A zero-shot subject-driven video generation framework that decomposes the task into identity injection from 200K subject-image pairs and motion preservation from 4K arbitrary videos, trained in 288 A100 GPU hours on C...

  21. Vera: Identity-Faithful Human Subject-to-Video Generation

    cs.CV 2026-07 conditional novelty 5.0

    Vera improves identity consistency in human subject-to-video generation using cross-clip identity-aligned data, face-weighted masked loss, and layer-aware reference attention.

  22. Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation

    cs.CV 2026-07 conditional novelty 5.0

    A keyframe-anchored, training-free pipeline—terminal-state prompts, chained keyframe generation, and identity-aware sampling—ranks third on the IPVG26 Track 2 leaderboard.

  23. DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation

    cs.CV 2026-06 unverdicted novelty 5.0

    DomainShuttle introduces domain-aware modeling and token separation techniques to achieve high subject fidelity with generative flexibility in open-domain subject-driven text-to-video tasks.

  24. EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis

    cs.CV 2026-06 unverdicted novelty 5.0

    EchoStyle is a text-driven framework for arbitrary-length video stylization that creates the V-Style20k dataset through reverse synthesis and adds init-follow-mode with sliding windows to reduce style drift and motion issues.

  25. HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation

    cs.CV 2026-06 unverdicted novelty 5.0

    HarmoView proposes Multi-level Feature Injection, learnable proxy tokens, Jump-RoPE, and Progressive View Curriculum plus a new multi-view dataset to achieve state-of-the-art identity-consistent video generation from ...

  26. Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation

    cs.CV 2026-05 unverdicted novelty 5.0

    Omni-Customizer proposes an end-to-end framework using Omni-Context Fusion, Masked TTS Cross-Attention, Semantic-Anchored Multimodal RoPE, and specialized training curricula to achieve precise multimodal identity bind...

  27. Movie Gen: A Cast of Media Foundation Models

    cs.CV 2024-10 unverdicted novelty 5.0

    A 30B-parameter transformer and related models generate high-quality videos and audio, claiming state-of-the-art results on text-to-video, video editing, personalization, and audio generation tasks.

  28. Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation

    cs.CV 2026-06 unverdicted novelty 4.0

    ST-DRC proposes latent in-context injection, TASS-RoPE, appearance-invariant augmentation, and three-stream guidance to improve identity preservation in text-to-video diffusion models built on LTX-2.3.

  29. Wan-Image: Pushing the Boundaries of Generative Visual Intelligence

    cs.CV 2026-04 unverdicted novelty 3.0

    Wan-Image is a unified multi-modal system that integrates LLMs and diffusion transformers to deliver professional-grade image generation features including complex typography, multi-subject consistency, and precise ed...

  30. Evolution of Video Generative Foundations

    cs.CV 2026-04 unverdicted novelty 2.0

    This survey traces video generation technology from GANs to diffusion models and then to autoregressive and multimodal approaches while analyzing principles, strengths, and future trends.