Pith. sign in

REVIEW 25 cited by

Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.17117 v3 pith:2XACAJKJ submitted 2023-11-28 cs.CV

Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation

classification cs.CV
keywords characteranimationimage-to-videoanimateapproachconsistencydiffusionensure
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Character Animation aims to generating character videos from still images through driving signals. Currently, diffusion models have become the mainstream in visual generation research, owing to their robust generative capabilities. However, challenges persist in the realm of image-to-video, especially in character animation, where temporally maintaining consistency with detailed information from character remains a formidable problem. In this paper, we leverage the power of diffusion models and propose a novel framework tailored for character animation. To preserve consistency of intricate appearance features from reference image, we design ReferenceNet to merge detail features via spatial attention. To ensure controllability and continuity, we introduce an efficient pose guider to direct character's movements and employ an effective temporal modeling approach to ensure smooth inter-frame transitions between video frames. By expanding the training data, our approach can animate arbitrary characters, yielding superior results in character animation compared to other image-to-video methods. Furthermore, we evaluate our method on benchmarks for fashion video and human dance synthesis, achieving state-of-the-art results.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance

    cs.CV 2026-05 unverdicted novelty 7.0

    iTryOn is a video diffusion Transformer that injects spatial 3D hand guidance and semantic action captions to enable interactive garment replacement in videos.

  2. iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance

    cs.CV 2026-05 unverdicted novelty 7.0

    iTryOn is a diffusion-based framework that adds spatial 3D hand guidance and semantic action-aware embeddings to handle complex garment deformations during human-clothing interactions in videos.

  3. ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos

    cs.CV 2026-04 unverdicted novelty 7.0

    ExpertEdit edits novice motions to expert skill levels by learning a motion prior from unpaired videos and infilling masked skill-critical spans.

  4. Screen, Cache, and Match: A Training-Free Causality-Consistent Reference Frame Framework for Human Animation

    cs.GR 2025-12 unverdicted novelty 7.0

    FrameCache uses a Screen-Cache-Match strategy and Trajectory-Aware Autoregressive Generation to convert past frames into causal guidance for temporally coherent human animation videos.

  5. GimbalDiffusion: Gravity-Aware Camera Control for Video Generation

    cs.CV 2025-12 conditional novelty 7.0

    A text-to-video method that conditions on absolute, gravity-aligned camera pitch/roll (trained from 360° video) achieves about a quarter lower absolute pitch error than the best baseline on a new extreme-angle benchmark.

  6. GimbalDiffusion: Gravity-Aware Camera Control for Video Generation

    cs.CV 2025-12 conditional novelty 7.0

    GimbalDiffusion adds gravity-referenced absolute camera control and null-pitch conditioning to text-to-video diffusion models, trained on full-sphere panoramic data, to support extreme trajectories and reduce prompt e...

  7. GimbalDiffusion: Gravity-Aware Camera Control for Video Generation

    cs.CV 2025-12 conditional novelty 7.0

    GimbalDiffusion lets text-to-video models follow absolute, gravity-aligned camera rotations by training on random crops from 360° video with forward-facing captions.

  8. Vid-Freeze: Protecting Images from Malicious Image-to-Video Generation via Temporal Freezing

    cs.CV 2025-09 unverdicted novelty 7.0

    Vid-Freeze immunizes images by adding perturbations that target attention dynamics in I2V models to enforce temporal freezing and suppress motion synthesis.

  9. MultiAnimate: A Unified Framework for Controllable Multi-Character Animation

    cs.CV 2026-07 conditional novelty 6.0

    A diffusion-based framework that animates multiple characters in one scene from separate reference images and pose sequences while preserving each character's identity.

  10. HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control

    cs.CV 2026-07 unverdicted novelty 6.0

    HandsOnWorld creates a hand-controlled egocentric video generator from unconstrained monocular video via a new EgoVid-Pro dataset from monocular reconstruction and a Plücker Hand Map that disentangles camera and hand motion.

  11. 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

    cs.CV 2026-06 unverdicted novelty 6.0

    Presents a scene-adaptive 3D human animation method using ground-adaptive motion retargeting and viewpoint-adaptive latent fusion to control human trajectories and camera views, reporting gains on two benchmarks.

  12. 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

    cs.CV 2026-06 conditional novelty 6.0

    A training-free 3D scene-adaptive human animation framework that controls human motion and camera trajectories via ground-adaptive retargeting and visibility-masked point-cloud fusion in a diffusion backbone.

  13. Error-Conditioned Neural Solvers

    cs.LG 2026-06 unverdicted novelty 6.0

    Error-Conditioned Neural Solvers improve PDE prediction accuracy by using the residual field as network input for learned corrections, outperforming residual-minimization methods by up to 10x on turbulent flows and ge...

  14. EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration

    cs.CV 2026-05 unverdicted novelty 6.0

    EverAnimate restores drifted latent flow trajectories in chunked video generation via persistent latent propagation and restorative flow matching, achieving measurable gains in PSNR, SSIM, LPIPS, and FID over prior lo...

  15. MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model

    cs.CV 2026-03 conditional novelty 6.0

    Using a 3D foundation model to produce viewpoint-aware anchors plus multi-view reference textures enables realistic human-object-interaction reenactment with large out-of-plane rotations.

  16. STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

    cs.CV 2025-12 conditional novelty 6.0

    STARCaster is a 2D spatio-temporal video diffusion model that unifies identity-conditioned audio-driven portrait animation and novel-view synthesis without explicit 3D reconstruction.

  17. CameraCtrl: Enabling Camera Control for Text-to-Video Generation

    cs.CV 2024-04 unverdicted novelty 6.0

    CameraCtrl enables accurate camera pose control in video diffusion models through a trained plug-and-play module and dataset choices emphasizing diverse camera trajectories with matching appearance.

  18. 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

    cs.CV 2026-06 unverdicted novelty 5.0

    Presents a scene-adaptive 3D human image animation framework using ground-adaptive motion retargeting and viewpoint-adaptive latent fusion to control human and camera trajectories, claiming improvements on two benchmarks.

  19. Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation

    cs.CV 2026-05 unverdicted novelty 5.0

    A controllable generative augmentation approach synthesizes diverse pose videos from indoor and outdoor datasets to improve model performance on unseen domains in 3D human pose estimation.

  20. InfinityHuman: Towards Long-Term Audio-Driven Human

    cs.CV 2025-08 conditional novelty 5.0

    A coarse-to-fine audio-driven animation framework that uses pose-guided refinement and hand-specific reward learning to generate long, identity-stable talking videos.

  21. DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment

    cs.RO 2025-04 unverdicted novelty 5.0

    DriVerse is a generative model that simulates driving scenes from an image and trajectory using multimodal prompting and motion alignment, achieving better performance on nuScenes and Waymo datasets with minimal training.

  22. Pose-dIVE: Pose-Diversified Augmentation with Diffusion Model for Person Re-Identification

    cs.CV 2024-06 unverdicted novelty 5.0

    Pose-dIVE augments Re-ID training sets with diffusion-generated images of diverse poses and viewpoints by conditioning on SMPL parameters.

  23. How to Build Digital Humans? From Priors to Photorealistic Avatars

    cs.GR 2026-07 accept novelty 4.0

    A taxonomy-driven state-of-the-art report that structures controllable 3D human avatar creation around prior learning and personalization, reviewing full-body, head, and layered (hair/hands/garments) methods.

  24. EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation

    cs.CV 2026-02 unverdicted novelty 4.0

    EchoTorrent combines multi-teacher distillation, adaptive CFG calibration, hybrid long-tail forcing, and VAE decoder refinement to enable few-pass autoregressive streaming video generation with improved temporal consi...

  25. Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

    cs.CV 2024-02 unverdicted novelty 2.0

    The paper reviews the background, technology, applications, limitations, and future directions of OpenAI's Sora text-to-video generative model based on public information.