Pith. sign in

REVIEW 3 cited by

RealisDance-DiT: Simple yet Strong Baseline towards Controllable Character Animation in the Wild

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.14977 v1 pith:OZO4CRT5 submitted 2025-04-21 cs.CV

classification cs.CV
keywords modelfoundationanimationcharactercontrollabledatasetrealisdance-ditbaseline
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Controllable character animation remains a challenging problem, particularly in handling rare poses, stylized characters, character-object interactions, complex illumination, and dynamic scenes. To tackle these issues, prior work has largely focused on injecting pose and appearance guidance via elaborate bypass networks, but often struggles to generalize to open-world scenarios. In this paper, we propose a new perspective that, as long as the foundation model is powerful enough, straightforward model modifications with flexible fine-tuning strategies can largely address the above challenges, taking a step towards controllable character animation in the wild. Specifically, we introduce RealisDance-DiT, built upon the Wan-2.1 video foundation model. Our sufficient analysis reveals that the widely adopted Reference Net design is suboptimal for large-scale DiT models. Instead, we demonstrate that minimal modifications to the foundation model architecture yield a surprisingly strong baseline. We further propose the low-noise warmup and "large batches and small iterations" strategies to accelerate model convergence during fine-tuning while maximally preserving the priors of the foundation model. In addition, we introduce a new test dataset that captures diverse real-world challenges, complementing existing benchmarks such as TikTok dataset and UBC fashion video dataset, to comprehensively evaluate the proposed method. Extensive experiments show that RealisDance-DiT outperforms existing methods by a large margin.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A visual proxy that renders human motion under the driving camera and overlays camera trajectory markers lets a video diffusion model control both body motion and camera movement from a single visual conditioning space.

  2. 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Presents a scene-adaptive 3D human image animation framework using ground-adaptive motion retargeting and viewpoint-adaptive latent fusion to control human and camera trajectories, claiming improvements on two benchmarks.

  3. AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Identity-finetuned video diffusion plus RF-Inversion can supply multi-view body supervision that lets 3D Gaussian avatars be completed and animated from heavily occluded monocular video.

Pith tools