Pith. sign in

REVIEW 11 cited by

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.04847 v5 pith:4UONZEG6 submitted 2025-02-07 cs.CV

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

classification cs.CV
keywords videohumanditsequencesgenerationposevideosacrossbody
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Human motion video generation has advanced significantly, while existing methods still struggle with accurately rendering detailed body parts like hands and faces, especially in long sequences and intricate motions. Current approaches also rely on fixed resolution and struggle to maintain visual consistency. To address these limitations, we propose HumanDiT, a pose-guided Diffusion Transformer (DiT)-based framework trained on a large and wild dataset containing 14,000 hours of high-quality video to produce high-fidelity videos with fine-grained body rendering. Specifically, (i) HumanDiT, built on DiT, supports numerous video resolutions and variable sequence lengths, facilitating learning for long-sequence video generation; (ii) we introduce a prefix-latent reference strategy to maintain personalized characteristics across extended sequences. Furthermore, during inference, HumanDiT leverages Keypoint-DiT to generate subsequent pose sequences, facilitating video continuation from static images or existing videos. It also utilizes a Pose Adapter to enable pose transfer with given sequences. Extensive experiments demonstrate its superior performance in generating long-form, pose-accurate videos across diverse scenarios.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Reshoot-Anything: A Self-Supervised Model for In-the-Wild Video Reshooting

    cs.CV 2026-04 unverdicted novelty 7.0

    Reshoot-Anything trains a diffusion transformer on pseudo multi-view triplets created by cropping and warping monocular videos to achieve temporally consistent video reshooting with robust camera control on dynamic scenes.

  2. CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos

    cs.CV 2026-01 unverdicted novelty 7.0

    CoMoVi co-generates 3D human motions and 2D videos synchronously in a single diffusion denoising loop using 3D-to-2D projection and dual-branch diffusion with 3D-2D cross attentions.

  3. One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer

    cs.CV 2025-11 unverdicted novelty 7.0

    One-to-All Animation enables alignment-free character animation and image pose transfer via self-supervised outpainting reformulation, reference extraction, hybrid fusion attention, identity-robust pose control, and t...

  4. ViDS: Video Diffusion Shader using 3D Face Tracking

    cs.CV 2026-07 conditional novelty 6.0

    Fine-tuning a video diffusion model on dense 3DMM normal maps from Pixel3DMM lets a single portrait photo be animated with a driving video's expressions and pose, surpassing landmark- and latent-based portrait animati...

  5. Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Video Synthesis

    cs.CV 2026-06 unverdicted novelty 6.0

    A new framework factorizes weather video synthesis into semantic appearance anchoring, physics-informed Gaussian particle simulation under gravity/wind/turbulence, and geometry-grounded alignment to produce diverse re...

  6. Beyond Skeletons: Learning Animation Directly from Driving Videos with Same2X Training Strategy

    cs.CV 2026-06 unverdicted novelty 6.0

    DirectAnimator bypasses pose extraction using a Driving Cue Triplet and Same2X training strategy to achieve state-of-the-art human animation quality and robustness from raw videos.

  7. Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization

    cs.CV 2026-06 unverdicted novelty 6.0

    Introduces mesh tokenization to condition DiT-based video diffusion models directly on 3D human meshes for motion control without 2D rendering.

  8. EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration

    cs.CV 2026-05 unverdicted novelty 6.0

    EverAnimate restores drifted latent flow trajectories in chunked video generation via persistent latent propagation and restorative flow matching, achieving measurable gains in PSNR, SSIM, LPIPS, and FID over prior lo...

  9. OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation

    cs.CV 2026-04 unverdicted novelty 6.0

    OmniShow unifies text, image, audio, and pose conditions into an end-to-end model for high-quality human-object interaction video generation and introduces the HOIVG-Bench benchmark, claiming state-of-the-art results.

  10. AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors

    cs.CV 2026-03 conditional novelty 6.0

    Identity-finetuned video diffusion plus RF-Inversion can supply multi-view body supervision that lets 3D Gaussian avatars be completed and animated from heavily occluded monocular video.

  11. Image-to-Video Diffusion: From Foundations to Open Frontiers

    cs.CV 2026-05 unverdicted novelty 3.0

    A survey that organizes diffusion image-to-video methods into a taxonomy, distills core designs in condition encoding, temporal modeling, noise prior, and upsampling, and discusses applications plus challenges.