Pith. sign in

REVIEW 27 cited by

MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.19680 v2 pith:T2NZJFS4 submitted 2024-06-28 cs.CV cs.AIcs.MM

MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

classification cs.CV cs.AIcs.MM
keywords generationmimicmotionvideoguidancelengthposevideosarbitrary
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In recent years, generative artificial intelligence has achieved significant advancements in the field of image generation, spawning a variety of applications. However, video generation still faces considerable challenges in various aspects, such as controllability, video length, and richness of details, which hinder the application and popularization of this technology. In this work, we propose a controllable video generation framework, dubbed MimicMotion, which can generate high-quality videos of arbitrary length mimicking specific motion guidance. Compared with previous methods, our approach has several highlights. Firstly, we introduce confidence-aware pose guidance that ensures high frame quality and temporal smoothness. Secondly, we introduce regional loss amplification based on pose confidence, which significantly reduces image distortion. Lastly, for generating long and smooth videos, we propose a progressive latent fusion strategy. By this means, we can produce videos of arbitrary length with acceptable resource consumption. With extensive experiments and user studies, MimicMotion demonstrates significant improvements over previous approaches in various aspects. Detailed results and comparisons are available on our project page: https://tencent.github.io/MimicMotion .

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 27 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis

    cs.CV 2026-04 unverdicted novelty 7.0

    ReImagine decouples human appearance from temporal consistency via pretrained image backbones, SMPL-X motion guidance, and training-free video diffusion refinement to generate high-quality controllable videos.

  2. HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation

    cs.CV 2026-04 unverdicted novelty 7.0

    HumANDiff improves motion consistency in human video generation by sampling diffusion noise on an articulated human body template and adding joint appearance-motion prediction plus a geometric consistency loss.

  3. CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos

    cs.CV 2026-01 unverdicted novelty 7.0

    CoMoVi co-generates 3D human motions and 2D videos synchronously in a single diffusion denoising loop using 3D-to-2D projection and dual-branch diffusion with 3D-2D cross attentions.

  4. One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer

    cs.CV 2025-11 unverdicted novelty 7.0

    One-to-All Animation enables alignment-free character animation and image pose transfer via self-supervised outpainting reformulation, reference extraction, hybrid fusion attention, identity-robust pose control, and t...

  5. AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment

    cs.CV 2026-07 conditional novelty 6.0

    AgentHOI generates human-object interaction videos from text plus one human image and one object image, using multi-agent action planning and implicit text-to-motion feature alignment inside a video diffusion model.

  6. MultiAnimate: A Unified Framework for Controllable Multi-Character Animation

    cs.CV 2026-07 conditional novelty 6.0

    A diffusion-based framework that animates multiple characters in one scene from separate reference images and pose sequences while preserving each character's identity.

  7. Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control

    cs.RO 2026-07 conditional novelty 6.0

    A DiT Mixture-of-Experts world model jointly learns locomotion, dual-arm manipulation, and egocentric hand control, with shared experts for world dynamics and progressive expert expansion for new modalities.

  8. 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

    cs.CV 2026-06 conditional novelty 6.0

    A training-free 3D scene-adaptive human animation framework that controls human motion and camera trajectories via ground-adaptive retargeting and visibility-masked point-cloud fusion in a diffusion backbone.

  9. 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

    cs.CV 2026-06 unverdicted novelty 6.0

    Presents a scene-adaptive 3D human animation method using ground-adaptive motion retargeting and viewpoint-adaptive latent fusion to control human trajectories and camera views, reporting gains on two benchmarks.

  10. EMOSH: Expressive Motion and Shape Disentanglement for Human Animation

    cs.CV 2026-06 unverdicted novelty 6.0

    EMOSH proposes an Expressive Human Model with disentangled parameters, coarse-to-fine motion injection, and spatially-aligned conditioning to generate high-fidelity expressive human videos without driving-subject shap...

  11. Cinematic Compositing Using Character-Environment-Harmonized Video Generation Models

    cs.CV 2026-06 conditional novelty 6.0

    An end-to-end diffusion model generates backgrounds, relights green-screen actors, and replaces or creates props in one pass, improving over cascaded inpainting plus relighting baselines in user preference and auto metrics.

  12. SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning

    cs.CV 2026-06 unverdicted novelty 6.0

    SCAIL-2 achieves end-to-end character animation via direct video concatenation, in-context mask conditioning, mode-specific RoPE, the synthetic MotionPair-60K dataset, and Bias-Aware DPO, outperforming prior methods o...

  13. Beyond Skeletons: Learning Animation Directly from Driving Videos with Same2X Training Strategy

    cs.CV 2026-06 unverdicted novelty 6.0

    DirectAnimator bypasses pose extraction using a Driving Cue Triplet and Same2X training strategy to achieve state-of-the-art human animation quality and robustness from raw videos.

  14. FreeAnimate: Training-Free Human Image Animation with Preview-Guided Denoising

    cs.CV 2026-06 unverdicted novelty 6.0

    A training-free framework for human image animation that uses preview-guided denoising plus Inversion-Boosted and Reference-Anchored Self-Attention modules to achieve temporal consistency and identity preservation in ...

  15. Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization

    cs.CV 2026-06 unverdicted novelty 6.0

    Introduces mesh tokenization to condition DiT-based video diffusion models directly on 3D human meshes for motion control without 2D rendering.

  16. EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration

    cs.CV 2026-05 unverdicted novelty 6.0

    EverAnimate restores drifted latent flow trajectories in chunked video generation via persistent latent propagation and restorative flow matching, achieving measurable gains in PSNR, SSIM, LPIPS, and FID over prior lo...

  17. Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing

    cs.RO 2026-05 unverdicted novelty 6.0

    A dual-contrastive disentanglement method factorizes videos into independent task and embodiment latents, then uses a parameter-efficient adapter on a frozen video diffusion model to synthesize robot executions from s...

  18. SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages

    cs.CV 2026-05 unverdicted novelty 6.0

    SignVerse-2M provides a 2-million-clip multilingual pose-native dataset for sign language derived from public videos via DWPose preprocessing to enable robust modeling in real-world conditions.

  19. Exploring the Role of Synthetic Data Augmentation in Controllable Human-Centric Video Generation

    cs.CV 2026-04 unverdicted novelty 6.0

    Synthetic data complements real data in diffusion-based controllable human video generation, with effective sample selection improving motion realism, temporal consistency, and identity preservation.

  20. HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis

    cs.CV 2026-03 unverdicted novelty 6.0

    HVG-3D uses a 3D-aware diffusion architecture with ControlNet to synthesize high-fidelity hand-object interaction videos from 3D control signals, achieving state-of-the-art spatial fidelity and temporal coherence on t...

  21. HairWeaver: Few-Shot Photorealistic Hair Motion Synthesis with Sim-to-Real Guided Video Diffusion

    cs.CV 2026-02 conditional novelty 6.0

    HairWeaver animates a single human photo with physically plausible hair motion by transferring simulated CG hair dynamics into a frozen video diffusion model via two lightweight LoRA adapters.

  22. SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation

    cs.CV 2025-11 unverdicted novelty 6.0

    SteadyDancer is an I2V framework using condition reconciliation, synergistic pose modulation, and staged training to achieve robust first-frame preservation and coherent motion control in human image animation.

  23. 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

    cs.CV 2026-06 unverdicted novelty 5.0

    Presents a scene-adaptive 3D human image animation framework using ground-adaptive motion retargeting and viewpoint-adaptive latent fusion to control human and camera trajectories, claiming improvements on two benchmarks.

  24. Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control

    cs.CV 2026-06 unverdicted novelty 5.0

    A decoupled-control autoregressive video model using Fast-Slow Memory training, dynamic projection, and staged camera control to produce stable long-horizon outputs with human and viewpoint guidance.

  25. Cinematic Compositing Using Character-Environment-Harmonized Video Generation Models

    cs.CV 2026-06 unverdicted novelty 5.0

    End-to-end video diffusion framework with tri-mask guidance and RGB-D denoising for joint modeling of character-to-environment physical interactions and environment-to-character lighting harmonization in cinematic com...

  26. Adding Thermal Awareness to Visual Systems in Real-Time via Distilled Diffusion Models

    cs.CV 2026-05 unverdicted novelty 5.0

    FusionProxy is a distilled diffusion-based fusion module that adds thermal awareness to RGB vision systems in real time as an independent plug-and-play component.

  27. LUNA: Learning Universal 3D Human Animation Beyond Skinning

    cs.CV 2026-06 unverdicted novelty 4.0

    LUNA is an LBS-free neural animation model that maps 2D controls to 3D Gaussian deformations via a transformer motion regressor and hybrid supervision for realistic motion and zero-shot generalization.