REVIEW 12 cited by
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Generating text-editable and pose-controllable character videos have an imperious demand in creating various digital human. Nevertheless, this task has been restricted by the absence of a comprehensive dataset featuring paired video-pose captions and the generative prior models for videos. In this work, we design a novel two-stage training scheme that can utilize easily obtained datasets (i.e.,image pose pair and pose-free video) and the pre-trained text-to-image (T2I) model to obtain the pose-controllable character videos. Specifically, in the first stage, only the keypoint-image pairs are used only for a controllable text-to-image generation. We learn a zero-initialized convolutional encoder to encode the pose information. In the second stage, we finetune the motion of the above network via a pose-free video dataset by adding the learnable temporal self-attention and reformed cross-frame self-attention blocks. Powered by our new designs, our method successfully generates continuously pose-controllable character videos while keeps the editing and concept composition ability of the pre-trained T2I model. The code and models will be made publicly available.
Forward citations
Cited by 12 Pith papers
-
FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors
Interactive image editing can be cast as image-to-video generation: initializing from Stable Video Diffusion plus a new matching attention mechanism yields high-quality sketch, drag, and coarse-edit results with far l...
-
MultiAnimate: A Unified Framework for Controllable Multi-Character Animation
A diffusion-based framework that animates multiple characters in one scene from separate reference images and pose sequences while preserving each character's identity.
-
3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement
Presents a scene-adaptive 3D human image animation framework using ground-adaptive motion retargeting and viewpoint-adaptive latent fusion to control human and camera trajectories, claiming improvements on two benchmarks.
-
STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
STARCaster is a 2D spatio-temporal video diffusion model that unifies identity-conditioned audio-driven portrait animation and novel-view synthesis without explicit 3D reconstruction.
-
HumanRAM: Feed-forward Human Reconstruction and Animation Model using Transformers
A feed-forward transformer model that adds SMPL-X neural-texture pose conditioning to LVSM, enabling single-pass human novel-view and novel-pose synthesis that surpasses prior generalizable methods on four benchmarks.
-
An Exploratory Study on Multi-modal Generative AI in AR Storytelling
The paper maps how storytellers prefer to use AI-generated text, audio, images, videos, and 3D content to augment AR stories, based on a 223-video analysis and two user studies with 30 participants.
-
Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss
A training-free method that matches sparse-point inter-frame feature correlations to transfer reference motion to generated videos with improved temporal consistency.
-
CFSynthesis: Controllable and Free-view 3D Human Video Synthesis
CFSynthesis generates free-view human videos from one reference image by conditioning a diffusion model on a textured SMPL body model and separately encoded foreground and background.
-
Wan-Animate-2: Pushing the Application Boundaries of Character Animation
Wan-Animate-2 animates a reference character from a driving video in one end-to-end diffusion transformer, adds text-driven viewpoint control, and distills a real-time streaming variant.
-
EnerVerse-AC: Envisioning Embodied Environments with Action Condition
EnerVerse-AC generates realistic multi-view robot videos conditioned on action sequences and shows early evidence it can augment training data and rank policy performance like a real robot.
-
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
VBench++ is a benchmark that scores text-to-video and image-to-video models on 16 quality dimensions plus trustworthiness, reporting human-alignment correlations for each.
-
Hierarchical Patch Compression for ColPali: Efficient Multi-Vector Document Retrieval with Dynamic Pruning and Quantization
HPC-ColPali compresses ColPali's patch embeddings via K-means quantization, attention-based pruning, and binary Hamming search, but all reported numbers are estimates, not measurements.
Discussion (0). Continue with ORCID to comment.