REVIEW 13 cited by
MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The growing demand for high-fidelity video generation from textual descriptions has catalyzed significant research in this field. In this work, we introduce MagicVideo-V2 that integrates the text-to-image model, video motion generator, reference image embedding module and frame interpolation module into an end-to-end video generation pipeline. Benefiting from these architecture designs, MagicVideo-V2 can generate an aesthetically pleasing, high-resolution video with remarkable fidelity and smoothness. It demonstrates superior performance over leading Text-to-Video systems such as Runway, Pika 1.0, Morph, Moon Valley and Stable Video Diffusion model via user evaluation at large scale.
Forward citations
Cited by 13 Pith papers
-
GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion
GeoMan predicts temporally consistent depth and normals for human videos by conditioning an image-to-video diffusion model on first-frame geometry and using a root-relative depth representation.
-
GUAVA: Generalizable Upper Body 3D Gaussian Avatar
From a single image, GUAVA builds an animatable upper-body 3D Gaussian avatar in one forward pass, using a new hybrid SMPLX/FLAME template and inverse texture mapping, then renders it in real time.
-
ProxyUp: Training-Free Proxy-Conditioned Video Generation for Controllable Dynamics
A training-free method preserves proxy-video dynamics via region-wise latent noising and Stochastic Flow Relaxation, outperforming editing and motion-transfer baselines on a custom 76-video set.
-
Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models
A video-diffusion pipeline conditioned on geometry maps, followed by component-wise UV inpainting, produces more coherent and seam-free textures for 3D meshes than Text2Tex, Paint3D, and Meshy in the reported tests.
-
Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
Phantom-Data provides around one million cross-context, identity-consistent reference-video pairs for subject-to-video generation, and training on it improves prompt following and visual quality.
-
ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos
ProphetDWM is a one-stage diffusion world model that jointly predicts future driving video and low-level actions from a current frame and a short action sequence.
-
StableAnimator: High-Quality Identity-Preserving Human Image Animation
StableAnimator uses a video diffusion model with face-embedding adapters and per-step latent optimization to generate pose-driven videos that preserve the reference person's identity end-to-end.
-
I2VControl: Disentangled and Unified Video Motion Synthesis Control
I2VControl unifies camera, drag, and brush controls into a single point-trajectory-based adapter for image-to-video diffusion models, enabling conflict-free combined motion control.
-
StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation
A diffusion-based avatar generator that uses a timestep-aware audio adapter, audio-adaptive guidance, and weighted sliding-window fusion to produce audio-synced talking-head videos several minutes long with reduced id...
-
StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation
StableAnimator++ combines learnable SVD-guided pose alignment, a distribution-aware ID Adapter, and an HJB-based inference-time face optimizer to preserve identity in human image animation under severe pose misalignment.
-
SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers
An audio-conditioned video diffusion transformer that animates portraits from image, video, text, and audio inputs with a sliding-window fusion for long videos.
-
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
VBench++ is a benchmark that scores text-to-video and image-to-video models on 16 quality dimensions plus trustworthiness, reporting human-alignment correlations for each.
-
Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook
The paper introduces BioDeepAV, a benchmark of real and fake talking-face videos, and reports that state-of-the-art deepfake detectors drop sharply when tested on deepfakes from unseen generators.
Discussion (0). Continue with ORCID to comment.