REVIEW 35 cited by
SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present Stable Video 4D (SV4D), a latent video diffusion model for multi-frame and multi-view consistent dynamic 3D content generation. Unlike previous methods that rely on separately trained generative models for video generation and novel view synthesis, we design a unified diffusion model to generate novel view videos of dynamic 3D objects. Specifically, given a monocular reference video, SV4D generates novel views for each video frame that are temporally consistent. We then use the generated novel view videos to optimize an implicit 4D representation (dynamic NeRF) efficiently, without the need for cumbersome SDS-based optimization used in most prior works. To train our unified novel view video generation model, we curate a dynamic 3D object dataset from the existing Objaverse dataset. Extensive experimental results on multiple datasets and user studies demonstrate SV4D's state-of-the-art performance on novel-view video synthesis as well as 4D generation compared to prior works.
Forward citations
Cited by 35 Pith papers
-
MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing
MV-Forcing composes temporal and view-sequential autoregression in a single diffusion model, using a recurrent 3D reconstruction model as a geometric bridge to generate arbitrarily long, multi-view consistent videos.
-
4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
4D-LRM is a transformer that maps sparse posed frames scattered across time to a cloud of 4D Gaussians and renders any query view at any query time in under 1.5 seconds.
-
Wonderland: Navigating 3D Scenes from a Single Image
A feed-forward pipeline reconstructs 3D Gaussian scenes from single images by regressing 3DGS directly from camera-conditioned video diffusion latents.
-
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.
-
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
Distributed video-generation agents can create a spatiotemporally consistent shared world by tiling four views, exchanging cross-agent attention, and querying a spatial memory cache.
-
WildSmoke: Ready-to-Use Dynamic 3D Smoke Assets from a Single Video in the Wild
WildSmoke reconstructs editable dynamic 3D smoke assets from a single in-the-wild video and reports a +2.22 dB average PSNR gain over prior fluid reconstruction methods on four real-world videos.
-
Learning an Implicit Physics Model for Image-based Fluid Simulation
Using a simplified physics loss and 3D Gaussians, a neural network animates a single fluid image into videos with novel views, beating earlier methods on quality and motion accuracy.
-
CharacterShot: Controllable and Consistent 4D Character Animation
A new pipeline generates pose-controlled, view-consistent 4D character animations from one reference image and a 2D pose sequence, backed by a new 13,115-character dataset and benchmark.
-
4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation
A two-stage cascaded video diffusion model generates 16-view consistent videos from a monocular video, enabling higher-quality 4D content reconstruction.
-
Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
A video-to-4D model that encodes mesh animations into compact Gaussian variation latents and diffuses them conditioned on the video and a canonical Gaussian splat.
-
Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models
A sliding iterative denoising scheme that alternates spatial and temporal passes, combined with skeleton conditioning, lets a diffusion model create spatio-temporally consistent multi-view human videos from sparse inp...
-
Voyaging into Perpetual Dynamic Scenes from a Single View
A single-view dynamic scene can be extended into an unbounded fly-through video by iteratively outpainting partial views of a learned 4D point cloud with ray distance guidance.
-
WonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene Exploration
A pipeline that restores corrupted novel-view videos with a video diffusion model and jointly denoises multiple viewpoints to improve 3D scene exploration from a single image.
-
CoCo4D: Comprehensive and Complex 4D Scene Generation
CoCo4D generates multi-view consistent 4D scenes from text or image prompts in about one hour by generating a reference video, reconstructing the foreground and background separately, and composing them with a learned...
-
Restereo: Diffusion stereo video generation and restoration
A diffusion model fine-tuned on synthetically degraded stereo videos simultaneously generates a consistent stereo pair and restores low-resolution or compressed input, outperforming prior stereo generators on low-qual...
-
TesserAct: Learning 4D Embodied World Models
A 4D embodied world model that generates RGB-depth-normal videos from an image and instruction, reconstructs the scene as point clouds, and uses those point clouds to train better robot manipulation policies.
-
Not All Frame Features Are Equal: Video-to-4D Generation via Decoupling Dynamic-Static Features
A video-to-4D generation method that decouples dynamic and static features in DINOv2 space and fuses similar dynamic information across views reports state-of-the-art scores on Consistent4D and Objaverse.
-
Articulate That Object Part (ATOP): 3D Part Articulation via Text and Motion Personalization
ATOP personalizes a pre-trained multi-view diffusion model with a few reference videos to generate part motion from text and masks, then lifts that motion to a 3D articulation axis via score distillation.
-
GAS: Generative Avatar Synthesis from a Single Image
GAS generates view-consistent, temporally coherent avatars from a single image by feeding NeRF renderings of the target view plus SMPL normal maps into a video diffusion model.
-
AR4D: Autoregressive 4D Generation from Monocular Videos
AR4D generates 4D content from monocular video by autoregressively deforming frame-wise 3D Gaussians, with progressive pseudo-view supervision from a pre-trained reconstruction model.
-
DreamDrive: Generative 4D Scene Modeling from Street View Images
DreamDrive generates 3D-consistent driving videos from a single image by lifting diffusion-generated reference frames into a hybrid static and dynamic 4D Gaussian scene.
-
Grid: Omni Visual Generation
GRID shows that fine-tuning an image diffusion model on videos arranged as grid images can generate coherent video and multi-view sequences with far less data and compute than specialized video models.
-
SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
A cross-view attention module and hybrid training recipe let a frozen text-to-video model generate synchronized multi-camera videos from arbitrary viewpoints given only a text prompt and camera poses.
-
4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion
4Real-Video generates consistent 4D video grids, frames across time and viewpoint, with a parallel two-stream diffusion transformer that synchronizes temporal and viewpoint token streams, achieving faster and higher-q...
-
Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis
By recasting text-to-image generation as multi-frame video generation and adding a differential camera encoder, the method achieves camera intrinsic control with scene consistency, outperforming current text-to-image ...
-
Gaussians-to-Life: Text-Driven Animation of 3D Gaussian Splatting Scenes
A text-driven pipeline that lifts 2D video diffusion motion into 3D Gaussian Splatting scenes via point tracking and depth estimation.
-
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
AC3D improves camera control in video diffusion transformers by conditioning only early denoising steps and the first 8 of 32 blocks, and by adding 20K static-camera dynamic videos to training.
-
Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation
A diffusion model generates dense 4D point trajectories from a single image, and a separate view-synthesis module renders them into novel-view videos.
-
DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective
A monocular human avatar reconstruction method generates pseudo back-view videos with a fine-tuned diffusion model and uses them as extra training data for a 3D Gaussian avatar.
-
MVG4D: Image Matrix-Based Multi-View and Motion Generation for 4D Content Creation from a Single Image
MVG4D generates a time-ordered multi-view image matrix from a single image and uses it to optimize a deformable 4D Gaussian Splatting scene, reporting higher quality and lower runtime than baselines.
-
Drive Any Mesh: 4D Latent Diffusion for Mesh Deformation from Video
A video-conditioned latent diffusion model generates mesh vertex trajectories that deform an input 3D asset into render-ready 4D animations.
-
DSplats: 3D Generation by Denoising Splats-Based Multiview Diffusion Models
DSplats is a single-stage multiview diffusion model that denoises latents through a 3D Gaussian reconstructor, achieving state-of-the-art single-image-to-3D reconstruction on the Google Scanned Objects benchmark.
-
PaintScene4D: Consistent 4D Scene Generation from Text Prompts
A training-free pipeline that turns one text-to-video clip into a multi-view 4D scene renderable along user-chosen camera paths.
-
Reconstructing 4D Spatial Intelligence: A Survey
A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.
-
Dynamic View Synthesis as an Inverse Problem
Dynamic view synthesis from a monocular video is achieved by redesigning the noise initialization of a pretrained video diffusion model using a recursive interpolation and a stochastic latent modulation.
Discussion (0). Continue with ORCID to comment.