Pith. sign in

REVIEW 35 cited by

SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.17470 v2 pith:RXWRTLYP submitted 2024-07-24 cs.CV

classification cs.CV
keywords videodynamicgenerationnovelsv4dviewmodelconsistent
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present Stable Video 4D (SV4D), a latent video diffusion model for multi-frame and multi-view consistent dynamic 3D content generation. Unlike previous methods that rely on separately trained generative models for video generation and novel view synthesis, we design a unified diffusion model to generate novel view videos of dynamic 3D objects. Specifically, given a monocular reference video, SV4D generates novel views for each video frame that are temporally consistent. We then use the generated novel view videos to optimize an implicit 4D representation (dynamic NeRF) efficiently, without the need for cumbersome SDS-based optimization used in most prior works. To train our unified novel view video generation model, we curate a dynamic 3D object dataset from the existing Objaverse dataset. Extensive experimental results on multiple datasets and user studies demonstrate SV4D's state-of-the-art performance on novel-view video synthesis as well as 4D generation compared to prior works.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 35 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

    cs.CV 2026-07 conditional novelty 7.0 of 10

    MV-Forcing composes temporal and view-sequential autoregression in a single diffusion model, using a recurrent 3D reconstruction model as a geometric bridge to generate arbitrarily long, multi-view consistent videos.

  2. 4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time

    cs.CV 2025-06 conditional novelty 7.0 of 10

    4D-LRM is a transformer that maps sparse posed frames scattered across time to a cloud of 4D Gaussians and renders any query view at any query time in under 1.5 seconds.

  3. Wonderland: Navigating 3D Scenes from a Single Image

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A feed-forward pipeline reconstructs 3D Gaussian scenes from single images by regressing 3DGS directly from camera-conditioned video diffusion latents.

  4. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  5. ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Distributed video-generation agents can create a spatiotemporally consistent shared world by tiling four views, exchanging cross-agent attention, and querying a spatial memory cache.

  6. WildSmoke: Ready-to-Use Dynamic 3D Smoke Assets from a Single Video in the Wild

    cs.CV 2025-09 conditional novelty 6.0 of 10

    WildSmoke reconstructs editable dynamic 3D smoke assets from a single in-the-wild video and reports a +2.22 dB average PSNR gain over prior fluid reconstruction methods on four real-world videos.

  7. Learning an Implicit Physics Model for Image-based Fluid Simulation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Using a simplified physics loss and 3D Gaussians, a neural network animates a single fluid image into videos with novel views, beating earlier methods on quality and motion accuracy.

  8. CharacterShot: Controllable and Consistent 4D Character Animation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A new pipeline generates pose-controlled, view-consistent 4D character animations from one reference image and a 2D pose sequence, backed by a new 13,115-character dataset and benchmark.

  9. 4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A two-stage cascaded video diffusion model generates 16-view consistent videos from a monocular video, enabling higher-quality 4D content reconstruction.

  10. Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A video-to-4D model that encodes mesh animations into compact Gaussian variation latents and diffuses them conditioned on the video and a canonical Gaussian splat.

  11. Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A sliding iterative denoising scheme that alternates spatial and temporal passes, combined with skeleton conditioning, lets a diffusion model create spatio-temporally consistent multi-view human videos from sparse inp...

  12. Voyaging into Perpetual Dynamic Scenes from a Single View

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single-view dynamic scene can be extended into an unbounded fly-through video by iteratively outpainting partial views of a learned 4D point cloud with ray distance guidance.

  13. WonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene Exploration

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A pipeline that restores corrupted novel-view videos with a video diffusion model and jointly denoises multiple viewpoints to improve 3D scene exploration from a single image.

  14. CoCo4D: Comprehensive and Complex 4D Scene Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CoCo4D generates multi-view consistent 4D scenes from text or image prompts in about one hour by generating a reference video, reconstructing the foreground and background separately, and composing them with a learned...

  15. Restereo: Diffusion stereo video generation and restoration

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A diffusion model fine-tuned on synthetically degraded stereo videos simultaneously generates a consistent stereo pair and restores low-resolution or compressed input, outperforming prior stereo generators on low-qual...

  16. TesserAct: Learning 4D Embodied World Models

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A 4D embodied world model that generates RGB-depth-normal videos from an image and instruction, reconstructs the scene as point clouds, and uses those point clouds to train better robot manipulation policies.

  17. Not All Frame Features Are Equal: Video-to-4D Generation via Decoupling Dynamic-Static Features

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A video-to-4D generation method that decouples dynamic and static features in DINOv2 space and fuses similar dynamic information across views reports state-of-the-art scores on Consistent4D and Objaverse.

  18. Articulate That Object Part (ATOP): 3D Part Articulation via Text and Motion Personalization

    cs.CV 2025-02 conditional novelty 6.0 of 10

    ATOP personalizes a pre-trained multi-view diffusion model with a few reference videos to generate part motion from text and masks, then lifts that motion to a 3D articulation axis via score distillation.

  19. GAS: Generative Avatar Synthesis from a Single Image

    cs.CV 2025-02 conditional novelty 6.0 of 10

    GAS generates view-consistent, temporally coherent avatars from a single image by feeding NeRF renderings of the target view plus SMPL normal maps into a video diffusion model.

  20. AR4D: Autoregressive 4D Generation from Monocular Videos

    cs.CV 2025-01 conditional novelty 6.0 of 10

    AR4D generates 4D content from monocular video by autoregressively deforming frame-wise 3D Gaussians, with progressive pseudo-view supervision from a pre-trained reconstruction model.

  21. DreamDrive: Generative 4D Scene Modeling from Street View Images

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DreamDrive generates 3D-consistent driving videos from a single image by lifting diffusion-generated reference frames into a hybrid static and dynamic 4D Gaussian scene.

  22. Grid: Omni Visual Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GRID shows that fine-tuning an image diffusion model on videos arranged as grid images can generate coherent video and multi-view sequences with far less data and compute than specialized video models.

  23. SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A cross-view attention module and hybrid training recipe let a frozen text-to-video model generate synchronized multi-camera videos from arbitrary viewpoints given only a text prompt and camera poses.

  24. 4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion

    cs.CV 2024-12 conditional novelty 6.0 of 10

    4Real-Video generates consistent 4D video grids, frames across time and viewpoint, with a parallel two-stream diffusion transformer that synchronizes temporal and viewpoint token streams, achieving faster and higher-q...

  25. Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    By recasting text-to-image generation as multi-frame video generation and adding a differential camera encoder, the method achieves camera intrinsic control with scene consistency, outperforming current text-to-image ...

  26. Gaussians-to-Life: Text-Driven Animation of 3D Gaussian Splatting Scenes

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A text-driven pipeline that lifts 2D video diffusion motion into 3D Gaussian Splatting scenes via point tracking and depth estimation.

  27. AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers

    cs.CV 2024-11 conditional novelty 6.0 of 10

    AC3D improves camera control in video diffusion transformers by conditioning only early denoising steps and the first 8 of 32 blocks, and by adding 20K static-camera dynamic videos to training.

  28. Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A diffusion model generates dense 4D point trajectories from a single image, and a separate view-synthesis module renders them into novel-view videos.

  29. DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective

    cs.CV 2025-08 reject novelty 5.0 of 10

    A monocular human avatar reconstruction method generates pseudo back-view videos with a fine-tuned diffusion model and uses them as extra training data for a 3D Gaussian avatar.

  30. MVG4D: Image Matrix-Based Multi-View and Motion Generation for 4D Content Creation from a Single Image

    cs.CV 2025-07 reject novelty 5.0 of 10

    MVG4D generates a time-ordered multi-view image matrix from a single image and uses it to optimize a deformable 4D Gaussian Splatting scene, reporting higher quality and lower runtime than baselines.

  31. Drive Any Mesh: 4D Latent Diffusion for Mesh Deformation from Video

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A video-conditioned latent diffusion model generates mesh vertex trajectories that deform an input 3D asset into render-ready 4D animations.

  32. DSplats: 3D Generation by Denoising Splats-Based Multiview Diffusion Models

    eess.IV 2024-12 conditional novelty 5.0 of 10

    DSplats is a single-stage multiview diffusion model that denoises latents through a 3D Gaussian reconstructor, achieving state-of-the-art single-image-to-3D reconstruction on the Google Scanned Objects benchmark.

  33. PaintScene4D: Consistent 4D Scene Generation from Text Prompts

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A training-free pipeline that turns one text-to-video clip into a multi-view 4D scene renderable along user-chosen camera paths.

  34. Reconstructing 4D Spatial Intelligence: A Survey

    cs.CV 2025-07 accept novelty 4.0 of 10

    A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.

  35. Dynamic View Synthesis as an Inverse Problem

    cs.CV 2025-06 reject novelty 3.0 of 10

    Dynamic view synthesis from a monocular video is achieved by redesigning the noise initialization of a pretrained video diffusion model using a recursive interpolation and a stochastic latent modulation.

Pith tools