Pith. sign in

REVIEW 24 cited by

Text-To-4D Dynamic Scene Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.11280 v1 pith:WSNOLUYY submitted 2023-01-26 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords dynamictextapproachmav3dmethodmodelscenescenes
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present MAV3D (Make-A-Video3D), a method for generating three-dimensional dynamic scenes from text descriptions. Our approach uses a 4D dynamic Neural Radiance Field (NeRF), which is optimized for scene appearance, density, and motion consistency by querying a Text-to-Video (T2V) diffusion-based model. The dynamic video output generated from the provided text can be viewed from any camera location and angle, and can be composited into any 3D environment. MAV3D does not require any 3D or 4D data and the T2V model is trained only on Text-Image pairs and unlabeled videos. We demonstrate the effectiveness of our approach using comprehensive quantitative and qualitative experiments and show an improvement over previously established internal baselines. To the best of our knowledge, our method is the first to generate 3D dynamic scenes given a text description.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 24 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A feed-forward VAE plus rectified-flow model animates arbitrary static meshes from text prompts in seconds, with a new 4M-sequence training dataset.

  2. STP4D: Spatio-Temporal-Prompt Consistent Modeling for Text-to-4D Gaussian Splatting

    cs.CV 2025-04 conditional novelty 7.0 of 10

    STP4D directly denoises 4D Gaussian splatting attributes with DDIM conditioned on time-varying text embeddings, producing high-fidelity 4D assets in about 4.6 seconds.

  3. ID-V2V: Identity-Preserving Video Restylization

    cs.CV 2026-07 conditional novelty 6.0 of 10

    ID-V2V restyles video by conditioning a diffusion model on edited keyframes, depth, relit faces, and face normals, so scene edits propagate while facial identity and performance are preserved.

  4. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  5. LivingWorld: Interactive 4D World Generation with Environmental Dynamics

    cs.CV 2026-04 conditional novelty 6.0 of 10

    An interactive pipeline generates expanding 4D worlds with globally coherent environmental dynamics from a single image in roughly 12 seconds per expansion step.

  6. SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-body Manipulation

    cs.RO 2026-02 conditional novelty 6.0 of 10

    SoMA couples robot joint actions, environmental forces, and learned Gaussian-splat dynamics into a single neural simulator, improving resimulation and generalization on real robot soft-body manipulation by about 20% o...

  7. CharacterShot: Controllable and Consistent 4D Character Animation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A new pipeline generates pose-controlled, view-consistent 4D character animations from one reference image and a 2D pose sequence, backed by a new 13,115-character dataset and benchmark.

  8. Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Video-rewinding joint training preserves geometry while re-animating a single-video scene with novel motion from a text prompt and an image-to-video model.

  9. Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A video-to-4D model that encodes mesh animations into compact Gaussian variation latents and diffuses them conditioned on the video and a canonical Gaussian splat.

  10. Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A sliding iterative denoising scheme that alternates spatial and temporal passes, combined with skeleton conditioning, lets a diffusion model create spatio-temporally consistent multi-view human videos from sparse inp...

  11. Advancing Text-to-3D Generation with Linearized Lookahead Variational Score Distillation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Updating the LoRA score model one step ahead of the 3D model and keeping only the first-order correction term yields L2-VSD, a stable and higher-quality variant of VSD for text-to-3D generation.

  12. CoCo4D: Comprehensive and Complex 4D Scene Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CoCo4D generates multi-view consistent 4D scenes from text or image prompts in about one hour by generating a reference video, reconstructing the foreground and background separately, and composing them with a learned...

  13. TesserAct: Learning 4D Embodied World Models

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A 4D embodied world model that generates RGB-depth-normal videos from an image and instruction, reconstructs the scene as point clouds, and uses those point clouds to train better robot manipulation policies.

  14. Not All Frame Features Are Equal: Video-to-4D Generation via Decoupling Dynamic-Static Features

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A video-to-4D generation method that decouples dynamic and static features in DINOv2 space and fuses similar dynamic information across views reports state-of-the-art scores on Consistent4D and Objaverse.

  15. DreamDrive: Generative 4D Scene Modeling from Street View Images

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DreamDrive generates 3D-consistent driving videos from a single image by lifting diffusion-generated reference frames into a hybrid static and dynamic 4D Gaussian scene.

  16. DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DrivingRecon predicts 4D Gaussians of street scenes from surround-view video in one forward pass, using a novel Prune and Dilate Block to reduce redundant overlapping points.

  17. 4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion

    cs.CV 2024-12 conditional novelty 6.0 of 10

    4Real-Video generates consistent 4D video grids, frames across time and viewpoint, with a parallel two-stream diffusion transformer that synchronizes temporal and viewpoint token streams, achieving faster and higher-q...

  18. CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    CAT4D uses a multi-view video diffusion model to convert monocular video into multi-view video and reconstruct a dynamic 3D Gaussian scene, with competitive results on 4D reconstruction benchmarks.

  19. Pathways on the Image Manifold: Image Editing via Video Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Frame2Frame performs text-based image editing by generating a short video transition from the source image and selecting the best resulting frame, achieving competitive or better benchmark scores than single-image dif...

  20. RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation

    cs.RO 2025-10 unverdicted novelty 5.0 of 10

    Abstract describes RoDyn but full text describes iMoWM; the record is internally inconsistent and the headline claims are absent from the body.

  21. TextMesh4D: Zero-shot Text-to-4D Mesh Generation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    TextMesh4D generates text-conditioned dynamic meshes by combining a Jacobian Deformation Field, video score distillation, and a local-global semantic regularizer in a zero-shot pipeline.

  22. Generative 4D Scene Gaussian Splatting with Object View-Synthesis Priors

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A test-time optimization method that jointly fits deformable per-object 3D Gaussians with object-centric diffusion priors to generate 4D scenes and point tracks from monocular multi-object videos.

  23. Drive Any Mesh: 4D Latent Diffusion for Mesh Deformation from Video

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A video-conditioned latent diffusion model generates mesh vertex trajectories that deform an input 3D asset into render-ready 4D animations.

  24. From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

    cs.RO 2026-07 conditional novelty 4.0 of 10

    Physical intelligence needs an embodied brain that reasons over interventions and emits capability requests, grounded by a physical harness and shared experience contracts rather than direct actuator policies.

Pith tools