Pith. sign in

REVIEW 43 cited by

DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.04928 v1 pith:TWG5MJUZ submitted 2024-11-07 cs.CV cs.AIcs.GR

classification cs.CVcs.AIcs.GR
keywords videodiffusiongenerationspatialtemporalscenescontrollabledimensionx
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we introduce \textbf{DimensionX}, a framework designed to generate photorealistic 3D and 4D scenes from just a single image with video diffusion. Our approach begins with the insight that both the spatial structure of a 3D scene and the temporal evolution of a 4D scene can be effectively represented through sequences of video frames. While recent video diffusion models have shown remarkable success in producing vivid visuals, they face limitations in directly recovering 3D/4D scenes due to limited spatial and temporal controllability during generation. To overcome this, we propose ST-Director, which decouples spatial and temporal factors in video diffusion by learning dimension-aware LoRAs from dimension-variant data. This controllable video diffusion approach enables precise manipulation of spatial structure and temporal dynamics, allowing us to reconstruct both 3D and 4D representations from sequential frames with the combination of spatial and temporal dimensions. Additionally, to bridge the gap between generated videos and real-world scenes, we introduce a trajectory-aware mechanism for 3D generation and an identity-preserving denoising strategy for 4D generation. Extensive experiments on various real-world and synthetic datasets demonstrate that DimensionX achieves superior results in controllable video generation, as well as in 3D and 4D scene generation, compared with previous methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 43 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    UniGeo unifies geometric guidance across three levels in video models to reduce geometric drift and improve consistency in camera-controllable image editing.

  2. SOPHY: Learning to Generate Simulation-Ready Objects with Physical Materials

    cs.GR 2025-04 conditional novelty 7.0 of 10

    A diffusion-based generative model jointly predicts shape, texture, and physics material parameters for 3D objects, using a new VLM-and-expert annotated dataset of 3,004 objects.

  3. Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation

    cs.CV 2025-01 conditional novelty 7.0 of 10

    A diffusion-based image animation method that uses rendered colored spheres and a checkerboard world envelope as 3D-aware motion control signals, enabling joint camera and object motion control with fine granularity.

  4. Wonderland: Navigating 3D Scenes from a Single Image

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A feed-forward pipeline reconstructs 3D Gaussian scenes from single images by regressing 3DGS directly from camera-conditioned video diffusion latents.

  5. StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization

    cs.CV 2026-08 conditional novelty 6.0 of 10

    StateFlow constructs, evolves, and accesses a persistent 3D world state for previsualization, reporting higher consistency and controllability than one-shot video generation on VBench and user studies.

  6. UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    UniWorld-View couples an occlusion-aware point cloud renderer with a dual-stream video diffusion model to synthesize large-baseline novel views from monocular video.

  7. 4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A diffusion model trained on 60,000 fitted 4D Gaussian Splatting human clips generates text-prompted, view-consistent dynamic humans directly in 4D, over 10x faster than video-first pipelines.

  8. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  9. PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A single pixel-space diffusion model jointly performs 3D scene reconstruction and generation by supervising flow matching on rendered multi-view images, matching SOTA reconstruction and outperforming latent-space generation.

  10. Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Diffusing in a unified 3D representation (geometry plus appearance and semantics) produces more cross-view-consistent 3D Gaussian scenes than 2D latent pipelines.

  11. VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation

    cs.CV 2026-01 unverdicted novelty 6.0 of 10

    VideoGPA distills geometry priors via self-supervised DPO to enhance 3D consistency, temporal stability, and motion coherence in video diffusion models.

  12. CustomX: Unified Character, Action, and Scene Customization in Video World Models

    cs.CV 2025-12 conditional novelty 6.0 of 10

    AniX generates controllable videos of a user-supplied character performing typed actions inside a user-supplied 3D scene by fine-tuning a pre-trained video generator on small locomotion datasets.

  13. WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling

    cs.CV 2025-12 unverdicted novelty 6.0 of 10

    WorldPlay uses dual action representation, reconstituted context memory, and context forcing distillation to produce consistent 720p streaming video at 24 FPS for interactive world modeling.

  14. InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A training-free, near-zero-overhead inverse solver for novel-view video generation and inpainting that projects masks into continuous multi-channel latent masks and applies DDS with conjugate gradient in latent space.

  15. GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation

    cs.CV 2025-11 conditional novelty 6.0 of 10

    GeoWorld improves image-to-3D scene generation by conditioning a video-diffusion model on full-frame geometry features extracted by a multi-view geometry model, yielding higher PSNR/SSIM/LPIPS than prior methods.

  16. CausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesis

    cs.CV 2025-09 conditional novelty 6.0 of 10

    An autoregressive multi-view diffusion model that generates novel views sequentially with flexible input-output configurations, using causal masking, relative pose encoding, and KV caching.

  17. 4DNeX: Feed-Forward 4D Generative Modeling Made Easy

    cs.CV 2025-08 conditional novelty 6.0 of 10

    4DNeX generates dynamic 3D point clouds and matching RGB video from a single image by fine-tuning a pretrained video diffusion model on a large pseudo-annotated 4D dataset.

  18. CharacterShot: Controllable and Consistent 4D Character Animation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A new pipeline generates pose-controlled, view-consistent 4D character animations from one reference image and a 2D pose sequence, backed by a new 13,115-character dataset and benchmark.

  19. 4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A two-stage cascaded video diffusion model generates 16-view consistent videos from a monocular video, enabling higher-quality 4D content reconstruction.

  20. Voyaging into Perpetual Dynamic Scenes from a Single View

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single-view dynamic scene can be extended into an unbounded fly-through video by iteratively outpainting partial views of a learned 4D point cloud with ray distance guidance.

  21. From Virtual Games to Real-World Play

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A chunk-wise video diffusion model trained on labeled game data plus unlabeled real footage transfers game-style control commands to real-world entities.

  22. Emergent Temporal Correspondences from Video Diffusion Transformers

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Video diffusion transformers encode temporal correspondences primarily in query-key similarities of a few specific attention layers, which can be extracted for zero-shot point tracking and used for training-free motio...

  23. EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance

    cs.CV 2025-05 conditional novelty 6.0 of 10

    EPiC trains a 30M-parameter visibility-aware ControlNet on mask-based anchor videos from 5,000 in-the-wild videos and 500 steps, reaching SOTA camera accuracy on RealEstate10K and MiraData.

  24. SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations

    cs.CV 2025-05 conditional novelty 6.0 of 10

    SpatialCrafter generates camera-controlled video from sparse input views and reconstructs a 3D Gaussian scene from the generated video latents, improving novel view synthesis.

  25. VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    VidCRAFT3 is a single image-to-video diffusion system that accepts camera, object, and lighting direction controls separately or jointly, trained in three stages with a new synthetic lighting dataset.

  26. Grid: Omni Visual Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GRID shows that fine-tuning an image diffusion model on videos arranged as grid images can generate coherent video and multi-view sequences with far less data and compute than specialized video models.

  27. You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale

    cs.CV 2024-12 reject novelty 6.0 of 10

    See3D proposes a pose-free visual condition for multi-view diffusion trained on web videos, claiming SOTA single- and sparse-view 3D generation, but the evaluation protocol leaks ground-truth information and mixes ben...

  28. CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    CAT4D uses a multi-view video diffusion model to convert monocular video into multi-view video and reconstruct a dynamic 3D Gaussian scene, with competitive results on 4D reconstruction benchmarks.

  29. AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers

    cs.CV 2024-11 conditional novelty 6.0 of 10

    AC3D improves camera control in video diffusion transformers by conditioning only early denoising steps and the first 8 of 32 blocks, and by adding 20K static-camera dynamic videos to training.

  30. High-Quality Exposure Correction with Diffusion-Based Image Generation Priors

    cs.CV 2026-08 conditional novelty 5.0 of 10

    DPEC fine-tunes a pretrained diffusion model for single-step exposure correction and fuses its low-frequency output into a regression network, improving perceptual metrics on LCDP, MSEC, and SICE.

  31. Quo Vadis, World Modeling?

    cs.CV 2026-08 conditional novelty 5.0 of 10

    An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.

  32. Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A diffusion model generates dense 4D point trajectories from a single image, and a separate view-synthesis module renders them into novel-view videos.

  33. PanoLora: Bridging Perspective and Panoramic Video Generation with LoRA Adaptation

    cs.CV 2025-09 reject novelty 5.0 of 10

    Fine-tuning a pretrained video diffusion model with LoRA rank 16 on about 1,000 synthetic videos produces panoramic video with good seam closure, but the claim that rank must exceed 8 degrees of freedom is not proven.

  34. Non-invasive Assessment of Pancreatic Duct Hypertension Using Computational Flow Modeling

    physics.med-ph 2025-08 unverdicted novelty 5.0 of 10

    A computational model estimates pancreatic duct pressure non-invasively from MRCP geometry, with reported agreement against ERCP pressure measurements.

  35. DIP-GS: Deep Image Prior For Gaussian Splatting Sparse View Recovery

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    DIP-GS applies a deep image prior in a coarse-to-fine manner to enable 3D Gaussian Splatting for sparse-view reconstruction.

  36. Impact-driven Context Filtering For Cross-file Code Completion

    cs.SE 2025-08 unverdicted novelty 5.0 of 10

    The manuscript's abstract claims a new code-completion filtering method, yet the body contains an unrelated 3D animation paper, leaving the claimed work unverifiable.

  37. HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A staged pipeline generates layered, mesh-based 3D worlds from text or images by combining panoramic diffusion, semantic layer decomposition, and video-based expansion.

  38. LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion

    cs.CV 2025-07 conditional novelty 5.0 of 10

    LiON-LoRA adds a learned scaling token to video-diffusion LoRA adapters, enabling linear and independent control of camera trajectory and object motion strength.

  39. LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion

    cs.CV 2025-07 conditional novelty 5.0 of 10

    LangScene-X generates RGB, normal, and semantic videos from sparse views to reconstruct 3D language-embedded Gaussian fields that support open-ended text queries.

  40. Follow-Your-Creation: Empowering 4D Creation through Video Inpainting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Follow-Your-Creation fine-tunes the Wan2.1 video inpainting model on composite point-cloud and editing masks so a single monocular video can be converted into editable 4D video with new camera motion.

  41. TwoSquared: 4D Generation from 2D Image Pairs

    cs.CV 2025-04 conditional novelty 5.0 of 10

    A method that generates a temporally consistent, textured 4D mesh sequence from only two RGB images showing an object's initial and final poses.

  42. Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A feed-forward system that generates object-level and scene-level 3D Gaussian scenes from text in about eight seconds by diffusing multi-view RGB-D latent codes and decoding them into pixel-aligned 3D Gaussians.

  43. 3D Scene Generation: A Survey

    cs.CV 2025-05 conditional

    The paper surveys 3D scene generation and organizes methods into four paradigms, with datasets, evaluation metrics, applications, and future directions.

Pith tools