Pith. sign in

REVIEW 23 cited by

RealmDreamer: Text-Driven 3D Scene Generation with Inpainting and Depth Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.07199 v2 pith:36CUHDLG submitted 2024-04-10 cs.CV cs.AIcs.GRcs.LG

classification cs.CVcs.AIcs.GRcs.LG
keywords diffusioninpaintingmethodrealmdreamercomplexconditioneddepthdistillation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce RealmDreamer, a technique for generating forward-facing 3D scenes from text descriptions. Our method optimizes a 3D Gaussian Splatting representation to match complex text prompts using pretrained diffusion models. Our key insight is to leverage 2D inpainting diffusion models conditioned on an initial scene estimate to provide low variance supervision for unknown regions during 3D distillation. In conjunction, we imbue high-fidelity geometry with geometric distillation from a depth diffusion model, conditioned on samples from the inpainting model. We find that the initialization of the optimization is crucial, and provide a principled methodology for doing so. Notably, our technique doesn't require video or multi-view data and can synthesize various high-quality 3D scenes in different styles with complex layouts. Further, the generality of our method allows 3D synthesis from a single image. As measured by a comprehensive user study, our method outperforms all existing approaches, preferred by 88-95%. Project Page: https://realmdreamer.github.io/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding

    cs.CV 2025-08 conditional novelty 7.0 of 10

    LSD-3D generates explicit, 3D-consistent driving scenes by combining a generated proxy mesh with geometry-grounded distillation from a 2D diffusion model.

  2. Towards Affordance-Aware Articulation Synthesis for Rigged Objects

    cs.CV 2025-01 conditional novelty 7.0 of 10

    A training-free optimization framework synthesizes affordance-aware articulation for arbitrary rigged 3D objects by aligning them with 2D diffusion-inpainted references.

  3. Wonderland: Navigating 3D Scenes from a Single Image

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A feed-forward pipeline reconstructs 3D Gaussian scenes from single images by regressing 3DGS directly from camera-conditioned video diffusion latents.

  4. From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A diffusion model trained on 1 million 360-degree videos synthesizes novel views with camera translation and enables 3D reconstruction from a single image.

  5. SimVS: Simulating World Inconsistencies for Robust View Synthesis

    cs.CV 2024-12 conditional novelty 7.0 of 10

    Video diffusion models simulate world inconsistencies, and a harmonization network trained on the simulated data reconciles sparse inconsistent multi-view images into consistent 3D scenes.

  6. SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control

    cs.GR 2026-07 conditional novelty 6.5 of 10

    Automatic view scheduling via a directed generation graph plus object-level identity and adherence conditioning enables high-quality outdoor 3DGS scenes from arbitrary input geometry without user camera paths.

  7. UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    UniWorld-View couples an occlusion-aware point cloud renderer with a dual-stream video diffusion model to synthesize large-baseline novel views from monocular video.

  8. SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SynCity 3000 generates large, coherent 3D scenes from text by fine-tuning an image-to-3D diffusion model to operate convolutionally on overlapping windows, trained on procedurally generated synthetic scene data.

  9. CustomX: Unified Character, Action, and Scene Customization in Video World Models

    cs.CV 2025-12 conditional novelty 6.0 of 10

    AniX generates controllable videos of a user-supplied character performing typed actions inside a user-supplied 3D scene by fine-tuning a pre-trained video generator on small locomotion datasets.

  10. EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    EarthCrafter generates 600-meter-scale 3D Earth scenes using separate latent diffusion models for structure and texture, conditioned on semantics, images, or nothing.

  11. DreamDance: Animating Character Art via Inpainting Stable Gaussian Worlds

    cs.CV 2025-05 conditional novelty 6.0 of 10

    DreamDance animates a single character artwork by reconstructing its background as a 3D Gaussian scene and then inpainting the animated character into the rendered video.

  12. SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations

    cs.CV 2025-05 conditional novelty 6.0 of 10

    SpatialCrafter generates camera-controlled video from sparse input views and reconstructs a 3D Gaussian scene from the generated video latents, improving novel view synthesis.

  13. DreamDrive: Generative 4D Scene Modeling from Street View Images

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DreamDrive generates 3D-consistent driving videos from a single image by lifting diffusion-generated reference frames into a hybrid static and dynamic 4D Gaussian scene.

  14. You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale

    cs.CV 2024-12 reject novelty 6.0 of 10

    See3D proposes a pose-free visual condition for multi-view diffusion trained on web videos, claiming SOTA single- and sparse-view 3D generation, but the evaluation protocol leaks ground-truth information and mixes ben...

  15. FiffDepth: Feed-forward Transformation of Diffusion-Based Generators for Detailed Depth Estimation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FiffDepth transforms a pre-trained diffusion image generator into a feed-forward monocular depth estimator that combines generative detail with DINOv2-based robustness.

  16. AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers

    cs.CV 2024-11 conditional novelty 6.0 of 10

    AC3D improves camera control in video diffusion transformers by conditioning only early denoising steps and the first 8 of 32 blocks, and by adding 20K static-camera dynamic videos to training.

  17. SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesis

    cs.CV 2024-11 conditional novelty 6.0 of 10

    SplatFlow jointly generates multi-view images, depths, and camera poses with a rectified flow model, then decodes them into editable 3D Gaussian Splatting scenes.

  18. MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A multi-view diffusion model trained on 1.6 million scenes uses warped depth-based 3D priors and a key-rescaling trick to synthesize up to 158 views in one forward pass.

  19. WorldClaw: Agentic 3D Open-World Generation at Scale

    cs.AI 2026-08 conditional novelty 4.0 of 10

    WorldClaw generates globally coherent, locally detailed, editable 3D worlds from open-ended text using a coarse-to-fine agentic pipeline.

  20. AI-powered Contextual 3D Environment Generation: A Systematic Review

    cs.GR 2025-06 conditional novelty 4.0 of 10

    A PRISMA-based systematic review of 136 papers finds diffusion models dominate high-quality AI 3D scene generation, with computational cost, data quality, and evaluation metrics as key limitations.

  21. A Critical Synthesis of Uncertainty Quantification and Foundation Models in Monocular Depth Estimation

    cs.CV 2025-01 conditional novelty 4.0 of 10

    Fine-tuning DepthAnythingV2 with a Gaussian negative log-likelihood loss yields the most reliable pixel-wise uncertainty estimates on indoor, street, and object scenes, but it fails on aerial large-depth data.

  22. Advancing Extended Reality with 3D Gaussian Splatting: Innovations and Prospects

    cs.CV 2024-12 conditional novelty 4.0 of 10

    3D Gaussian Splatting research relevant to Extended Reality is organized into a five-part taxonomy with suggested future directions.

  23. Survey on Monocular Metric Depth Estimation

    cs.CV 2025-01 unverdicted novelty 1.0 of 10

    A survey of monocular metric depth estimation methods, datasets, and open challenges, with no new experimental results.

Pith tools