REVIEW 23 cited by
RealmDreamer: Text-Driven 3D Scene Generation with Inpainting and Depth Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce RealmDreamer, a technique for generating forward-facing 3D scenes from text descriptions. Our method optimizes a 3D Gaussian Splatting representation to match complex text prompts using pretrained diffusion models. Our key insight is to leverage 2D inpainting diffusion models conditioned on an initial scene estimate to provide low variance supervision for unknown regions during 3D distillation. In conjunction, we imbue high-fidelity geometry with geometric distillation from a depth diffusion model, conditioned on samples from the inpainting model. We find that the initialization of the optimization is crucial, and provide a principled methodology for doing so. Notably, our technique doesn't require video or multi-view data and can synthesize various high-quality 3D scenes in different styles with complex layouts. Further, the generality of our method allows 3D synthesis from a single image. As measured by a comprehensive user study, our method outperforms all existing approaches, preferred by 88-95%. Project Page: https://realmdreamer.github.io/
Forward citations
Cited by 23 Pith papers
-
LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding
LSD-3D generates explicit, 3D-consistent driving scenes by combining a generated proxy mesh with geometry-grounded distillation from a 2D diffusion model.
-
Towards Affordance-Aware Articulation Synthesis for Rigged Objects
A training-free optimization framework synthesizes affordance-aware articulation for arbitrary rigged 3D objects by aligning them with 2D diffusion-inpainted references.
-
Wonderland: Navigating 3D Scenes from a Single Image
A feed-forward pipeline reconstructs 3D Gaussian scenes from single images by regressing 3DGS directly from camera-conditioned video diffusion latents.
-
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
A diffusion model trained on 1 million 360-degree videos synthesizes novel views with camera translation and enables 3D reconstruction from a single image.
-
SimVS: Simulating World Inconsistencies for Robust View Synthesis
Video diffusion models simulate world inconsistencies, and a harmonization network trained on the simulated data reconciles sparse inconsistent multi-view images into consistent 3D scenes.
-
SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control
Automatic view scheduling via a directed generation graph plus object-level identity and adherence conditioning enables high-quality outdoor 3DGS scenes from arbitrary input geometry without user camera paths.
-
UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
UniWorld-View couples an occlusion-aware point cloud renderer with a dual-stream video diffusion model to synthesize large-baseline novel views from monocular video.
-
SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
SynCity 3000 generates large, coherent 3D scenes from text by fine-tuning an image-to-3D diffusion model to operate convolutionally on overlapping windows, trained on procedurally generated synthetic scene data.
-
CustomX: Unified Character, Action, and Scene Customization in Video World Models
AniX generates controllable videos of a user-supplied character performing typed actions inside a user-supplied 3D scene by fine-tuning a pre-trained video generator on small locomotion datasets.
-
EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion
EarthCrafter generates 600-meter-scale 3D Earth scenes using separate latent diffusion models for structure and texture, conditioned on semantics, images, or nothing.
-
DreamDance: Animating Character Art via Inpainting Stable Gaussian Worlds
DreamDance animates a single character artwork by reconstructing its background as a 3D Gaussian scene and then inpainting the animated character into the rendered video.
-
SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations
SpatialCrafter generates camera-controlled video from sparse input views and reconstructs a 3D Gaussian scene from the generated video latents, improving novel view synthesis.
-
DreamDrive: Generative 4D Scene Modeling from Street View Images
DreamDrive generates 3D-consistent driving videos from a single image by lifting diffusion-generated reference frames into a hybrid static and dynamic 4D Gaussian scene.
-
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
See3D proposes a pose-free visual condition for multi-view diffusion trained on web videos, claiming SOTA single- and sparse-view 3D generation, but the evaluation protocol leaks ground-truth information and mixes ben...
-
FiffDepth: Feed-forward Transformation of Diffusion-Based Generators for Detailed Depth Estimation
FiffDepth transforms a pre-trained diffusion image generator into a feed-forward monocular depth estimator that combines generative detail with DINOv2-based robustness.
-
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
AC3D improves camera control in video diffusion transformers by conditioning only early denoising steps and the first 8 of 32 blocks, and by adding 20K static-camera dynamic videos to training.
-
SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesis
SplatFlow jointly generates multi-view images, depths, and camera poses with a rectified flow model, then decodes them into editable 3D Gaussian Splatting scenes.
-
MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model
A multi-view diffusion model trained on 1.6 million scenes uses warped depth-based 3D priors and a key-rescaling trick to synthesize up to 158 views in one forward pass.
-
WorldClaw: Agentic 3D Open-World Generation at Scale
WorldClaw generates globally coherent, locally detailed, editable 3D worlds from open-ended text using a coarse-to-fine agentic pipeline.
-
AI-powered Contextual 3D Environment Generation: A Systematic Review
A PRISMA-based systematic review of 136 papers finds diffusion models dominate high-quality AI 3D scene generation, with computational cost, data quality, and evaluation metrics as key limitations.
-
A Critical Synthesis of Uncertainty Quantification and Foundation Models in Monocular Depth Estimation
Fine-tuning DepthAnythingV2 with a Gaussian negative log-likelihood loss yields the most reliable pixel-wise uncertainty estimates on indoor, street, and object scenes, but it fails on aerial large-depth data.
-
Advancing Extended Reality with 3D Gaussian Splatting: Innovations and Prospects
3D Gaussian Splatting research relevant to Extended Reality is organized into a five-part taxonomy with suggested future directions.
-
Survey on Monocular Metric Depth Estimation
A survey of monocular metric depth estimation methods, datasets, and open challenges, with no new experimental results.
Discussion (0). Continue with ORCID to comment.