REVIEW 9 cited by
Diffusion Probabilistic Models for Scene-Scale 3D Categorical Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this paper, we learn a diffusion model to generate 3D data on a scene-scale. Specifically, our model crafts a 3D scene consisting of multiple objects, while recent diffusion research has focused on a single object. To realize our goal, we represent a scene with discrete class labels, i.e., categorical distribution, to assign multiple objects into semantic categories. Thus, we extend discrete diffusion models to learn scene-scale categorical distributions. In addition, we validate that a latent diffusion model can reduce computation costs for training and deploying. To the best of our knowledge, our work is the first to apply discrete and latent diffusion for 3D categorical data on a scene-scale. We further propose to perform semantic scene completion (SSC) by learning a conditional distribution using our diffusion model, where the condition is a partial observation in a sparse point cloud. In experiments, we empirically show that our diffusion models not only generate reasonable scenes, but also perform the scene completion task better than a discriminative model. Our code and models are available at https://github.com/zoomin-lee/scene-scale-diffusion
Forward citations
Cited by 9 Pith papers
-
OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation
A joint diffusion framework trains a Stable Diffusion generator and a semantic occupancy perception model together, so each task improves the other, producing text-conditional RGB-occupancy pairs.
-
SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
SynCity 3000 generates large, coherent 3D scenes from text by fine-tuning an image-to-3D diffusion model to operate convolutionally on overlapping windows, trained on procedurally generated synthetic scene data.
-
PVNet: Point-Voxel Interaction LiDAR Scene Upsampling Via Diffusion Models
PVNet uses diffusion models with point-voxel interaction to upsample LiDAR scenes at arbitrary rates without dense supervision.
-
Distilling Diffusion Models to Efficient 3D LiDAR Scene Completion
ScoreLiDAR distills a 50-step LiDAR scene completion diffusion model into an 8-step student that runs about 5x faster and matches or improves completion quality on SemanticKITTI and KITTI-360.
-
3D and 4D World Modeling: A Survey
A survey that defines 3D/4D world modeling, organizes methods into VideoGen, OccGen, and LiDARGen categories, and compiles datasets, metrics, and benchmark numbers.
-
Diffusion-Based Generative Models for 3D Occupancy Prediction in Autonomous Driving
Diffusion-based generative models, using discrete categorical diffusion conditioned on BEV features, improve 3D occupancy prediction and downstream planning for autonomous driving.
-
Map Imagination Like Blind Humans: Group Diffusion Model for Robotic Map Generation
A group diffusion model can generate LiDAR-style 3D maps from path-only odometry data, and adding 50 LiDAR points improves the maps.
-
SSEditor: Controllable Mask-to-Scene Generation with Diffusion Model
SSEditor generates controllable 3D semantic urban scenes from mask conditions using a triplane autoencoder and a mask-conditional diffusion model, avoiding multi-step resampling.
-
3D Scene Generation: A Survey
The paper surveys 3D scene generation and organizes methods into four paradigms, with datasets, evaluation metrics, applications, and future directions.
Discussion (0). Continue with ORCID to comment.