Pith. sign in

REVIEW 9 cited by

Diffusion Probabilistic Models for Scene-Scale 3D Categorical Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.00527 v1 pith:2TE373IT submitted 2023-01-02 cs.CV

classification cs.CV
keywords diffusionmodelcategoricalmodelsscenescene-scaledatadiscrete
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we learn a diffusion model to generate 3D data on a scene-scale. Specifically, our model crafts a 3D scene consisting of multiple objects, while recent diffusion research has focused on a single object. To realize our goal, we represent a scene with discrete class labels, i.e., categorical distribution, to assign multiple objects into semantic categories. Thus, we extend discrete diffusion models to learn scene-scale categorical distributions. In addition, we validate that a latent diffusion model can reduce computation costs for training and deploying. To the best of our knowledge, our work is the first to apply discrete and latent diffusion for 3D categorical data on a scene-scale. We further propose to perform semantic scene completion (SSC) by learning a conditional distribution using our diffusion model, where the condition is a partial observation in a sparse point cloud. In experiments, we empirically show that our diffusion models not only generate reasonable scenes, but also perform the scene completion task better than a discriminative model. Our code and models are available at https://github.com/zoomin-lee/scene-scale-diffusion

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A joint diffusion framework trains a Stable Diffusion generator and a semantic occupancy perception model together, so each task improves the other, producing text-conditional RGB-occupancy pairs.

  2. SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SynCity 3000 generates large, coherent 3D scenes from text by fine-tuning an image-to-3D diffusion model to operate convolutionally on overlapping windows, trained on procedurally generated synthetic scene data.

  3. PVNet: Point-Voxel Interaction LiDAR Scene Upsampling Via Diffusion Models

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    PVNet uses diffusion models with point-voxel interaction to upsample LiDAR scenes at arbitrary rates without dense supervision.

  4. Distilling Diffusion Models to Efficient 3D LiDAR Scene Completion

    cs.CV 2024-12 conditional novelty 6.0 of 10

    ScoreLiDAR distills a 50-step LiDAR scene completion diffusion model into an 8-step student that runs about 5x faster and matches or improves completion quality on SemanticKITTI and KITTI-360.

  5. 3D and 4D World Modeling: A Survey

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A survey that defines 3D/4D world modeling, organizes methods into VideoGen, OccGen, and LiDARGen categories, and compiles datasets, metrics, and benchmark numbers.

  6. Diffusion-Based Generative Models for 3D Occupancy Prediction in Autonomous Driving

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Diffusion-based generative models, using discrete categorical diffusion conditioned on BEV features, improve 3D occupancy prediction and downstream planning for autonomous driving.

  7. Map Imagination Like Blind Humans: Group Diffusion Model for Robotic Map Generation

    cs.RO 2024-12 conditional novelty 5.0 of 10

    A group diffusion model can generate LiDAR-style 3D maps from path-only odometry data, and adding 50 LiDAR points improves the maps.

  8. SSEditor: Controllable Mask-to-Scene Generation with Diffusion Model

    cs.CV 2024-11 conditional novelty 5.0 of 10

    SSEditor generates controllable 3D semantic urban scenes from mask conditions using a triplane autoencoder and a mask-conditional diffusion model, avoiding multi-step resampling.

  9. 3D Scene Generation: A Survey

    cs.CV 2025-05 conditional

    The paper surveys 3D scene generation and organizes methods into four paradigms, with datasets, evaluation metrics, applications, and future directions.

Pith tools