REVIEW 5 cited by
Compositional 3D Scene Generation using Locally Conditioned Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Designing complex 3D scenes has been a tedious, manual process requiring domain expertise. Emerging text-to-3D generative models show great promise for making this task more intuitive, but existing approaches are limited to object-level generation. We introduce \textbf{locally conditioned diffusion} as an approach to compositional scene diffusion, providing control over semantic parts using text prompts and bounding boxes while ensuring seamless transitions between these parts. We demonstrate a score distillation sampling--based text-to-3D synthesis pipeline that enables compositional 3D scene generation at a higher fidelity than relevant baselines.
Forward citations
Cited by 5 Pith papers
-
PartGen: Part-level 3D Generation and Reconstruction with Multi-View Diffusion Models
A multi-view diffusion pipeline that segments 3D objects into parts, completes occluded or invisible parts, and reconstructs them into a compositional 3D asset.
-
SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
SynCity 3000 generates large, coherent 3D scenes from text by fine-tuning an image-to-3D diffusion model to operate convolutionally on overlapping windows, trained on procedurally generated synthetic scene data.
-
SemLayoutDiff: Semantic Layout Generation with Diffusion Model for Indoor Scene Synthesis
SemLayoutDiff uses a categorical diffusion model over top-down semantic maps, conditioned on architectural room masks, to generate coherent 3D indoor layouts across room types.
-
PhyCAGE: Physically Plausible Compositional 3D Asset Generation from a Single Image
A single-image pipeline that generates physically plausible compositional 3D Gaussian Splatting assets by using a physics simulator as a gradient-driven optimizer.
-
DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation
A pipeline that generates editable 3D scenes from natural language by combining LLM-based layout planning, multi-timestep diffusion distillation, and staged camera sampling.
Discussion (0). Continue with ORCID to comment.