Pith. sign in

REVIEW 5 cited by

Compositional 3D Scene Generation using Locally Conditioned Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.12218 v2 pith:UAQFXHMS submitted 2023-03-21 cs.CV

classification cs.CV
keywords compositionaldiffusiongenerationsceneconditionedlocallypartstext-to-3d
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Designing complex 3D scenes has been a tedious, manual process requiring domain expertise. Emerging text-to-3D generative models show great promise for making this task more intuitive, but existing approaches are limited to object-level generation. We introduce \textbf{locally conditioned diffusion} as an approach to compositional scene diffusion, providing control over semantic parts using text prompts and bounding boxes while ensuring seamless transitions between these parts. We demonstrate a score distillation sampling--based text-to-3D synthesis pipeline that enables compositional 3D scene generation at a higher fidelity than relevant baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PartGen: Part-level 3D Generation and Reconstruction with Multi-View Diffusion Models

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A multi-view diffusion pipeline that segments 3D objects into parts, completes occluded or invisible parts, and reconstructs them into a compositional 3D asset.

  2. SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SynCity 3000 generates large, coherent 3D scenes from text by fine-tuning an image-to-3D diffusion model to operate convolutionally on overlapping windows, trained on procedurally generated synthetic scene data.

  3. SemLayoutDiff: Semantic Layout Generation with Diffusion Model for Indoor Scene Synthesis

    cs.GR 2025-08 conditional novelty 6.0 of 10

    SemLayoutDiff uses a categorical diffusion model over top-down semantic maps, conditioned on architectural room masks, to generate coherent 3D indoor layouts across room types.

  4. PhyCAGE: Physically Plausible Compositional 3D Asset Generation from a Single Image

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A single-image pipeline that generates physically plausible compositional 3D Gaussian Splatting assets by using a physics simulator as a gradient-driven optimizer.

  5. DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A pipeline that generates editable 3D scenes from natural language by combining LLM-based layout planning, multi-timestep diffusion distillation, and staged camera sampling.

Pith tools