Pith. sign in

REVIEW 2 cited by

SemLayoutDiff: Semantic Layout Generation with Diffusion Model for Indoor Scene Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.18597 v2 pith:5BMJZXVQ submitted 2025-08-26 cs.GR cs.CV

SemLayoutDiff: Semantic Layout Generation with Diffusion Model for Indoor Scene Synthesis

classification cs.GR cs.CV
keywords modelsemlayoutdifflayoutscenesemanticarchitecturalcoherentdiffusion
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We present SemLayoutDiff, a unified model for synthesizing diverse 3D indoor scenes across multiple room types. The model introduces a scene layout representation combining a top-down semantic map and attributes for each object. Unlike prior approaches, which cannot condition on architectural constraints, SemLayoutDiff employs a categorical diffusion model capable of conditioning scene synthesis explicitly on room masks. It first generates a coherent semantic map, followed by a cross-attention-based network to predict furniture placements that respect the synthesized layout. Our method also accounts for architectural elements such as doors and windows, ensuring that generated furniture arrangements remain practical and unobstructed. Experiments on the 3D-FRONT dataset show that SemLayoutDiff produces spatially coherent, realistic, and varied scenes, outperforming previous methods.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. VoxScene: Anchor-Conditioned Voxel Diffusion for Indoor Scene Arrangement

    cs.GR 2026-05 unverdicted novelty 6.0

    VoxScene is a new anchor-conditioned voxel diffusion model that synthesizes collision-free 3D indoor scene arrangements via discrete volumetric occupancies and uses the grids for asset retrieval.

  2. Tokenizing Buildings: A Transformer for Layout Synthesis

    cs.CV 2025-12 unverdicted novelty 5.0

    SBM tokenizes building rooms via a sparse attribute-feature matrix and trains a Transformer for high-fidelity embeddings plus autoregressive layout generation, yielding better retrieval and fewer layout errors than baselines.