Pith. sign in

REVIEW 4 cited by

SceneFormer: Indoor Scene Generation with Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.09793 v2 pith:MS6ZYHM7 submitted 2020-12-17 cs.CV

SceneFormer: Indoor Scene Generation with Transformers

classification cs.CV
keywords scenessceneindoorroomgenerationtransformersappearanceconditioned
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We address the task of indoor scene generation by generating a sequence of objects, along with their locations and orientations conditioned on a room layout. Large-scale indoor scene datasets allow us to extract patterns from user-designed indoor scenes, and generate new scenes based on these patterns. Existing methods rely on the 2D or 3D appearance of these scenes in addition to object positions, and make assumptions about the possible relations between objects. In contrast, we do not use any appearance information, and implicitly learn object relations using the self-attention mechanism of transformers. We show that our model design leads to faster scene generation with similar or improved levels of realism compared to previous methods. Our method is also flexible, as it can be conditioned not only on the room layout but also on text descriptions of the room, using only the cross-attention mechanism of transformers. Our user study shows that our generated scenes are preferred to the state-of-the-art FastSynth scenes 53.9% and 56.7% of the time for bedroom and living room scenes, respectively. At the same time, we generate a scene in 1.48 seconds on average, 20% faster than FastSynth.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Learning to Place Objects with Programs and Iterative Self Training

    cs.GR 2025-03 unverdicted novelty 7.0

    A generative model writes programs in a relational constraint DSL and uses bootstrapping to learn object placement distributions that align more closely with human annotations than data-driven or LLM baselines.

  2. ShellMaker: Language-Guided Exterior Completion under Structural Constraints

    cs.CV 2026-06 unverdicted novelty 6.0

    ShellMaker generates complete building exteriors from scaffolds and style prompts via parametric roofs, LLM prompt refinement, material retrieval, and geometry-aware assembly while preserving structural constraints.

  3. HetScene: Heterogeneity-Aware Diffusion for Dense Indoor Scene Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    HetScene proposes a two-stage heterogeneous diffusion framework that decomposes scenes into primary structural objects and secondary contextual objects to generate denser, more plausible indoor layouts.

  4. GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation

    cs.CV 2026-03 conditional novelty 6.0

    A causal transformer with 3D RoPE generates vector-quantized 3D Gaussian latent grids autoregressively, enabling unconditional synthesis, completion, and open-ended outpainting of indoor scenes.