Pith. sign in

hub

Instructscene: Instruction- driven 3d indoor scene synthesis with semantic graph prior

13 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.

13 Pith papers citing it
1 external citations · Pith
abstract

Comprehending natural language instructions is a charming property for 3D indoor scene synthesis systems. Existing methods directly model object joint distributions and express object relations implicitly within a scene, thereby hindering the controllability of generation. We introduce InstructScene, a novel generative framework that integrates a semantic graph prior and a layout decoder to improve controllability and fidelity for 3D scene synthesis. The proposed semantic graph prior jointly learns scene appearances and layout distributions, exhibiting versatility across various downstream tasks in a zero-shot manner. To facilitate the benchmarking for text-driven 3D scene synthesis, we curate a high-quality dataset of scene-instruction pairs with large language and multimodal models. Extensive experimental results reveal that the proposed method surpasses existing state-of-the-art approaches by a large margin. Thorough ablation studies confirm the efficacy of crucial design components. Project page: https://chenguolin.github.io/projects/InstructScene.

hub tools

citation-role summary

baseline 1

citation-polarity summary

years

2026 11 2025 2

roles

baseline 1

polarities

baseline 1

representative citing papers

SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

cs.CV · 2026-07-06 · conditional · novelty 6.0

SynCity 3000 generates large, coherent 3D scenes from text by fine-tuning an image-to-3D diffusion model to operate convolutionally on overlapping windows, trained on procedurally generated synthetic scene data.

Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments

cs.AI · 2026-07-02 · unverdicted · novelty 3.0

SPG-Layout combines statistical object priors with hierarchical large-object-first placement to produce physically plausible text-driven 3D scenes in non-Manhattan rooms and outperforms baselines on a new 500-scene benchmark.

citing papers explorer

Showing 13 of 13 citing papers.