REVIEW 18 cited by
InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Comprehending natural language instructions is a charming property for 3D indoor scene synthesis systems. Existing methods directly model object joint distributions and express object relations implicitly within a scene, thereby hindering the controllability of generation. We introduce InstructScene, a novel generative framework that integrates a semantic graph prior and a layout decoder to improve controllability and fidelity for 3D scene synthesis. The proposed semantic graph prior jointly learns scene appearances and layout distributions, exhibiting versatility across various downstream tasks in a zero-shot manner. To facilitate the benchmarking for text-driven 3D scene synthesis, we curate a high-quality dataset of scene-instruction pairs with large language and multimodal models. Extensive experimental results reveal that the proposed method surpasses existing state-of-the-art approaches by a large margin. Thorough ablation studies confirm the efficacy of crucial design components. Project page: https://chenguolin.github.io/projects/InstructScene.
Forward citations
Cited by 18 Pith papers
-
D3D-GEN: Robot-Aware Domain-Grounded Interactive 3D World Generation for Social Robotics
D3D-GEN automatically builds a domain knowledge base from web research and uses it to generate interactive 3D robot simulation worlds for residential, office, and hospital settings.
-
Agentic Designer: Progressive Multi-Agent Collaboration for Structure-Aware Interior Layout Generation
A progressive generator-evaluator-refiner loop, trained on a new 18,853-room benchmark, reduces furniture-boundary violations in generated interior layouts from more than 90% to single digits while keeping layout density.
-
CR-Refiner: An Object-Centric Optimal Transport Reranker for Edit-Conditioned 3D Scene Retrieval
An unbalanced optimal-transport reranker with structural priors and an LLM verifier improves hard-subset 3D scene retrieval, evaluated on the new synthetic 3D-CER benchmark.
-
SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
SynCity 3000 generates large, coherent 3D scenes from text by fine-tuning an image-to-3D diffusion model to operate convolutionally on overlapping windows, trained on procedurally generated synthetic scene data.
-
3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds
A self-improving vision-language-model policy iteratively crafts 3D environments from text, and renderings of those environments serve as effective synthetic pretraining data for vision models.
-
PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers
PartCrafter generates several separable 3D part meshes at once from a single image by fine-tuning a pretrained 3D diffusion transformer with part identity tokens and local-global attention.
-
FreeScene: Mixed Graph Diffusion for 3D Scene Synthesis from Free Prompts
FreeScene parses free-form text and image prompts into scene graphs via a VLM-based Graph Designer, then generates 3D indoor layouts with a mixed graph diffusion transformer, reporting improved quality and controllabi...
-
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
ReSpace is an autoregressive LLM framework for text-driven 3D indoor scene editing and synthesis, using a structured JSON scene representation and a voxelization-based layout metric.
-
ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary
A training-free pipeline that generates editable 3D scenes from text by using a generated 2D image as an intermediary to extract object shapes, appearances, positions, and poses.
-
Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation
Scenethesis integrates LLM planning, vision-guided layout refinement, and SDF-based collision and stability optimization to generate physically plausible interactive 3D scenes from text.
-
ScanEdit: Hierarchically-Guided Functional 3D Scan Editing
ScanEdit uses hierarchical scene graphs and LLM-based planning, placement, and optimization to rearrange objects in real-world 3D scans from text instructions.
-
HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
HiScene generates compositional 3D scenes by treating a room as an object under isometric view, then decomposing and regenerating each instance with video-diffusion amodal completion.
-
LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models
LayoutVLM couples VLM-generated pose estimates and spatial relations with differentiable optimization to create physically plausible, instruction-aligned 3D layouts.
-
Diorama: Unleashing Zero-shot Single-view 3D Indoor Scene Modeling
Diorama produces a structured, CAD-based 3D scene model from one RGB image using pretrained foundation models and staged layout optimization, with no end-to-end training.
-
SG-Layout: Structured Scene Graph-Guided Layout Generation with LLMs
Feeding scene-graph embeddings into a frozen LLM through a two-stage alignment and LoRA pipeline improves layout accuracy in relation-dense scenes, at the cost of small losses on simple two-object layouts.
-
MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene Generation
MMGDreamer generates 3D indoor scenes from a mixed-modality scene graph whose nodes can be text, images, or both, and it predicts missing object relationships for more coherent layouts.
-
SSEditor: Controllable Mask-to-Scene Generation with Diffusion Model
SSEditor generates controllable 3D semantic urban scenes from mask conditions using a triplane autoencoder and a mask-conditional diffusion model, avoiding multi-step resampling.
-
Global Graph-Validated Optimization for VLM-based 3D Indoor Scene Generation
A graph-verified, evolution-plus-gradient pipeline for text-driven 3D indoor layout generation reports improved GPT-4o-judged semantic and physical quality over four prior methods.
Discussion (0). Continue with ORCID to comment.