FreeScene parses free-form text and image prompts into scene graphs via a VLM-based Graph Designer, then generates 3D indoor layouts with a mixed graph diffusion transformer, reporting improved quality and controllability over prior baselines.
Text to 3D Scene Generation with Rich Lexical Grounding
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The ability to map descriptions of scenes to 3D geometric representations has many applications in areas such as art, education, and robotics. However, prior work on the text to 3D scene generation task has used manually specified object categories and language that identifies them. We introduce a dataset of 3D scenes annotated with natural language descriptions and learn from this data how to ground textual descriptions to physical objects. Our method successfully grounds a variety of lexical terms to concrete referents, and we show quantitatively that our method improves 3D scene generation over previous work using purely rule-based methods. We evaluate the fidelity and plausibility of 3D scenes generated with our grounding approach through human judgments. To ease evaluation on this task, we also introduce an automated metric that strongly correlates with human judgments.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
FreeScene: Mixed Graph Diffusion for 3D Scene Synthesis from Free Prompts
FreeScene parses free-form text and image prompts into scene graphs via a VLM-based Graph Designer, then generates 3D indoor layouts with a mixed graph diffusion transformer, reporting improved quality and controllability over prior baselines.