Pith. sign in

REVIEW 2 cited by

Text to 3D Scene Generation with Rich Lexical Grounding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1505.06289 v2 pith:YRJNWNBB submitted 2015-05-23 cs.CL cs.GR

Text to 3D Scene Generation with Rich Lexical Grounding

classification cs.CL cs.GR
keywords descriptionsgenerationscenescenesgroundinghumanintroducejudgments
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The ability to map descriptions of scenes to 3D geometric representations has many applications in areas such as art, education, and robotics. However, prior work on the text to 3D scene generation task has used manually specified object categories and language that identifies them. We introduce a dataset of 3D scenes annotated with natural language descriptions and learn from this data how to ground textual descriptions to physical objects. Our method successfully grounds a variety of lexical terms to concrete referents, and we show quantitatively that our method improves 3D scene generation over previous work using purely rule-based methods. We evaluate the fidelity and plausibility of 3D scenes generated with our grounding approach through human judgments. To ease evaluation on this task, we also introduce an automated metric that strongly correlates with human judgments.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. GS-Agent: Creating 4D Physical Worlds With Generative Simulation

    cs.RO 2026-07 conditional novelty 6.0

    Three LLM agents write physics-engine code from text, review rendered frames, and correct errors, turning prompts into physically simulated 4D worlds with camera control.

  2. ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning

    cs.CV 2026-07 conditional novelty 6.0

    A progressive reasoning framework where a VLM generates or edits 3D layouts one reasoned object placement at a time, trained on 224,757 GPT-4o-annotated placement pairs plus tier-decoupled GDPO.