Pith. sign in

REVIEW 2 cited by

SceneTeller: Language-to-3D Scene Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.20727 v1 pith:25ATY7LL submitted 2024-07-30 cs.CV

classification cs.CV
keywords roomscenedesignhigh-qualityproducespromptscenessceneteller
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Designing high-quality indoor 3D scenes is important in many practical applications, such as room planning or game development. Conventionally, this has been a time-consuming process which requires both artistic skill and familiarity with professional software, making it hardly accessible for layman users. However, recent advances in generative AI have established solid foundation for democratizing 3D design. In this paper, we propose a pioneering approach for text-based 3D room design. Given a prompt in natural language describing the object placement in the room, our method produces a high-quality 3D scene corresponding to it. With an additional text prompt the users can change the appearance of the entire scene or of individual objects in it. Built using in-context learning, CAD model retrieval and 3D-Gaussian-Splatting-based stylization, our turnkey pipeline produces state-of-the-art 3D scenes, while being easy to use even for novices. Our project page is available at https://sceneteller.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can You Move These Over There? An LLM-based VR Mover for Supporting Object Manipulation

    cs.HC 2025-02 conditional novelty 6.0 of 10

    A large language model interprets pointing and speech in VR, and in a 24-person user study this interface significantly speeds up multi-object manipulation, lowers workload and arm fatigue, and improves reported user ...

  2. AccioScene: Compositional 3D Scene Generation via Graph Diffusion and Interaction-driven Critics

    cs.LG 2025-02 conditional novelty 4.0 of 10

    A text-to-3D scene pipeline that adds LLM-predicted human-object actions to a graph-diffusion scene generator and then removes or shifts objects that intersect a placed human body, yielding modest gains over InstructS...

Pith tools