Pith. sign in

REVIEW 5 cited by

Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.11989 v2 pith:B2YVL7KR submitted 2023-03-21 cs.CV

classification cs.CV
keywords geometrymethodtexttexturedapproachexistinggeneratinginput
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Text2Room, a method for generating room-scale textured 3D meshes from a given text prompt as input. To this end, we leverage pre-trained 2D text-to-image models to synthesize a sequence of images from different poses. In order to lift these outputs into a consistent 3D scene representation, we combine monocular depth estimation with a text-conditioned inpainting model. The core idea of our approach is a tailored viewpoint selection such that the content of each image can be fused into a seamless, textured 3D mesh. More specifically, we propose a continuous alignment strategy that iteratively fuses scene frames with the existing geometry to create a seamless mesh. Unlike existing works that focus on generating single objects or zoom-out trajectories from text, our method generates complete 3D scenes with multiple objects and explicit 3D geometry. We evaluate our approach using qualitative and quantitative metrics, demonstrating it as the first method to generate room-scale 3D geometry with compelling textures from only text as input.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A causal transformer with 3D RoPE generates vector-quantized 3D Gaussian latent grids autoregressively, enabling unconditional synthesis, completion, and open-ended outpainting of indoor scenes.

  2. Conversational Interfaces for Parametric Conceptual Architectural Design: Integrating Mixed Reality with LLM-driven Interaction

    cs.HC 2025-06 conditional novelty 6.0 of 10

    A conversational MR interface with three LLM agents converts speech and gestures into parametric modeling code; a 27-person user study reports lower barriers and successful compilation, though quantitative evidence is...

  3. Quo Vadis, World Modeling?

    cs.CV 2026-08 conditional novelty 5.0 of 10

    An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.

  4. DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A pipeline that generates editable 3D scenes from natural language by combining LLM-based layout planning, multi-timestep diffusion distillation, and staged camera sampling.

  5. AI-powered Contextual 3D Environment Generation: A Systematic Review

    cs.GR 2025-06 conditional novelty 4.0 of 10

    A PRISMA-based systematic review of 136 papers finds diffusion models dominate high-quality AI 3D scene generation, with computational cost, data quality, and evaluation metrics as key limitations.

Pith tools