Pith. sign in

REVIEW 11 cited by

Text2Immersion: Generative Immersive Scene with 3D Gaussians

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.09242 v1 pith:MQADG3PF submitted 2023-12-14 cs.CV cs.GR

classification cs.CVcs.GR
keywords scenesscenetext2immersioncloudcreationgaussianimmersivemethods
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce Text2Immersion, an elegant method for producing high-quality 3D immersive scenes from text prompts. Our proposed pipeline initiates by progressively generating a Gaussian cloud using pre-trained 2D diffusion and depth estimation models. This is followed by a refining stage on the Gaussian cloud, interpolating and refining it to enhance the details of the generated scene. Distinct from prevalent methods that focus on single object or indoor scenes, or employ zoom-out trajectories, our approach generates diverse scenes with various objects, even extending to the creation of imaginary scenes. Consequently, Text2Immersion can have wide-ranging implications for various applications such as virtual reality, game development, and automated content creation. Extensive evaluations demonstrate that our system surpasses other methods in rendering quality and diversity, further progressing towards text-driven 3D scene generation. We will make the source code publicly accessible at the project page.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PrITTI: Primitive-based Generation of Controllable and Editable 3D Semantic Urban Scenes

    cs.CV 2025-06 conditional novelty 7.0 of 10

    PrITTI generates controllable 3D semantic urban scenes from a hybrid primitive/raster representation and reports state-of-the-art generation quality over voxel-based baselines on KITTI-360.

  2. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  3. SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SynCity 3000 generates large, coherent 3D scenes from text by fine-tuning an image-to-3D diffusion model to operate convolutionally on overlapping windows, trained on procedurally generated synthetic scene data.

  4. CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-Centric 3D Scene Generation

    cs.GR 2026-07 conditional novelty 6.0 of 10

    CGGS generates viewpoint-consistent, text-aligned ego-centric 3D scenes via consistency-augmented multi-view diffusion, flow-guided layout initialization, and mutual-information depth-refined Gaussian optimization.

  5. LivingWorld: Interactive 4D World Generation with Environmental Dynamics

    cs.CV 2026-04 conditional novelty 6.0 of 10

    An interactive pipeline generates expanding 4D worlds with globally coherent environmental dynamics from a single image in roughly 12 seconds per expansion step.

  6. Text2Stereo: Repurposing Stable Diffusion for Stereo Generation with Consistency Rewards

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Text2Stereo adapts Stable Diffusion to generate wide-baseline stereo image pairs from text by fine-tuning with LoRA and a disparity-correlation consistency reward.

  7. HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A framework that generates panoramic videos from one image and reconstructs them into 4D Gaussian scenes, with a new panoramic video dataset and improved depth alignment.

  8. BloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    BloomScene generates 3D scenes from text or images by combining progressive point cloud construction, depth-prior regularization, and hash-grid compression, cutting storage about 5.8x versus LucidDreamer.

  9. RoomPainter: View-Integrated Diffusion for Consistent Indoor Scene Texturing

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A two-stage, zero-shot diffusion pipeline that textures room-scale meshes with global style consistency and per-instance repainting.

  10. DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A pipeline that generates editable 3D scenes from natural language by combining LLM-based layout planning, multi-timestep diffusion distillation, and staged camera sampling.

  11. Advancing Extended Reality with 3D Gaussian Splatting: Innovations and Prospects

    cs.CV 2024-12 conditional novelty 4.0 of 10

    3D Gaussian Splatting research relevant to Extended Reality is organized into a five-part taxonomy with suggested future directions.

Pith tools