Pith. sign in

REVIEW 18 cited by

GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.07207 v2 pith:Q6JG62XV submitted 2024-02-11 cs.CV

classification cs.CV
keywords gala3dgenerationlayout-guidedscenecompositionalcontentgaussiangenerate
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present GALA3D, generative 3D GAussians with LAyout-guided control, for effective compositional text-to-3D generation. We first utilize large language models (LLMs) to generate the initial layout and introduce a layout-guided 3D Gaussian representation for 3D content generation with adaptive geometric constraints. We then propose an instance-scene compositional optimization mechanism with conditioned diffusion to collaboratively generate realistic 3D scenes with consistent geometry, texture, scale, and accurate interactions among multiple objects while simultaneously adjusting the coarse layout priors extracted from the LLMs to align with the generated scene. Experiments show that GALA3D is a user-friendly, end-to-end framework for state-of-the-art scene-level 3D content generation and controllable editing while ensuring the high fidelity of object-level entities within the scene. The source codes and models will be available at gala3d.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  2. SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SynCity 3000 generates large, coherent 3D scenes from text by fine-tuning an image-to-3D diffusion model to operate convolutionally on overlapping windows, trained on procedurally generated synthetic scene data.

  3. 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

    cs.GR 2025-07 conditional novelty 6.0 of 10

    A self-improving vision-language-model policy iteratively crafts 3D environments from text, and renderings of those environments serve as effective synthetic pretraining data for vision models.

  4. Sat2City: 3D City Generation from A Single Satellite Image with Cascaded Latent Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Sat2City generates explicit 3D city geometry and appearance from a height-map condition using cascaded latent diffusion on sparse voxel grids, beating prior methods on a new synthetic city dataset.

  5. Efficient multi-view training for 3D Gaussian Splatting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Training 3D Gaussian Splatting with multiple images per iteration, using partial rendering and a 3D-aware SSIM loss, improves novel-view synthesis quality over single-view training.

  6. ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A training-free pipeline that generates editable 3D scenes from text by using a generated 2D image as an intermediary to extract object shapes, appearances, positions, and poses.

  7. Apply Hierarchical-Chain-of-Generation to Complex Attributes Text-to-3D Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    HCoG uses an LLM to sort object parts from inside out and sequentially optimizes 3D Gaussian splats, improving attribute binding for complex text-to-3D prompts.

  8. Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Scenethesis integrates LLM planning, vision-guided layout refinement, and SDF-based collision and stability optimization to generate physically plausible interactive 3D scenes from text.

  9. CoherenDream: Boosting Holistic Text Coherence in 3D Generation via Multimodal Large Language Models Feedback

    cs.CV 2025-04 conditional novelty 6.0 of 10

    CoherenDream injects MLLM-based semantic feedback into the SDS loop and improves text-3D alignment on multi-object prompts.

  10. ScanEdit: Hierarchically-Guided Functional 3D Scan Editing

    cs.CV 2025-04 conditional novelty 6.0 of 10

    ScanEdit uses hierarchical scene graphs and LLM-based planning, placement, and optimization to rearrange objects in real-world 3D scans from text instructions.

  11. HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation

    cs.GR 2025-04 conditional novelty 6.0 of 10

    HiScene generates compositional 3D scenes by treating a room as an object under isometric view, then decomposing and regenerating each instance with video-diffusion amodal completion.

  12. LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LayoutVLM couples VLM-generated pose estimates and spatial relations with differentiable optimization to create physically plausible, instruction-aligned 3D layouts.

  13. MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    MARVEL-40M+ provides multi-level captions for over 8.9 million 3D assets and a two-stage text-to-3D pipeline that generates textured meshes in 15 seconds.

  14. Toward Scene Graph and Layout Guided Complex 3D Scene Generation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    GraLa3D builds a graph with single-object nodes and super-nodes plus layout boxes, then generates 3D Gaussian scenes from those structures to preserve spatial and interaction relations.

  15. GSEditPro: 3D Gaussian Splatting Editing with Attention-based Progressive Localization

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A text-driven 3D editing framework that tags 3D Gaussian points via cross-attention and uses SDS plus pseudo-GT guidance to edit only the target region.

  16. DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A pipeline that generates editable 3D scenes from natural language by combining LLM-based layout planning, multi-timestep diffusion distillation, and staged camera sampling.

  17. AccioScene: Compositional 3D Scene Generation via Graph Diffusion and Interaction-driven Critics

    cs.LG 2025-02 conditional novelty 4.0 of 10

    A text-to-3D scene pipeline that adds LLM-predicted human-object actions to a graph-diffusion scene generator and then removes or shifts objects that intersect a placed human body, yielding modest gains over InstructS...

  18. Advancing Extended Reality with 3D Gaussian Splatting: Innovations and Prospects

    cs.CV 2024-12 conditional novelty 4.0 of 10

    3D Gaussian Splatting research relevant to Extended Reality is organized into a five-part taxonomy with suggested future directions.

Pith tools