Pith. sign in

REVIEW 13 cited by

DreamScene: 3D Gaussian-based Text-to-3D Scene Generation via Formation Pattern Sampling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.03575 v2 pith:37RCG3TJ submitted 2024-04-04 cs.CV

classification cs.CV
keywords dreamscenesamplingsceneformationgenerationtext-to-3dconsistencyediting
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Text-to-3D scene generation holds immense potential for the gaming, film, and architecture sectors. Despite significant progress, existing methods struggle with maintaining high quality, consistency, and editing flexibility. In this paper, we propose DreamScene, a 3D Gaussian-based novel text-to-3D scene generation framework, to tackle the aforementioned three challenges mainly via two strategies. First, DreamScene employs Formation Pattern Sampling (FPS), a multi-timestep sampling strategy guided by the formation patterns of 3D objects, to form fast, semantically rich, and high-quality representations. FPS uses 3D Gaussian filtering for optimization stability, and leverages reconstruction techniques to generate plausible textures. Second, DreamScene employs a progressive three-stage camera sampling strategy, specifically designed for both indoor and outdoor settings, to effectively ensure object-environment integration and scene-wide 3D consistency. Last, DreamScene enhances scene editing flexibility by integrating objects and environments, enabling targeted adjustments. Extensive experiments validate DreamScene's superiority over current state-of-the-art techniques, heralding its wide-ranging potential for diverse applications. Code and demos will be released at https://dreamscene-project.github.io .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding

    cs.CV 2025-08 conditional novelty 7.0 of 10

    LSD-3D generates explicit, 3D-consistent driving scenes by combining a generated proxy mesh with geometry-grounded distillation from a 2D diffusion model.

  2. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  3. Efficient multi-view training for 3D Gaussian Splatting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Training 3D Gaussian Splatting with multiple images per iteration, using partial rendering and a 3D-aware SSIM loss, improves novel-view synthesis quality over single-view training.

  4. BloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    BloomScene generates 3D scenes from text or images by combining progressive point cloud construction, depth-prior regularization, and hash-grid compression, cutting storage about 5.8x versus LucidDreamer.

  5. Text-to-3D Generation by 2D Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GE3D generates 3D objects by aligning latents across multi-step 2D diffusion editing trajectories, replacing the single-step SDS noise comparison and yielding more photorealistic results.

  6. InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A three-stage pipeline generates up to 100,000 square meters of dynamic 3D driving scenes with 200-frame videos, controlled by HD maps, bounding boxes, and text.

  7. SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesis

    cs.CV 2024-11 conditional novelty 6.0 of 10

    SplatFlow jointly generates multi-view images, depths, and camera poses with a rectified flow model, then decodes them into editable 3D Gaussian Splatting scenes.

  8. Quo Vadis, World Modeling?

    cs.CV 2026-08 conditional novelty 5.0 of 10

    An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.

  9. Hybrid Mesh-Gaussian Representation for Efficient Indoor Scene Reconstruction

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A hybrid representation routes texture-rich flat indoor regions to a textured mesh and keeps Gaussians only for complex geometry, reducing Gaussian counts by 18-50% with roughly comparable rendering quality.

  10. GSEditPro: 3D Gaussian Splatting Editing with Attention-based Progressive Localization

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A text-driven 3D editing framework that tags 3D Gaussian points via cross-attention and uses SDS plus pseudo-GT guidance to edit only the target region.

  11. DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A pipeline that generates editable 3D scenes from natural language by combining LLM-based layout planning, multi-timestep diffusion distillation, and staged camera sampling.

  12. Advancing Extended Reality with 3D Gaussian Splatting: Innovations and Prospects

    cs.CV 2024-12 conditional novelty 4.0 of 10

    3D Gaussian Splatting research relevant to Extended Reality is organized into a five-part taxonomy with suggested future directions.

  13. MS2Mesh-XR: Multi-modal Sketch-to-Mesh Generation in XR Environments

    cs.CV 2024-12 conditional novelty 3.0 of 10

    A system that turns mid-air sketches plus voice into textured 3D meshes in XR by chaining ControlNet image generation with convolutional mesh reconstruction.

Pith tools