Pith. sign in

REVIEW 17 cited by

GaussianCube: A Structured and Explicit Radiance Representation for 3D Generative Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.19655 v4 pith:G3QMS5Q2 submitted 2024-03-28 cs.CV

classification cs.CV
keywords gaussiancubemodelingrepresentationgenerativeradiancestructuredfittinggaussians
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce a radiance representation that is both structured and fully explicit and thus greatly facilitates 3D generative modeling. Existing radiance representations either require an implicit feature decoder, which significantly degrades the modeling power of the representation, or are spatially unstructured, making them difficult to integrate with mainstream 3D diffusion methods. We derive GaussianCube by first using a novel densification-constrained Gaussian fitting algorithm, which yields high-accuracy fitting using a fixed number of free Gaussians, and then rearranging these Gaussians into a predefined voxel grid via Optimal Transport. Since GaussianCube is a structured grid representation, it allows us to use standard 3D U-Net as our backbone in diffusion modeling without elaborate designs. More importantly, the high-accuracy fitting of the Gaussians allows us to achieve a high-quality representation with orders of magnitude fewer parameters than previous structured representations for comparable quality, ranging from one to two orders of magnitude. The compactness of GaussianCube greatly eases the difficulty of 3D generative modeling. Extensive experiments conducted on unconditional and class-conditioned object generation, digital avatar creation, and text-to-3D synthesis all show that our model achieves state-of-the-art generation results both qualitatively and quantitatively, underscoring the potential of GaussianCube as a highly accurate and versatile radiance representation for 3D generative modeling. Project page: https://gaussiancube.github.io/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    A conditional diffusion model using proprioception and multi-contact touch produces metric-scale, physically consistent 3D object reconstructions under hand occlusion.

  2. Structured 3D Latents for Scalable and Versatile 3D Generation

    cs.CV 2024-12 unverdicted novelty 7.0 of 10

    SLAT provides a unified 3D latent representation enabling versatile high-quality generation across multiple output formats from text or image inputs.

  3. MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hierarchical point-shuffle densification plus local AVS-Conv multi-scale decoding lets compact VecSet VAEs approach voxel-level 3D reconstruction fidelity at much lower token and query cost.

  4. PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    PixGS is a single-stage pixel-space diffusion model that directly produces high-quality 3D Gaussian Splats from text or images in ~1s, outperforming multi-stage latent methods on standard benchmarks.

  5. PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    A single-stage pixel-space diffusion model for direct 3D Gaussian Splat generation that bypasses latent compression and adds geometric supervisions to outperform prior multi-stage methods.

  6. Generative 3D Gaussians with Learned Density Control

    cs.GR 2026-05 unverdicted novelty 6.0 of 10

    DeG models 3D Gaussians via learned octree density and uses VecSeq Sobol re-indexing to turn set generation into sequence modeling, claiming SOTA quality in single-image-to-3D.

  7. Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Proprioception and multi-contact touch, fused with a physics-guided conditional diffusion model over a Structure-VAE SDF latent space, improve metric amodal object reconstruction under severe hand occlusion versus vis...

  8. UniRecGen: Unifying Multi-View 3D Reconstruction and Generation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    UniRecGen unifies reconstruction and generation via shared canonical space and disentangled cooperative learning to produce complete, consistent 3D models from sparse views.

  9. Native and Compact Structured Latents for 3D Generation

    cs.CV 2025-12 unverdicted novelty 6.0 of 10

    Introduces O-Voxel omni-voxel representation and Sparse Compression VAE for structured native 3D latents, enabling efficient training of large flow-matching models that produce higher-quality geometry and materials th...

  10. Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Can3Tok tokenizes scene-level 3D Gaussian splats into canonical latent tokens with normalization and saliency filtering, enabling reconstruction and text/image-to-3D generation.

  11. Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A video-to-4D model that encodes mesh animations into compact Gaussian variation latents and diffuses them conditioned on the video and a canonical Gaussian splat.

  12. GASE: Gaussian Splatting-Based Automated System for Reconstructing Embodied-Simulation Environments

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    GASE automates high-fidelity simulation scene reconstruction from multi-view panoramic videos via Gaussian splatting, object extraction, and inpainting, yielding robot policies with under 10% performance gap versus re...

  13. Predicting 3D structure by latent posterior sampling

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    A two-stage latent-variable model uses diffusion-based score matching to sample 3D scenes from posteriors conditioned on varied observations via volumetric rendering likelihoods.

  14. Predicting 3D structure by latent posterior sampling

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    A two-stage method trains NeRF latents then a diffusion prior to sample posteriors for 3D reconstruction from varied observations including single-view, multi-view, noisy, sparse pixels, and sparse depth.

  15. Predicting 3D structure by latent posterior sampling

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    A latent-variable approach uses diffusion models on NeRF-encoded scene representations to perform posterior sampling for 3D reconstruction from single-view, multi-view, noisy, sparse-pixel, or sparse-depth inputs.

  16. Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation

    cs.GR 2025-08 unverdicted novelty 4.0 of 10

    A visual prompt engineering system for text-to-3D generation uses multi-view MLLM scoring and interactive visualizations to help designers create models faster, with 70.5% time reduction and higher quality ratings (4....

  17. A Survey on 3D Gaussian Splatting

    cs.CV 2024-01 unverdicted novelty 2.0 of 10

    A survey compiling principles, applications, benchmarks, and challenges of 3D Gaussian Splatting for explicit 3D scene representation.

Pith tools