REVIEW 10 cited by
Triplane Meets Gaussian Splatting: Fast and Generalizable Single-View 3D Reconstruction with Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent advancements in 3D reconstruction from single images have been driven by the evolution of generative models. Prominent among these are methods based on Score Distillation Sampling (SDS) and the adaptation of diffusion models in the 3D domain. Despite their progress, these techniques often face limitations due to slow optimization or rendering processes, leading to extensive training and optimization times. In this paper, we introduce a novel approach for single-view reconstruction that efficiently generates a 3D model from a single image via feed-forward inference. Our method utilizes two transformer-based networks, namely a point decoder and a triplane decoder, to reconstruct 3D objects using a hybrid Triplane-Gaussian intermediate representation. This hybrid representation strikes a balance, achieving a faster rendering speed compared to implicit representations while simultaneously delivering superior rendering quality than explicit representations. The point decoder is designed for generating point clouds from single images, offering an explicit representation which is then utilized by the triplane decoder to query Gaussian features for each point. This design choice addresses the challenges associated with directly regressing explicit 3D Gaussian attributes characterized by their non-structural nature. Subsequently, the 3D Gaussians are decoded by an MLP to enable rapid rendering through splatting. Both decoders are built upon a scalable, transformer-based architecture and have been efficiently trained on large-scale 3D datasets. The evaluations conducted on both synthetic datasets and real-world images demonstrate that our method not only achieves higher quality but also ensures a faster runtime in comparison to previous state-of-the-art techniques. Please see our project page at https://zouzx.github.io/TriplaneGaussian/.
Forward citations
Cited by 10 Pith papers
-
MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data Augmentation
A single photo is converted into a 3D mesh with PBR textures using a render-enhanced auto-encoder, two data-augmentation schemes, and a multi-view texturing pipeline, with the claimed result being the best quality amo...
-
Zero-1-to-G: Taming Pretrained 2D Diffusion Model for Direct 3D Generation
A single image can be turned into a 3D Gaussian splat model by fine-tuning a pretrained 2D diffusion model to output decomposed multi-view splatter attribute images.
-
GaussianPainter: Painting Point Cloud into 3D Gaussians with Normal Guidance
GaussianPainter produces 3D Gaussians from a point cloud and reference image in one forward pass by constraining Gaussian rotations with predicted surface normals.
-
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
See3D proposes a pose-free visual condition for multi-view diffusion trained on web videos, claiming SOTA single- and sparse-view 3D generation, but the evaluation protocol leaks ground-truth information and mixes ben...
-
LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models
Fine-tuning LLaMA-3.1-8B on an OBJ-as-text dataset lets one chat model both answer questions and generate simple 3D meshes, with no vocabulary expansion.
-
Few-step Flow for 3D Generation via Marginal-Data Transport Distillation
MDT-dist distills a pretrained 3D flow model into a 1-2 step generator using velocity matching plus velocity distillation, cutting TRELLIS inference from 6.1s to 0.68s while approximately preserving generation quality.
-
UnCommon Objects in 3D
uCO3D is a large, diverse, high-quality real-object video dataset with 3D annotations that improves training of feedforward 3D reconstruction and text-to-3D models.
-
MVBoost: Boost 3D Reconstruction with Multi-View Refinement
MVBoost generates pseudo-ground-truth multi-view images by diffusing renderings of a base 3D model, trains a boosted reconstruction model on them, and reports SOTA on GSO.
-
M3D: Dual-Stream Selective State Spaces and Depth-Driven Framework for High-Fidelity Single-View 3D Reconstruction
A dual-stream Mamba plus depth network for single-view 3D reconstruction reports state-of-the-art 3D-FRONT scores, but its experimental evidence is inconsistent and missing controlled baselines.
-
FlexSplat: Flexible Feed-Forward 3D Gaussian Splatting without Point Cloud Correspondence
FlexSplat jointly trains a geometry estimator with a query-based Gaussian decoder, reconstructing objects from uncalibrated photos within 0.7 dB PSNR of posed baselines on GSO.
Discussion (0). Continue with ORCID to comment.