REVIEW 12 cited by
latentSplat: Autoencoding Variational Gaussians for Fast Generalizable 3D Reconstruction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present latentSplat, a method to predict semantic Gaussians in a 3D latent space that can be splatted and decoded by a light-weight generative 2D architecture. Existing methods for generalizable 3D reconstruction either do not scale to large scenes and resolutions, or are limited to interpolation of close input views. latentSplat combines the strengths of regression-based and generative approaches while being trained purely on readily available real video data. The core of our method are variational 3D Gaussians, a representation that efficiently encodes varying uncertainty within a latent space consisting of 3D feature Gaussians. From these Gaussians, specific instances can be sampled and rendered via efficient splatting and a fast, generative decoder. We show that latentSplat outperforms previous works in reconstruction quality and generalization, while being fast and scalable to high-resolution data.
Forward citations
Cited by 12 Pith papers
-
STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes
A feed-forward Transformer turns sparse multi-view video frames into 3D Gaussians with velocities, reconstructing dynamic driving scenes in 0.2 seconds and estimating scene flow without motion labels.
-
MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation Models
A feed-forward architecture that reuses a frozen depth foundation model to predict 3D Gaussian primitives, improving novel view synthesis and cross-dataset generalization.
-
GSsplat: Generalizable Semantic Gaussian Splatting for Novel-view Synthesis in 3D Scenes
GSsplat is a feed-forward generalizable 3D Gaussian Splatting model that renders novel-view colors and semantic maps from multi-view inputs without per-scene training, claiming state-of-the-art semantic accuracy at th...
-
Seeing World Dynamics in a Nutshell
NutWorld is a feed-forward model that represents a monocular video as structured dynamic 3D Gaussians in a canonical orthographic space, trained with depth and flow priors.
-
GSemSplat: Generalizable Semantic 3D Gaussian Splatting from Uncalibrated Image Pairs
GSemSplat predicts open-vocabulary semantic features attached to 3D Gaussians from two uncalibrated images and generalizes across scenes with a single feed-forward pass.
-
MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data
A 700K-scene procedural, non-semantic synthetic dataset improves large reconstruction models by 1.2 to 1.8 dB PSNR when combined with real data.
-
CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from A Single-View Image
CATSplat uses image captions and back-projected depth points to improve single-image 3D Gaussian scene reconstruction and novel view synthesis.
-
PanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian Splatting
A feed-forward Gaussian splatting system that synthesizes novel 4K panoramic views from two wide-baseline inputs, using Fibonacci-lattice Gaussians and memory-efficient training.
-
FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D Reconstruction
A feed-forward transformer that jointly predicts pixel-aligned 3D Gaussians and camera poses from uncalibrated sparse views.
-
Extrapolated Urban View Synthesis Benchmark
A new benchmark with 90,810 frames from public AV datasets quantifies a large performance drop of neural rendering methods on extrapolated urban views.
-
PhysMotion: Physics-Grounded Dynamics From a Single Image
PhysMotion generates physically plausible videos from a single image by simulating 3D object motion with a material point method, then enhancing the rendering with a diffusion model.
-
Direct and Explicit 3D Generation from a Single Image
A modified Stable Diffusion model generates six views of depth, color, and 3D Gaussian features from one image, then lifts them into a textured mesh or splatted scene in 15 to 25 seconds.
Discussion (0). Continue with ORCID to comment.