Pith. sign in

REVIEW 12 cited by

latentSplat: Autoencoding Variational Gaussians for Fast Generalizable 3D Reconstruction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.16292 v2 pith:6BZOVFEP submitted 2024-03-24 cs.CV

classification cs.CV
keywords gaussianslatentsplatfastgenerativereconstructiondatageneralizablelatent
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present latentSplat, a method to predict semantic Gaussians in a 3D latent space that can be splatted and decoded by a light-weight generative 2D architecture. Existing methods for generalizable 3D reconstruction either do not scale to large scenes and resolutions, or are limited to interpolation of close input views. latentSplat combines the strengths of regression-based and generative approaches while being trained purely on readily available real video data. The core of our method are variational 3D Gaussians, a representation that efficiently encodes varying uncertainty within a latent space consisting of 3D feature Gaussians. From these Gaussians, specific instances can be sampled and rendered via efficient splatting and a fast, generative decoder. We show that latentSplat outperforms previous works in reconstruction quality and generalization, while being fast and scalable to high-resolution data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A feed-forward Transformer turns sparse multi-view video frames into 3D Gaussians with velocities, reconstructing dynamic driving scenes in 0.2 seconds and estimating scene flow without motion labels.

  2. MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A feed-forward architecture that reuses a frozen depth foundation model to predict 3D Gaussian primitives, improving novel view synthesis and cross-dataset generalization.

  3. GSsplat: Generalizable Semantic Gaussian Splatting for Novel-view Synthesis in 3D Scenes

    cs.GR 2025-05 conditional novelty 6.0 of 10

    GSsplat is a feed-forward generalizable 3D Gaussian Splatting model that renders novel-view colors and semantic maps from multi-view inputs without per-scene training, claiming state-of-the-art semantic accuracy at th...

  4. Seeing World Dynamics in a Nutshell

    cs.CV 2025-02 conditional novelty 6.0 of 10

    NutWorld is a feed-forward model that represents a monocular video as structured dynamic 3D Gaussians in a canonical orthographic space, trained with depth and flow priors.

  5. GSemSplat: Generalizable Semantic 3D Gaussian Splatting from Uncalibrated Image Pairs

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GSemSplat predicts open-vocabulary semantic features attached to 3D Gaussians from two uncalibrated images and generalizes across scenes with a single feed-forward pass.

  6. MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A 700K-scene procedural, non-semantic synthetic dataset improves large reconstruction models by 1.2 to 1.8 dB PSNR when combined with real data.

  7. CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from A Single-View Image

    cs.CV 2024-12 conditional novelty 6.0 of 10

    CATSplat uses image captions and back-projected depth points to improve single-image 3D Gaussian scene reconstruction and novel view synthesis.

  8. PanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian Splatting

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A feed-forward Gaussian splatting system that synthesizes novel 4K panoramic views from two wide-baseline inputs, using Fibonacci-lattice Gaussians and memory-efficient training.

  9. FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D Reconstruction

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A feed-forward transformer that jointly predicts pixel-aligned 3D Gaussians and camera poses from uncalibrated sparse views.

  10. Extrapolated Urban View Synthesis Benchmark

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new benchmark with 90,810 frames from public AV datasets quantifies a large performance drop of neural rendering methods on extrapolated urban views.

  11. PhysMotion: Physics-Grounded Dynamics From a Single Image

    cs.CV 2024-11 conditional novelty 6.0 of 10

    PhysMotion generates physically plausible videos from a single image by simulating 3D object motion with a material point method, then enhancing the rendering with a diffusion model.

  12. Direct and Explicit 3D Generation from a Single Image

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A modified Stable Diffusion model generates six views of depth, color, and 3D Gaussian features from one image, then lifts them into a textured mesh or splatted scene in 15 to 25 seconds.

Pith tools