Pith. sign in

REVIEW 21 cited by

MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.14627 v2 pith:RUPLCYDH submitted 2024-03-21 cs.CV

MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images

classification cs.CV
keywords gaussianmvsplatcostfeed-forwardvolumecentersefficientgaussians
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We introduce MVSplat, an efficient model that, given sparse multi-view images as input, predicts clean feed-forward 3D Gaussians. To accurately localize the Gaussian centers, we build a cost volume representation via plane sweeping, where the cross-view feature similarities stored in the cost volume can provide valuable geometry cues to the estimation of depth. We also learn other Gaussian primitives' parameters jointly with the Gaussian centers while only relying on photometric supervision. We demonstrate the importance of the cost volume representation in learning feed-forward Gaussians via extensive experimental evaluations. On the large-scale RealEstate10K and ACID benchmarks, MVSplat achieves state-of-the-art performance with the fastest feed-forward inference speed (22~fps). More impressively, compared to the latest state-of-the-art method pixelSplat, MVSplat uses $10\times$ fewer parameters and infers more than $2\times$ faster while providing higher appearance and geometry quality as well as better cross-dataset generalization.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction

    cs.CV 2026-05 unverdicted novelty 7.0

    GenRecon lifts object-level generative priors to scene-scale reconstruction by chunking scenes and using projection-based conditioning on multi-view features, claiming 16% better results than prior methods.

  2. No Pose, No Problem in 4D: Feed-Forward Dynamic Gaussians from Unposed Multi-View Videos

    cs.CV 2026-05 unverdicted novelty 7.0

    NoPo4D is the first feed-forward system for dynamic 4D Gaussian splatting from unposed multi-view videos, using velocity decomposition supervised by optical flow and a bidirectional motion encoder.

  3. Splats in Splats++: Robust and Generalizable 3D Gaussian Splatting Steganography

    cs.CV 2026-04 conditional novelty 7.0

    Splats in Splats++ embeds messages into 3DGS via importance-graded SH encryption, hash-grid opacity mapping, and a gradient-gated consistency loss, achieving higher fidelity and robustness than prior methods.

  4. SparseSplat: Towards Applicable Feed-Forward 3D Gaussian Splatting with Pixel-Unaligned Prediction

    cs.CV 2026-04 unverdicted novelty 7.0

    SparseSplat uses entropy-based probabilistic sampling and a specialized point cloud network to generate compact 3D Gaussian maps that retain high rendering quality with far fewer Gaussians than prior feed-forward methods.

  5. AdaptiveSplat:Texture Aware Controllable 3D Gaussian Allocation for Feed-Forward Reconstruction

    cs.CV 2026-07 conditional novelty 6.0

    Texture-aware SuperCluster pruning plus an adaptive Gaussian head lets feed-forward 3DGS models hit a user budget β while outperforming post-hoc pruners on RE10K, ACID, DL3DV and DTU.

  6. PointSplat: Compact Gaussian Splatting via Human-Centric Prediction

    cs.CV 2026-06 unverdicted novelty 6.0

    PointSplat infers compact Gaussian splats directly in 3D space from input point sets via ray casting and Point-Image Transformer to reduce inter-view redundancy and improve novel-view quality for humans.

  7. Learning Efficient 4D Gaussian Representations from Monocular Videos with Flow Splatting

    cs.CV 2026-06 unverdicted novelty 6.0

    Flow Splatting extends 4D Gaussian volumes with time-varying means and covariances, approximates a velocity field, and splats it to render optical flow for supervising dynamic reconstruction from monocular video.

  8. Hand-4DGS: Feed-Forward 3D Gaussian Splatting for 4D Hand Reconstruction from Egocentric Videos

    cs.CV 2026-06 unverdicted novelty 6.0

    Hand-4DGS introduces the first feed-forward 3D Gaussian Splatting framework for 4D hand reconstruction from egocentric videos, achieving ~60 FPS inference and generalization on H2O and ARCTIC datasets.

  9. Self-Learning Expression Deformations for Data-Efficient Gaussian Avatars

    cs.CV 2026-06 unverdicted novelty 6.0

    SAGE self-learns Gaussian expression deformations via joint surfel-SDF optimization and self-supervised consistency, enabling comparable avatar quality from single frames, monocular rotations, or one-shot inputs.

  10. DelowlightSplat: Feed-Forward Gaussian Splatting for Lowlight 3D Scene Reconstruction

    cs.CV 2026-05 unverdicted novelty 6.0

    DelowlightSplat adds a lightweight Lowlight Adapter and cost-volume multi-view inference to feed-forward Gaussian splatting, enabling direct prediction of clean 3D Gaussians from degraded lowlight context views.

  11. You Only Gaussian Once: Controllable 3D Gaussian Splatting for Ultra-Densely Sampled Scenes

    cs.CV 2026-04 reject novelty 6.0

    YOGO enforces a fixed Gaussian budget during training via a deterministic controller, and the dense Immersion dataset shifts evaluation from sparse-view interpolation to physical fidelity.

  12. You Only Gaussian Once: Controllable 3D Gaussian Splatting for Ultra-Densely Sampled Scenes

    cs.CV 2026-04 unverdicted novelty 6.0

    YOGO reformulates stochastic 3D Gaussian Splatting into a deterministic budget-aware system and supplies an ultra-dense dataset to enforce physical fidelity over viewpoint interpolation.

  13. You Only Gaussian Once: Controllable 3D Gaussian Splatting for Ultra-Densely Sampled Scenes

    cs.CV 2026-04 conditional novelty 6.0

    YOGO delivers deterministic budget-controlled 3D Gaussian Splatting that matches or exceeds prior methods on a new ultra-dense multi-sensor indoor benchmark while keeping primitive counts strictly fixed.

  14. GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation

    cs.CV 2026-03 conditional novelty 6.0

    A causal transformer with 3D RoPE generates vector-quantized 3D Gaussian latent grids autoregressively, enabling unconditional synthesis, completion, and open-ended outpainting of indoor scenes.

  15. FLEG: Feed-Forward Language Embedded Gaussian Splatting from Any Views via Compact Semantic Representation

    cs.CV 2025-12 unverdicted novelty 6.0

    FLEG reconstructs language-embedded 3D Gaussians from arbitrary input views using a dual-branch distillation framework and a sparse set of semantic Gaussians that requires only 5% of prior embeddings.

  16. The Less You Depend, The More You Learn: Synthesizing Novel Views from Sparse, Unposed Images with Minimal 3D Knowledge

    cs.CV 2025-06 unverdicted novelty 6.0

    Data-centric novel view synthesis models with minimal 3D knowledge and no pose annotations scale better with data volume and outperform traditional bias-driven methods.

  17. RoadVGGT: Road-Structure-Aware Feed-Forward Road Surface Reconstruction

    cs.CV 2026-07 conditional novelty 5.5

    A feed-forward Gaussian head on OmniVGGT plus road-plane grid fusion and structure-aware grouping reconstructs compact road surfaces that beat RoGS and AnySplat on Waymo and zero-shot nuScenes.

  18. $\text{VG}^2$GT: Voxel-Gaussian Splatting Visual Geometry Grounded Transformer

    cs.CV 2026-06 unverdicted novelty 5.0

    VG²GT regresses Gaussian primitive parameters from multi-scale voxel features of a frozen VFM and uses stochastic solid volume rendering for depth supervision to produce geometrically accurate reconstructions that out...

  19. Long-LRM++: Preserving Fine Details in Feed-Forward Wide-Coverage Reconstruction

    cs.CV 2025-12 unverdicted novelty 5.0

    Long-LRM++ achieves real-time 14 FPS high-fidelity 360-degree scene reconstruction from 32-64 views by using semi-explicit Gaussians plus a light decoder, matching LaCT quality on DL3DV and improving depth prediction.

  20. Turbo-GS: Accelerating 3D Gaussian Fitting for High-Quality Radiance Fields

    cs.CV 2024-12 unverdicted novelty 5.0

    Turbo-GS accelerates 3D Gaussian Splatting training via dilated rendering of pixel subsets, convergence-aware Gaussian budget allocation, and combined positional-appearance error densification to enable faster 4K fitt...

  21. DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion

    cs.CV 2025-10 conditional novelty 4.0

    DriveGen3D makes long driving-video synthesis and 3D scene reconstruction practical by caching only the conditional diffusion branch, quantizing cross-view attention, and fusing temporal context into a feed-forward Ga...