REVIEW 47 cited by
InstantSplat: Sparse-view Gaussian Splatting in Seconds
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
While neural 3D reconstruction has advanced substantially, its performance significantly degrades with sparse-view data, which limits its broader applicability, since SfM is often unreliable in sparse-view scenarios where feature matches are scarce. In this paper, we introduce InstantSplat, a novel approach for addressing sparse-view 3D scene reconstruction at lightning-fast speed. InstantSplat employs a self-supervised framework that optimizes 3D scene representation and camera poses by unprojecting 2D pixels into 3D space and aligning them using differentiable neural rendering. The optimization process is initialized with a large-scale trained geometric foundation model, which provides dense priors that yield initial points through model inference, after which we further optimize all scene parameters using photometric errors. To mitigate redundancy introduced by the prior model, we propose a co-visibility-based geometry initialization, and a Gaussian-based bundle adjustment is employed to rapidly adapt both the scene representation and camera parameters without relying on a complex adaptive density control process. Overall, InstantSplat is compatible with multiple point-based representations for view synthesis and surface reconstruction. It achieves an acceleration of over 30x in reconstruction and improves visual quality (SSIM) from 0.3755 to 0.7624 compared to traditional SfM with 3D-GS.
Forward citations
Cited by 47 Pith papers
-
SimVS: Simulating World Inconsistencies for Robust View Synthesis
Video diffusion models simulate world inconsistencies, and a harmonization network trained on the simulated data reconciles sparse inconsistent multi-view images into consistent 3D scenes.
-
DynSUP: Dynamic Gaussian Splatting from An Unposed Image Pair
A pose-free two-image pipeline decomposes a dynamic scene into rigid objects and fits per-Gaussian SE(3) motions to synthesize novel views of moving scenes.
-
UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
UniWorld-View couples an occlusion-aware point cloud renderer with a dual-stream video diffusion model to synthesize large-baseline novel views from monocular video.
-
D^2-4DGS: Dual-Depth Guided Sparse-Camera 4D Gaussian Splatting
D²-4DGS aligns monocular and multi-view depths, uses their agreement as verified anchors to guide densification, pruning, and depth supervision, and reports the best PSNR in all nine sparse-camera settings tested.
-
Swimm3R: Splatting with Medium-aware SfM for Underwater 3D Reconstruction
Swimm3R couples a scattering-aware, feed-forward structure-from-motion backbone with underwater Beta splatting to reconstruct and render 3D scenes from turbid underwater video, improving rendering PSNR and localizatio...
-
CORF-GS: Real-Time Wireless Radiance Field Reconstruction via Coupled Optical-RF Gaussian Splatting
CORF-GS reconstructs wireless radiance fields online by sharing Gaussian geometry between optical images and RF spectra, cutting reconstruction time ~6.4x versus RF-3DGS while improving RF spectrum PSNR.
-
GEAR: Reconstruction of Classical Paintings via Geometry Grounding and Appearance Restitution
GeAR recovers more geometrically plausible and painterly-faithful 3D Gaussian reconstructions from classical paintings by separating geometry grounding from appearance restitution.
-
MAC-Splat: Multi-Attribute Consistency for High-Fidelity Sparse-View Reconstruction
Semantically enriched MASt3R correspondences plus a multi-attribute 3D consistency loss raise sparse-view ScanNet++ PSNR by >4.5 dB over Splatt3R and preserve quality under wide baselines.
-
NoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction
Anchoring Gaussian centers to predicted raymaps and jointly optimizing RGB, raymap, and camera losses with a dual-frequency curriculum suppresses pose drift and improves pose-free 3D reconstruction on long sequences.
-
AdaptiveSplat:Texture Aware Controllable 3D Gaussian Allocation for Feed-Forward Reconstruction
Texture-aware SuperCluster pruning plus an adaptive Gaussian head lets feed-forward 3DGS models hit a user budget β while outperforming post-hoc pruners on RE10K, ACID, DL3DV and DTU.
-
The Role of Initialization in 3D Gaussian Splatting
Dense initialization of 3DGS does not consistently beat sparse SfM initialization for standard novel views, but improves off-trajectory generalization; no densification method wins everywhere.
-
GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis
Feature-space Gaussian Splat Feature Adapter (GS-Adapter) grounds camera-controlled video diffusion in 3D Gaussians, improving geometric consistency and controllability over SEVA and CameraCtrl without retraining geom...
-
DAV-GSWT: Diffusion-Active-View Sampling for Data-Efficient Gaussian Splatting Wang Tiles
DAV-GSWT uses diffusion priors and active view sampling to synthesize high-fidelity Gaussian Splatting Wang Tiles from minimal observations while preserving visual quality and tile transitions.
-
Hybrid Foveated Path Tracing with Peripheral Gaussians for Immersive Anatomy
A hybrid VR renderer combines foveated volumetric path tracing with a rapidly regenerated Gaussian-splatting periphery for interactive medical anatomy visualization.
-
E-4DGS: High-Fidelity Dynamic Reconstruction from the Multi-view Event Cameras
E-4DGS is a deformable 3D Gaussian Splatting method that reconstructs dynamic scenes directly from multi-view event camera streams, outperforming event-to-image baseline approaches.
-
Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians
Can3Tok tokenizes scene-level 3D Gaussian splats into canonical latent tokens with normalization and saliency filtering, enabling reconstruction and text/image-to-3D generation.
-
NeRF Is a Valuable Assistant for 3D Gaussian Splatting
NeRF-GS jointly optimizes a NeRF and a 3D Gaussian Splatting model in one scene, using shared features, residual corrections, and mutual loss constraints to beat both standalone methods.
-
UFV-Splatter: Pose-Free Feed-Forward 3D Gaussian Splatting Adapted to Unfavorable Views
Re-centering inputs, LoRA fine-tuning, and a Gaussian refiner let a pretrained pose-free 3D Gaussian Splatting model reconstruct objects from off-center, unknown camera views without any unfavorable-view training data.
-
PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction
PanoSplatt3R adapts a perspective pretrained stereo model to unposed wide-baseline panorama reconstruction with per-head rolled rotary positional embeddings, achieving SOTA on HM3D and Replica.
-
GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting
A camera-only Gaussian-surfel pipeline reconstructs full Waymo scenes, converts them to binary occupancy labels, and trains CVT-Occ to generalize on Occ3D-Waymo and Occ3D-nuScenes at a level close to or above LiDAR-la...
-
TRAN-D: 2D Gaussian Splatting-based Sparse-view Transparent Object Depth Reconstruction via Physics Simulation for Scene Update
TRAN-D reconstructs transparent-object depth from sparse views via segmentation-conditioned 2D Gaussian Splatting with an object-aware loss, and updates scenes after object removal using one image and physics simulation.
-
VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment
VideoLifter is a fragment-based, SfM-free video-to-3D pipeline using learned stereo priors and hierarchical Gaussian merging to cut training time by over 80% while matching or beating CF-3DGS quality.
-
EasySplat: View-Adaptive Learning makes 3D Gaussian Splatting Easy
EasySplat combines DUSt3R pointmap initialization with a KNN-based Gaussian splitting rule to improve 3D Gaussian Splatting for dense-view novel view synthesis.
-
What Makes for a Good Stereoscopic Image?
SCOPE, a VR-annotated stereo preference dataset with 2,400 comparisons, and iSQoE, a model trained on it, outperform existing image-quality metrics on mono-to-stereo conversion ranking.
-
DAS3R: Dynamics-Aware Gaussian Splatting for Static Scene Reconstruction
A dynamics-aware Gaussian splatting method reconstructs clean static backgrounds from unposed videos with large dynamic objects by training dynamic masks on image pairs and optimizing a per-Gaussian staticness score.
-
Dust to Tower: Coarse-to-Fine Photo-Realistic Scene Reconstruction from Sparse Uncalibrated Images
A coarse-to-fine pipeline jointly optimizes 3D Gaussian Splatting and camera poses from sparse, uncalibrated images, using warped and inpainted pseudo-views for supervision.
-
Learning Radiance Fields from a Single Snapshot Compressive Image
SCINeRF and SCISplat recover a view-consistent 3D scene representation from a single snapshot compressive image, with SCISplat reaching 35.94 dB PSNR and 205 FPS rendering.
-
GSemSplat: Generalizable Semantic 3D Gaussian Splatting from Uncalibrated Image Pairs
GSemSplat predicts open-vocabulary semantic features attached to 3D Gaussians from two uncalibrated images and generalizes across scenes with a single feed-forward pass.
-
FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D Reconstruction
A feed-forward transformer that jointly predicts pixel-aligned 3D Gaussians and camera poses from uncalibrated sparse views.
-
LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation Models
LoRA3D specializes pretrained 3D foundation models to target scenes via confidence-calibrated pseudo-labels from multi-view robust optimization and LoRA fine-tuning, improving performance by up to 88%.
-
MAtCha Gaussians: Atlas of Charts for High-Quality Geometry and Photorealism From Sparse Views
MAtCha models a scene as an atlas of per-view depth charts, aligns and refines them with Gaussian surfel rendering, and extracts high-quality meshes from sparse images.
-
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
See3D proposes a pose-free visual condition for multi-view diffusion trained on web videos, claiming SOTA single- and sparse-view 3D generation, but the evaluation protocol leaks ground-truth information and mixes ben...
-
Dynamics-Aware Gaussian Splatting Streaming Towards Fast On-the-Fly 4D Reconstruction
A dynamics-aware three-stage Gaussian splatting pipeline achieves the fastest reported on-the-fly 4D reconstruction training with competitive quality on N3DV and MeetRoom benchmarks.
-
SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D Reconstruction
SPARS3R aligns dense depth-prior points to SfM points in two stages, including a SAM-based semantic outlier alignment, and uses the result to initialize Gaussian Splatting, improving sparse-view rendering.
-
DeblurSplat: SfM-free 3D Gaussian Splatting with Event Camera for Robust Deblurring
A pose-free deblurring 3D Gaussian Splatting pipeline using DUSt3R point clouds, confidence-balanced sampling, and event-decoded latent image supervision.
-
Non-invasive Assessment of Pancreatic Duct Hypertension Using Computational Flow Modeling
A computational model estimates pancreatic duct pressure non-invasively from MRCP geometry, with reported agreement against ERCP pressure measurements.
-
Reconstruction Using the Invisible: Intuition from NIR and Metadata for Enhanced 3D Gaussian Splatting
NIRSplat fuses NIR imagery and vegetation-index metadata with 3D Gaussian Splatting via cross-attention and positional encoding, outperforming 3DGS, CoR-GS, and InstantSplat on the new multimodal agriculture dataset NIRPlant.
-
DIP-GS: Deep Image Prior For Gaussian Splatting Sparse View Recovery
DIP-GS applies a deep image prior in a coarse-to-fine manner to enable 3D Gaussian Splatting for sparse-view reconstruction.
-
LocalDyGS: Multi-view Global Dynamic Scene Modeling via Adaptive Local Implicit Feature Decoupling
LocalDyGS reconstructs dynamic scenes by decomposing space into seed-based local regions and generating time-varying Temporal Gaussians, though its claim of being first for large-scale scenes omits the existing Swift4...
-
RGE-GS: Reward-Guided Expansive Driving Scene Reconstruction via Diffusion Priors
RGE-GS uses a learned pixel-wise reward filter and convergence-aware Gaussian training to improve cross-lane novel view synthesis for driving scenes.
-
Intern-GS: Vision Model Guided Sparse-View 3D Gaussian Splatting
Intern-GS improves sparse-view 3D Gaussian Splatting by initializing from DUSt3R dense point clouds and regularizing with depth and diffusion-refined pseudo-views, achieving state-of-the-art results on three benchmarks.
-
Improving Novel view synthesis of 360$^\circ$ Scenes in Extremely Sparse Views by Jointly Training Hemisphere Sampled Synthetic Images
A pipeline that samples and enhances synthetic upper-hemisphere views from a DUSt3R point cloud to train 3D Gaussian Splatting, improving four-view 360-degree novel view synthesis.
-
SuperGS: Consistent and Detailed 3D Super-Resolution Scene Reconstruction via Gaussian Splatting
SuperGS outperforms prior Gaussian-splatting methods on high-resolution novel view synthesis by combining a latent feature field, multi-view voting densification, and variational uncertainty weighting.
-
Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person Retrieval
GCKD combines graph-based cross-domain propagation with momentum knowledge distillation to improve text-to-image person retrieval without target-domain annotations.
-
SmileSplat: Generalizable Gaussian Splats for Unconstrained Sparse Images
SmileSplat predicts Gaussian surfels from sparse unposed image pairs and jointly optimizes scene geometry and camera intrinsics and extrinsics, reporting state-of-the-art novel view synthesis on Re10K, ACID, Replica, ...
-
Unifying Scale-Aware Depth Prediction and Perceptual Priors for Monocular Endoscope Pose Estimation and Tissue Reconstruction
A monocular endoscopy framework fuses Depth Pro and Depth Anything depth with RAFT-LPIPS temporal refinement and dog-leg pose optimization to reconstruct tissue surfaces and camera trajectories.
-
Sparse-View 3D Reconstruction: Recent Advances and Open Challenges
A comprehensive survey that organizes sparse-view 3D reconstruction methods into geometry-based, NeRF, 3DGS, and diffusion-based categories, with benchmarks and open challenges.
Discussion (0). Continue with ORCID to comment.