REVIEW 19 cited by
DepthSplat: Connecting Gaussian Splatting and Depth
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Gaussian splatting and single-view depth estimation are typically studied in isolation. In this paper, we present DepthSplat to connect Gaussian splatting and depth estimation and study their interactions. More specifically, we first contribute a robust multi-view depth model by leveraging pre-trained monocular depth features, leading to high-quality feed-forward 3D Gaussian splatting reconstructions. We also show that Gaussian splatting can serve as an unsupervised pre-training objective for learning powerful depth models from large-scale multi-view posed datasets. We validate the synergy between Gaussian splatting and depth estimation through extensive ablation and cross-task transfer experiments. Our DepthSplat achieves state-of-the-art performance on ScanNet, RealEstate10K and DL3DV datasets in terms of both depth estimation and novel view synthesis, demonstrating the mutual benefits of connecting both tasks. In addition, DepthSplat enables feed-forward reconstruction from 12 input views (512x960 resolutions) in 0.6 seconds.
Forward citations
Cited by 19 Pith papers
-
TinySplat: Feedforward Approach for Generating Compact 3D Scene Representation
TinySplat compresses feedforward 3D Gaussian scenes by 105-199x on two-view benchmarks (about 50x on DL3DV) while keeping rendered quality close to the uncompressed model.
-
OmniSplat: Taming Feed-Forward 3D Gaussian Splatting for Omnidirectional Images with Editable Capabilities
A training-free feed-forward framework that uses Yin-Yang grid decomposition to make pretrained perspective-image 3D Gaussian splatting models work on omnidirectional images, achieving state-of-the-art feed-forward no...
-
LongSplat: Online Generalizable 3D Gaussian Splatting from Long Sequence Images
A feed-forward 3D Gaussian Splatting pipeline that incrementally fuses and compresses historical Gaussians using a 2D image-like representation.
-
RegGS: Unposed Sparse Views Gaussian Splatting with 3DGS Registration
RegGS aligns locally generated 3D Gaussian maps using a Sinkhorn-approximated mixture Wasserstein distance, improving pose estimation and novel view synthesis from sparse unposed views.
-
JointSplat: Probabilistic Joint Flow-Depth Optimization for Sparse-View Gaussian Splatting
A feed-forward 3D Gaussian splatting method fuses depth and optical flow via a learned reliability mask, improving novel-view PSNR on RealEstate10K by 0.19 dB over its backbone.
-
RadarSplat: Radar Gaussian Splatting for High-Fidelity Data Synthesis and 3D Reconstruction of Autonomous Driving Scenes
RadarSplat brings Gaussian Splatting to automotive radar, explicitly modeling multipath and receiver noise to synthesize realistic radar images and estimate occupancy.
-
MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation Models
A feed-forward architecture that reuses a frozen depth foundation model to predict 3D Gaussian primitives, improving novel view synthesis and cross-dataset generalization.
-
Not All Frame Features Are Equal: Video-to-4D Generation via Decoupling Dynamic-Static Features
A video-to-4D generation method that decouples dynamic and static features in DINOv2 space and fuses similar dynamic information across views reports state-of-the-art scores on Consistent4D and Objaverse.
-
PanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian Splatting
A feed-forward Gaussian splatting system that synthesizes novel 4K panoramic views from two wide-baseline inputs, using Fibonacci-lattice Gaussians and memory-efficient training.
-
SLGaussian: Fast Language Gaussian Splatting in Sparse Views
SLGaussian builds a 3D semantic field from two photos in a single forward pass, stores CLIP features in a memory bank for fast open-vocabulary queries, and reports higher IoU than LangSplat and LERF on the LERF and 3D...
-
MAtCha Gaussians: Atlas of Charts for High-Quality Geometry and Photorealism From Sparse Views
MAtCha models a scene as an atlas of per-view depth charts, aligns and refines them with Gaussian surfel rendering, and extracts high-quality meshes from sparse images.
-
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
See3D proposes a pose-free visual condition for multi-view diffusion trained on web videos, claiming SOTA single- and sparse-view 3D generation, but the evaluation protocol leaks ground-truth information and mixes ben...
-
Splatter-360: Generalizable 360$^{\circ}$ Gaussian Splatting for Wide-baseline Panoramic Images
Splatter-360 is an end-to-end generalizable 3D Gaussian splatting model that builds a spherical cost volume to improve geometry and rendering from wide-baseline panoramic images.
-
Generative Densification: Learning to Densify Gaussians for High-Fidelity Generalizable 3D Reconstruction
Generative Densification improves feed-forward Gaussian 3D reconstruction by learning to generate fine Gaussians for detailed regions in one forward pass, and it beats baselines on object and scene datasets.
-
SelfSplat: Pose-Free and 3D Prior-Free Generalizable 3D Gaussian Splatting
SelfSplat jointly predicts depth, camera poses and 3D Gaussians from unposed image triplets, and outperforms prior pose-free baselines on RealEstate10K, ACID and DL3DV.
-
VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation
VistaVLA lifts multi-view vision-language features into 3D Gaussians, compresses them 99% via Merge-then-Query, and improves real-robot manipulation success by ~23% over baselines.
-
FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion Distillation
FVGen uses GAN-based adversarial distillation and softened reverse KL divergence to compress a video diffusion teacher for novel-view synthesis into a four-step student with comparable quality.
-
Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation
A feed-forward system that generates object-level and scene-level 3D Gaussian scenes from text in about eight seconds by diffusing multi-view RGB-D latent codes and decoding them into pixel-aligned 3D Gaussians.
-
Align3R: Aligned Monocular Depth Estimation for Dynamic Videos
Align3R injects monocular depth estimates into a fine-tuned DUSt3R model to recover temporally consistent video depth and camera poses for dynamic monocular videos.
Discussion (0). Continue with ORCID to comment.