REVIEW 10 cited by
MoDGS: Dynamic Gaussian Splatting from Casually-captured Monocular Videos with Depth Priors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this paper, we propose MoDGS, a new pipeline to render novel views of dy namic scenes from a casually captured monocular video. Previous monocular dynamic NeRF or Gaussian Splatting methods strongly rely on the rapid move ment of input cameras to construct multiview consistency but struggle to recon struct dynamic scenes on casually captured input videos whose cameras are either static or move slowly. To address this challenging task, MoDGS adopts recent single-view depth estimation methods to guide the learning of the dynamic scene. Then, a novel 3D-aware initialization method is proposed to learn a reasonable deformation field and a new robust depth loss is proposed to guide the learning of dynamic scene geometry. Comprehensive experiments demonstrate that MoDGS is able to render high-quality novel view images of dynamic scenes from just a casually captured monocular video, which outperforms state-of-the-art meth ods by a significant margin. The code will be publicly available.
Forward citations
Cited by 10 Pith papers
-
RelayGS: Reconstructing Dynamic Scenes with Large-Scale and Complex Motions via Relay Gaussians
RelayGS improves dynamic 3D Gaussian reconstruction of large-scale motions by decoupling foreground from background with a learnable mask and decomposing trajectories into per-segment Relay Gaussians, gaining about 1 ...
-
4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans
A diffusion model trained on 60,000 fitted 4D Gaussian Splatting human clips generates text-prompted, view-consistent dynamic humans directly in 4D, over 10x faster than video-first pipelines.
-
VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow Prediction
A training-only Gaussian splatting loss, which renders predicted 3D semantics and motion into 2D camera views, improves semantic occupancy and scene flow prediction across several camera-based models.
-
Monocular Dynamic Gaussian Splatting: Fast, Brittle, and Scene Complexity Rules
A comprehensive benchmark shows monocular dynamic Gaussian splatting methods are fast and brittle, with scene complexity dominating method differences.
-
Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering
UGSDF achieves state-of-the-art novel-view rendering of dynamic urban objects without LiDAR or 3D motion annotations by jointly optimizing SDFs and 3D Gaussians under 2D depth and point-tracking priors.
-
3D Gaussian Representations with Motion Trajectory Field for Dynamic Scene Reconstruction
A 3D Gaussian Splatting model whose Gaussian centers are represented as a learned combination of shared global motion bases recovers dynamic scenes and motion trajectories from monocular video.
-
Enhanced Velocity Field Modeling for Gaussian Video Reconstruction
Velocity field rendering with flow-based losses and flow-assisted densification lifts dynamic Gaussian novel-view PSNR by about 2.5 dB on Nvidia-long and Neu3D.
-
Align3R: Aligned Monocular Depth Estimation for Dynamic Videos
Align3R injects monocular depth estimates into a fine-tuned DUSt3R model to recover temporally consistent video depth and camera poses for dynamic monocular videos.
-
Advances in 4D Representation: Geometry, Motion, and Interaction
A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.
-
Advancing Extended Reality with 3D Gaussian Splatting: Innovations and Prospects
3D Gaussian Splatting research relevant to Extended Reality is organized into a five-part taxonomy with suggested future directions.
Discussion (0). Continue with ORCID to comment.