Pith. sign in

REVIEW 8 cited by

MoDGS: Dynamic Gaussian Splatting from Casually-captured Monocular Videos with Depth Priors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.00434 v3 pith:6RHHQDO4 submitted 2024-06-01 cs.CV

classification cs.CV
keywords dynamicmodgsmonocularcapturedcasuallydepthnovelscenes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we propose MoDGS, a new pipeline to render novel views of dy namic scenes from a casually captured monocular video. Previous monocular dynamic NeRF or Gaussian Splatting methods strongly rely on the rapid move ment of input cameras to construct multiview consistency but struggle to recon struct dynamic scenes on casually captured input videos whose cameras are either static or move slowly. To address this challenging task, MoDGS adopts recent single-view depth estimation methods to guide the learning of the dynamic scene. Then, a novel 3D-aware initialization method is proposed to learn a reasonable deformation field and a new robust depth loss is proposed to guide the learning of dynamic scene geometry. Comprehensive experiments demonstrate that MoDGS is able to render high-quality novel view images of dynamic scenes from just a casually captured monocular video, which outperforms state-of-the-art meth ods by a significant margin. The code will be publicly available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Realizing Immersive Volumetric Video: A Multimodal Framework for 6-DoF VR Engagement

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    The paper presents a multimodal framework, dataset, and reconstruction pipeline to create immersive volumetric videos supporting large 6-DoF audiovisual interaction from real multi-view captures.

  2. 4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A diffusion model trained on 60,000 fitted 4D Gaussian Splatting human clips generates text-prompted, view-consistent dynamic humans directly in 4D, over 10x faster than video-first pipelines.

  3. PD-4DGS:Progressive Decomposition of 4D Gaussian Splatting for Bandwidth-Adaptive Dynamic Scene Streaming

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    PD-4DGS decomposes 4DGS into static scaffold, global deformation, and local refinement layers using hierarchical decomposition and custom losses, achieving over 60% bitstream reduction and reducing first-frame latency...

  4. Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering

    cs.CV 2025-10 conditional novelty 5.0 of 10

    UGSDF achieves state-of-the-art novel-view rendering of dynamic urban objects without LiDAR or 3D motion annotations by jointly optimizing SDFs and 3D Gaussians under 2D depth and point-tracking priors.

  5. 3D Gaussian Representations with Motion Trajectory Field for Dynamic Scene Reconstruction

    cs.RO 2025-08 conditional novelty 5.0 of 10

    A 3D Gaussian Splatting model whose Gaussian centers are represented as a learned combination of shared global motion bases recovers dynamic scenes and motion trajectories from monocular video.

  6. Enhanced Velocity Field Modeling for Gaussian Video Reconstruction

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Velocity field rendering with flow-based losses and flow-assisted densification lifts dynamic Gaussian novel-view PSNR by about 2.5 dB on Nvidia-long and Neu3D.

  7. MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion

    cs.CV 2024-10 unverdicted novelty 5.0 of 10

    By fine-tuning DUST3R to output per-timestep pointmaps on scarce dynamic video datasets, MonST3R achieves stronger video depth and pose estimation without explicit motion modeling.

  8. Advances in 4D Representation: Geometry, Motion, and Interaction

    cs.CV 2025-10 conditional novelty 4.0 of 10

    A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.

Pith tools