Pith. sign in

REVIEW 20 cited by

Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.10642 v3 pith:AOABQOXY submitted 2023-10-16 cs.CV

Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting

classification cs.CV
keywords scenedynamiccomplexmodelingrenderingtimeappearancedeformation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Reconstructing dynamic 3D scenes from 2D images and generating diverse views over time is challenging due to scene complexity and temporal dynamics. Despite advancements in neural implicit models, limitations persist: (i) Inadequate Scene Structure: Existing methods struggle to reveal the spatial and temporal structure of dynamic scenes from directly learning the complex 6D plenoptic function. (ii) Scaling Deformation Modeling: Explicitly modeling scene element deformation becomes impractical for complex dynamics. To address these issues, we consider the spacetime as an entirety and propose to approximate the underlying spatio-temporal 4D volume of a dynamic scene by optimizing a collection of 4D primitives, with explicit geometry and appearance modeling. Learning to optimize the 4D primitives enables us to synthesize novel views at any desired time with our tailored rendering routine. Our model is conceptually simple, consisting of a 4D Gaussian parameterized by anisotropic ellipses that can rotate arbitrarily in space and time, as well as view-dependent and time-evolved appearance represented by the coefficient of 4D spherindrical harmonics. This approach offers simplicity, flexibility for variable-length video and end-to-end training, and efficient real-time rendering, making it suitable for capturing complex dynamic scene motions. Experiments across various benchmarks, including monocular and multi-view scenarios, demonstrate our 4DGS model's superior visual quality and efficiency.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video

    cs.CV 2026-04 unverdicted novelty 8.0

    ReconPhys is the first feedforward neural network that jointly reconstructs 3D geometry and appearance via Gaussian Splatting while estimating physical attributes from a single monocular video using self-supervised training.

  2. No Pose, No Problem in 4D: Feed-Forward Dynamic Gaussians from Unposed Multi-View Videos

    cs.CV 2026-05 unverdicted novelty 7.0

    NoPo4D is the first feed-forward system for dynamic 4D Gaussian splatting from unposed multi-view videos, using velocity decomposition supervised by optical flow and a bidirectional motion encoder.

  3. MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics

    cs.CV 2026-05 unverdicted novelty 7.0

    MoCam unifies static and dynamic novel view synthesis by temporally decoupling geometric alignment and appearance refinement within the diffusion denoising process.

  4. MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics

    cs.CV 2026-05 unverdicted novelty 7.0

    MoCam uses structured denoising dynamics in diffusion models to temporally decouple geometric alignment from appearance refinement, enabling unified novel view synthesis that outperforms prior methods on imperfect poi...

  5. PaMoSplat: Part-Aware Motion-Guided Gaussian Splatting for Dynamic Scene Reconstruction

    cs.CV 2026-05 unverdicted novelty 7.0

    PaMoSplat reconstructs dynamic scenes by lifting 2D segmentations to coherent 3D Gaussian parts and estimating their motions via optical flow-guided differential evolution for higher quality rendering and faster training.

  6. ClipGStream: Clip-Stream Gaussian Splatting for Any Length and Any Motion Multi-View Dynamic Scene Reconstruction

    cs.CV 2026-04 unverdicted novelty 7.0

    ClipGStream enables scalable flicker-free reconstruction of long dynamic multi-view videos by performing stream optimization at the clip level with clip-independent spatio-temporal fields, residual anchor compensation...

  7. Realizing Immersive Volumetric Video: A Multimodal Framework for 6-DoF VR Engagement

    cs.CV 2026-04 unverdicted novelty 7.0

    The paper presents a multimodal framework, dataset, and reconstruction pipeline to create immersive volumetric videos supporting large 6-DoF audiovisual interaction from real multi-view captures.

  8. GS-Surrogate: Deformable Gaussian Splatting for Parameter Space Exploration of Ensemble Simulations

    cs.GR 2026-04 unverdicted novelty 7.0

    GS-Surrogate creates a canonical Gaussian field that is sequentially deformed by simulation parameters to enable real-time, controllable 3D exploration of ensemble data while separating simulation variations from visu...

  9. SparseCam4D: Spatio-Temporally Consistent 4D Reconstruction from Sparse Cameras

    cs.CV 2026-03 unverdicted novelty 7.0

    SparseCam4D achieves spatio-temporally consistent high-fidelity 4D reconstruction from sparse cameras via a Spatio-Temporal Distortion Field that corrects inconsistencies in generative observations.

  10. Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping

    cs.CV 2026-02 unverdicted novelty 7.0

    MoGaF groups Gaussians by motion in 4D splatting representations to enable stable long-term forecasting of dynamic scenes.

  11. Robust Dreamer: Deviation-Aware Latent Gaussian Memory for Action-Controlled AR Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    Robust Dreamer uses Latent Gaussian Memory anchored to diffusion latents and Deviation Learning with a Dynamic Deviation Archive to reduce drift in long-horizon action-controlled image-to-video generation, reporting S...

  12. Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized Distillation

    cs.CV 2026-05 unverdicted novelty 6.0

    DeGO decouples rigid and nonrigid motion in Gaussian occupancy prediction via factorized 4D distillation from VGGT, reporting SOTA results on Occ3D-NuScenes with 13.5% gains on human-centric cases.

  13. PD-4DGS:Progressive Decomposition of 4D Gaussian Splatting for Bandwidth-Adaptive Dynamic Scene Streaming

    cs.CV 2026-05 unverdicted novelty 6.0

    PD-4DGS decomposes 4DGS into static scaffold, global deformation, and local refinement layers using hierarchical decomposition and custom losses, achieving over 60% bitstream reduction and reducing first-frame latency...

  14. HOIGS: Human-Object Interaction Gaussian Splatting

    cs.CV 2026-04 unverdicted novelty 6.0

    HOIGS adds a cross-attention HOI module to Gaussian Splatting that combines HexPlane human features with Cubic Hermite Spline object features to model interaction-induced deformations.

  15. Splatography: Sparse multi-view dynamic Gaussian Splatting for filmmaking challenges

    cs.CV 2025-11 unverdicted novelty 6.0

    Splatography improves dynamic 3D reconstruction from sparse multi-view videos by splitting foreground and background Gaussian representations and applying tailored deformation learning for each.

  16. Boosting Zero-Shot 3D Style Transfer with 2D Pre-trained Priors

    cs.CV 2026-05 unverdicted novelty 5.0

    DS-StyleGaussian integrates a 2D-pretrained decoder with feature Gaussian splatting and deferred stylization to achieve view-consistent zero-shot 3D style transfer from a single style image.

  17. Generator-Refiner-Examiner: A Tri-Module Data Augmentation Framework for 3D Human Avatar Learning from Monocular Videos

    cs.CV 2026-05 unverdicted novelty 5.0

    TrioMan is a tri-module data augmentation framework using a Generator for pose/camera perturbations, a Refiner with one-step diffusion, and an Examiner with dual-branch attention to improve 3D avatar learning from mon...

  18. Flow4DGS-SLAM: Optical Flow-Guided 4D Gaussian Splatting SLAM

    cs.CV 2026-04 unverdicted novelty 5.0

    Flow4DGS-SLAM uses optical flow to generate motion masks, initialize poses, and guide 4D Gaussian modeling with scene flow and GMM for temporal properties, claiming SOTA results in dynamic tracking and reconstruction.

  19. GeoRect4D: Geometry-Compatible Generative Rectification for Dynamic Sparse-View 3D Reconstruction

    cs.CV 2026-04 unverdicted novelty 5.0

    GeoRect4D couples 3D Gaussian splatting with a single-step diffusion rectifier via degradation-aware feedback and progressive optimization to improve fidelity and consistency in sparse-view dynamic 3D reconstruction.

  20. Beyond Static Gaussians: An Empirical Investigation of Architectural Paradigms for Dynamic 3D Scene Reconstruction

    cs.CV 2026-05 unverdicted novelty 4.0

    Structure-guided dynamic 3DGS methods deliver superior reconstruction fidelity and compactness on D-NeRF while gaussian-centric methods provide higher rendering speeds at the cost of quality variability and storage.