REVIEW 26 cited by
4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
Representing and rendering dynamic scenes has been an important but challenging task. Especially, to accurately model complex motions, high efficiency is usually hard to guarantee. To achieve real-time dynamic scene rendering while also enjoying high training and storage efficiency, we propose 4D Gaussian Splatting (4D-GS) as a holistic representation for dynamic scenes rather than applying 3D-GS for each individual frame. In 4D-GS, a novel explicit representation containing both 3D Gaussians and 4D neural voxels is proposed. A decomposed neural voxel encoding algorithm inspired by HexPlane is proposed to efficiently build Gaussian features from 4D neural voxels and then a lightweight MLP is applied to predict Gaussian deformations at novel timestamps. Our 4D-GS method achieves real-time rendering under high resolutions, 82 FPS at an 800$\times$800 resolution on an RTX 3090 GPU while maintaining comparable or better quality than previous state-of-the-art methods. More demos and code are available at https://guanjunwu.github.io/4dgs/.
Forward citations
Cited by 26 Pith papers
-
Representing Long Volumetric Video with Temporal Gaussian Hierarchy
A hierarchical temporal arrangement of 4D Gaussian primitives achieves near-constant GPU memory, compact storage, and real-time rendering for long volumetric videos.
-
3D Gaussian Splatting for Scientific Particle Data Compression and Rendering
ParticleGS uses 3D Gaussian splats to mimic ParaView renderings of 281M-particle data, reaching 30 dB PSNR at 65x compression and rendering at 662 FPS.
-
Future Rendering $\neq$ Future Surface: A Benchmark and Dataset for Dynamic Surface Reconstruction Beyond the Observed Window
FutureSurf, a new benchmark for held-out future surface reconstruction, shows deformation-MLP methods leave a 2-6.6× future-surface gap while rendering quality stays flat.
-
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.
-
Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments
Flow equivariant world models use a latent memory that shifts with the agent and with inferred object motion, giving stable long-horizon prediction under partial observability.
-
Learning Efficient and Generalizable Human Representation with Human Gaussian Model
A graph over SMPL mesh vertices aggregates per-frame 3D Gaussians from a video, yielding a reposable human avatar in a single feed-forward pass.
-
HuSc3D: Human Sculpture dataset for 3D object reconstruction
HuSc3D provides six real-world scenes of white, low-texture sculptures with varied capture conditions, and benchmarks show Gaussian-splatting methods clearly outperform NeRF-based methods.
-
FreeTimeGS: Free Gaussian Primitives at Anytime and Anywhere for Dynamic Scene Reconstruction
A dynamic-scene representation where Gaussian primitives live freely in 4D space-time with linear motion and Gaussian time windows achieves state-of-the-art novel-view quality on complex-motion benchmarks.
-
MaintaAvatar: A Maintainable Avatar Based on Neural Radiance Fields by Continual Learning
MaintaAvatar continually adds new appearances to a NeRF human avatar from a few images per task and retains old appearances via replay, per-appearance triplanes, and pose distillation.
-
LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors
LiftImage3D generates small-motion video clips from one image, registers them with MASt3R, and fits a distortion-aware 3D Gaussian field whose canonical scene renders new views.
-
DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving
DrivingRecon predicts 4D Gaussians of street scenes from surround-view video in one forward pass, using a novel Prune and Dilate Block to reduce redundant overlapping points.
-
4D Gaussian Splatting with Scale-aware Residual Field and Adaptive Optimization for Real-time Rendering of Temporally Complex Dynamic Scenes
SaRO-GS models dynamic scenes with 4D Gaussians plus a scale-aware residual field and adaptive per-Gaussian optimization, achieving state-of-the-art PSNR at real-time frame rates on D-NeRF and Plenoptic Video datasets.
-
Monocular Dynamic Gaussian Splatting: Fast, Brittle, and Scene Complexity Rules
A comprehensive benchmark shows monocular dynamic Gaussian splatting methods are fast and brittle, with scene complexity dominating method differences.
-
ChatSplat: 3D Conversational Gaussian Splatting
ChatSplat learns a 3D conversational field in Gaussian Splatting that supports object-, view-, and scene-level chat with an LLM at real-time speeds.
-
4D Scaffold Gaussian Splatting with Dynamic-Aware Anchor Growing for Efficient and High-Fidelity Dynamic Scene Reconstruction
A 4D anchor-based Gaussian splatting method with dynamic-aware anchor growing achieves state-of-the-art dynamic-region quality on N3DV and Technicolor while using far less storage than 4DGS.
-
Quo Vadis, World Modeling?
An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.
-
Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering
UGSDF achieves state-of-the-art novel-view rendering of dynamic urban objects without LiDAR or 3D motion annotations by jointly optimizing SDFs and 3D Gaussians under 2D depth and point-tracking priors.
-
Gaussian kernel-based motion measurement
A Gaussian kernel representation with motion consistency and super-resolution constraints measures sub-pixel displacement without per-sample tuning, as shown on synthetic and one experimental target.
-
CTRL-GS: Cascaded Temporal Residue Learning for 4D Gaussian Splatting
CTRL-GS represents dynamic Gaussian scenes as cascaded video-segment-frame residuals, improving reconstruction quality over 4D-GS on several dynamic-view benchmarks.
-
GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs
A training-free pipeline uses SAM segmentation and GPT-4V material recognition, then votes across views to attach density, elasticity, and friction values to 3D Gaussians for simulation and grasping.
-
SmileSplat: Generalizable Gaussian Splats for Unconstrained Sparse Images
SmileSplat predicts Gaussian surfels from sparse unposed image pairs and jointly optimizes scene geometry and camera intrinsics and extrinsics, reporting state-of-the-art novel view synthesis on Re10K, ACID, Replica, ...
-
DrivingGaussian++: Towards Realistic Reconstruction and Editable Simulation for Surrounding Dynamic Driving Scenes
DrivingGaussian++ reconstructs dynamic surround-view driving scenes and performs training-free multi-task editing (weather, texture, object manipulation) using Gaussians, diffusion models, and LLM-generated trajectories.
-
Cooperative Perception: A Resource-Efficient Framework for Multi-Drone 3D Scene Reconstruction Using Federated Diffusion and NeRF
The framework claims drone swarms can reconstruct 3D scenes by sharing semantic labels and poses, with a federated diffusion model generating missing views for NeRF training.
-
SparSplat: Fast Multi-View Reconstruction with Generalizable 2D Gaussian Splatting
A feed-forward model regressing 2D Gaussian splat parameters from three views reports the best Chamfer distance on DTU sparse reconstruction in its comparison table, and competitive novel view synthesis, at roughly 80...
-
EventSplat: 3D Gaussian Splatting from Moving Event Cameras for Real-time Rendering
EventSplat achieves real-time novel view synthesis from event-only camera streams by supervising 3D Gaussian Splatting with accumulated event differences, event-to-video-guided initialization, and spline-interpolated poses.
-
Sparse Input View Synthesis: 3D Representations and Reliable Priors
Regularizing sparse-input radiance fields with visibility priors, simpler-solution depth supervision, and sparse flow priors improves novel view synthesis and depth estimation on multiple benchmarks.
Discussion (0). Continue with ORCID to comment.