Pith. sign in

REVIEW 32 cited by

EmerNeRF: Emergent Spatial-Temporal Scene Decomposition via Self-Supervision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.02077 v1 pith:5RUDAJNA submitted 2023-11-03 cs.CV

classification cs.CV
keywords dynamicemernerffieldsflowscenesfieldstaticdecomposition
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present EmerNeRF, a simple yet powerful approach for learning spatial-temporal representations of dynamic driving scenes. Grounded in neural fields, EmerNeRF simultaneously captures scene geometry, appearance, motion, and semantics via self-bootstrapping. EmerNeRF hinges upon two core components: First, it stratifies scenes into static and dynamic fields. This decomposition emerges purely from self-supervision, enabling our model to learn from general, in-the-wild data sources. Second, EmerNeRF parameterizes an induced flow field from the dynamic field and uses this flow field to further aggregate multi-frame features, amplifying the rendering precision of dynamic objects. Coupling these three fields (static, dynamic, and flow) enables EmerNeRF to represent highly-dynamic scenes self-sufficiently, without relying on ground truth object annotations or pre-trained models for dynamic object segmentation or optical flow estimation. Our method achieves state-of-the-art performance in sensor simulation, significantly outperforming previous methods when reconstructing static (+2.93 PSNR) and dynamic (+3.70 PSNR) scenes. In addition, to bolster EmerNeRF's semantic generalization, we lift 2D visual foundation model features into 4D space-time and address a general positional bias in modern Transformers, significantly boosting 3D perception performance (e.g., 37.50% relative improvement in occupancy prediction accuracy on average). Finally, we construct a diverse and challenging 120-sequence dataset to benchmark neural fields under extreme and highly-dynamic settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 32 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Omni-Scene: Omni-Gaussian Representation for Ego-Centric Sparse-View Scene Reconstruction

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A hybrid pixel-plus-volume Gaussian representation with triplane transformer and depth-guided training yields state-of-the-art feed-forward sparse-view reconstruction for ego-centric driving scenes.

  2. Point as Skeleton: Accumulated Point Cloud Enhanced Autoregressive Generation for Closed-Loop Autonomous Driving Simulation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Point-cloud skeleton conditions and a Reset-and-Roll inference scheme enable stable frame-wise autoregressive driving video generation for closed-loop autonomous driving simulation.

  3. Robust 4D Driving Scene Reconstruction from Imperfect Visual Priors

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    A self-correcting Gaussian scene graph uses semantic attention and adaptive topology updates to reconstruct dynamic driving scenes from noisy video-only priors.

  4. WeatherCity: Urban Scene Reconstruction with Controllable Multi-Weather Transformation

    cs.CV 2026-02 conditional novelty 6.0 of 10

    WeatherCity turns a driving video into an editable 4D scene that can be re-rendered in consistent, controllable rain, snow, and fog with stable geometry.

  5. SIMSplat: Language-Aligned 4D Gaussian Splatting for Driving Scenario Generation

    cs.RO 2025-10 conditional novelty 6.0 of 10

    SIMSplat embeds appearance, motion, and location semantics into scene-graph 4D Gaussian Splatting so driving scenes can be queried and edited via natural language, with multi-agent trajectories refined by a learned mo...

  6. One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Given one RGB-D photo of an unseen object, an AI-generated 3D mesh, aligned jointly in metric scale and pose, yields state-of-the-art one-shot 6D pose estimation on YCBInEOAT, TOYL, and LM-O.

  7. DoGFlow: Self-Supervised LiDAR Scene Flow via Cross-Modal Doppler Guidance

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Radar Doppler velocities, clustered under rigidity assumptions, can be propagated to LiDAR as pseudo scene flow labels, outperforming self-supervised baselines on TruckScenes and improving label efficiency.

  8. ExtraGS: Geometric-Aware Trajectory Extrapolation with Uncertainty-Guided Generative Priors

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    ExtraGS combines Gaussian-SDF road surfaces, far-field Gaussians, and spherical-harmonics uncertainty gating to generate geometrically consistent extrapolated driving views.

  9. AD-GS: Object-Aware B-Spline Gaussian Splatting for Self-Supervised Autonomous Driving

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A self-supervised Gaussian splatting method for driving scenes models object motion with learnable B-spline and quaternion B-spline curves plus bidirectional temporal visibility masks, achieving state-of-the-art rende...

  10. Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Post-trained Cosmos world models generate controllable multi-view driving videos and LiDAR; augmenting real AV training data with these synthetic clips improves downstream perception and policy metrics, especially in ...

  11. RadarSplat: Radar Gaussian Splatting for High-Fidelity Data Synthesis and 3D Reconstruction of Autonomous Driving Scenes

    cs.CV 2025-06 conditional novelty 6.0 of 10

    RadarSplat brings Gaussian Splatting to automotive radar, explicitly modeling multipath and receiver noise to synthesize realistic radar images and estimate occupancy.

  12. UrbanCraft: Urban View Extrapolation via Hierarchical Sem-Geometric Priors

    cs.CV 2025-05 reject novelty 6.0 of 10

    UrbanCraft uses hierarchical semantic-geometric priors to condition diffusion-based score distillation, enabling extrapolated view synthesis for urban 3D Gaussian Splatting scenes.

  13. Gaussian Splatting is an Effective Data Generator for 3D Object Detection

    cs.CV 2025-04 conditional novelty 6.0 of 10

    Inserting Gaussian-splat reconstructed 3D objects into reconstructed driving scenes is a more effective augmentation for camera-based 3D object detection than diffusion-based image synthesis.

  14. Predicting 3D representations for Dynamic Scenes

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A self-supervised model predicts future 3D radiance-field representations (triplanes) from monocular video and renders future views better than two adapted baselines on held-out scenes.

  15. DreamDrive: Generative 4D Scene Modeling from Street View Images

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DreamDrive generates 3D-consistent driving videos from a single image by lifting diffusion-generated reference frames into a hybrid static and dynamic 4D Gaussian scene.

  16. Extrapolated Urban View Synthesis Benchmark

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new benchmark with 90,810 frames from public AV datasets quantifies a large performance drop of neural rendering methods on extrapolated urban views.

  17. InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A three-stage pipeline generates up to 100,000 square meters of dynamic 3D driving scenes with 200-frame videos, controlled by HD maps, bounding boxes, and text.

  18. FreeSim: Toward Free-viewpoint Camera Simulation in Driving Scenes

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A generation-reconstruction pipeline with a diffusion enhancer trained on simulated degradations enables off-trajectory camera rendering in driving scenes.

  19. SplatFlow: Self-Supervised Dynamic Gaussian Splatting in Neural Motion Flow Field for Autonomous Driving

    cs.CV 2024-11 conditional novelty 6.0 of 10

    SplatFlow learns 4D Gaussian scene representations inside a neural motion flow field, achieving state-of-the-art novel view synthesis on Waymo and KITTI without tracked 3D boxes.

  20. DeSiRe-GS: 4D Street Gaussians for Static-Dynamic Decomposition and Surface Reconstruction for Urban Driving Scenes

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A self-supervised 4D Gaussian splatting pipeline that extracts motion masks from rendered-versus-real feature differences and uses them, together with geometric and cross-view constraints, to decompose and reconstruct...

  21. Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering

    cs.CV 2025-10 conditional novelty 5.0 of 10

    UGSDF achieves state-of-the-art novel-view rendering of dynamic urban objects without LiDAR or 3D motion annotations by jointly optimizing SDFs and 3D Gaussians under 2D depth and point-tracking priors.

  22. VAD-GS: Visibility-Aware Densification for 3D Gaussian Splatting in Dynamic Urban Scenes

    cs.CV 2025-10 conditional novelty 5.0 of 10

    Visibility-aware densification with multi-view stereo fills missing geometry in dynamic urban 3D Gaussian splatting, improving reconstruction on Waymo and nuScenes.

  23. OcRFDet: Object-Centric Radiance Fields for Multi-View 3D Object Detection in Autonomous Driving

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Adding object-centric radiance-field rendering and height-aware opacity attention to the DualBEV detector improves camera-only 3D object detection on nuScenes by up to 2.0 mAP points.

  24. Parallel Sequence Modeling via Generalized Spatial Propagation Network

    cs.CV 2025-01 conditional novelty 5.0 of 10

    GSPN is a 2D line-scan propagation mechanism for vision that reports SOTA ImageNet accuracy, strong class-conditional generation FID, and large high-resolution text-to-image speedups.

  25. Street Gaussians without 3D Object Tracker

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Replacing 3D object trackers with a 2D foundation model plus LiDAR and a motion-learning correction produces state-of-the-art street-scene reconstructions without ground-truth object poses.

  26. UrbanGS: Semantic-Guided Gaussian Splatting for Urban Scene Reconstruction

    cs.CV 2024-12 conditional novelty 5.0 of 10

    UrbanGS improves urban scene reconstruction by separating static and dynamic Gaussians with 2D semantic maps and adding static-invariance, ground-consistency, and learnable time-embedding losses.

  27. ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration

    cs.CV 2024-11 conditional novelty 5.0 of 10

    ReconDreamer fine-tunes a driving world model as an online restorer and progressively expands novel-trajectory training data, reporting first-time effective rendering of multi-lane shifts in driving scenes.

  28. DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion

    cs.CV 2025-10 conditional novelty 4.0 of 10

    DriveGen3D makes long driving-video synthesis and 3D scene reconstruction practical by caching only the conditional diffusion branch, quantizing cross-view attention, and fusing temporal context into a feed-forward Ga...

  29. DrivingGaussian++: Towards Realistic Reconstruction and Editable Simulation for Surrounding Dynamic Driving Scenes

    cs.CV 2025-08 conditional novelty 4.0 of 10

    DrivingGaussian++ reconstructs dynamic surround-view driving scenes and performs training-free multi-task editing (weather, texture, object manipulation) using Gaussians, diffusion models, and LLM-generated trajectories.

  30. Reconstructing 4D Spatial Intelligence: A Survey

    cs.CV 2025-07 accept novelty 4.0 of 10

    A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.

  31. Editing Implicit and Explicit Representations of Radiance Fields: A Survey

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A review that classifies radiance field editing into explicit, latent space, text-guided, compositional, and other categories, with application and dataset tables.

  32. EMD: Explicit Motion Modeling for High-Quality Street Gaussian Splatting

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A learnable motion embedding and two-level deformation module improve dynamic street-scene rendering for several Gaussian splatting baselines.

Pith tools