REVIEW 35 cited by
OmniRe: Omni Urban Scene Reconstruction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce OmniRe, a comprehensive system for efficiently creating high-fidelity digital twins of dynamic real-world scenes from on-device logs. Recent methods using neural fields or Gaussian Splatting primarily focus on vehicles, hindering a holistic framework for all dynamic foregrounds demanded by downstream applications, e.g., the simulation of human behavior. OmniRe extends beyond vehicle modeling to enable accurate, full-length reconstruction of diverse dynamic objects in urban scenes. Our approach builds scene graphs on 3DGS and constructs multiple Gaussian representations in canonical spaces that model various dynamic actors, including vehicles, pedestrians, cyclists, and others. OmniRe allows holistically reconstructing any dynamic object in the scene, enabling advanced simulations (~60Hz) that include human-participated scenarios, such as pedestrian behavior simulation and human-vehicle interaction. This comprehensive simulation capability is unmatched by existing methods. Extensive evaluations on the Waymo dataset show that our approach outperforms prior state-of-the-art methods quantitatively and qualitatively by a large margin. We further extend our results to 5 additional popular driving datasets to demonstrate its generalizability on common urban scenes.
Forward citations
Cited by 35 Pith papers
-
Rig3R: Rig-Aware Conditioning for Learned 3D Reconstruction
Rig3R conditions learned 3D reconstruction on optional rig metadata and predicts rig-relative raymaps, enabling state-of-the-art pose estimation and rig calibration discovery from images.
-
STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes
A feed-forward Transformer turns sparse multi-view video frames into 3D Gaussians with velocities, reconstructing dynamic driving scenes in 0.2 seconds and estimating scene flow without motion labels.
-
SplatAD: Real-Time Lidar and Camera Rendering with 3D Gaussian Splatting for Autonomous Driving
A 3D Gaussian Splatting method renders dynamic driving scenes for both cameras and lidar in real time, using spherical-coordinate lidar rasterization and rolling shutter compensation.
-
WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence
A real multi-city, multi-kilometer surround-view driving dataset plus an urban-tailored 3DGS baseline shows that city-scale reconstruction still degrades with scale, off-trajectory views, and real-world noise.
-
Point as Skeleton: Accumulated Point Cloud Enhanced Autoregressive Generation for Closed-Loop Autonomous Driving Simulation
Point-cloud skeleton conditions and a Reset-and-Roll inference scheme enable stable frame-wise autoregressive driving video generation for closed-loop autonomous driving simulation.
-
Realistic and Controllable 3D Gaussian-Guided Object Editing for Driving Video Generation
G2Editor combines rendered 3D Gaussians, scene-level boxes, and reference-image features in a diffusion inpainting model to enable controllable object repositioning, insertion, and deletion in driving videos.
-
ExtraGS: Geometric-Aware Trajectory Extrapolation with Uncertainty-Guided Generative Priors
ExtraGS combines Gaussian-SDF road surfaces, far-field Gaussians, and spherical-harmonics uncertainty gating to generate geometrically consistent extrapolated driving views.
-
GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting
A camera-only Gaussian-surfel pipeline reconstructs full Waymo scenes, converts them to binary occupancy labels, and trains CVT-Occ to generalize on Occ3D-Waymo and Occ3D-nuScenes at a level close to or above LiDAR-la...
-
InvRGB+L: Inverse Rendering of Complex Scenes with Unified Color and LiDAR Reflectance Modeling
InvRGB+L jointly estimates visible and LiDAR albedo with a physics-based specular LiDAR model and cross-modal consistency losses, improving inverse rendering and LiDAR intensity simulation for urban and indoor scenes.
-
AD-GS: Object-Aware B-Spline Gaussian Splatting for Self-Supervised Autonomous Driving
A self-supervised Gaussian splatting method for driving scenes models object motion with learnable B-spline and quaternion B-spline curves plus bidirectional temporal visibility masks, achieving state-of-the-art rende...
-
Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
Post-trained Cosmos world models generate controllable multi-view driving videos and LiDAR; augmenting real AV training data with these synthetic clips improves downstream perception and policy metrics, especially in ...
-
VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow Prediction
A training-only Gaussian splatting loss, which renders predicted 3D semantics and motion into 2D camera views, improves semantic occupancy and scene flow prediction across several camera-based models.
-
Gaussian Splatting is an Effective Data Generator for 3D Object Detection
Inserting Gaussian-splat reconstructed 3D objects into reconstructed driving scenes is a more effective augmentation for camera-based 3D object detection than diffusion-based image synthesis.
-
PINGS: Gaussian Splatting Meets Distance Fields within a Point-Based Implicit Neural Map
PINGS jointly builds a signed distance field and a Gaussian splatting radiance field in one point-based neural map, using geometric consistency to improve both.
-
sshELF: Single-Shot Hierarchical Extrapolation of Latent Features for 3D Reconstruction from Sparse-Views
sshELF reconstructs full 360-degree outdoor scenes from six sparse views in 0.18 seconds by generating intermediate virtual views before decoding 3D Gaussian primitives.
-
LiDAR-RT: Gaussian-based Ray Tracing for Dynamic LiDAR Re-simulation
LiDAR-RT re-simulates LiDAR views of dynamic driving scenes in real time by ray tracing Gaussian primitives with learnable intensity and ray-drop properties.
-
StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
StreetCrafter conditions a video diffusion model on LiDAR point cloud renderings to synthesize controllable street views, and distills it into a real-time 3D Gaussian representation.
-
3DGUT: Enabling Distorted Cameras and Secondary Rays in Gaussian Splatting
Replacing EWA splatting's linearized projection with an Unscented Transform lets 3DGS handle fisheye and rolling-shutter cameras and enables hybrid rasterization plus traced secondary rays.
-
DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving
DrivingRecon predicts 4D Gaussians of street scenes from surround-view video in one forward pass, using a novel Prune and Dilate Block to reduce redundant overlapping points.
-
Extrapolated Urban View Synthesis Benchmark
A new benchmark with 90,810 frames from public AV datasets quantifies a large performance drop of neural rendering methods on extrapolated urban views.
-
InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models
A three-stage pipeline generates up to 100,000 square meters of dynamic 3D driving scenes with 200-frame videos, controlled by HD maps, bounding boxes, and text.
-
FreeSim: Toward Free-viewpoint Camera Simulation in Driving Scenes
A generation-reconstruction pipeline with a diffusion enhancer trained on simulated degradations enables off-trajectory camera rendering in driving scenes.
-
DeSiRe-GS: 4D Street Gaussians for Static-Dynamic Decomposition and Surface Reconstruction for Urban Driving Scenes
A self-supervised 4D Gaussian splatting pipeline that extracts motion masks from rendered-versus-real feature differences and uses them, together with geometric and cross-view constraints, to decompose and reconstruct...
-
Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering
UGSDF achieves state-of-the-art novel-view rendering of dynamic urban objects without LiDAR or 3D motion annotations by jointly optimizing SDFs and 3D Gaussians under 2D depth and point-tracking priors.
-
Decomposing Densification in Gaussian Splatting for Faster 3D Scene Reconstruction
A split-then-clone densification schedule with energy-guided multi-resolution training roughly halves 3D Gaussian Splatting training time while keeping reconstruction quality.
-
OmniIndoor3D: Comprehensive Indoor 3D Reconstruction
OmniIndoor3D jointly optimizes appearance, geometry, and panoptic labels in a single set of 3D Gaussians initialized from RGB-D camera depth, reporting state-of-the-art numbers on ScanNet and ScanNet++.
-
Bench2Drive-R: Turning Real World Data into Reactive Closed-Loop Autonomous Driving Benchmark by Generative Model
A reactive closed-loop driving simulator that uses a diffusion renderer with retrieval from real recordings, plus a nuPlan behavioral controller, to generate sensor images in response to an end-to-end driving model's actions.
-
UrbanGS: Semantic-Guided Gaussian Splatting for Urban Scene Reconstruction
UrbanGS improves urban scene reconstruction by separating static and dynamic Gaussians with 2D semantic maps and adding static-invariance, ground-consistency, and learnable time-embedding losses.
-
ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration
ReconDreamer fine-tunes a driving world model as an online restorer and progressively expands novel-trajectory training data, reporting first-time effective rendering of multi-lane shifts in driving scenes.
-
Beyond Gaussians: Fast and High-Fidelity 3D Splatting with Linear Kernels
3DLS replaces Gaussian kernels with bounded linear kernels plus distribution alignment and gradient scaling, yielding slightly better fidelity and faster rendering than 3DGS on some scenes.
-
Advances in 4D Representation: Geometry, Motion, and Interaction
A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.
-
Editing Implicit and Explicit Representations of Radiance Fields: A Survey
A review that classifies radiance field editing into explicit, latent space, text-guided, compositional, and other categories, with application and dataset tables.
-
DynamicAvatars: Accurate Dynamic Facial Avatars Reconstruction and Precise Editing with Diffusion Models
DynamicAvatars reconstructs dynamic 3D head avatars from video and enables prompt-based editing via dual Gaussian tracking, semantic masks, and LLM-guided diffusion editing.
-
EMD: Explicit Motion Modeling for High-Quality Street Gaussian Splatting
A learnable motion embedding and two-level deformation module improve dynamic street-scene rendering for several Gaussian splatting baselines.
-
Generative AI for Autonomous Driving: Frontiers and Opportunities
A comprehensive, structured survey of generative AI for autonomous driving, covering model families, sensor modalities, real-world applications, and open research challenges.
Discussion (0). Continue with ORCID to comment.