Pith. sign in

REVIEW 1 cited by

CARFF: Conditional Auto-encoded Radiance Field for 3D Scene Forecasting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.18075 v2 pith:JULBQKK3 submitted 2024-01-31 cs.CV

classification cs.CV
keywords scenecarffmethodscenarioscomplexdrivingfieldlatent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose CARFF, a method for predicting future 3D scenes given past observations. Our method maps 2D ego-centric images to a distribution over plausible 3D latent scene configurations and predicts the evolution of hypothesized scenes through time. Our latents condition a global Neural Radiance Field (NeRF) to represent a 3D scene model, enabling explainable predictions and straightforward downstream planning. This approach models the world as a POMDP and considers complex scenarios of uncertainty in environmental states and dynamics. Specifically, we employ a two-stage training of Pose-Conditional-VAE and NeRF to learn 3D representations, and auto-regressively predict latent scene representations utilizing a mixture density network. We demonstrate the utility of our method in scenarios using the CARLA driving simulator, where CARFF enables efficient trajectory and contingency planning in complex multi-agent autonomous driving scenarios involving occlusions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-View Pedestrian Occupancy Prediction with a Novel Synthetic Dataset

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new synthetic multi-view dataset and baseline model for predicting voxel-level pedestrian occupancy and panoptic labels in dense urban scenes.

Pith tools