A neural network learns obstacle shapes as Signed Distance Functions from LiDAR, but only training-loss results are reported and the collision-safety claim is not tested.
Multi-View Stereo by Temporal Nonparametric Fusion
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We propose a novel idea for depth estimation from multi-view image-pose pairs, where the model has capability to leverage information from previous latent-space encodings of the scene. This model uses pairs of images and poses, which are passed through an encoder--decoder model for disparity estimation. The novelty lies in soft-constraining the bottleneck layer by a nonparametric Gaussian process prior. We propose a pose-kernel structure that encourages similar poses to have resembling latent spaces. The flexibility of the Gaussian process (GP) prior provides adapting memory for fusing information from previous views. We train the encoder--decoder and the GP hyperparameters jointly end-to-end. In addition to a batch method, we derive a lightweight estimation scheme that circumvents standard pitfalls in scaling Gaussian process inference, and demonstrate how our scheme can run in real-time on smart devices.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Implicit 3D scene reconstruction using deep learning towards efficient collision understanding in autonomous driving
A neural network learns obstacle shapes as Signed Distance Functions from LiDAR, but only training-loss results are reported and the collision-safety claim is not tested.