REVIEW 3 major objections 5 minor 77 references
Mono-SF: Multi-View Geometry Meets Single-View Depth for Monocular Scene Flow Estimation of Dynamic Traffic Scenes
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A monocular camera plus probabilistic depth can estimate 3D scene flow in traffic scenes, reaching a combined scene flow error of 21.60 on KITTI.
desk verdict Solid monocular scene flow paper with a fixable sign error in the core homography that needs correcting before acceptance. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the energy function over scaled plane normals $n_i$ and rigid-body motions $T_j \in SE(3)$. A pixel's 3D position is fixed by the scaled normal through $n_i^T X = 1$, and its motion is a homography $p_1 = K(R_j - t_j n_i^T) K^{-1} p_0$. This homography lets the same variables explain image warping and depth: the photometric term compares Census descriptors at the warped location, and the depth term evaluates the implied inverse depth $d_t$ under the ProbDepthNet density at both times. The optimization runs particle belief propagation over the normals and motions, with rigid-body poses initialized by jointly fitting sparse flow correspondences and ProbDepthNet depth.
What would settle it
Run Mono-SF on a KITTI sequence where a moving vehicle is deliberately not detected by the segmentation (for example, by withholding one object's ground-truth mask during evaluation); the optimization will assign that object's pixels to the background rigid body and output zero 3D motion, and the foreground scene flow error will jump by roughly the object's image-area fraction.
Extended reading notes
Core claim
Mono-SF models a traffic scene as piecewise 3D planes, each associated with a rigid body (the background or an instance detected by segmentation), and minimizes an energy with three terms: a Census-based photometric distance from warping the reference image into the next frame using the plane normal and the rigid motion; a negative log-likelihood term that scores the plane's implied inverse depth against ProbDepthNet's mixture-of-Gaussians depth densities at both timestamps; and pairwise smoothness priors on depth and orientation. ProbDepthNet is trained to output per-pixel inverse-depth distributions; its CalibNet subnetwork, trained on a separate split, rescales the variance and mixture weights to counter overconfident estimates. The paper reports a scene flow error of 21.60 on the KITTI scene flow training set, the best among the monocular methods compared, with ablations showing that each energy term and the recalibration step contributes to the final result.
Load-bearing premise
The method assumes every moving object is detected by an instance segmentation network and moves as one rigid body; a missed object is treated as static, so the optimization cannot recover its motion.
Editorial extensions
If this is right
- A monocular camera can produce 3D scene flow competitive with stereo-based methods, provided the single-view depth uncertainty is well calibrated.
- The warp-consistency energy could serve as a training signal for a network that predicts depth and motion, potentially replacing the iterative optimization at test time.
- The reported ablation shows that probabilistic depth distributions are not a minor refinement: using the distribution instead of the mean lowers the combined scene flow error by several points.
- The plane-plus-rigid-body representation yields per-object 6D motions and planar surfaces that downstream planning modules can consume directly.
- The CalibNet recalibration approach applies to other probabilistic regression networks, including multi-hypothesis and probabilistic-layer alternatives, improving their calibration on the same data.
Reading between the lines
- The recalibration step is generic: any regression network that outputs a distribution and overfits its training split could adopt a separately trained rescaling subnetwork, not just depth networks.
- Running at about 41 seconds per image on a single CPU, Mono-SF is a proof of concept rather than a real-time system; a natural next step is distilling the optimized scene flow into a feed-forward network to remove the iterative loop.
- Because a missed segmentation is structurally fatal, a testable extension is to add a fallback that re-estimates pixels with high photometric residual as new rigid bodies rather than letting them remain in the background.
- The same formulation could transfer to stereo inputs by replacing the monocular depth term with a stereo disparity consistency term, although the paper does not test this variant.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Mono-SF, a monocular scene flow method for dynamic street scenes. The scene is modeled as superpixel-based 3D planes, each assigned to a rigid body (the background or a Mask R-CNN instance). Mono-SF minimizes an energy with (i) a Census-based photometric term that warps the reference image into the next frame via a plane-induced homography, (ii) a term penalizing inconsistency with pixel-wise depth distributions from a proposed ProbDepthNet, and (iii) pairwise smoothness priors. ProbDepthNet outputs a mixture of Gaussians over inverse depth and includes CalibNet, a small network trained on a hold-out split to recalibrate variances and weights. Experiments on the KITTI scene flow training set report the best SF-all score (21.60) among the monocular baselines compared, and ablations support the benefit of probabilistic, recalibrated depth and of each energy term. The paper also reports the first monocular submission to the KITTI scene flow benchmark.
Significance. If the results withstand scrutiny, the contribution is solid and useful: a first monocular entry on the KITTI scene flow benchmark, a clean probabilistic integration of single-view depth into a plane-and-rigid-body scene flow framework, and a simple recalibration idea (CalibNet) that is shown to help across several probabilistic depth formulations on KITTI. The ablation studies in Tables 3 and 4 are a strength because they isolate the contributions of the proposed components. Nevertheless, the published derivation of the central homography is inconsistent with the stated plane convention, and the comparative evaluation lacks error bars and sensitivity analysis, so I cannot regard the claims as fully established in the present form.
major comments (3)
- [Sec. 3.2, Eq. (5)] The homography is printed as K(R_j - t_j n_i^T)K^{-1}, but the model states that each plane satisfies n_i^T X = 1. For a point X_0 on that plane, X_1 = R_j X_0 + t_j = (R_j + t_j n_i^T)X_0, so the induced homography is K(R_j + t_j n_i^T)K^{-1}. The minus sign is inconsistent with the stated convention and would change the warp for any nontrivial translation, including the fronto-parallel translation case. Because Eq. (5) defines the correspondences used in both Phi_pho and the t=1 part of Phi_svd, this is a load-bearing error in the published derivation. Please correct the equation or explicitly introduce a different plane convention, and confirm that the experiments were run with the corrected form.
- [Sec. 4.2, Table 1] The central empirical claim that Mono-SF outperforms state-of-the-art monocular baselines on scene flow is supported only by single point estimates on one dataset split. No error bars, per-sequence variance, or significance tests are reported. Given that the energy in Eq. (3) depends on hand-set weights Theta_0..Theta_4 and truncation thresholds tau_0..tau_2 (Sec. 3.2), the authors should provide a sensitivity analysis or scene-wise statistics to demonstrate that the reported margins are not an artifact of parameter tuning.
- [Sec. 3.2, Initialization and Table 1] The method assumes that every moving object is detected by Mask R-CNN and successfully paired across frames by sparse-flow voting; any missed instance is treated as static and its motion is not estimated. The paper does not report how often this occurs on the KITTI scene flow set or how the foreground (fg) and scene-flow (SF) metrics depend on segmentation and pairing quality. Please add such an analysis, or qualify the claims to make this dependency explicit.
minor comments (5)
- [Eq. (9)] Please use an explicit dot product notation, e.g., n_k · n_l, instead of |n_k n_l|, which is ambiguous.
- [Sec. 4.2, Table 2] Table 2 lists stereo-based methods without a clear statement that the comparison is not like-for-like because the sensor input differs; the text should state this explicitly.
- [General] The paper refers to supplementary material at several points, but no supplementary material is included in the submission; either include it or remove these references.
- [General] A statement on code availability would materially help reproducibility, especially in light of the Eq. (5) sign issue.
- [Sec. 4.2, Table 1] In Table 1, DMDE, S. Soup, and MFA are only evaluated on MRE; the statement that Mono-SF shows the best rating on most metrics should be narrowed to the metrics actually reported for all compared methods.
Circularity Check
No significant circularity; the central scene flow claim is benchmarked against independent KITTI ground truth and the depth network is trained on LiDAR/SGM data, not on the scene flow output.
full rationale
The paper's central claim, that Mono-SF outperforms monocular baselines on KITTI scene flow, is supported by comparisons to independent ground-truth data (Table 1) and to external published baselines. The single-view depth network, ProbDepthNet, is trained with a negative log-likelihood loss against SGM-completed LiDAR depth (Eq. 2), not against the scene flow result, so the scene flow accuracy is not an input recycled as an output. The ablations in Tables 3 and 4 compare variants of the same pipeline against the same external ground truth, which is a standard ablation and not a self-definitional prediction. The self-citations to prior Mono-Stixels work are used as a baseline and as methodological inspiration, not as an unverified load-bearing premise. No fitted parameter is renamed as a prediction; the CalibNet recalibration is trained on a hold-out split and evaluated by calibration curves and NLL. The apparent sign inconsistency in Eq. (5) noted by the skeptic is a mathematical correctness or reproducibility concern, not a circularity: it does not reduce the derived scene flow to the method's own inputs by construction. Overall, the derivation chain is self-contained against external benchmarks, and no circular step can be exhibited.
Assumptions & free parameters
free parameters (5)
- Number of Gaussian components K =
8
- Energy weights Theta_0..Theta_4 =
not reported
- Truncation thresholds tau_0, tau_1, tau_2 =
not reported
- Belief propagation particles/iterations =
5 motion particles, 10 plane particles, 10 iterations
- Training schedule for ProbDepthNet =
15 epochs, LR 1e-4 halved every 5, batch 4, image 512x256
assumptions (5)
- domain assumption Traffic scenes are approximated by piecewise planar surface elements and rigid bodies.
- domain assumption Each detected object instance moves with a single 6D rigid body motion.
- domain assumption Mask R-CNN instance segmentation and sparse flow correspondences correctly identify and pair moving objects.
- domain assumption The depth distribution estimated by ProbDepthNet generalizes from its training split (KITTI raw) to the evaluation set.
- standard math Standard pinhole camera and SE(3) motion models apply.
Cite this review
Pith. "Pith review of Mono-SF: Multi-View Geometry Meets Single-View Depth for Monocular Scene Flow Estimation of Dynamic Traffic Scenes." pith.science (2026). https://pith.science/paper/E4GQEWHW
@misc{pith2026190806316,
author = {Pith},
title = {Pith review of: Mono-SF: Multi-View Geometry Meets Single-View Depth for Monocular Scene Flow Estimation of Dynamic Traffic Scenes},
year = {2026},
howpublished = {\url{https://pith.science/paper/E4GQEWHW}},
note = {Machine review of arXiv:1908.06316}
}
read the original abstract
Existing 3D scene flow estimation methods provide the 3D geometry and 3D motion of a scene and gain a lot of interest, for example in the context of autonomous driving. These methods are traditionally based on a temporal series of stereo images. In this paper, we propose a novel monocular 3D scene flow estimation method, called Mono-SF. Mono-SF jointly estimates the 3D structure and motion of the scene by combining multi-view geometry and single-view depth information. Mono-SF considers that the scene flow should be consistent in terms of warping the reference image in the consecutive image based on the principles of multi-view geometry. For integrating single-view depth in a statistical manner, a convolutional neural network, called ProbDepthNet, is proposed. ProbDepthNet estimates pixel-wise depth distributions from a single image rather than single depth values. Additionally, as part of ProbDepthNet, a novel recalibration technique for regression problems is proposed to ensure well-calibrated distributions. Our experiments show that Mono-SF outperforms state-of-the-art monocular baselines and ablation studies support the Mono-SF approach and ProbDepthNet design.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Exploiting semantic information and deep matching for op- tical flow
Min Bai, Wenjie Luo, Kaustav Kundu, and Raquel Urtasun. Exploiting semantic information and deep matching for op- tical flow. In Proc. of European Conference on Computer Vision (ECCV), pages 154–170. Springer, 2016. 5
work page 2016
-
[2]
Dan Barnes, Will Maddern, Geoffrey Pascoe, and Ingmar Posner. Driven to distraction: Self-supervised distractor learning for robust monocular visual odometry in urban en- vironments. In Proc. of IEEE International Conference on Robotics and Automation (ICRA) , pages 1894–1900. IEEE,
work page 1900
-
[3]
Multi-view scene flow estimation: A view centered variational ap- proach
Tali Basha, Yael Moses, and Nahum Kiryati. Multi-view scene flow estimation: A view centered variational ap- proach. International Journal of Computer Vision, 101(1):6– 21, 2013. 2
work page 2013
-
[4]
Aseem Behl, Omid Hosseini Jafari, Siva Karthik Mustikovela, Hassan Abu Alhaija, Carsten Rother, and Andreas Geiger. Bounding Boxes, Segmentations and Object Coordinates: How Important is Recognition for 3D Scene Flow Estimation in Autonomous Driving Scenarios? In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2574–2583, ...
work page 2017
-
[5]
Exploiting Single Image Depth Prediction for Mono-Stixel Estimation
Fabian Brickwedde, Steffen Abraham, and Rudolf Mester. Exploiting Single Image Depth Prediction for Mono-Stixel Estimation. In Proc. of European Conference of Computer Vision Workshops (ECCV Workshops). IEEE, 2018. 3, 7, 8
work page 2018
-
[6]
Mono-Stixels: Monocular Depth Reconstruction of Dy- namic Street Scenes
Fabian Brickwedde, Steffen Abraham, and Rudolf Mester. Mono-Stixels: Monocular Depth Reconstruction of Dy- namic Street Scenes. In Proc. of IEEE International Confer- ence on Robotics and Automation (ICRA), pages 1–7. IEEE,
-
[7]
Depth and scene flow from a single moving camera
Neil Brikbeck, Dana Cobzas, and Martin J ¨agersand. Depth and scene flow from a single moving camera. In Proc. of International Symposium on 3D Data Processing, Visualiza- tion and Transmission (3DPVT), 2010. 2
work page 2010
-
[8]
Sebastian Bullinger, Christoph Bodensteiner, Michael Arens, and Rainer Stiefelhagen. 3D Vehicle Trajectory Re- construction in Monocular Video Data Using Environment Structure Constraints. In Proc. of European Conference on Computer Vision (ECCV), pages 35–50, 2018. 1, 2
work page 2018
Show all 77 references
-
[9]
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognitio...
2016
-
[10]
Depth map prediction from a single image using a multi-scale deep net- work
David Eigen, Christian Puhrsch, and Rob Fergus. Depth map prediction from a single image using a multi-scale deep net- work. In Proc. of Advances in neural information processing systems (NeurIPS), pages 2366–2374, 2014. 1, 2
2014
-
[11]
LSD- SLAM: Large-scale direct monocular SLAM
Jakob Engel, Thomas Sch ¨ops, and Daniel Cremers. LSD- SLAM: Large-scale direct monocular SLAM. In Proc. of European Conference on Computer Vision (ECCV) , pages 834–849. Springer, 2014. 2
2014
-
[12]
Single-View and Multi-View Depth Fusion
F ´acil, Jos ´e M and Concha, Alejo and Montesano, Luis and Civera, Javier. Single-View and Multi-View Depth Fusion. IEEE Robotics and Automation Letters , 2(4):1994–2001, Oct 2017. 1, 2
1994
-
[13]
Predictive monocular odometry (PMO): What is possible without RANSAC and multiframe bundle adjustment? Image and Vision Computing, 2017
Nolang Fanani, Alina St ¨urck, Matthias Ochs, Henry Bradler, and Rudolf Mester. Predictive monocular odometry (PMO): What is possible without RANSAC and multiframe bundle adjustment? Image and Vision Computing, 2017. 2
2017
-
[14]
Deep Ordinal Regression Network for Monocular Depth Estimation
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Bat- manghelich, and Dacheng Tao. Deep Ordinal Regression Network for Monocular Depth Estimation. In Proc. of the IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), pages 2002–2011, 2018. 1, 2, 6, 7
2002
-
[15]
Unsupervised CNN for single view depth estimation: Geometry to the res- cue
Ravi Garg, Gustavo Carneiro, and Ian Reid. Unsupervised CNN for single view depth estimation: Geometry to the res- cue. In Proc. of European Conference on Computer Vision (ECCV), pages 740–756. Springer, 2016. 2
2016
-
[16]
Dense variational reconstruction of non-rigid surfaces from monoc- ular video
Ravi Garg, Anastasios Roussos, and Lourdes Agapito. Dense variational reconstruction of non-rigid surfaces from monoc- ular video. In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1272–1279, 2013. 2
2013
-
[17]
Lightweight Probabilistic Deep Networks
Jochen Gast and Stefan Roth. Lightweight Probabilistic Deep Networks. In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 3369–3378,
-
[18]
Stere- oscan: Dense 3D reconstruction in real-time
Andreas Geiger, Julius Ziegler, and Christoph Stiller. Stere- oscan: Dense 3D reconstruction in real-time. In Proc. of IEEE Intelligent Vehicles Symposium (IV) , pages 963–968,
-
[19]
Clement Godard, Oisin Mac Aodha, and Gabriel J. Brostow. Unsupervised Monocular Depth Estimation With Left-Right Consistency. In Proc. of IEEE Conference on Computer Vi- sion and Pattern Recognition (CVPR), July 2017. 1, 2, 6, 7, 8
2017
-
[20]
NRSfM-Flow: Recovering Non-Rigid Scene Flow from Monocular Image Sequences
Vladislav Golyanik, Aman S Mathur, and Didier Stricker. NRSfM-Flow: Recovering Non-Rigid Scene Flow from Monocular Image Sequences. In Proc. of British Machine Vision Conference (BMVC), 2016. 2
2016
-
[21]
On Calibration of Modern Neural Networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. On Calibration of Modern Neural Networks. In Proc. of In- ternational Conference on Machine Learning (ICML), pages 1321–1330, 2017. 2, 3
2017
-
[22]
Multiple view ge- ometry in computer vision
Richard Hartley and Andrew Zisserman. Multiple view ge- ometry in computer vision . Cambridge university press,
-
[23]
Mask R-CNN
Kaiming He, Georgia Gkioxari, Piotr Doll ´ar, and Ross Gir- shick. Mask R-CNN. In Proc. of IEEE International Confer- ence on Computer Vision (ICCV) , pages 2980–2988. IEEE,
-
[24]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proc. of the IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), pages 770–778, 2016. 3
2016
-
[25]
RGB-D flow: Dense 3-D motion estimation using color and depth
Evan Herbst, Xiaofeng Ren, and Dieter Fox. RGB-D flow: Dense 3-D motion estimation using color and depth. InProc. of IEEE International Conference on Robotics and Automa- tion (ICRA), pages 2276–2282, May 2013. 2
2013
-
[26]
Accurate and efficient stereo processing by semi-global matching and mutual information
Heiko Hirschmuller. Accurate and efficient stereo processing by semi-global matching and mutual information. In Proc. of IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR), pages 807–814, 2005. 3, 4, 8
2005
-
[27]
Auto- matic photo pop-up
Derek Hoiem, Alexei A Efros, and Martial Hebert. Auto- matic photo pop-up. In Proc. of ACM transactions on graph- ics (TOG), volume 24, pages 577–584. ACM, 2005. 2
2005
-
[28]
SphereFlow: 6 DoF scene flow from RGB-D pairs
Michael Hornacek, Andrew Fitzgibbon, and Carsten Rother. SphereFlow: 6 DoF scene flow from RGB-D pairs. In Proc. of IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR), pages 3526–3533, 2014. 8
2014
-
[29]
A variational method for scene flow estimation from stereo sequences
Fr ´ed´eric Huguet and Fr ´ed´eric Devernay. A variational method for scene flow estimation from stereo sequences. In Proc. of IEEE International Conference on Computer Vision, pages 1–7. IEEE, 2007. 2
2007
-
[30]
MirrorFlow: Exploiting symmetries in joint optical flow and occlusion estimation
Junhwa Hur and Stefan Roth. MirrorFlow: Exploiting symmetries in joint optical flow and occlusion estimation. In Proc. of International Conference on Computer Vision (ICCV), 2017. 7, 8
2017
-
[31]
Uncertainty Esti- mates and Multi-Hypotheses Networks for Optical Flow
Eddy Ilg, Ozgun Cicek, Silvio Galesso, Aaron Klein, Osama Makansi, Frank Hutter, and Thomas Brox. Uncertainty Esti- mates and Multi-Hypotheses Networks for Optical Flow. In Proc. of European Conference on Computer Vision (ECCV),
-
[32]
What uncertainties do we need in bayesian deep learning for computer vision? In Proc
Alex Kendall and Yarin Gal. What uncertainties do we need in bayesian deep learning for computer vision? In Proc. of Advances in neural information processing systems (NeurIPS), pages 5574–5584, 2017. 2, 3, 4, 6
2017
-
[33]
Adam: Amethod for stochastic optimization
Diederik P Kingma and Jimmy Lei Ba. Adam: Amethod for stochastic optimization. In Proc. of International Conference for Learning Representations (ICLR), 2014. 5
2014
-
[34]
Supervising the new with the old: learning SFM from SFM
Maria Klodt and Andrea Vedaldi. Supervising the new with the old: learning SFM from SFM. InProc. of European Con- ference on Computer Vision (ECCV), pages 698–713, 2018. 2, 3, 4
2018
-
[35]
Accurate Uncertainties for Deep Learning Using Calibrated Regression
V olodymyr Kuleshov, Nathan Fenner, and Stefano Ermon. Accurate Uncertainties for Deep Learning Using Calibrated Regression. In Proc. of International Conference on Ma- chine Learning (ICML), pages 2801–2809, 2018. 3
2018
-
[36]
Monoc- ular dense 3D reconstruction of a complex dynamic scene from two perspective frames
Suryansh Kumar, Yuchao Dai, and Hongdong Li. Monoc- ular dense 3D reconstruction of a complex dynamic scene from two perspective frames. In Proc. of IEEE International Conference on Computer Vision (ICCV), pages 4649–4657,
-
[37]
A Motion Free Approach to Dense Depth Estimation in Complex Dynamic Scene
Suryansh Kumar, Ram Srivatsav Ghorakavi, Yuchao Dai, and Hongdong Li. A Motion Free Approach to Dense Depth Estimation in Complex Dynamic Scene. arXiv preprint arXiv:1902.03791, 2019. 2, 7, 8
1902 arXiv
-
[38]
g 2 o: A general framework for graph optimization
Rainer K ¨ummerle, Giorgio Grisetti, Hauke Strasdat, Kurt Konolige, and Wolfram Burgard. g 2 o: A general framework for graph optimization. In Proc. of IEEE International Con- ference on Robotics and Automation (ICRA) , pages 3607–
-
[39]
Semi- supervised deep learning for monocular depth map predic- tion
Yevhen Kuznietsov, J ¨org St¨uckler, and Bastian Leibe. Semi- supervised deep learning for monocular depth map predic- tion. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 6647–6655, 2017. 6
2017
-
[40]
Pulling things out of perspective
Lubor Ladicky, Jianbo Shi, and Marc Pollefeys. Pulling things out of perspective. InProceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR) , pages 89–96, 2014. 2
2014
-
[41]
Deep rigid instance scene flow
Wei-Chiu Ma, Shenlong Wang, Rui Hu, Yuwen Xiong, and Raquel Urtasun. Deep rigid instance scene flow. In Proc. of IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), June 2019. 8
2019
-
[42]
Un- supervised Learning of Depth and Ego-Motion from Monoc- ular Video Using 3D Geometric Constraints
Reza Mahjourian, Martin Wicke, and Anelia Angelova. Un- supervised Learning of Depth and Ego-Motion from Monoc- ular Video Using 3D Geometric Constraints. In Proc. of IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR), pages 5667–5675, 2018. 2
2018
-
[43]
Predictive uncertainty esti- mation via prior networks
Andrey Malinin and Mark Gales. Predictive uncertainty esti- mation via prior networks. InProc. of Advances in Neural In- formation Processing Systems (NeurIPS), pages 7047–7058,
-
[44]
Object Scene Flow for Autonomous Vehicles
Moritz Menze and Andreas Geiger. Object Scene Flow for Autonomous Vehicles. InProc. of IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), 2015. 1, 2, 4, 5, 6, 7
2015
-
[45]
Ob- ject Scene Flow
Moritz Menze, Christian Heipke, and Andreas Geiger. Ob- ject Scene Flow. ISPRS Journal of Photogrammetry and Remote Sensing, 140:60 – 76, 2018. Geospatial Computer Vision. 2, 4, 5, 7
2018
-
[46]
Monocular Concurrent Recovery of Structure and Motion Scene Flow
Amar Mitiche, Yosra Mathlouthi, and Ismail Ben Ayed. Monocular Concurrent Recovery of Structure and Motion Scene Flow. Frontiers in ICT, 2:16, 2015. 1, 2
2015
-
[47]
ORB-SLAM: a versatile and accurate monocular SLAM system
Raul Mur-Artal, Jose Maria Martinez Montiel, and Juan D Tardos. ORB-SLAM: a versatile and accurate monocular SLAM system. IEEE Transactions on Robotics, 31(5):1147– 1163, 2015. 2
2015
-
[48]
Monocular Visual Odometry with Cyclic Estimation
Fabio Irigon Pereira, Gustavo Ilha, Joel Luft, Marcelo Ne- greiros, and Altamiro Susin. Monocular Visual Odometry with Cyclic Estimation. In Graphics, Patterns and Images (SIBGRAPI), 2017 30th SIBGRAPI Conference on, pages 1–
2017
-
[49]
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
John Platt. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Ad- vances in large margin classifiers, 10(3):61–74, 1999. 3
1999
-
[50]
Multi-view stereo reconstruction and scene flow estimation with a global image-based matching score
Jean-Philippe Pons, Renaud Keriven, and Olivier Faugeras. Multi-view stereo reconstruction and scene flow estimation with a global image-based matching score. International Journal of Computer Vision, 72(2):179–193, 2007. 2
2007
-
[51]
Dense monocular depth estimation in complex dy- namic scenes
Ren ´e Ranftl, Vibhav Vineet, Qifeng Chen, and Vladlen Koltun. Dense monocular depth estimation in complex dy- namic scenes. In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 4058–4066,
-
[52]
Learn- ing depth from single monocular images
Ashutosh Saxena, Sung H Chung, and Andrew Y Ng. Learn- ing depth from single monocular images. In Proc. of Ad- vances in Neural Information Processing Systems (NeurIPS), volume 18, pages 1–8, 2005. 2
2005
-
[53]
Make3D: Learning 3D Scene Structure from a Single Still Image
Ashutosh Saxena, Min Sun, and Andrew Y Ng. Make3D: Learning 3D Scene Structure from a Single Still Image. IEEE Transactions on Pattern Analysis and Machine Intel- ligence, 31(5):824–840, May 2009. 6
2009
-
[54]
CNN-SLAM: Real-time dense monocular SLAM with learned depth prediction
Keisuke Tateno, Federico Tombari, Iro Laina, and Nassir Navab. CNN-SLAM: Real-time dense monocular SLAM with learned depth prediction. In Proc. of the IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR) , volume 2, 2017. 1, 2
2017
-
[55]
Occlusion- Aware Unsupervised Learning of Monocular Depth, Optical Flow and Camera Pose with Geometric Constraints
Qianru Teng, Yimin Chen, and Chen Huang. Occlusion- Aware Unsupervised Learning of Monocular Depth, Optical Flow and Camera Pose with Geometric Constraints. Future Internet, 10(10):92, 2018. 2
2018
-
[56]
Sparsity Invariant CNNs
Jonas Uhrig, N Schneider, L Schneider, U Franke, Thomas Brox, and A Geiger. Sparsity Invariant CNNs. In Proc. of IEEE International Conference on 3D Vision (3DV), 2017. 2
2017
-
[57]
DeMoN: Depth and Motion Network for Learning Monocular Stereo
Benjamin Ummenhofer, Huizhong Zhou, Jonas Uhrig, Niko- laus Mayer, Eddy Ilg, Alexey Dosovitskiy, and Thomas Brox. DeMoN: Depth and Motion Network for Learning Monocular Stereo. In Proc. of IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), July 2017. 2
2017
-
[58]
Joint es- timation of motion, structure and geometry from stereo se- quences
Levi Valgaerts, Andr ´es Bruhn, Henning Zimmer, Joachim Weickert, Carsten Stoll, and Christian Theobalt. Joint es- timation of motion, structure and geometry from stereo se- quences. In Proc. of European Conference on Computer Vi- sion (ECCV), pages 568–581. Springer, 2010. 2
2010
-
[59]
Three-dimensional scene flow
Sundar Vedula, Simon Baker, Peter Rander, Robert Collins, and Takeo Kanade. Three-dimensional scene flow. In Proc. of IEEE International Conference on Computer Vision (ICCV), volume 2, pages 722–729. IEEE, 1999. 1, 2
1999
-
[60]
Three-dimensional scene flow
Sundar Vedula, Peter Rander, Robert Collins, and Takeo Kanade. Three-dimensional scene flow. IEEE transactions on pattern analysis and machine intelligence , 27(3):475– 480, 2005. 1, 2
2005
-
[61]
Piece- wise rigid scene flow
Christoph V ogel, Konrad Schindler, and Stefan Roth. Piece- wise rigid scene flow. In Proc. of the IEEE International Conference on Computer Vision (ICCV), pages 1377–1384,
-
[62]
Learning Depth from Monocular Videos using Direct Methods
Chaoyang Wang, Jos ´e Miguel Buenaposada, Rui Zhu, and Simon Lucey. Learning Depth from Monocular Videos using Direct Methods. In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 2022–2030,
2022
-
[63]
An MXNet implementation of Mask R-CNN
Naiyan Wang. An MXNet implementation of Mask R-CNN. https://github.com/TuSimple/mx-maskrcnn, 2018. [accessed June, 25 2018]. 5
2018
-
[64]
Stereoscopic scene flow computation for 3D motion understanding
Andreas Wedel, Thomas Brox, Tobi Vaudrey, Clemens Rabe, Uwe Franke, and Daniel Cremers. Stereoscopic scene flow computation for 3D motion understanding. International Journal of Computer Vision, 95(1):29–51, 2011. 2
2011
-
[65]
Efficient dense scene flow from sparse or dense stereo data
Andreas Wedel, Clemens Rabe, Tobi Vaudrey, Thomas Brox, Uwe Franke, and Daniel Cremers. Efficient dense scene flow from sparse or dense stereo data. In Proc. of Euro- pean Conference on Computer Vision (ECCV) , pages 739–
-
[66]
Monoc- ular scene flow estimation via variational method
Degui Xiao, Qiuwei Yang, Bing Yang, and Wei Wei. Monoc- ular scene flow estimation via variational method. Multime- dia Tools and Applications, 76(8):10575–10597, Apr 2017. 1, 2
2017
-
[67]
Robust monocular epipolar flow estimation
Koichiro Yamaguchi, David McAllester, and Raquel Urta- sun. Robust monocular epipolar flow estimation. In Proc. of IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), pages 1862–1869, 2013. 5
2013
-
[68]
Efficient Joint Segmentation, Occlusion Labeling, Stereo and Flow Estimation
Koichiro Yamaguchi, David McAllester, and Raquel Ur- tasun. Efficient Joint Segmentation, Occlusion Labeling, Stereo and Flow Estimation. In David Fleet, Tomas Pajdla, Bernt Schiele, and Tinne Tuytelaars, editors, Proc. of Euro- pean Conference on Computer Vision (ECCV) , pages...
2014
-
[69]
Deep virtual stereo odometry: Leveraging deep depth pre- diction for monocular direct sparse odometry
Nan Yang, Rui Wang, J ¨org St ¨uckler, and Daniel Cremers. Deep virtual stereo odometry: Leveraging deep depth pre- diction for monocular direct sparse odometry. In Proc. of European Conference on Computer Vision (ECCV) , pages 835–852. Springer, 2018. 2, 5
2018
-
[70]
Every Pixel Counts: Unsupervised Geom- etry Learning with Holistic 3D Motion Understanding
Zhenheng Yang, Peng Wang, Yang Wang, Wei Xu, and Ram Nevatia. Every Pixel Counts: Unsupervised Geom- etry Learning with Holistic 3D Motion Understanding. In Proc. of European Conference on Computer Vision Work- shops (ECCV Workshops), 2018. 2, 7
2018
-
[71]
Scale recovery for monocular visual odometry using depth estimated with deep convolutional neural fields
Xiaochuan Yin, Xiangwei Wang, Xiaoguo Du, and Qijun Chen. Scale recovery for monocular visual odometry using depth estimated with deep convolutional neural fields. In Proceedings of the IEEE International Conference on Com- puter Vision, pages 5870–5878, 2017. 1, 2, 5
2017
-
[72]
Hierarchical discrete distribution decomposition for match density esti- mation
Zhichao Yin, Trevor Darrell, and Fisher Yu. Hierarchical discrete distribution decomposition for match density esti- mation. In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2019. 7
2019
-
[73]
GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose
Zhichao Yin and Jianping Shi. GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose. In Proc. of IEEE Conference on Computer Vision and Pat- tern Recognition (CVPR), volume 2, 2018. 2, 7
2018
-
[74]
Non-parametric local trans- forms for computing visual correspondence
Ramin Zabih and John Woodfill. Non-parametric local trans- forms for computing visual correspondence. In Proc. of Eu- ropean Conference on Computer Vision (ECCV), pages 151–
-
[75]
Unsupervised Learn- ing of Monocular Depth Estimation and Visual Odometry with Deep Feature Reconstruction
Huangying Zhan, Ravi Garg, Chamara Saroj Weerasekera, Kejie Li, Harsh Agarwal, and Ian Reid. Unsupervised Learn- ing of Monocular Depth Estimation and Visual Odometry with Deep Feature Reconstruction. In Proc. of IEEE Confer- ence on Computer Vision and Pattern Recognition (CV...
2018
-
[76]
Tinghui Zhou, Matthew Brown, Noah Snavely, and David G. Lowe. Unsupervised Learning of Depth and Ego-Motion from Video. In Proc. of IEEE Conference on Computer Vi- sion and Pattern Recognition (CVPR), 2017. 2
2017
-
[77]
DF-Net: Un- supervised Joint Learning of Depth and Flow using Cross- Network Consistency
Yuliang Zou, Zelun Luo, and Jia-Bin Huang. DF-Net: Un- supervised Joint Learning of Depth and Flow using Cross- Network Consistency. In Proc. of European Conference on Computer Vision (ECCV), pages 36–53, 2018. 2, 7
2018
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.