Pith. sign in

REVIEW 3 major objections 5 minor 77 references

Mono-SF: Multi-View Geometry Meets Single-View Depth for Monocular Scene Flow Estimation of Dynamic Traffic Scenes

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A monocular camera plus probabilistic depth can estimate 3D scene flow in traffic scenes, reaching a combined scene flow error of 21.60 on KITTI.

desk verdict Solid monocular scene flow paper with a fixable sign error in the core homography that needs correcting before acceptance. read the letter →

arxiv 1908.06316 v1 pith:E4GQEWHW submitted 2019-08-17 cs.CV

classification cs.CV
keywords monocularsceneflowprobabilisticdepthestimationmixtureofGaussiansrecalibrationrigid-bodymotion3DplanesegmentationKITTIbenchmarkautonomousdriving
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Monocular scene flow — recovering each pixel's 3D position and 3D motion from two consecutive images of one camera — is ill-posed because scale is ambiguous. Mono-SF resolves this by fusing two sources: multi-view geometry, which warps the first image into the second via plane homographies, and single-view depth, supplied as per-pixel probability distributions by a network called ProbDepthNet. The paper claims that this joint probabilistic optimization, over 3D planes attached to rigid bodies, gives the best monocular scene flow result on the KITTI scene flow training set, with a combined scene flow error of 21.60. A second claim is that well-calibrated depth distributions matter: the ablation replacing probabilistic depth with point estimates raises the error, and removing recalibration raises it from 21.60 to 26.91. The reader should care because monocular cameras are cheaper and simpler than stereo rigs, so a working monocular scene flow pipeline is directly relevant to autonomous driving.

What carries the argument

The central object is the energy function over scaled plane normals $n_i$ and rigid-body motions $T_j \in SE(3)$. A pixel's 3D position is fixed by the scaled normal through $n_i^T X = 1$, and its motion is a homography $p_1 = K(R_j - t_j n_i^T) K^{-1} p_0$. This homography lets the same variables explain image warping and depth: the photometric term compares Census descriptors at the warped location, and the depth term evaluates the implied inverse depth $d_t$ under the ProbDepthNet density at both times. The optimization runs particle belief propagation over the normals and motions, with rigid-body poses initialized by jointly fitting sparse flow correspondences and ProbDepthNet depth.

What would settle it

Run Mono-SF on a KITTI sequence where a moving vehicle is deliberately not detected by the segmentation (for example, by withholding one object's ground-truth mask during evaluation); the optimization will assign that object's pixels to the background rigid body and output zero 3D motion, and the foreground scene flow error will jump by roughly the object's image-area fraction.

Watch

Extended reading notes

Core claim

Mono-SF models a traffic scene as piecewise 3D planes, each associated with a rigid body (the background or an instance detected by segmentation), and minimizes an energy with three terms: a Census-based photometric distance from warping the reference image into the next frame using the plane normal and the rigid motion; a negative log-likelihood term that scores the plane's implied inverse depth against ProbDepthNet's mixture-of-Gaussians depth densities at both timestamps; and pairwise smoothness priors on depth and orientation. ProbDepthNet is trained to output per-pixel inverse-depth distributions; its CalibNet subnetwork, trained on a separate split, rescales the variance and mixture weights to counter overconfident estimates. The paper reports a scene flow error of 21.60 on the KITTI scene flow training set, the best among the monocular methods compared, with ablations showing that each energy term and the recalibration step contributes to the final result.

Load-bearing premise

The method assumes every moving object is detected by an instance segmentation network and moves as one rigid body; a missed object is treated as static, so the optimization cannot recover its motion.

Editorial extensions

If this is right

  • A monocular camera can produce 3D scene flow competitive with stereo-based methods, provided the single-view depth uncertainty is well calibrated.
  • The warp-consistency energy could serve as a training signal for a network that predicts depth and motion, potentially replacing the iterative optimization at test time.
  • The reported ablation shows that probabilistic depth distributions are not a minor refinement: using the distribution instead of the mean lowers the combined scene flow error by several points.
  • The plane-plus-rigid-body representation yields per-object 6D motions and planar surfaces that downstream planning modules can consume directly.
  • The CalibNet recalibration approach applies to other probabilistic regression networks, including multi-hypothesis and probabilistic-layer alternatives, improving their calibration on the same data.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The recalibration step is generic: any regression network that outputs a distribution and overfits its training split could adopt a separately trained rescaling subnetwork, not just depth networks.
  • Running at about 41 seconds per image on a single CPU, Mono-SF is a proof of concept rather than a real-time system; a natural next step is distilling the optimized scene flow into a feed-forward network to remove the iterative loop.
  • Because a missed segmentation is structurally fatal, a testable extension is to add a fallback that re-estimates pixels with high photometric residual as new rigid bodies rather than letting them remain in the background.
  • The same formulation could transfer to stereo inputs by replacing the monocular depth term with a stereo disparity consistency term, although the paper does not test this variant.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents Mono-SF, a monocular scene flow method for dynamic street scenes. The scene is modeled as superpixel-based 3D planes, each assigned to a rigid body (the background or a Mask R-CNN instance). Mono-SF minimizes an energy with (i) a Census-based photometric term that warps the reference image into the next frame via a plane-induced homography, (ii) a term penalizing inconsistency with pixel-wise depth distributions from a proposed ProbDepthNet, and (iii) pairwise smoothness priors. ProbDepthNet outputs a mixture of Gaussians over inverse depth and includes CalibNet, a small network trained on a hold-out split to recalibrate variances and weights. Experiments on the KITTI scene flow training set report the best SF-all score (21.60) among the monocular baselines compared, and ablations support the benefit of probabilistic, recalibrated depth and of each energy term. The paper also reports the first monocular submission to the KITTI scene flow benchmark.

Significance. If the results withstand scrutiny, the contribution is solid and useful: a first monocular entry on the KITTI scene flow benchmark, a clean probabilistic integration of single-view depth into a plane-and-rigid-body scene flow framework, and a simple recalibration idea (CalibNet) that is shown to help across several probabilistic depth formulations on KITTI. The ablation studies in Tables 3 and 4 are a strength because they isolate the contributions of the proposed components. Nevertheless, the published derivation of the central homography is inconsistent with the stated plane convention, and the comparative evaluation lacks error bars and sensitivity analysis, so I cannot regard the claims as fully established in the present form.

major comments (3)
  1. [Sec. 3.2, Eq. (5)] The homography is printed as K(R_j - t_j n_i^T)K^{-1}, but the model states that each plane satisfies n_i^T X = 1. For a point X_0 on that plane, X_1 = R_j X_0 + t_j = (R_j + t_j n_i^T)X_0, so the induced homography is K(R_j + t_j n_i^T)K^{-1}. The minus sign is inconsistent with the stated convention and would change the warp for any nontrivial translation, including the fronto-parallel translation case. Because Eq. (5) defines the correspondences used in both Phi_pho and the t=1 part of Phi_svd, this is a load-bearing error in the published derivation. Please correct the equation or explicitly introduce a different plane convention, and confirm that the experiments were run with the corrected form.
  2. [Sec. 4.2, Table 1] The central empirical claim that Mono-SF outperforms state-of-the-art monocular baselines on scene flow is supported only by single point estimates on one dataset split. No error bars, per-sequence variance, or significance tests are reported. Given that the energy in Eq. (3) depends on hand-set weights Theta_0..Theta_4 and truncation thresholds tau_0..tau_2 (Sec. 3.2), the authors should provide a sensitivity analysis or scene-wise statistics to demonstrate that the reported margins are not an artifact of parameter tuning.
  3. [Sec. 3.2, Initialization and Table 1] The method assumes that every moving object is detected by Mask R-CNN and successfully paired across frames by sparse-flow voting; any missed instance is treated as static and its motion is not estimated. The paper does not report how often this occurs on the KITTI scene flow set or how the foreground (fg) and scene-flow (SF) metrics depend on segmentation and pairing quality. Please add such an analysis, or qualify the claims to make this dependency explicit.
minor comments (5)
  1. [Eq. (9)] Please use an explicit dot product notation, e.g., n_k · n_l, instead of |n_k n_l|, which is ambiguous.
  2. [Sec. 4.2, Table 2] Table 2 lists stereo-based methods without a clear statement that the comparison is not like-for-like because the sensor input differs; the text should state this explicitly.
  3. [General] The paper refers to supplementary material at several points, but no supplementary material is included in the submission; either include it or remove these references.
  4. [General] A statement on code availability would materially help reproducibility, especially in light of the Eq. (5) sign issue.
  5. [Sec. 4.2, Table 1] In Table 1, DMDE, S. Soup, and MFA are only evaluated on MRE; the statement that Mono-SF shows the best rating on most metrics should be narrowed to the metrics actually reported for all compared methods.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity; the central scene flow claim is benchmarked against independent KITTI ground truth and the depth network is trained on LiDAR/SGM data, not on the scene flow output.

full rationale

The paper's central claim, that Mono-SF outperforms monocular baselines on KITTI scene flow, is supported by comparisons to independent ground-truth data (Table 1) and to external published baselines. The single-view depth network, ProbDepthNet, is trained with a negative log-likelihood loss against SGM-completed LiDAR depth (Eq. 2), not against the scene flow result, so the scene flow accuracy is not an input recycled as an output. The ablations in Tables 3 and 4 compare variants of the same pipeline against the same external ground truth, which is a standard ablation and not a self-definitional prediction. The self-citations to prior Mono-Stixels work are used as a baseline and as methodological inspiration, not as an unverified load-bearing premise. No fitted parameter is renamed as a prediction; the CalibNet recalibration is trained on a hold-out split and evaluated by calibration curves and NLL. The apparent sign inconsistency in Eq. (5) noted by the skeptic is a mathematical correctness or reproducibility concern, not a circularity: it does not reduce the derived scene flow to the method's own inputs by construction. Overall, the derivation chain is self-contained against external benchmarks, and no circular step can be exhibited.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The Mono-SF result depends on a number of hand-set hyperparameters (energy weights, thresholds, particle counts) for which no sensitivity analysis is provided. The domain assumptions are the piecewise-planar scene model and the reliance on instance segmentation and single-view depth generalization. No new physical entities are introduced, so the invented_entities list is empty; CalibNet is a learned module, not a new entity with independent falsifiable evidence.

free parameters (5)
  • Number of Gaussian components K = 8
    Chosen architecture hyperparameter; not justified by an experiment.
  • Energy weights Theta_0..Theta_4 = not reported
    Hand-tuned weights for photometric, depth, and smoothness terms; no sensitivity analysis.
  • Truncation thresholds tau_0, tau_1, tau_2 = not reported
    Hand-tuned clipping values in the energy terms.
  • Belief propagation particles/iterations = 5 motion particles, 10 plane particles, 10 iterations
    Selected to balance runtime and accuracy; no ablation.
  • Training schedule for ProbDepthNet = 15 epochs, LR 1e-4 halved every 5, batch 4, image 512x256
    Standard training choices; not the central claim.
assumptions (5)
  • domain assumption Traffic scenes are approximated by piecewise planar surface elements and rigid bodies.
    Sec. 3.2 model description.
  • domain assumption Each detected object instance moves with a single 6D rigid body motion.
    Sec. 3.2 and Intro; pedestrians are assumed to be approximated by their dominant rigid motion.
  • domain assumption Mask R-CNN instance segmentation and sparse flow correspondences correctly identify and pair moving objects.
    Sec. 3.2 Initialization; if segmentation fails, object motion is not estimated.
  • domain assumption The depth distribution estimated by ProbDepthNet generalizes from its training split (KITTI raw) to the evaluation set.
    Sec. 4.1; training and evaluation are both KITTI-derived.
  • standard math Standard pinhole camera and SE(3) motion models apply.
    Used in homography warping Eq. 5.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Mono-SF: Multi-View Geometry Meets Single-View Depth for Monocular Scene Flow Estimation of Dynamic Traffic Scenes." pith.science (2026). https://pith.science/paper/E4GQEWHW

@misc{pith2026190806316,
  author       = {Pith},
  title        = {Pith review of: Mono-SF: Multi-View Geometry Meets Single-View Depth for Monocular Scene Flow Estimation of Dynamic Traffic Scenes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E4GQEWHW}},
  note         = {Machine review of arXiv:1908.06316}
}
read the original abstract

Existing 3D scene flow estimation methods provide the 3D geometry and 3D motion of a scene and gain a lot of interest, for example in the context of autonomous driving. These methods are traditionally based on a temporal series of stereo images. In this paper, we propose a novel monocular 3D scene flow estimation method, called Mono-SF. Mono-SF jointly estimates the 3D structure and motion of the scene by combining multi-view geometry and single-view depth information. Mono-SF considers that the scene flow should be consistent in terms of warping the reference image in the consecutive image based on the principles of multi-view geometry. For integrating single-view depth in a statistical manner, a convolutional neural network, called ProbDepthNet, is proposed. ProbDepthNet estimates pixel-wise depth distributions from a single image rather than single depth values. Additionally, as part of ProbDepthNet, a novel recalibration technique for regression problems is proposed to ensure well-calibrated distributions. Our experiments show that Mono-SF outperforms state-of-the-art monocular baselines and ablation studies support the Mono-SF approach and ProbDepthNet design.

Figures

Figures reproduced from arXiv: 1908.06316 by the authors.

Figure 1
Figure 1. Overview of Mono-SF for monocular scene flow estima [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of ProbDepthNet for probabilistic single-view [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Variables of Mono-SF model and energy minimization [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Exemplary estimates of ProbDepthNet on KITTI scene flow set [ [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Generalization of ProbDepthNet (trained on KITTI) on [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Top: Mean negative log-likelihood (NLL) of ProbDepth [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: Exemplary qualitative results of monocular scene flow estimation methods on the KITTI scene flow training set [ [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: Exemplary qualitative result of Mono-SF on a crop of [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

77 extracted references · 76 canonical work pages

  1. [1]

    Exploiting semantic information and deep matching for op- tical flow

    Min Bai, Wenjie Luo, Kaustav Kundu, and Raquel Urtasun. Exploiting semantic information and deep matching for op- tical flow. In Proc. of European Conference on Computer Vision (ECCV), pages 154–170. Springer, 2016. 5

  2. [2]

    Driven to distraction: Self-supervised distractor learning for robust monocular visual odometry in urban en- vironments

    Dan Barnes, Will Maddern, Geoffrey Pascoe, and Ingmar Posner. Driven to distraction: Self-supervised distractor learning for robust monocular visual odometry in urban en- vironments. In Proc. of IEEE International Conference on Robotics and Automation (ICRA) , pages 1894–1900. IEEE,

  3. [3]

    Multi-view scene flow estimation: A view centered variational ap- proach

    Tali Basha, Yael Moses, and Nahum Kiryati. Multi-view scene flow estimation: A view centered variational ap- proach. International Journal of Computer Vision, 101(1):6– 21, 2013. 2

  4. [4]

    Bounding Boxes, Segmentations and Object Coordinates: How Important is Recognition for 3D Scene Flow Estimation in Autonomous Driving Scenarios? In Proc

    Aseem Behl, Omid Hosseini Jafari, Siva Karthik Mustikovela, Hassan Abu Alhaija, Carsten Rother, and Andreas Geiger. Bounding Boxes, Segmentations and Object Coordinates: How Important is Recognition for 3D Scene Flow Estimation in Autonomous Driving Scenarios? In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2574–2583, ...

  5. [5]

    Exploiting Single Image Depth Prediction for Mono-Stixel Estimation

    Fabian Brickwedde, Steffen Abraham, and Rudolf Mester. Exploiting Single Image Depth Prediction for Mono-Stixel Estimation. In Proc. of European Conference of Computer Vision Workshops (ECCV Workshops). IEEE, 2018. 3, 7, 8

  6. [6]

    Mono-Stixels: Monocular Depth Reconstruction of Dy- namic Street Scenes

    Fabian Brickwedde, Steffen Abraham, and Rudolf Mester. Mono-Stixels: Monocular Depth Reconstruction of Dy- namic Street Scenes. In Proc. of IEEE International Confer- ence on Robotics and Automation (ICRA), pages 1–7. IEEE,

  7. [7]

    Depth and scene flow from a single moving camera

    Neil Brikbeck, Dana Cobzas, and Martin J ¨agersand. Depth and scene flow from a single moving camera. In Proc. of International Symposium on 3D Data Processing, Visualiza- tion and Transmission (3DPVT), 2010. 2

  8. [8]

    3D Vehicle Trajectory Re- construction in Monocular Video Data Using Environment Structure Constraints

    Sebastian Bullinger, Christoph Bodensteiner, Michael Arens, and Rainer Stiefelhagen. 3D Vehicle Trajectory Re- construction in Monocular Video Data Using Environment Structure Constraints. In Proc. of European Conference on Computer Vision (ECCV), pages 35–50, 2018. 1, 2

Show all 77 references
  1. [9]

    The cityscapes dataset for semantic urban scene understanding

    Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognitio...

  2. [10]

    Depth map prediction from a single image using a multi-scale deep net- work

    David Eigen, Christian Puhrsch, and Rob Fergus. Depth map prediction from a single image using a multi-scale deep net- work. In Proc. of Advances in neural information processing systems (NeurIPS), pages 2366–2374, 2014. 1, 2

  3. [11]

    LSD- SLAM: Large-scale direct monocular SLAM

    Jakob Engel, Thomas Sch ¨ops, and Daniel Cremers. LSD- SLAM: Large-scale direct monocular SLAM. In Proc. of European Conference on Computer Vision (ECCV) , pages 834–849. Springer, 2014. 2

  4. [12]

    Single-View and Multi-View Depth Fusion

    F ´acil, Jos ´e M and Concha, Alejo and Montesano, Luis and Civera, Javier. Single-View and Multi-View Depth Fusion. IEEE Robotics and Automation Letters , 2(4):1994–2001, Oct 2017. 1, 2

  5. [13]

    Predictive monocular odometry (PMO): What is possible without RANSAC and multiframe bundle adjustment? Image and Vision Computing, 2017

    Nolang Fanani, Alina St ¨urck, Matthias Ochs, Henry Bradler, and Rudolf Mester. Predictive monocular odometry (PMO): What is possible without RANSAC and multiframe bundle adjustment? Image and Vision Computing, 2017. 2

  6. [14]

    Deep Ordinal Regression Network for Monocular Depth Estimation

    Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Bat- manghelich, and Dacheng Tao. Deep Ordinal Regression Network for Monocular Depth Estimation. In Proc. of the IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), pages 2002–2011, 2018. 1, 2, 6, 7

  7. [15]

    Unsupervised CNN for single view depth estimation: Geometry to the res- cue

    Ravi Garg, Gustavo Carneiro, and Ian Reid. Unsupervised CNN for single view depth estimation: Geometry to the res- cue. In Proc. of European Conference on Computer Vision (ECCV), pages 740–756. Springer, 2016. 2

  8. [16]

    Dense variational reconstruction of non-rigid surfaces from monoc- ular video

    Ravi Garg, Anastasios Roussos, and Lourdes Agapito. Dense variational reconstruction of non-rigid surfaces from monoc- ular video. In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1272–1279, 2013. 2

  9. [17]

    Lightweight Probabilistic Deep Networks

    Jochen Gast and Stefan Roth. Lightweight Probabilistic Deep Networks. In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 3369–3378,

  10. [18]

    Stere- oscan: Dense 3D reconstruction in real-time

    Andreas Geiger, Julius Ziegler, and Christoph Stiller. Stere- oscan: Dense 3D reconstruction in real-time. In Proc. of IEEE Intelligent Vehicles Symposium (IV) , pages 963–968,

  11. [19]

    Clement Godard, Oisin Mac Aodha, and Gabriel J. Brostow. Unsupervised Monocular Depth Estimation With Left-Right Consistency. In Proc. of IEEE Conference on Computer Vi- sion and Pattern Recognition (CVPR), July 2017. 1, 2, 6, 7, 8

  12. [20]

    NRSfM-Flow: Recovering Non-Rigid Scene Flow from Monocular Image Sequences

    Vladislav Golyanik, Aman S Mathur, and Didier Stricker. NRSfM-Flow: Recovering Non-Rigid Scene Flow from Monocular Image Sequences. In Proc. of British Machine Vision Conference (BMVC), 2016. 2

  13. [21]

    On Calibration of Modern Neural Networks

    Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. On Calibration of Modern Neural Networks. In Proc. of In- ternational Conference on Machine Learning (ICML), pages 1321–1330, 2017. 2, 3

  14. [22]

    Multiple view ge- ometry in computer vision

    Richard Hartley and Andrew Zisserman. Multiple view ge- ometry in computer vision . Cambridge university press,

  15. [23]

    Mask R-CNN

    Kaiming He, Georgia Gkioxari, Piotr Doll ´ar, and Ross Gir- shick. Mask R-CNN. In Proc. of IEEE International Confer- ence on Computer Vision (ICCV) , pages 2980–2988. IEEE,

  16. [24]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proc. of the IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), pages 770–778, 2016. 3

  17. [25]

    RGB-D flow: Dense 3-D motion estimation using color and depth

    Evan Herbst, Xiaofeng Ren, and Dieter Fox. RGB-D flow: Dense 3-D motion estimation using color and depth. InProc. of IEEE International Conference on Robotics and Automa- tion (ICRA), pages 2276–2282, May 2013. 2

  18. [26]

    Accurate and efficient stereo processing by semi-global matching and mutual information

    Heiko Hirschmuller. Accurate and efficient stereo processing by semi-global matching and mutual information. In Proc. of IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR), pages 807–814, 2005. 3, 4, 8

  19. [27]

    Auto- matic photo pop-up

    Derek Hoiem, Alexei A Efros, and Martial Hebert. Auto- matic photo pop-up. In Proc. of ACM transactions on graph- ics (TOG), volume 24, pages 577–584. ACM, 2005. 2

  20. [28]

    SphereFlow: 6 DoF scene flow from RGB-D pairs

    Michael Hornacek, Andrew Fitzgibbon, and Carsten Rother. SphereFlow: 6 DoF scene flow from RGB-D pairs. In Proc. of IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR), pages 3526–3533, 2014. 8

  21. [29]

    A variational method for scene flow estimation from stereo sequences

    Fr ´ed´eric Huguet and Fr ´ed´eric Devernay. A variational method for scene flow estimation from stereo sequences. In Proc. of IEEE International Conference on Computer Vision, pages 1–7. IEEE, 2007. 2

  22. [30]

    MirrorFlow: Exploiting symmetries in joint optical flow and occlusion estimation

    Junhwa Hur and Stefan Roth. MirrorFlow: Exploiting symmetries in joint optical flow and occlusion estimation. In Proc. of International Conference on Computer Vision (ICCV), 2017. 7, 8

  23. [31]

    Uncertainty Esti- mates and Multi-Hypotheses Networks for Optical Flow

    Eddy Ilg, Ozgun Cicek, Silvio Galesso, Aaron Klein, Osama Makansi, Frank Hutter, and Thomas Brox. Uncertainty Esti- mates and Multi-Hypotheses Networks for Optical Flow. In Proc. of European Conference on Computer Vision (ECCV),

  24. [32]

    What uncertainties do we need in bayesian deep learning for computer vision? In Proc

    Alex Kendall and Yarin Gal. What uncertainties do we need in bayesian deep learning for computer vision? In Proc. of Advances in neural information processing systems (NeurIPS), pages 5574–5584, 2017. 2, 3, 4, 6

  25. [33]

    Adam: Amethod for stochastic optimization

    Diederik P Kingma and Jimmy Lei Ba. Adam: Amethod for stochastic optimization. In Proc. of International Conference for Learning Representations (ICLR), 2014. 5

  26. [34]

    Supervising the new with the old: learning SFM from SFM

    Maria Klodt and Andrea Vedaldi. Supervising the new with the old: learning SFM from SFM. InProc. of European Con- ference on Computer Vision (ECCV), pages 698–713, 2018. 2, 3, 4

  27. [35]

    Accurate Uncertainties for Deep Learning Using Calibrated Regression

    V olodymyr Kuleshov, Nathan Fenner, and Stefano Ermon. Accurate Uncertainties for Deep Learning Using Calibrated Regression. In Proc. of International Conference on Ma- chine Learning (ICML), pages 2801–2809, 2018. 3

  28. [36]

    Monoc- ular dense 3D reconstruction of a complex dynamic scene from two perspective frames

    Suryansh Kumar, Yuchao Dai, and Hongdong Li. Monoc- ular dense 3D reconstruction of a complex dynamic scene from two perspective frames. In Proc. of IEEE International Conference on Computer Vision (ICCV), pages 4649–4657,

  29. [37]

    A Motion Free Approach to Dense Depth Estimation in Complex Dynamic Scene

    Suryansh Kumar, Ram Srivatsav Ghorakavi, Yuchao Dai, and Hongdong Li. A Motion Free Approach to Dense Depth Estimation in Complex Dynamic Scene. arXiv preprint arXiv:1902.03791, 2019. 2, 7, 8

  30. [38]

    g 2 o: A general framework for graph optimization

    Rainer K ¨ummerle, Giorgio Grisetti, Hauke Strasdat, Kurt Konolige, and Wolfram Burgard. g 2 o: A general framework for graph optimization. In Proc. of IEEE International Con- ference on Robotics and Automation (ICRA) , pages 3607–

  31. [39]

    Semi- supervised deep learning for monocular depth map predic- tion

    Yevhen Kuznietsov, J ¨org St¨uckler, and Bastian Leibe. Semi- supervised deep learning for monocular depth map predic- tion. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 6647–6655, 2017. 6

  32. [40]

    Pulling things out of perspective

    Lubor Ladicky, Jianbo Shi, and Marc Pollefeys. Pulling things out of perspective. InProceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR) , pages 89–96, 2014. 2

  33. [41]

    Deep rigid instance scene flow

    Wei-Chiu Ma, Shenlong Wang, Rui Hu, Yuwen Xiong, and Raquel Urtasun. Deep rigid instance scene flow. In Proc. of IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), June 2019. 8

  34. [42]

    Un- supervised Learning of Depth and Ego-Motion from Monoc- ular Video Using 3D Geometric Constraints

    Reza Mahjourian, Martin Wicke, and Anelia Angelova. Un- supervised Learning of Depth and Ego-Motion from Monoc- ular Video Using 3D Geometric Constraints. In Proc. of IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR), pages 5667–5675, 2018. 2

  35. [43]

    Predictive uncertainty esti- mation via prior networks

    Andrey Malinin and Mark Gales. Predictive uncertainty esti- mation via prior networks. InProc. of Advances in Neural In- formation Processing Systems (NeurIPS), pages 7047–7058,

  36. [44]

    Object Scene Flow for Autonomous Vehicles

    Moritz Menze and Andreas Geiger. Object Scene Flow for Autonomous Vehicles. InProc. of IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), 2015. 1, 2, 4, 5, 6, 7

  37. [45]

    Ob- ject Scene Flow

    Moritz Menze, Christian Heipke, and Andreas Geiger. Ob- ject Scene Flow. ISPRS Journal of Photogrammetry and Remote Sensing, 140:60 – 76, 2018. Geospatial Computer Vision. 2, 4, 5, 7

  38. [46]

    Monocular Concurrent Recovery of Structure and Motion Scene Flow

    Amar Mitiche, Yosra Mathlouthi, and Ismail Ben Ayed. Monocular Concurrent Recovery of Structure and Motion Scene Flow. Frontiers in ICT, 2:16, 2015. 1, 2

  39. [47]

    ORB-SLAM: a versatile and accurate monocular SLAM system

    Raul Mur-Artal, Jose Maria Martinez Montiel, and Juan D Tardos. ORB-SLAM: a versatile and accurate monocular SLAM system. IEEE Transactions on Robotics, 31(5):1147– 1163, 2015. 2

  40. [48]

    Monocular Visual Odometry with Cyclic Estimation

    Fabio Irigon Pereira, Gustavo Ilha, Joel Luft, Marcelo Ne- greiros, and Altamiro Susin. Monocular Visual Odometry with Cyclic Estimation. In Graphics, Patterns and Images (SIBGRAPI), 2017 30th SIBGRAPI Conference on, pages 1–

  41. [49]

    Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods

    John Platt. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Ad- vances in large margin classifiers, 10(3):61–74, 1999. 3

  42. [50]

    Multi-view stereo reconstruction and scene flow estimation with a global image-based matching score

    Jean-Philippe Pons, Renaud Keriven, and Olivier Faugeras. Multi-view stereo reconstruction and scene flow estimation with a global image-based matching score. International Journal of Computer Vision, 72(2):179–193, 2007. 2

  43. [51]

    Dense monocular depth estimation in complex dy- namic scenes

    Ren ´e Ranftl, Vibhav Vineet, Qifeng Chen, and Vladlen Koltun. Dense monocular depth estimation in complex dy- namic scenes. In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 4058–4066,

  44. [52]

    Learn- ing depth from single monocular images

    Ashutosh Saxena, Sung H Chung, and Andrew Y Ng. Learn- ing depth from single monocular images. In Proc. of Ad- vances in Neural Information Processing Systems (NeurIPS), volume 18, pages 1–8, 2005. 2

  45. [53]

    Make3D: Learning 3D Scene Structure from a Single Still Image

    Ashutosh Saxena, Min Sun, and Andrew Y Ng. Make3D: Learning 3D Scene Structure from a Single Still Image. IEEE Transactions on Pattern Analysis and Machine Intel- ligence, 31(5):824–840, May 2009. 6

  46. [54]

    CNN-SLAM: Real-time dense monocular SLAM with learned depth prediction

    Keisuke Tateno, Federico Tombari, Iro Laina, and Nassir Navab. CNN-SLAM: Real-time dense monocular SLAM with learned depth prediction. In Proc. of the IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR) , volume 2, 2017. 1, 2

  47. [55]

    Occlusion- Aware Unsupervised Learning of Monocular Depth, Optical Flow and Camera Pose with Geometric Constraints

    Qianru Teng, Yimin Chen, and Chen Huang. Occlusion- Aware Unsupervised Learning of Monocular Depth, Optical Flow and Camera Pose with Geometric Constraints. Future Internet, 10(10):92, 2018. 2

  48. [56]

    Sparsity Invariant CNNs

    Jonas Uhrig, N Schneider, L Schneider, U Franke, Thomas Brox, and A Geiger. Sparsity Invariant CNNs. In Proc. of IEEE International Conference on 3D Vision (3DV), 2017. 2

  49. [57]

    DeMoN: Depth and Motion Network for Learning Monocular Stereo

    Benjamin Ummenhofer, Huizhong Zhou, Jonas Uhrig, Niko- laus Mayer, Eddy Ilg, Alexey Dosovitskiy, and Thomas Brox. DeMoN: Depth and Motion Network for Learning Monocular Stereo. In Proc. of IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), July 2017. 2

  50. [58]

    Joint es- timation of motion, structure and geometry from stereo se- quences

    Levi Valgaerts, Andr ´es Bruhn, Henning Zimmer, Joachim Weickert, Carsten Stoll, and Christian Theobalt. Joint es- timation of motion, structure and geometry from stereo se- quences. In Proc. of European Conference on Computer Vi- sion (ECCV), pages 568–581. Springer, 2010. 2

  51. [59]

    Three-dimensional scene flow

    Sundar Vedula, Simon Baker, Peter Rander, Robert Collins, and Takeo Kanade. Three-dimensional scene flow. In Proc. of IEEE International Conference on Computer Vision (ICCV), volume 2, pages 722–729. IEEE, 1999. 1, 2

  52. [60]

    Three-dimensional scene flow

    Sundar Vedula, Peter Rander, Robert Collins, and Takeo Kanade. Three-dimensional scene flow. IEEE transactions on pattern analysis and machine intelligence , 27(3):475– 480, 2005. 1, 2

  53. [61]

    Piece- wise rigid scene flow

    Christoph V ogel, Konrad Schindler, and Stefan Roth. Piece- wise rigid scene flow. In Proc. of the IEEE International Conference on Computer Vision (ICCV), pages 1377–1384,

  54. [62]

    Learning Depth from Monocular Videos using Direct Methods

    Chaoyang Wang, Jos ´e Miguel Buenaposada, Rui Zhu, and Simon Lucey. Learning Depth from Monocular Videos using Direct Methods. In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 2022–2030,

  55. [63]

    An MXNet implementation of Mask R-CNN

    Naiyan Wang. An MXNet implementation of Mask R-CNN. https://github.com/TuSimple/mx-maskrcnn, 2018. [accessed June, 25 2018]. 5

  56. [64]

    Stereoscopic scene flow computation for 3D motion understanding

    Andreas Wedel, Thomas Brox, Tobi Vaudrey, Clemens Rabe, Uwe Franke, and Daniel Cremers. Stereoscopic scene flow computation for 3D motion understanding. International Journal of Computer Vision, 95(1):29–51, 2011. 2

  57. [65]

    Efficient dense scene flow from sparse or dense stereo data

    Andreas Wedel, Clemens Rabe, Tobi Vaudrey, Thomas Brox, Uwe Franke, and Daniel Cremers. Efficient dense scene flow from sparse or dense stereo data. In Proc. of Euro- pean Conference on Computer Vision (ECCV) , pages 739–

  58. [66]

    Monoc- ular scene flow estimation via variational method

    Degui Xiao, Qiuwei Yang, Bing Yang, and Wei Wei. Monoc- ular scene flow estimation via variational method. Multime- dia Tools and Applications, 76(8):10575–10597, Apr 2017. 1, 2

  59. [67]

    Robust monocular epipolar flow estimation

    Koichiro Yamaguchi, David McAllester, and Raquel Urta- sun. Robust monocular epipolar flow estimation. In Proc. of IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), pages 1862–1869, 2013. 5

  60. [68]

    Efficient Joint Segmentation, Occlusion Labeling, Stereo and Flow Estimation

    Koichiro Yamaguchi, David McAllester, and Raquel Ur- tasun. Efficient Joint Segmentation, Occlusion Labeling, Stereo and Flow Estimation. In David Fleet, Tomas Pajdla, Bernt Schiele, and Tinne Tuytelaars, editors, Proc. of Euro- pean Conference on Computer Vision (ECCV) , pages...

  61. [69]

    Deep virtual stereo odometry: Leveraging deep depth pre- diction for monocular direct sparse odometry

    Nan Yang, Rui Wang, J ¨org St ¨uckler, and Daniel Cremers. Deep virtual stereo odometry: Leveraging deep depth pre- diction for monocular direct sparse odometry. In Proc. of European Conference on Computer Vision (ECCV) , pages 835–852. Springer, 2018. 2, 5

  62. [70]

    Every Pixel Counts: Unsupervised Geom- etry Learning with Holistic 3D Motion Understanding

    Zhenheng Yang, Peng Wang, Yang Wang, Wei Xu, and Ram Nevatia. Every Pixel Counts: Unsupervised Geom- etry Learning with Holistic 3D Motion Understanding. In Proc. of European Conference on Computer Vision Work- shops (ECCV Workshops), 2018. 2, 7

  63. [71]

    Scale recovery for monocular visual odometry using depth estimated with deep convolutional neural fields

    Xiaochuan Yin, Xiangwei Wang, Xiaoguo Du, and Qijun Chen. Scale recovery for monocular visual odometry using depth estimated with deep convolutional neural fields. In Proceedings of the IEEE International Conference on Com- puter Vision, pages 5870–5878, 2017. 1, 2, 5

  64. [72]

    Hierarchical discrete distribution decomposition for match density esti- mation

    Zhichao Yin, Trevor Darrell, and Fisher Yu. Hierarchical discrete distribution decomposition for match density esti- mation. In Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2019. 7

  65. [73]

    GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose

    Zhichao Yin and Jianping Shi. GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose. In Proc. of IEEE Conference on Computer Vision and Pat- tern Recognition (CVPR), volume 2, 2018. 2, 7

  66. [74]

    Non-parametric local trans- forms for computing visual correspondence

    Ramin Zabih and John Woodfill. Non-parametric local trans- forms for computing visual correspondence. In Proc. of Eu- ropean Conference on Computer Vision (ECCV), pages 151–

  67. [75]

    Unsupervised Learn- ing of Monocular Depth Estimation and Visual Odometry with Deep Feature Reconstruction

    Huangying Zhan, Ravi Garg, Chamara Saroj Weerasekera, Kejie Li, Harsh Agarwal, and Ian Reid. Unsupervised Learn- ing of Monocular Depth Estimation and Visual Odometry with Deep Feature Reconstruction. In Proc. of IEEE Confer- ence on Computer Vision and Pattern Recognition (CV...

  68. [76]

    Tinghui Zhou, Matthew Brown, Noah Snavely, and David G. Lowe. Unsupervised Learning of Depth and Ego-Motion from Video. In Proc. of IEEE Conference on Computer Vi- sion and Pattern Recognition (CVPR), 2017. 2

  69. [77]

    DF-Net: Un- supervised Joint Learning of Depth and Flow using Cross- Network Consistency

    Yuliang Zou, Zelun Luo, and Jia-Bin Huang. DF-Net: Un- supervised Joint Learning of Depth and Flow using Cross- Network Consistency. In Proc. of European Conference on Computer Vision (ECCV), pages 36–53, 2018. 2, 7

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.