Pith. sign in

REVIEW 4 major objections 5 minor 44 references

Unsupervised Learning of Depth and Deep Representation for Visual Odometry from Monocular Videos in a Metric Space

T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A direct feature-alignment solver, not a pose regressor, is what keeps unsupervised monocular odometry scale-consistent.

desk verdict A promising hybrid of learned features and direct pose optimization, but the scale-consistency claim and pose evaluation protocol need to be pinned down before the headline numbers can be taken at face value. read the letter →

arxiv 1908.01367 v1 pith:O57FLEYD submitted 2019-08-04 cs.CV cs.ROeess.IV

classification cs.CVcs.ROeess.IV
keywords unsupervisedlearningmonocularvisualodometrydepthestimationdirectmethodmetricspacefeaturescaleconsistencypyramidKITTI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to show that a monocular visual odometry system can learn depth and a hierarchical feature representation without supervision, then recover camera motion by directly aligning the learned features in a metric space instead of regressing the pose from a neural network. It presents DFO, which feeds depth and feature pyramids into an iterative Gauss-Newton pose solver, refining the motion from coarse, object-level features down to fine, local features while a self-learned attention mask selects reliable points. The claimed payoff is that translation scale stays consistent across frames and pyramid levels, because the scale is tied to predicted depth rather than to an unconstrained pose regressor. On KITTI, the paper reports pose accuracy better than direct and feature-based baselines and depth accuracy close to supervised single-view depth estimators.

What carries the argument

The engine of the method is the metric-space direct feature odometry objective $E(\xi)=\sum_i w_t(u_i)\left\|\varphi_t(u_i)-\varphi_s(\omega(u_i,d_i,\xi))\right\|_2^2$, minimized in a Gauss-Newton loop with an inverse compositional update. The same objective is coupled with two scale-consistency devices: three-frame snippets, which tie translation scale to the depth of the middle frame, and the inter-level initialization $t^{(l-1)} = (\bar{z}^{(l-1)}/\bar{z}^{(l)})\,t^{(l)}$, which rescales translation between pyramid levels.

What would settle it

Run DFO on a long monocular sequence with large depth-range changes and compare translation scale drift against GPS ground truth: if scale drifts whenever the depth network's mean depth scale changes, the central claim fails.

Watch

Extended reading notes

Core claim

The central discovery claimed is that replacing the pose regression network with a direct feature-alignment solver makes unsupervised monocular odometry scale-consistent while still allowing all representations to be learned end-to-end from view synthesis. In DFO, two networks produce a depth pyramid and feature pyramids, and the relative camera pose $\xi$ minimizes a weighted Euclidean distance between warped feature patches, with features z-score normalized so that Euclidean distance, cosine similarity, and Pearson correlation agree. The pose is initialized at the coarsest pyramid level and refined downward, with translation at each finer level rescaled by the ratio of mean depth values, while three-frame snippets constrain depth scale through the middle frame. The paper reports that this scheme outperforms direct and feature-based visual odometry baselines on KITTI and matches supervised depth estimation, without assuming forward motion.

Load-bearing premise

The scale-consistency mechanism assumes the depth prediction network keeps relative depth scale stable through the dataset and through the pyramid; if depth scale drifts between frames or levels, the translation scale factor drifts with it.

Editorial extensions

If this is right

  • On KITTI odometry sequences 09 and 10, DFO with 3-frame snippets reports absolute trajectory error 0.017 plus or minus 0.017 and 0.011 plus or minus 0.010, lower than DSO (full), ORB-SLAM (short), and pose-regression baselines on the same split.
  • The depth output retains scale rather than normalized inverse depth, which is what lets the estimated depth and pose feed classical SLAM components such as keyframe selection, marginalization, and bundle adjustment.
  • Because the pose is computed by explicit feature alignment, the framework does not need the forward-motion prior that the paper shows pose networks pick up from KITTI-style training data.
  • In the Eigen-split depth evaluation, the full method reaches Abs Rel 0.152, comparable to the supervised baseline of 0.148 and better than the unsupervised baselines, with the ablation attributing a large share of the gain to learned point selection.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A test the authors do not run: feed DFO a sequence with artificial depth-scale shifts or a reinitialized depth network and check whether the three-frame snippet truly corrects translation scale rather than inheriting the drift.
  • Because the pose solver is iterative and feature-aligned, DFO is naturally extendable to sliding-window optimization; the reported accuracy over two short test sequences is plausibly a lower bound for a version with keyframe selection and bundle adjustment.
  • The ablation suggests learned point selection carries much of the depth gain; whether those masks encode geometric salience or merely KITTI's forward-driving regularities is worth testing on high-rotation or non-urban sequences.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This manuscript proposes DFO, an unsupervised framework for monocular depth estimation and visual odometry. It learns a hierarchical feature representation by training a feature pyramid with view-synthesis and autoencoder losses while a separate network predicts depth. The camera pose is obtained by direct alignment of the learned feature maps in a coarse-to-fine scheme, using iterative Gauss-Newton optimization, rather than by a pose-regression network. The paper claims that this direct feature odometry formulation constrains the translation scale factor through 3-frame snippets and pyramid-level rescaling, and it reports depth and pose results on the KITTI dataset.

Significance. If the claims are verified, the framework would be a valuable bridge between learned feature representations and classical direct visual odometry, potentially enabling integration with bundle adjustment, keyframe selection, and loop closure. The paper is commendable for formulating pose estimation as an explicit least-squares alignment with learned feature selection and outlier removal, for evaluating on held-out KITTI sequences, and for providing an ablation of the point-selection mask. However, the evaluation as presented does not yet establish the central contributions: the scale-consistency mechanism rests on an unverified assumption, and the pose comparisons in Table 2 do not clearly support the stated superiority over feature-based and direct methods.

major comments (4)
  1. [Section 4.4, Table 2] Table 2 does not support the claim that DFO 'outperforms both the feature-based and the direct visual odometry methods.' On Seq. 09, ORB-SLAM (full) achieves ATE 0.014 +/- 0.008 versus 0.017 +/- 0.017 for 'Our method (3-frame)', and the error bars overlap on Seq. 10 as well. The table also does not state whether the ATE was computed after a 7-DOF Sim(3) alignment (standard for monocular VO), after a 6-DOF metric alignment, or after no alignment; without this detail, the ATE numbers cannot validate the metric-scale claim. Please state the alignment protocol explicitly and compare against the configuration that matches the claim, e.g., ORB-SLAM (short) or DSO without full-sequence postprocessing.
  2. [Section 3.2, Eq. (10)] The scale-consistency mechanism is not sufficient for the claimed contribution. Eq. (10) only rescales the translation initialization between pyramid levels of the same frame; it does not couple different snippets of a test sequence. The paper explicitly assumes that 'the depth prediction network can maintain the relative depths through the dataset.' Because the view-synthesis loss is invariant to a joint rescaling of depth and translation, as the paper itself notes in the discussion of Wang et al. [36], nothing in the training objective or in Eq. (10) prevents depth scale from drifting over a test trajectory, and such drift directly changes the scale of the integrated translation. Please provide evidence that this assumption holds--for example, the per-sequence scale factor between the predicted trajectory and the GPS/IMU ground truth, or a comparison of ATE with and without scale alignment--or remove the absolute-scale claim.
  3. [Section 4.3, Table 1] The text claims that 'our method performs better than supervised methods,' but Table 1 shows Godard et al. [14] with Pose K achieving Abs Rel 0.148 versus DFO's 0.152, so DFO is not better on that metric. The Wang et al. [36] DN baseline also attains 0.151 Abs Rel, essentially matching DFO, which means the depth experiments do not demonstrate a scale-consistency advantage. Please correct the claim and report uncertainty estimates or significance tests for the depth metrics.
  4. [Section 4.4, Table 2] The contribution attributed to 3-frame snippets is not isolated by any ablation. The paper claims that the 3-frame constraint preserves scale consistency, but no experiment compares training with 2-frame pairs against 3-frame snippets in terms of ATE or depth-scale drift. Since the scale-consistency claim is one of the three stated contributions, such an ablation (or a direct measurement of trajectory scale error) should be added.
minor comments (5)
  1. [Section 4.4] The text refers to 'DOS (full)' where Table 2 uses DSO; please correct this typo.
  2. [Section 3.2, Eq. (8)] The Gumbel-softmax is written as a softmax over a single scalar log(p); for a Bernoulli selection variable, a two-logit form or a Gumbel-sigmoid should be specified.
  3. [Section 4.4] The sentence 'Our method falls short of the pose regression methods in Seq. 9' is inconsistent with Table 2, where DFO (3-frame) has lower ATE than Zhou et al. on Seq. 09; please reword.
  4. [Section 4.1] The statement that 'the two subnetworks are initialized with [15]' should clarify that [15] refers to a weight initialization scheme, not pretrained model weights.
  5. [Figure 3] The caption lists 'ResblockDeconvConv' without spacing; please typeset as 'Resblock/Deconv/Conv' for readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the scale-consistency claim depends on an explicit assumption, not on a circular derivation.

full rationale

The derivation is self-contained with respect to the benchmark claims. The depth and feature networks are trained on the KITTI train split with a view-synthesis loss, and the test-time pose is obtained by Gauss-Newton optimization of Eq. (5) on the fixed learned features and depths, with no pose-regression network and no test-time fitting of a scale parameter, so the reported depth metrics and ATE values are not fitted to the test trajectories. The scale-consistency contribution is the weakest point, but it is not circular: Sec. 3.2 explicitly assumes that 'the depth prediction network can maintain the relative depths through the dataset,' and the 3-frame snippet plus Eq. (10) only couple depths and poses within a snippet and across pyramid levels. This is an unverified premise and should be read as a correctness risk, not as a reduction of the conclusion to the input; Table 2 also does not state whether ATE was Sim(3)-aligned, which is an evaluation-protocol concern rather than a circularity. No load-bearing self-citation or imported uniqueness theorem is present, and the 'metric space' feature normalization is a standard z-score/cosine equivalence, not a renamed known result.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The method rests on standard assumptions about camera projection and brightness constancy, plus several tuning choices. The key unvalidated assumption is that the learned depth network maintains consistent relative depth scale across sequences, which is critical for the claimed scale consistency of the translation. The z-score normalization equivalence to cosine similarity is only approximate for local patches.

free parameters (4)
  • Outlier removal threshold = 0.5 * (median + max) of feature residuals
    Used to remove approximately 10% of feature points during pose optimization; chosen ad hoc without ablation.
  • Appearance loss thresholds = epsilon = 0.15 (DSSIM), 0.3 (L1)
    Hard thresholds for robust loss; chosen without sensitivity analysis.
  • Loss weights = lambda_sm=0.1, lambda_sp=0.01, lambda_ae=0.01
    Weights of smoothness, sparsity, and autoencoder terms; set manually.
  • Gumbel temperature = initialized 1, decreased to 0.1
    From prior work [21]; affects feature selection mask hardness.
assumptions (4)
  • standard math Euclidean distance, Pearson correlation and cosine similarity are equivalent after z-score normalization
    Invoked in Section 3.2 to justify comparing normalized features with Euclidean distance. The equivalence holds for whole vectors with zero mean and unit variance, but the residuals are computed on local patches, so it is approximate.
  • domain assumption Depth prediction network maintains relative depth scale through the dataset
    Used for scale consistency (Section 3.2). Not validated experimentally.
  • domain assumption Camera intrinsics K are known
    Used in projection equation (Eq. 3) to warp features and images.
  • domain assumption The scene is rigid except for outliers removed by thresholding
    The direct alignment assumes a single rigid transformation; moving objects are handled by the outlier removal threshold.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Unsupervised Learning of Depth and Deep Representation for Visual Odometry from Monocular Videos in a Metric Space." pith.science (2026). https://pith.science/paper/O57FLEYD

@misc{pith2026190801367,
  author       = {Pith},
  title        = {Pith review of: Unsupervised Learning of Depth and Deep Representation for Visual Odometry from Monocular Videos in a Metric Space},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/O57FLEYD}},
  note         = {Machine review of arXiv:1908.01367}
}
read the original abstract

For ego-motion estimation, the feature representation of the scenes is crucial. Previous methods indicate that both the low-level and semantic feature-based methods can achieve promising results. Therefore, the incorporation of hierarchical feature representation may benefit from both methods. From this perspective, we propose a novel direct feature odometry framework, named DFO, for depth estimation and hierarchical feature representation learning from monocular videos. By exploiting the metric distance, our framework is able to learn the hierarchical feature representation without supervision. The pose is obtained with a coarse-to-fine approach from high-level to low-level features in enlarged feature maps. The pixel-level attention mask can be self-learned to provide the prior information. In contrast to the previous methods, our proposed method calculates the camera motion with a direct method rather than regressing the ego-motion from the pose network. With this approach, the consistency of the scale factor of translation can be constrained. Additionally, the proposed method is thus compatible with the traditional SLAM pipeline. Experiments on the KITTI dataset demonstrate the effectiveness of our method.

Figures

Figures reproduced from arXiv: 1908.01367 by the authors.

Figure 1
Figure 1. Pipeline of our method for pose and depth estimation. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Illustration of our method for the pose calculation process using the [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. An overview of the deep feature representation network. [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Illustration of the scale consistency constraints using 3- [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Feature pyramid for pose estimation. The ego-motion is [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Visualization of the feature maps involved in the calcu [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

44 extracted references · 40 canonical work pages

  1. [36]

    C. Wang, J. M. Buenaposada, R. Zhu, and S. Lucey. Learning depth from monocular videos using direct methods. In Pro- ceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2022–2030, 2018. 2, 5, 6, 7, 8

  2. [14]

    Godard, O

    C. Godard, O. Mac Aodha, and G. J. Brostow. Unsuper- vised monocular depth estimation with left-right consistency. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, volume 2, page 7, 2017. 7, 8

  3. [1]

    Agrawal, J

    P. Agrawal, J. Carreira, and J. Malik. Learning to see by moving. In Proceedings of the IEEE International Confer- ence on Computer Vision, pages 37–45, 2015. 2

  4. [2]

    Andrew, R

    G. Andrew, R. Arora, J. Bilmes, and K. Livescu. Deep canonical correlation analysis. InProceedings of the Interna- tional Conference on Machine Learning , pages 1247–1255,

  5. [3]

    Baker and I

    S. Baker and I. Matthews. Lucas-kanade 20 years on: A uni- fying framework. International journal of computer vision , 56(3):221–255, 2004. 4

  6. [4]

    Belagiannis, C

    V . Belagiannis, C. Rupprecht, G. Carneiro, and N. Navab. Robust optimization for deep regression. In Proceedings of the IEEE International Conference on Computer Vision , pages 2830–2838, 2015. 6

  7. [5]

    M. R. Berthold and F. H ¨oppner. On clustering time se- ries using euclidean distance and pearson correlation. arXiv preprint arXiv:1601.02213, 2016. 4

  8. [6]

    Bloesch, J

    M. Bloesch, J. Czarnowski, R. Clark, S. Leutenegger, and A. J. Davison. Codeslam-learning a compact, optimisable representation for dense visual slam. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion, pages 2560–2568, 2018. 2

Show all 44 references
  1. [7]

    S. L. Bowman, N. Atanasov, K. Daniilidis, and G. J. Pappas. Probabilistic data association for semantic slam. In Proceed- ings of IEEE International Conference on Robotics and Au- tomation, pages 1722–1729. IEEE, 2017. 2

  2. [8]

    Chopra, R

    S. Chopra, R. Hadsell, and Y . LeCun. Learning a similarity metric discriminatively, with application to face verification. In Proceedings of the Computer Vision and Pattern Recogni- tion, 2005. CVPR 2005. IEEE Computer Society Conference on, volume 1, pages 539–546, 2005. 3

  3. [9]

    Cordts, M

    M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele. The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3213–3223, 2016. 8

  4. [10]

    Eigen, C

    D. Eigen, C. Puhrsch, and R. Fergus. Depth map prediction from a single image using a multi-scale deep network. In Advances in neural information processing systems , pages 2366–2374, 2014. 2, 7, 8

  5. [11]

    Engel, V

    J. Engel, V . Koltun, and D. Cremers. Direct sparse odometry. IEEE transactions on pattern analysis and machine intelli- gence, 4, 2017. 5, 7

  6. [12]

    Engel, T

    J. Engel, T. Sch ¨ops, and D. Cremers. Lsd-slam: Large-scale direct monocular slam. In Proceedings of the European Con- ference on Computer Vision, pages 834–849. Springer, 2014. 7

  7. [13]

    Geiger, P

    A. Geiger, P. Lenz, and R. Urtasun. Are we ready for au- tonomous driving? the kitti vision benchmark suite. In Pro- ceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2012. 8

  8. [15]

    K. He, X. Zhang, S. Ren, and J. Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international con- ference on computer vision, pages 1026–1034, 2015. 6

  9. [16]

    J. Hu, J. Lu, and Y .-P. Tan. Discriminative deep metric learn- ing for face verification in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion, pages 1875–1882, 2014. 4

  10. [17]

    Jaderberg, K

    M. Jaderberg, K. Simonyan, A. Zisserman, et al. Spatial transformer networks. In Advances in neural information processing systems, pages 2017–2025, 2015. 3

  11. [18]

    E. Jang, S. Gu, and B. Poole. Categorical reparameterization with gumbel-softmax. In Proceedings of the International Conference for Learning Representations, 2017. 5

  12. [19]

    D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. In Proceedings of the International Conference for Learning Representations, 2014. 6

  13. [20]

    Klodt and A

    M. Klodt and A. Vedaldi. Supervising the new with the old: learning sfm from sfm. InProceedings of the European Con- ference on Computer Vision, pages 698–713, 2018. 8

  14. [21]

    Kong and C

    S. Kong and C. Fowlkes. Pixel-wise attentional gat- ing for parsimonious pixel labeling. arXiv preprint arXiv:1805.01556, 2018. 5

  15. [22]

    F. Liu, C. Shen, G. Lin, and I. D. Reid. Learning depth from single monocular images using deep convolutional neural fields. IEEE Trans. Pattern Anal. Mach. Intell., 38(10):2024– 2039, 2016. 2, 7, 8

  16. [23]

    Y . Ma, S. Soatto, J. Kosecka, and S. S. Sastry. An invitation to 3-d vision: from images to geometric models , volume 26. Springer Science & Business Media, 2012. 4

  17. [24]

    C. J. Maddison, A. Mnih, and Y . W. Teh. The concrete dis- tribution: A continuous relaxation of discrete random vari- ables. In Proceedings of the International Conference for Learning Representations, 2017. 5

  18. [25]

    Mahjourian, M

    R. Mahjourian, M. Wicke, and A. Angelova. Unsupervised learning of depth and ego-motion from monocular video us- ing 3d geometric constraints. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 5667–5675, 2018. 2, 7, 8

  19. [26]

    Mayer, E

    N. Mayer, E. Ilg, P. Hausser, P. Fischer, D. Cremers, A. Dosovitskiy, and T. Brox. A large dataset to train con- volutional networks for disparity, optical flow, and scene flow estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 4...

  20. [27]

    Menze and A

    M. Menze and A. Geiger. Object scene flow for autonomous vehicles. In Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition , pages 3061–3070,

  21. [28]

    Mur-Artal and J

    R. Mur-Artal and J. D. Tard ´os. Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras.IEEE Transactions on Robotics, 33(5):1255–1262, 2017. 2, 7

  22. [29]

    Paszke, S

    A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. De- Vito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer. Auto- matic differentiation in pytorch. In Proceedings of the Ad- vances in neural information processing systems Workshop,

  23. [30]

    C. Pinard. Sfmlearner-pytorch. https://github.com/ ClementPinard/SfmLearner-Pytorch. Accessed Oct. 2018. 6, 8

  24. [31]

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, D. Batra, et al. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Pro- ceedings of the IEEE international conference on computer vision, pages 618–626, 2017. 3, 7

  25. [32]

    Szeliski

    R. Szeliski. Prediction error as a quality metric for motion and stereo. In Proceedings of the Seventh IEEE International Conference on Computer Vision, volume 2, pages 781–788. IEEE, 1999. 3

  26. [33]

    Tateno, F

    K. Tateno, F. Tombari, I. Laina, and N. Navab. Cnn-slam: Real-time dense monocular slam with learned depth predic- tion. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, volume 2, 2017. 2

  27. [34]

    Ummenhofer, H

    B. Ummenhofer, H. Zhou, J. Uhrig, N. Mayer, E. Ilg, A. Dosovitskiy, and T. Brox. Demon: Depth and motion network for learning monocular stereo. In Proceedings of the IEEE Conference on computer vision and pattern recog- nition, volume 5, page 6, 2017. 2

  28. [35]

    Vijayanarasimhan, S

    S. Vijayanarasimhan, S. Ricco, C. Schmid, R. Sukthankar, and K. Fragkiadaki. Sfm-net: Learning of structure and mo- tion from video. arXiv preprint arXiv:1704.07804, 2017. 2, 8

  29. [37]

    F. Wang, M. Jiang, C. Qian, S. Yang, C. Li, H. Zhang, X. Wang, and X. Tang. Residual attention network for im- age classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 3156– 3164, 2017. 6

  30. [38]

    W. Wang, R. Arora, K. Livescu, and J. Bilmes. On deep multi-view representation learning. In Proceedings of the In- ternational Conference on Machine Learning , pages 1083– 1092, 2015. 3, 4

  31. [39]

    Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simon- celli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image process- ing, 13(4):600–612, 2004. 5

  32. [40]

    K. Q. Weinberger and L. K. Saul. Unsupervised learning of image manifolds by semidefinite programming. Interna- tional journal of computer vision, 70(1):77–90, 2006. 1

  33. [41]

    N. Yang, R. Wang, J. St ¨uckler, and D. Cremers. Deep vir- tual stereo odometry: Leveraging deep depth prediction for monocular direct sparse odometry. In Proceedings of the European Conference on Computer Vision, pages 835–852. Springer, 2018. 2

  34. [42]

    Yin and J

    Z. Yin and J. Shi. Geonet: Unsupervised learning of dense depth, optical flow and camera pose. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion, pages 1983–1992, 2018. 2, 6, 7, 8

  35. [43]

    M. D. Zeiler and R. Fergus. Visualizing and understand- ing convolutional networks. In Proceedings of the European conference on computer vision , pages 818–833. Springer,

  36. [44]

    T. Zhou, M. Brown, N. Snavely, and D. G. Lowe. Unsuper- vised learning of depth and ego-motion from video. In Pro- ceedings of the IEEE Conference on Computer Vision and Pattern Recognition, volume 2, page 7, 2017. 2, 3, 6, 7, 8

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.