Pith. sign in

REVIEW 4 cited by

DF-VO: What Should Be Learnt for Visual Odometry?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.00933 v1 pith:BXUZNZGU submitted 2021-03-01 cs.CV

classification cs.CV
keywords deepdf-vodepthsmethodsodometrycameradynamicgeometric
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-view geometry-based methods dominate the last few decades in monocular Visual Odometry for their superior performance, while they have been vulnerable to dynamic and low-texture scenes. More importantly, monocular methods suffer from scale-drift issue, i.e., errors accumulate over time. Recent studies show that deep neural networks can learn scene depths and relative camera in a self-supervised manner without acquiring ground truth labels. More surprisingly, they show that the well-trained networks enable scale-consistent predictions over long videos, while the accuracy is still inferior to traditional methods because of ignoring geometric information. Building on top of recent progress in computer vision, we design a simple yet robust VO system by integrating multi-view geometry and deep learning on Depth and optical Flow, namely DF-VO. In this work, a) we propose a method to carefully sample high-quality correspondences from deep flows and recover accurate camera poses with a geometric module; b) we address the scale-drift issue by aligning geometrically triangulated depths to the scale-consistent deep depths, where the dynamic scenes are taken into account. Comprehensive ablation studies show the effectiveness of the proposed method, and extensive evaluation results show the state-of-the-art performance of our system, e.g., Ours (1.652%) v.s. ORB-SLAM (3.247%}) in terms of translation error in KITTI Odometry benchmark. Source code is publicly available at: \href{https://github.com/Huangying-Zhan/DF-VO}{DF-VO}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ZeroVO: Visual Odometry with Minimal Assumptions

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A two-frame visual odometry model using estimated depth, language priors, and semi-supervised pseudo-label filtering achieves zero-shot metric-scale pose estimation across multiple driving datasets.

  2. DVLO4D: Deep Visual-Lidar Odometry with Sparse Spatial-temporal Fusion

    cs.CV 2025-09 conditional novelty 5.0 of 10

    DVLO4D fuses sparse LiDAR queries with camera features, adds temporal memory and a sequence-level loss, and improves visual-LiDAR odometry accuracy to 0.73% translation error on KITTI 07-10.

  3. UNO: Unified Self-Supervised Monocular Odometry for Platform-Agnostic Deployment

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A self-supervised monocular visual odometry pipeline combines mixture-of-experts pose decoders, Gumbel-Softmax graph selection and scale-aware bundle adjustment to reach state-of-the-art reported errors on KITTI, EuRo...

  4. Dense-depth map guided deep Lidar-Visual Odometry with Sparse Point Clouds and Images

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A LiDAR-visual odometry network that uses completed dense depth maps to guide optical flow and hierarchical pose refinement reports strong results on the KITTI benchmark.

Pith tools