Pith. sign in

REVIEW 2 cited by

Transformer-Based Model for Monocular Visual Odometry: A Video Understanding Approach

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.06121 v3 pith:2AUEHVP6 submitted 2023-05-10 cs.CV cs.AIcs.RO

classification cs.CVcs.AIcs.RO
keywords odometryvisualcameramonocularunderstandingvideoapproachavailable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Estimating the camera's pose given images from a single camera is a traditional task in mobile robots and autonomous vehicles. This problem is called monocular visual odometry and often relies on geometric approaches that require considerable engineering effort for a specific scenario. Deep learning methods have been shown to be generalizable after proper training and with a large amount of available data. Transformer-based architectures have dominated the state-of-the-art in natural language processing and computer vision tasks, such as image and video understanding. In this work, we deal with the monocular visual odometry as a video understanding task to estimate the 6 degrees of freedom of a camera's pose. We contribute by presenting the TSformer-VO model based on spatio-temporal self-attention mechanisms to extract features from clips and estimate the motions in an end-to-end manner. Our approach achieved competitive state-of-the-art performance compared with geometry-based and deep learning-based methods on the KITTI visual odometry dataset, outperforming the DeepVO implementation highly accepted in the visual odometry community. The code is publicly available at https://github.com/aofrancani/TSformer-VO.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BRIGHT-VO: Brightness-Guided Hybrid Transformer for Visual Odometry with Multi-modality Refinement Module

    cs.CV 2025-01 reject novelty 5.0 of 10

    A brightness-guided Transformer with IMU-based pose graph optimization improves monocular VO accuracy on synthetic low-light and KITTI benchmarks, though the evaluation has significant flaws.

  2. VKFPos: A Learning-Based Monocular Positioning with Variational Bayesian Extended Kalman Filter Integration

    cs.CV 2025-01 conditional novelty 4.0 of 10

    A monocular positioning system that combines absolute and relative pose regression with a Kalman filter, using learned covariances, reports improved temporal accuracy on 7-Scenes and Oxford RobotCar.

Pith tools