Pith. sign in

REVIEW 1 cited by

Globally Consistent Video Depth and Pose Estimation with Efficient Test-Time Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.02709 v1 pith:E6ZLS6YF submitted 2022-08-04 cs.CV

classification cs.CV
keywords depthestimationposeconsistentgcvdgloballylearning-basedmethods
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Dense depth and pose estimation is a vital prerequisite for various video applications. Traditional solutions suffer from the robustness of sparse feature tracking and insufficient camera baselines in videos. Therefore, recent methods utilize learning-based optical flow and depth prior to estimate dense depth. However, previous works require heavy computation time or yield sub-optimal depth results. We present GCVD, a globally consistent method for learning-based video structure from motion (SfM) in this paper. GCVD integrates a compact pose graph into the CNN-based optimization to achieve globally consistent estimation from an effective keyframe selection mechanism. It can improve the robustness of learning-based methods with flow-guided keyframes and well-established depth prior. Experimental results show that GCVD outperforms the state-of-the-art methods on both depth and pose estimation. Besides, the runtime experiments reveal that it provides strong efficiency in both short- and long-term videos with global consistency provided.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RaCalNet: Radar Calibration Network for Sparse-Supervised Metric Depth Estimation

    cs.CV 2025-06 reject novelty 6.0 of 10

    A radar-camera depth estimation framework that recalibrates sparse radar points and aligns a frozen monocular depth model using sparse LiDAR labels, claiming state-of-the-art accuracy with roughly 1% supervision density.

Pith tools