Pith. sign in

REVIEW 4 cited by

Self-supervised Depth Estimation Leveraging Global Perception and Geometric Smoothness Using On-board Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.03505 v1 pith:O6PWMHV5 submitted 2021-06-07 cs.CV

Self-supervised Depth Estimation Leveraging Global Perception and Geometric Smoothness Using On-board Videos

classification cs.CV
keywords depthestimationglobalperformanceproposedsmoothnessblockdlnet
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Self-supervised depth estimation has drawn much attention in recent years as it does not require labeled data but image sequences. Moreover, it can be conveniently used in various applications, such as autonomous driving, robotics, realistic navigation, and smart cities. However, extracting global contextual information from images and predicting a geometrically natural depth map remain challenging. In this paper, we present DLNet for pixel-wise depth estimation, which simultaneously extracts global and local features with the aid of our depth Linformer block. This block consists of the Linformer and innovative soft split multi-layer perceptron blocks. Moreover, a three-dimensional geometry smoothness loss is proposed to predict a geometrically natural depth map by imposing the second-order smoothness constraint on the predicted three-dimensional point clouds, thereby realizing improved performance as a byproduct. Finally, we explore the multi-scale prediction strategy and propose the maximum margin dual-scale prediction strategy for further performance improvement. In experiments on the KITTI and Make3D benchmarks, the proposed DLNet achieves performance competitive to those of the state-of-the-art methods, reducing time and space complexities by more than $62\%$ and $56\%$, respectively. Extensive testing on various real-world situations further demonstrates the strong practicality and generalization capability of the proposed model.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SS3D: End2End Self-Supervised 3D from Web Videos

    cs.CV 2026-04 unverdicted novelty 6.0

    SS3D pretrains an end-to-end 3D estimator on filtered YouTube-8M videos via SfM self-supervision, achieving improved zero-shot transfer and fine-tuning over prior baselines.

  2. SS3D: End2End Self-Supervised 3D from Web Videos

    cs.CV 2026-04 unverdicted novelty 6.0

    SS3D pretrains an end-to-end feed-forward 3D estimator on filtered YouTube-8M videos via SfM self-supervision, MVS filtering, and expert distillation, delivering stronger zero-shot transfer and fine-tuning than prior ...

  3. Rebalancing gradient to improve self-supervised co-training of depth, odometry and optical flow predictions

    cs.CV 2026-05 unverdicted novelty 5.0

    CoopNet improves co-training of depth, odometry, and optical flow networks via a hybrid loss that rebalances gradients by modeling photometric error disagreements to identify and ignore moving objects.

  4. SS3D: End2End Self-Supervised 3D from Web Videos

    cs.CV 2026-04 unverdicted novelty 5.0

    SS3D scales SfM-based self-supervision to ~100M frames from YouTube-8M using a multi-view signal proxy for filtering and a two-stage training schedule, achieving strong zero-shot transfer and better fine-tuning than p...