Pith. sign in

REVIEW 6 cited by

Semantically-Guided Representation Learning for Self-Supervised Monocular Depth

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.12319 v1 pith:NZJTFBML submitted 2020-02-27 cs.CV cs.LG

Semantically-Guided Representation Learning for Self-Supervised Monocular Depth

classification cs.CV cs.LG
keywords learningself-supervisedsemanticdepthmonocularrepresentationguideleveraging
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Self-supervised learning is showing great promise for monocular depth estimation, using geometry as the only source of supervision. Depth networks are indeed capable of learning representations that relate visual appearance to 3D properties by implicitly leveraging category-level patterns. In this work we investigate how to leverage more directly this semantic structure to guide geometric representation learning, while remaining in the self-supervised regime. Instead of using semantic labels and proxy losses in a multi-task approach, we propose a new architecture leveraging fixed pretrained semantic segmentation networks to guide self-supervised representation learning via pixel-adaptive convolutions. Furthermore, we propose a two-stage training process to overcome a common semantic bias on dynamic objects via resampling. Our method improves upon the state of the art for self-supervised monocular depth prediction over all pixels, fine-grained details, and per semantic categories.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Improved monocular depth prediction using distance transform over pre-semantic contours with self-supervised neural networks

    eess.IV 2026-05 unverdicted novelty 7.0

    Self-supervised monocular depth estimation improves in low-texture regions by using distance transforms on jointly estimated pre-semantic contours to create more informative loss signals.

  2. Adaptive Depth-converted-Scale Convolution for Self-supervised Monocular Depth Estimation

    cs.CV 2026-04 unverdicted novelty 7.0

    DcSConv adapts convolution filter scales based on object depth to reduce size-depth ambiguity in self-supervised monocular depth estimation, improving performance on KITTI by up to 11.6% in SqRel.

  3. SS3D: End2End Self-Supervised 3D from Web Videos

    cs.CV 2026-04 unverdicted novelty 6.0

    SS3D pretrains an end-to-end 3D estimator on filtered YouTube-8M videos via SfM self-supervision, achieving improved zero-shot transfer and fine-tuning over prior baselines.

  4. SS3D: End2End Self-Supervised 3D from Web Videos

    cs.CV 2026-04 unverdicted novelty 6.0

    SS3D pretrains an end-to-end feed-forward 3D estimator on filtered YouTube-8M videos via SfM self-supervision, MVS filtering, and expert distillation, delivering stronger zero-shot transfer and fine-tuning than prior ...

  5. SS3D: End2End Self-Supervised 3D from Web Videos

    cs.CV 2026-04 unverdicted novelty 5.0

    SS3D scales SfM-based self-supervision to ~100M frames from YouTube-8M using a multi-view signal proxy for filtering and a two-stage training schedule, achieving strong zero-shot transfer and better fine-tuning than p...

  6. NAIMA: Semantics Aware RGB Guided Depth Super-Resolution

    eess.IV 2026-04 unverdicted novelty 5.0

    NAIMA distills global semantic context from DINOv2 token embeddings into RGB-guided depth super-resolution using cross-attention blocks, reporting gains over prior GDSR methods on multiple datasets and scales.