Pith. sign in

REVIEW 2 cited by

Digging Into Self-Supervised Monocular Depth Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1806.01260 v4 pith:SJ5F5YAJ submitted 2018-06-04 cs.CV stat.ML

classification cs.CVstat.ML
keywords depthself-supervisedlossmonoculartrainingestimationmethodsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Per-pixel ground-truth depth data is challenging to acquire at scale. To overcome this limitation, self-supervised learning has emerged as a promising alternative for training models to perform monocular depth estimation. In this paper, we propose a set of improvements, which together result in both quantitatively and qualitatively improved depth maps compared to competing self-supervised methods. Research on self-supervised monocular training usually explores increasingly complex architectures, loss functions, and image formation models, all of which have recently helped to close the gap with fully-supervised methods. We show that a surprisingly simple model, and associated design choices, lead to superior predictions. In particular, we propose (i) a minimum reprojection loss, designed to robustly handle occlusions, (ii) a full-resolution multi-scale sampling method that reduces visual artifacts, and (iii) an auto-masking loss to ignore training pixels that violate camera motion assumptions. We demonstrate the effectiveness of each component in isolation, and show high quality, state-of-the-art results on the KITTI benchmark.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing

    cs.CV 2026-03 conditional novelty 6.0 of 10

    RL3DEdit fine-tunes FLUX-Kontext with GRPO using VGGT confidence and pose rewards to produce multi-view consistent 3D scene edits in a single pass.

  2. StixelNExT++: Lightweight Monocular Scene Segmentation and Representation for Collective Perception

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A monocular neural network predicts 3D Stixels directly from RGB images in about 10 ms, with a self-defined Waymo evaluation showing competitive performance within 30 m.

Pith tools