Pith. sign in

REVIEW 3 cited by

Stealing Stable Diffusion Prior for Robust Monocular Depth Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.05056 v1 pith:NA7ZHF5W submitted 2024-03-08 cs.CV

classification cs.CV
keywords depthdiffusionestimationstableapproachchallengingconditionsmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Monocular depth estimation is a crucial task in computer vision. While existing methods have shown impressive results under standard conditions, they often face challenges in reliably performing in scenarios such as low-light or rainy conditions due to the absence of diverse training data. This paper introduces a novel approach named Stealing Stable Diffusion (SSD) prior for robust monocular depth estimation. The approach addresses this limitation by utilizing stable diffusion to generate synthetic images that mimic challenging conditions. Additionally, a self-training mechanism is introduced to enhance the model's depth estimation capability in such challenging environments. To enhance the utilization of the stable diffusion prior further, the DINOv2 encoder is integrated into the depth model architecture, enabling the model to leverage rich semantic priors and improve its scene understanding. Furthermore, a teacher loss is introduced to guide the student models in acquiring meaningful knowledge independently, thus reducing their dependency on the teacher models. The effectiveness of the approach is evaluated on nuScenes and Oxford RobotCar, two challenging public datasets, with the results showing the efficacy of the method. Source code and weights are available at: https://github.com/hitcslj/SSD.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Depth Anything at Any Condition

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A fine-tuned Depth Anything V2 model using perturbation consistency and spatial distance constraints improves monocular depth estimation under adverse conditions without any labeled data.

  2. Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Driver gaze and accident-reason text are used to train a video diffusion model that can edit and generate egocentric crash videos with the correct causal participants, with a new large gaze dataset for accidents.

  3. GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A 2D Gaussian Splatting method that uses foundation-model depth/normal priors plus deferred shading to improve reconstruction and relighting of reflective objects.

Pith tools