Pith. sign in

REVIEW 28 cited by

MiDaS v3.1 -- A Model Zoo for Robust Monocular Relative Depth Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.14460 v1 pith:OVJQSIK2 submitted 2023-07-26 cs.CV

classification cs.CV
keywords midasvisiondepthestimationmodelstransformersqualityrelease
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We release MiDaS v3.1 for monocular depth estimation, offering a variety of new models based on different encoder backbones. This release is motivated by the success of transformers in computer vision, with a large variety of pretrained vision transformers now available. We explore how using the most promising vision transformers as image encoders impacts depth estimation quality and runtime of the MiDaS architecture. Our investigation also includes recent convolutional approaches that achieve comparable quality to vision transformers in image classification tasks. While the previous release MiDaS v3.0 solely leverages the vanilla vision transformer ViT, MiDaS v3.1 offers additional models based on BEiT, Swin, SwinV2, Next-ViT and LeViT. These models offer different performance-runtime tradeoffs. The best model improves the depth estimation quality by 28% while efficient models enable downstream tasks requiring high frame rates. We also describe the general process for integrating new backbones. A video summarizing the work can be found at https://youtu.be/UjaeNNFf9sE and the code is available at https://github.com/isl-org/MiDaS.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 28 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth

    cs.CV 2026-07 conditional novelty 7.0 of 10

    EpiDistill uses depth-guided epipolar attention and learnable rectified stereo tokens to distill multi-view scale knowledge into single-view monocular depth models.

  2. Evaluating Robustness of Monocular Depth Estimation with Procedural Scene Perturbations

    cs.CV 2025-07 conditional novelty 7.0 of 10

    A procedural-generation benchmark (PDE) shows that depth models are surprisingly vulnerable to camera changes and occlusion, while resisting lighting changes.

  3. Relative Pose Estimation through Affine Corrections of Monocular Depth Priors

    cs.CV 2025-01 conditional novelty 7.0 of 10

    New solvers and a hybrid RANSAC pipeline that estimate relative pose while correcting scale and shift in monocular depth priors, improving pose accuracy for calibrated and uncalibrated cameras.

  4. Breaking the Horizontal Prior: From Long-Tailed Orientation Bias to Roll-Robust Monocular Depth Estimation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Auxiliary rotation-stable depth cues, added during fine-tuning, reduce monocular depth-error degradation under camera roll across five benchmarks.

  5. DGSfM: Depth-Guided Scale-Aware Global Structure-from-Motion

    cs.CV 2026-07 accept novelty 6.0 of 10

    Monocular depth priors turn scale-ambiguous global SfM into a scale-aware pipeline that measurably improves camera pose accuracy on ETH3D and IMC2021.

  6. AerialMetric: Benchmarking and Adapting UAV Monocular Metric Depth Estimation in the Real World

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    AerialMetric is a new benchmark dataset and evaluation suite for adapting monocular metric depth estimation models to real-world UAV aerial views.

  7. Towards Consistent Video Geometry Estimation

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    ViGeo is a feed-forward transformer for video geometry that introduces dynamic chunking attention and a completion-based data refinement framework to achieve SOTA on depth, normals, and point map estimation.

  8. Accuracy Does Not Guarantee Human-Likeness: Cross-Domain Human-Centered Benchmark in Monocular Depth Estimation

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Across 69 monocular depth estimators, human-likeness of error patterns peaks near human-level accuracy and declines for the most accurate models: accuracy does not guarantee human-like depth perception.

  9. SpatialTrackerV2: 3D Point Tracking Made Easy

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single feed-forward model jointly estimates video depth, camera poses, and 3D point trajectories from monocular video, setting a new state of the art on TAPVid-3D.

  10. RaCalNet: Radar Calibration Network for Sparse-Supervised Metric Depth Estimation

    cs.CV 2025-06 reject novelty 6.0 of 10

    A radar-camera depth estimation framework that recalibrates sparse radar points and aligns a frozen monocular depth model using sparse LiDAR labels, claiming state-of-the-art accuracy with roughly 1% supervision density.

  11. Repurposing Marigold for Zero-Shot Metric Depth Estimation via Defocus Blur Cues

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Two differently blurred images plus a pretrained diffusion depth prior are optimized together at inference time to recover metric depth without retraining.

  12. Flow Distillation Sampling: Regularizing 3D Gaussians with Pre-trained Matching Priors

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A new loss that matches an optical-flow model's predictions against analytically computed flows from 3D Gaussians, improving geometric reconstruction on sparse indoor scenes.

  13. Enhancing Monocular Depth Estimation with Multi-Source Auxiliary Tasks

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Training a frozen DINOv2 backbone with a shared decoder on auxiliary multi-label dense classification improves monocular depth accuracy by about 11 percent on in-domain benchmarks.

  14. Video Depth Anything: Consistent Depth Estimation for Super-Long Videos

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Video Depth Anything adapts Depth Anything V2 to produce temporally consistent depth for arbitrarily long videos using a temporal attention head, an optical-flow-free gradient loss, and key-frame-based stitching.

  15. RDG-GS: Relative Depth Guidance with Gaussian Splatting for Real-time Sparse-View 3D Rendering

    cs.CV 2025-01 conditional novelty 6.0 of 10

    RDG-GS combines refined monocular depth priors, a relative depth similarity loss, and adaptive point densification to improve sparse-view 3D Gaussian Splatting rendering.

  16. DEFOM-Stereo: Depth Foundation Model Based Stereo Matching

    cs.CV 2025-01 conditional novelty 6.0 of 10

    DEFOM-Stereo combines a depth foundation model's features and depth estimates with RAFT-Stereo's recurrent updates to improve zero-shot stereo matching and set top benchmark numbers.

  17. RePoseD: Efficient Relative Pose Estimation With Known Depth Information

    cs.CV 2025-01 conditional novelty 6.0 of 10

    New efficient minimal solvers estimate relative camera pose jointly with unknown depth scale and shift, improving speed and often accuracy over prior depth-aware solvers.

  18. LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation

    cs.CV 2026-08 conditional novelty 5.0 of 10

    LiteMVS improves efficient multi-view stereo depth estimation by injecting semantic descriptors, MoE cost aggregation, and pseudo-labels from monocular foundation models.

  19. ER-LoRA: Effective-Rank Guided Adaptation for Weather-Generalized Depth Estimation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Tuning only 8.7M parameters of a frozen DINOv2 on daytime data is reported to beat prior PEFT, full fine-tuning, synthetic-data depth methods, and Depth Anything V2 on zero-shot adverse-weather benchmarks.

  20. Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Concatenating monocular depth maps as an extra input channel improves video instance segmentation and reaches 56.2 AP, a new state of the art on OVIS.

  21. Depth Anything at Any Condition

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A fine-tuned Depth Anything V2 model using perturbation consistency and spatial distance constraints improves monocular depth estimation under adverse conditions without any labeled data.

  22. Vision without Images: End-to-End Computer Vision from Single Compressive Measurements

    cs.CV 2025-01 conditional novelty 5.0 of 10

    An end-to-end compressive sensing system, CompDAE, runs edge, depth, and segmentation tasks directly on single 8x8-mask measurements and reports strong low-light results without image reconstruction.

  23. Rethinking Encoder-Decoder Flow Through Shared Structures

    cs.CV 2025-01 conditional novelty 5.0 of 10

    Shared feature and sampling banks fed to every decoder block improve monocular depth estimation accuracy slightly for ViT and RepViT encoders at under 1% parameter overhead for ViTs.

  24. Joint Learning of Depth and Appearance for Portrait Image Animation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A single diffusion model jointly generates portrait RGB images and aligned depth maps, and its fine-tuned variants can estimate depth, edit from depth, relight, and produce audio-driven talking heads with depth.

  25. URS-Stereo: Uncertainty-Guided Residual Search for Real-Time Stereo Matching

    cs.CV 2026-07 conditional novelty 4.5 of 10

    Uncertainty-modulated residual offsets relocate local cost-volume centers in coarse-to-fine stereo matching, improving zero-shot disparity accuracy while keeping real-time speed.

  26. Spatial RoboGrasp: Generalized Robotic Grasping Control Policy

    cs.RO 2025-05 conditional novelty 4.0 of 10

    Spatial RoboGrasp combines AugFusion, monocular depth, and grasp prompts in a diffusion policy, claiming large gains under exposure change, without released artifacts or error bars.

  27. Visual question answering: from early developments to recent advances -- a survey

    cs.CV 2025-01 conditional novelty 2.0 of 10

    A survey that classifies VQA architectures by encoder, fusion, and decoder, reviews datasets and metrics, and discusses applications and future directions.

  28. Survey on Monocular Metric Depth Estimation

    cs.CV 2025-01 unverdicted novelty 1.0 of 10

    A survey of monocular metric depth estimation methods, datasets, and open challenges, with no new experimental results.

Pith tools