Pith. sign in

REVIEW 3 major objections 5 minor 76 references

A 6–33M-parameter model can carry foundation-level monocular depth generalization and metric accuracy if trained with bias-resistant sampling and encoder-frozen, camera-conditioned fine-tuning.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 19:01 UTC pith:CFFTP3JL

load-bearing objection Solid tiny-model depth recipe with convincing relative-depth gains; metric transfer rests on a scale-only mapping that needs an explicit shift control before I'd trust it fully. the 3 major comments →

arxiv 2607.17099 v1 pith:CFFTP3JL submitted 2026-07-19 cs.CV cs.AI

DepthART: Scaling Foundation Monocular Depth to Tiny Models

classification cs.CV cs.AI
keywords monocular depth estimationtiny modelsfoundation model distillationdataset biascamera-conditioned fine-tuningzero-shot generalizationmetric scaleon-device deployment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper sets out to prove that tiny monocular depth models need not sacrifice the cross-scene generalization and metric-scale accuracy that large foundation models deliver. It identifies two reasons small models fail when naively trained: they overfit dataset-specific biases, and full fine-tuning to metric scale destroys learned geometry. DepthART attacks both with a bias-resistant data sampling scheme that rebalances long-tailed training distributions, and a camera-conditioned fine-tuning protocol that freezes the distilled encoder and adds small camera adapters plus a multi-query scale head. Across five zero-shot benchmarks and multiple metric-transfer settings, DepthART surpasses all prior tiny baselines and in some cases approaches models tens of times larger, while running at real-time speeds on embedded GPUs.

Core claim

DepthART claims that the two capacity-driven bottlenecks for compact depth models are overfitting to dominant dataset-specific distributions and forgetting of transferable geometry during metric fine-tuning. The paper's central claim is that both can be fixed without changing the tiny architecture: bias-resistant sampling performs density-aware stratified selection in a compressed visual feature space to suppress overrepresented scene modes, and camera-conditioned fine-tuning freezes the encoder, injects multi-resolution camera prompts through small cross-attention adapters, and regresses a single per-image scale factor to convert relative depth to metric depth. With this recipe, DepthART-S

What carries the argument

BRDS (bias-resistant data sampling): density-aware stratified sampling in a visual feature space, which caps the number of samples drawn from dense, dominant scene modes and guarantees coverage of sparse long-tail modes, producing a 1.7M-image bias-resistant pool. CamFT (camera-conditioned fine-tuning): a frozen distilled encoder, pyramid camera-ray prompts injected through small cross-attention adapters, and a multi-query scale head that outputs one positive scale factor; metric depth is computed as the relative depth multiplied by that scale and a fixed maximum depth. The critical mechanism is that freezing the encoder preserves the transferred relative-depth structure, while the adapters

Load-bearing premise

The claim rests on the premise that metric depth can be obtained from the frozen relative-depth encoder by a single per-image scalar multiplied onto the relative depth with a fixed maximum, with no per-image offset or non-scalar correction.

What would settle it

Run DepthART, fine-tuned on KITTI, on a collection of handheld, drone, and telephoto images. Fit per-image optimal scale and shift to ground truth; if the optimal shift is systematically large and varies with scene content, the scale-only metric parameterization is falsified.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If correct, real-time on-device depth estimation can retain a large share of foundation-model generalization, closing the gap between research models and deployed systems.
  • The recipe yields specific speedups: DepthART-S/B/L run at 347/298/191 FPS at 224² and 245/204/124 FPS at 448² on an RTX A6000 in strict FP32, with >15 FPS on a Jetson Nano.
  • Zero-shot metric depth becomes practical without post-hoc scale-and-shift alignment, since metric predictions are directly usable as 3D point clouds.
  • Density-aware sampling is more sample-efficient than random or k-means selection for distillation, improving zero-shot transfer even at fixed data budgets.
  • Encoder-frozen, camera-conditioned fine-tuning substantially limits catastrophic forgetting compared to full fine-tuning, as measured by affine-invariant RMSE on out-of-distribution scenes.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The result implies that data diversity, not raw data volume, is the binding constraint for tiny models; the same density-aware sampling could improve distillation of other dense-prediction tasks (e.g., normals, segmentation) where small backbones overfit to dataset-specific cues.
  • The freezing strategy suggests that catastrophic forgetting in tiny models can be largely avoided by keeping the feature extractor fixed; this principle may generalize to other metric-regression problems that build on pretrained representations.
  • The scale-only metric parameterization is a testable bet; if scenes with strong offset variations (high viewpoint changes, telephoto crops) show systematic nonzero shift, an additional shift head would be needed—the paper's limitations acknowledge this.
  • Combining BRDS with a larger teacher or multi-teacher pseudo-labeling could push tiny-model zero-shot accuracy further while preserving speed.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes DepthART, a family of compact monocular depth models (6M/11M/32M parameters) built from a TinyViM encoder and a DPT decoder, distilled from Depth Anything v2 Large. Two contributions are presented: (i) BRDS, a density-aware stratified sub-sampling scheme that reduces dataset-specific distribution bias when constructing a 1.7M-image distillation pool, and (ii) CamFT, an encoder-frozen, camera-conditioned fine-tuning protocol with camera adapters and a multi-query scale head that converts relative depth to metric depth using a single per-image scale factor. The paper reports zero-shot affine-invariant results on NYUD v2, KITTI, ETH3D, DIODE, and DDAD, metric fine-tuning results on NYUD v2 and KITTI with zero-shot transfer to several unseen datasets, and detailed latency/power measurements on RTX A6000 and Jetson Orin NX. The central claim is that DepthART consistently outperforms prior tiny MDE baselines in zero-shot generalization and metric accuracy while retaining real-time on-device efficiency.

Significance. If the empirical claims hold, this is a practically useful result: it shows that a substantial fraction of foundation-model monocular depth generalization can be transferred to models small enough for embedded deployment, and it provides a reproducible recipe (data rebalancing plus frozen-encoder calibration) that other lightweight depth systems could adopt. The paper is strong in breadth of evaluation: five zero-shot relative-depth benchmarks, four cross-dataset metric-transfer benchmarks, two fine-tuning domains, three model sizes, and careful latency/memory measurements. The ablations in Tables 3-5 are informative and the authors explicitly release code. No circular derivation is present; the gains are measured against external baselines. However, several load-bearing aspects need strengthening before the claims can be fully accepted: the scale-only metric mapping in Eq. (5) is tested only indirectly, the BRDS reduction factors are undisclosed, all results are single-run, and the zero-shot evaluation is intertwined with hyperparameter selection.

major comments (3)
  1. [Sec. 3.3, Eq. (5), Table 2] The central metric-transfer claim rests on the scale-only calibration d_metric = d_max * (s ⊙ d_rel), with no shift term. The paper itself states in Sec. 3.3 that shift correction is not universally unnecessary, and the zero-shot sets in Table 2 (iBims-1, SUN RGB-D, DIODE, DDAD, ETH3D) are exactly where per-scene shift varies with camera, indoor/outdoor content, and truncated depth range. A fixed d_max (10 m for NYUD-v2-trained, 80 m for KITTI-trained) also imposes a hard output range, which can truncate outdoor zero-shot depth. Fully fine-tuned baselines can absorb arbitrary scale and shift, so the comparison is not fully apples-to-apples. Please add a scale+shift control (or per-image affine alignment at test time) and/or report per-image scale-shift statistics on the zero-shot benchmarks to show that the claimed transfer does not depend on this restriction.
  2. [Sec. 4.3, Eq. (2)] The BRDS source-specific reduction factors γ are never disclosed. Eq. (2) defines the sampling only in terms of γ, and the text merely says that source-specific factors are applied to five large-scale datasets. These values determine the actual training set and are essential for reproducibility. In addition, BRDS hyperparameters (J=8, B=10), CamFT components (Tables 4-5), d_max, α, λ, and the top-10% truncation are selected using the same NYUD v2 and iBims-1 numbers later reported as zero-shot evidence. This makes the 'zero-shot' framing less clean than it appears. Please disclose the γ values, state which hyperparameters were fixed a priori and which were tuned, or present a held-out selection protocol.
  3. [Tables 1-5] All accuracy numbers are single-run with no seeds or error bars. Several comparisons are close enough that optimization noise could matter, e.g., Table 3 at the 0.5M budget, Table 4 (freeze-only vs. full fine-tuning), and Table 5 (adapter variants). Without repeated runs or interval estimates, the claims of 'consistently surpasses' and the module-level ablation conclusions are not statistically grounded. Add at least three seeds with mean±std for the main tables and the key ablations, or an equivalent uncertainty analysis.
minor comments (5)
  1. [Eq. (5)] s is a scalar, so the symbol ⊙ is misleading; use ordinary scalar multiplication or define broadcasting explicitly.
  2. [Eq. (2)] The notation U(D_c, m_c) is not defined. Specify that it denotes uniform sampling without replacement.
  3. [Table 2 caption] The caption refers to pink and cyan shading; if the paper is read or printed in grayscale, the distinction is hard to see. Please use letters or distinct markers in addition to color.
  4. [Fig. 3(b)] The 'scale-shift distribution' is described only by standard deviations. Define how the per-image scale and shift alignment parameters are estimated before reporting their statistics.
  5. [Abstract / Sec. 4.3] The manuscript says code is released on GitHub but gives only a project page URL. Include the repository URL in the text.

Circularity Check

0 steps flagged

No significant circularity: the core derivation is empirical, externally benchmarked, and does not reduce to its own inputs.

full rationale

The paper's derivation chain is not self-referential in a load-bearing way. Stage-1 distillation uses an external foundation teacher (Depth Anything v2) on a BRDS-selected multi-source corpus, and zero-shot affine-invariant depth is evaluated by the standard per-image least-squares scale/shift protocol on external datasets, so the claimed relative-depth generalization is not defined in terms of the method's own outputs. The metric calibration in Eq. (5), d_metric = d_max * (s ⊙ d_rel), is an explicit modeling choice rather than a fitted prediction being reported as a result: the scale head is trained with SiLog supervision against metric ground truth, and d_max is a disclosed per-dataset constant. The paper even acknowledges that 'shift correction is not universally unnecessary' (§3.3), which undercuts the universality of the scale-only assumption but does not make the claim circular. BRDS hyperparameters and CamFT design choices may have been selected using NYUD v2/KITTI/iBims-1 zero-shot numbers, which is a test-set selection concern and somewhat weakens the 'zero-shot' framing, but no equation or fitted constant is being renamed as a prediction. The only self-citations ([22], [40], [56]) are contextual or baseline references and do not carry the central construction. Therefore, no specific circular reduction can be exhibited, and the paper should be scored as free of constructional circularity.

Axiom & Free-Parameter Ledger

7 free parameters · 6 axioms · 0 invented entities

The free parameters are model-selection hyperparameters, most chosen on downstream zero-shot benchmarks, which adds tuning burden but not a mathematical derivation. The load-bearing premises are scale-only metric calibration, known intrinsics, reliability of teacher pseudo-labels, and the assumption that VAE/PCA binning captures depth-relevant bias.

free parameters (7)
  • BRDS source-specific reduction factors γ = not disclosed (five values)
    Referenced in §4.3 but never enumerated; tuning these per large-scale dataset is part of the sampling recipe and affects the 1.7M pool.
  • PCA dimension J = 8
    §4.3; dimensionality of the VAE embedding used for cell binning; chosen by hand.
  • Bins per dimension B = 10
    §4.3; grid resolution for density-aware stratification; chosen by hand.
  • Scale-head query count Q = 8
    §3.3, Eq. 7; number of learnable queries in the multi-query scale head; chosen by hand.
  • Max depth d_max = 10 (NYUD v2), 80 (KITTI)
    §4.3; fixed per fine-tuning dataset and used in Eq. 5; an input to the metric mapping rather than a learned output.
  • Distillation loss weight α and top-10% truncation = α=0.5; top-10% highest-error pixels ignored
    Eq. 4; hand-set hyperparameters of the distillation objective.
  • Fine-tuning edge loss weight λ = not given
    Eq. 8; used in CamFT but the value is omitted from the manuscript.
axioms (6)
  • domain assumption Depth Anything v2 Large teacher pseudo-depth is a reliable distillation target.
    Stage 1 depends on teacher pseudo-labels; if the teacher is biased, the student inherits the bias. Invoked in §3.2.
  • domain assumption Stable Diffusion VAE + PCA binning captures image-level distribution bias relevant to depth learning.
    BRDS relies on this embedding for density-aware stratification; no independent validation is given beyond downstream results. Invoked in §3.2.
  • domain assumption A per-image global scale s and fixed d_max are sufficient to convert relative depth to metric depth.
    Eq. 5 defines metric depth as d_max * (s ⊙ d_rel); shift correction is deliberately omitted. Invoked in §3.3.
  • domain assumption Camera intrinsics are known at training and inference.
    Pyramid camera prompts are constructed from K=(fx,fy,cx,cy); the paper lists this dependence as a limitation. Invoked in §3.3 and Limitations.
  • domain assumption The evaluation benchmarks NYUD v2, KITTI, ETH3D, DIODE, and DDAD are not part of the 1.7M BRDS training pool.
    The 11 listed sources exclude these benchmarks, which underpins the zero-shot claims. Invoked in §4.1.
  • standard math Per-image least-squares scale-shift alignment is the accepted affine-invariant depth evaluation protocol.
    Used for zero-shot affine-invariant depth in Table 1; standard in the MiDaS/Depth-Anything evaluation lineage. Invoked in §4.4.

pith-pipeline@v1.3.0-alltime-deepseek · 18467 in / 13501 out tokens · 121851 ms · 2026-08-01T19:01:17.079799+00:00 · methodology

0 comments
read the original abstract

Recent geometric foundation models (e.g., Metric3D, Depth Anything and UniDepth) have substantially improved monocular depth estimation (MDE) in both cross-scene generalization and metric-scale prediction, yet these gains have not translated to tiny models. We bridge this gap with DepthART (Depth Anything Rethought for Tiny Models), which is a compact MDE model for on-device deployment across diverse scenes. We first identify two capacity-driven bottlenecks in tiny models: (i) overfitting to dataset-specific distribution bias and (ii) unstable metric adaptation under camera shift, where full fine-tuning easily damages transferable geometry. Accordingly, DepthART combines two simple but effective strategies: a bias-resistant data sampling scheme to reduce distribution bias under the same training budget, and a camera-conditioned fine-tuning protocol that freezes the distilled encoder and adjusts metric scale conditioned on intrinsics while better preserving cross-dataset generalization. Across datasets, DepthART consistently surpasses previous tiny baselines in both zero-shot generalization and metric accuracy (e.g., zero-shot $\delta_1$=0.964 for DepthART-S on NYUD v2), and in some cases approaches heavy models. We further provide a scalable model family, with DepthART-S reaching 347/245 FPS (strict FP32) on an RTX A6000 at $224^2/448^2$, 102 FPS (TF32) on a Orin NX 8GB, and over 15 FPS (FP32) on a Jetson Nano 4GB.

Figures

Figures reproduced from arXiv: 2607.17099 by Anlong Ming, Dianqiao Lei, Feng Xue, Guofeng Zhong, Haiyang Zhang, Haozhe Wang, Mingshuai Zhao, Nicu Sebe, Wu Chen, Zhaowen Lin.

Figure 1
Figure 1. Figure 1: DepthART brings Depth Anything-style capability to tiny models, achieving strong zero-shot generalization, reliable [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Bias-resistant data sampling (BRDS) and distillation pipeline. A 44M multi-source corpus is encoded (VAE) and [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Geometry and scale stability after fine-tuning on [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Camera-conditioned fine-tuning. We freeze the distilled encoder and inject lightweight camera adapters conditioned [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Qualitative comparison of affine-invariant depth and metric 3D reconstruction. Top-left: predictions on five zero [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Model-only inference latency of DepthART on a Jetson Orin NX 8GB under different power modes. The three panels [PITH_FULL_IMAGE:figures/full_fig_p009_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

76 extracted references · 8 linked inside Pith

  1. [1]

    Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka. 2021. Adabins: Depth estimation using adaptive bins. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  2. [2]

    Shariq Farooq Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, and Matthias Müller

  3. [3]

    Lucy Chai, Richard Tucker, Zhengqi Li, Phillip Isola, and Noah Snavely. 2023. Persistent Nature: A Generative Model of Unbounded 3D Worlds. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  4. [4]

    Jaehoon Cho, Dongbo Min, Youngjung Kim, and Kwanghoon Sohn. 2021. DIML/CVL RGB-D Dataset: 2M RGB-D Images of Natural Indoor and Outdoor Scenes.arXiv preprint arXiv:2110.11590(2021). arXiv:2110.11590

  5. [5]

    Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus En- zweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. 2016. The cityscapes dataset for semantic urban scene understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  6. [6]

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Ima- geNet: A large-scale hierarchical image database. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  7. [7]

    David Eigen and Rob Fergus. 2015. Predicting depth, surface normals and se- mantic labels with a common multi-scale convolutional architecture. InIEEE international conference on computer vision

  8. [8]

    Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao. 2018. Deep ordinal regression network for monocular depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion

  9. [9]

    Xiao Fu, Wei Yin, Mu Hu, Kaixuan Wang, Yuexin Ma, Ping Tan, Shaojie Shen, Dahua Lin, and Xiaoxiao Long. 2024. Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image. InECCV

  10. [10]

    Ming Gui, Johannes Schusterbauer, Ulrich Prestel, Pingchuan Ma, Dmytro Ko- tovenko, Olga Grebenkova, Stefan Andreas Baumann, Vincent Tao Hu, and Björn Ommer. 2025. DepthFM: Fast Generative Monocular Depth Estimation with Flow Matching. InProceedings of the AAAI Conference on Artificial Intelligence

  11. [11]

    Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Allan Raventos, and Adrien Gaidon

  12. [12]

    Jing He, Haodong Li, Mingzhi Sheng, and Ying-Cong Chen. 2025. Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model. arXiv:2512.01030 [cs.CV] https://arxiv.org/abs/2512.01030

  13. [13]

    Mu Hu, Wei Yin, Chi Zhang, Zhipeng Cai, Xiaoxiao Long, Hao Chen, Kaixuan Wang, Gang Yu, Chunhua Shen, and Shaojie Shen. 2024. Metric3D v2: A Versatile Monocular Geometric Foundation Model for Zero-Shot Metric Depth and Surface Normal Estimation.IEEE Transactions on Pattern Analysis and Machine Intelligence 46, 12 (2024), 10579–10596

  14. [14]

    Yiwen Hua, Puneet Kohli, Pritish Uplavikar, Anand Ravi, Saravana Gunaseelan, Jason Orozco, and Edward Li. 2020. Holopix50k: A large-scale in-the-wild stereo image dataset.arXiv preprint arXiv:2003.11172(2020)

  15. [15]

    Xinyu Huang, Xinjing Cheng, Qichuan Geng, Binbin Cao, Dingfu Zhou, Peng Wang, Yuanqing Lin, and Ruigang Yang. 2018. The apolloscape dataset for autonomous driving. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition workshop

  16. [16]

    Xueyang Kang, Zhengkang Xiang, Zezheng Zhang, and Kourosh Khoshelham

  17. [17]

    Fatemeh Baran Karimi, Amir Mehrpanah, and Reza Rawassizadeh. 2024. Light- Depth: A resource efficient depth estimation approach for dealing with ground truth sparsity via curriculum learning.Robotics and Autonomous Systems181 (2024), 104784

  18. [18]

    Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Metzger, Rodrigo Caye Daudt, and Konrad Schindler. 2024. Repurposing diffusion-based image genera- tors for monocular depth estimation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  19. [19]

    Krishnateja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Abir De, and Rishabh Iyer. 2021. GRAD-MATCH: Gradient Matching based Data Subset Selection for Efficient Deep Model Training. InProceedings of the International Conference on Machine Learning

  20. [20]

    Tobias Koch, Lukas Liebel, Friedrich Fraundorfer, and Marco Korner. 2018. Eval- uation of cnn-based single-image depth estimation methods. InProceedings of the European Conference on Computer Vision Workshops

  21. [21]

    Jaewook Lee, Filippo Aleotti, Diego Mazala, Guillermo Garcia-Hernando, Sara Vicente, Oliver James Johnston, Isabel Kraus-Liang, Jakub Powierza, Donghoon Shin, Jon E Froehlich, et al. 2025. Imaginatear: Ai-assisted in-situ authoring in augmented reality. InProceedings of the Annual ACM Symposium on User Interface Software and Technology

  22. [22]

    Yihao Liu, Feng Xue, Anlong Ming, Mingshuai Zhao, Huadong Ma, and Nicu Sebe

  23. [23]

    Jiahuan Long and Xin Zhou. 2025. LMDepth: Lightweight Mamba-based Monocu- lar Depth Estimation for Real-World Deployment.arXiv preprint arXiv:2505.00980 (2025). arXiv:2505.00980

  24. [24]

    Xiaowen Ma, Zhenliang Ni, and Xinghao Chen. 2025. Tinyvim: Frequency decou- pling for tiny hybrid vision mamba. InProceedings of the IEEE/CVF International Conference on Computer Vision

  25. [25]

    Nicholas Merrill, Patrick Geneva, Michael Paton, Zixin Li, Yulin Yang, and Guo- quan Huang. 2024. Fast and Robust Learned Single-View Depth-aided Monocular Visual-Inertial Initialization.The International Journal of Robotics Research43, 2 (2024), 237–257

  26. [26]

    Baharan Mirzasoleiman, Jeff Bilmes, and Jure Leskovec. 2020. Coresets for Data- efficient Training of Machine Learning Models. InProceedings of the International Conference on Machine Learning

  27. [27]

    Lorenzo Papa, Paolo Russo, and Irene Amerini. 2023. METER: A mobile vision transformer architecture for monocular depth estimation.IEEE Transactions on Circuits and Systems for Video Technology33, 10 (2023), 5882–5893

  28. [28]

    Luigi Piccinelli, Christos Sakaridis, Mattia Segu, Yung-Hsu Yang, Siyuan Li, Wim Abbeloos, and Luc Van Gool. 2025. UniK3D: Universal Camera Monocular 3D Estimation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  29. [29]

    Luigi Piccinelli, Christos Sakaridis, Yung-Hsu Yang, Mattia Segu, Siyuan Li, Wim Abbeloos, and Luc Van Gool. 2025. UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler.IEEE Transactions on Pattern Analysis and Machine Intelligence(2025), 1–14

  30. [30]

    Luigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis, Mattia Segu, Siyuan Li, Luc Van Gool, and Fisher Yu. 2024. UniDepth: Universal Monocular Metric Depth Estimation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  31. [31]

    René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. 2021. Vision Transformers for Dense Prediction. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 12179–12188. doi:10.1109/ICCV48922.2021.01196

  32. [32]

    René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun

  33. [33]

    Xuanchi Ren, Tianchang Shen, Jiahui Huang, Huan Ling, Yifan Lu, Merlin Nimier- David, Thomas Müller, Alexander Keller, Sanja Fidler, and Jun Gao. 2025. GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Con- trol. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  34. [34]

    Zeyu Ren, Zeyu Zhang, Wukai Li, Qingxiang Liu, and Hao Tang. 2026. Any- Depth: Depth Estimation Made Easy.arXiv preprint arXiv:2601.02760(2026). arXiv:2601.02760

  35. [35]

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion

  36. [36]

    Michael Rudolph, Youssef Dawoud, Ronja Güldenring, Lazaros Nalpantidis, and Vasileios Belagiannis. 2022. Lightweight monocular depth estimation through guided decoding. InProceedings of the IEEE International Conference on Robotics and Automation

  37. [37]

    Thomas Schops, Johannes L Schonberger, Silvano Galliani, Torsten Sattler, Kon- rad Schindler, Marc Pollefeys, and Andreas Geiger. 2017. A multi-view stereo benchmark with high-resolution images and multi-camera videos. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  38. [38]

    Ozan Sener and Silvio Savarese. 2018. Active Learning for Convolutional Neu- ral Networks: A Core-Set Approach. InInternational Conference on Learning Representations. DepthART: Scaling Foundation Monocular Depth to Tiny Models Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

  39. [39]

    Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. 2019. Objects365: A large-scale, high-quality dataset for object detection. InProceedings of the IEEE/CVF International Conference on Com- puter Vision

  40. [40]

    Fei Sheng, Feng Xue, Yicong Chang, Wenteng Liang, and Anlong Ming. 2022. Monocular depth distribution alignment with low computation. InProceedings of the IEEE International Conference on Robotics and Automation

  41. [41]

    Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. 2012. Indoor segmentation and support inference from rgbd images. InProceedings of the European Conference on Computer Vision

  42. [42]

    Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao. 2015. Sun rgb-d: A rgb-d scene understanding benchmark suite. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  43. [43]

    Wen Su, Haifeng Zhang, Jia Li, Wenzhen Yang, and Zengfu Wang. 2019. Monocu- lar depth estimation as regression of classification using piled residual networks. InProceedings of the ACM International Conference on Multimedia. 2161–2169

  44. [44]

    Xiaohan Tu, Cheng Xu, Siping Liu, Renfa Li, Guoqi Xie, Jing Huang, and Lau- rence Tianruo Yang. 2021. Efficient Monocular Depth Estimation for Edge Devices in Internet of Things.IEEE Transactions on Industrial Informatics (TII)17, 4 (2021), 2821–2832

  45. [45]

    Jonas Uhrig, Nick Schneider, Lukas Schneider, Uwe Franke, Thomas Brox, and Andreas Geiger. 2017. Sparsity Invariant CNNs. InProceedings of the International Conference on 3D Vision

  46. [46]

    Igor Vasiljevic, Nick Kolkin, Shanyi Zhang, Ruotian Luo, Haochen Wang, Falcon Z Dai, Andrea F Daniele, Mohammadreza Mostajabi, Steven Basart, Matthew R Walter, et al. 2019. Diode: A dense indoor and outdoor depth dataset.arXiv preprint arXiv:1908.00463(2019). arXiv:1908.00463

  47. [47]

    Alexander Veicht, Paul-Edouard Sarlin, Philipp Lindenberger, and Marc Polle- feys. 2024. GeoCalib: Single-image Calibration with Geometric Optimization. In Proceedings of the European Conference on Computer Vision

  48. [48]

    Ruicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang, Yu Deng, Xin Tong, and Jiaolong Yang. 2025. MoGe: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  49. [49]

    Ruicheng Wang, Sicheng Xu, Yue Dong, Yu Deng, Jianfeng Xiang, Zelong Lv, Guangzhong Sun, Xin Tong, and Jiaolong Yang. 2025. MoGe-2: Accurate Monoc- ular Geometry with Metric Scale and Sharp Details. arXiv:2507.02546 [cs.CV] https://arxiv.org/abs/2507.02546

  50. [50]

    Tobias Weyand, Andre Araujo, Bingyi Cao, and Jack Sim. 2020. Google land- marks dataset v2-a large-scale benchmark for instance-level recognition and retrieval. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  51. [51]

    Diana Wofk, Fangchang Ma, Tien-Ju Yang, Sertac Karaman, and Vivienne Sze

  52. [52]

    Zizhang Wu, Zhuozheng Li, Zhi-Gang Fan, Yunzhe Wu, Jian Pu, and Xianzhi Li. 2023. V2Depth: Monocular depth estimation via feature-Level virtual-view simulation and refinement. InProceedings of the ACM International Conference on Multimedia. 688–697

  53. [53]

    Ke Xian, Chunhua Shen, Zhiguo Cao, Hao Lu, Yang Xiao, Ruibo Li, and Zhenbo Luo. 2018. Monocular relative depth perception with web stereo data supervi- sion. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  54. [54]

    Ke Xian, Jianming Zhang, Oliver Wang, Long Mai, Zhe Lin, and Zhiguo Cao. 2020. Structure-Guided Ranking Loss for Single Image Depth Prediction. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  55. [55]

    Guangkai Xu, Yongtao Ge, Mingyu Liu, Chengxiang Fan, Kangyang Xie, Zhiyue Zhao, Hao Chen, and Chunhua Shen. 2025. What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?. InInternational Conference on Learning Representations

  56. [56]

    Feng Xue, Junfeng Cao, Yu Zhou, Fei Sheng, Yankai Wang, and Anlong Ming

  57. [57]

    Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Heng- shuang Zhao. 2024. Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  58. [58]

    Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. 2024. Depth Anything V2. InAdvances in Neural Information Processing Systems, Vol. 37. 21875–21911

  59. [59]

    Wei Yin, Yifan Liu, and Chunhua Shen. 2021. Virtual normal: Enforcing geometric constraints for accurate and robust depth prediction.TPAMI44, 10 (2021), 7282– 7295

  60. [60]

    Wei Yin, Chi Zhang, Hao Chen, Zhipeng Cai, Gang Yu, Kaixuan Wang, Xiaozhi Chen, and Chunhua Shen. 2023. Metric3D: Towards Zero-shot Metric 3D Predic- tion from A Single Image. InProceedings of the IEEE/CVF International Conference on Computer Vision. 9009–9019. doi:10.1109/ICCV51070.2023.00830

  61. [61]

    Naoki Yokoyama, Sehoon Ha, Dhruv Batra, Jiuguang Wang, and Bernadette Bucher. 2024. Vision-Language Frontier Maps for Zero-Shot Semantic Navigation. InProceedings of the IEEE International Conference on Robotics and Automation

  62. [62]

    Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell. 2020. Bdd100k: A diverse driving dataset for heterogeneous multitask learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  63. [63]

    Fisher Yu, Yinda Zhang, Shuran Song, Ari Seff, and Jianxiong Xiao. 2015. LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop.arXiv preprint arXiv:1506.03365(2015)

  64. [64]

    Songsong Yu, Yifan Wang, Yunzhi Zhuge, Lijun Wang, and Huchuan Lu. 2024. DME: Unveiling the Bias for Better Generalized Monocular Depth Estimation. In Proceedings of the AAAI Conference on Artificial Intelligence

  65. [65]

    Denis Zavadski, Johann-Friedrich Feiden, and Carsten Rother. 2024. ControlNet- XS: Rethinking the Control of Text-to-Image Diffusion Models as Feedback- Control Systems. InProceedings of the European Conference on Computer Vision

  66. [66]

    Longjian Zeng, Zunjie Zhu, Rongfeng Lu, Ming Lu, Bolun Zheng, Chenggang Yan, and Anke Xue. 2025. DepthDark: robust monocular depth estimation for low-light environments. InProceedings of the ACM International Conference on Multimedia. 11239–11248

  67. [67]

    Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding conditional con- trol to text-to-image diffusion models. InProceedings of the IEEE/CVF International Conference on Computer Vision

  68. [68]

    Wendong Zhang, Feng Gao, Bingbing Ni, Lingyu Duan, Yichao Yan, Jingwei Xu, and Xiaokang Yang. 2018. Depth structure preserving scene image generation. InProceedings of the ACM International Conference on Multimedia. 727–736

  69. [69]

    Kecheng Zheng, Zheng-Jun Zha, Yang Cao, Xuejin Chen, and Feng Wu. 2018. La- net: Layout-aware dense network for monocular depth estimation. InProceedings of the ACM International Conference on Multimedia. 1381–1388

  70. [2019]

    In Proceedings of the IEEE International Conference on Robotics and Automation

    Fastdepth: Fast monocular depth estimation on embedded systems. In Proceedings of the IEEE International Conference on Robotics and Automation

  71. [2020]

    InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    3d packing for self-supervised monocular depth estimation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  72. [2021]

    Boundary-induced and scene-aggregated network for monocular depth prediction.Pattern Recognition115 (2021), 107901

  73. [2022]

    doi:10.1109/TPAMI.2020.3019967

    Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero- Shot Cross-Dataset Transfer.IEEE Transactions on Pattern Analysis and Machine Intelligence44, 3 (2022), 1623–1637. doi:10.1109/TPAMI.2020.3019967

  74. [2023]

    arXiv:2302.12288 [cs.CV] https://arxiv.org/abs/2302.12288

    ZoeDepth: Zero-Shot Transfer by Combining Relative and Metric Depth. arXiv:2302.12288 [cs.CV] https://arxiv.org/abs/2302.12288

  75. [2024]

    InProceedings of the ACM International Conference on Multimedia

    SM4Depth: Seamless Monocular Metric Depth Estimation across Multiple Cameras and Scenes by One Model. InProceedings of the ACM International Conference on Multimedia

  76. [2025]

    InProceedings of the ACM International Conference on Multimedia

    Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion. InProceedings of the ACM International Conference on Multimedia. 9375–9384