REVIEW 3 major objections 5 minor 76 references
A 6–33M-parameter model can carry foundation-level monocular depth generalization and metric accuracy if trained with bias-resistant sampling and encoder-frozen, camera-conditioned fine-tuning.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 19:01 UTC pith:CFFTP3JL
load-bearing objection Solid tiny-model depth recipe with convincing relative-depth gains; metric transfer rests on a scale-only mapping that needs an explicit shift control before I'd trust it fully. the 3 major comments →
DepthART: Scaling Foundation Monocular Depth to Tiny Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
DepthART claims that the two capacity-driven bottlenecks for compact depth models are overfitting to dominant dataset-specific distributions and forgetting of transferable geometry during metric fine-tuning. The paper's central claim is that both can be fixed without changing the tiny architecture: bias-resistant sampling performs density-aware stratified selection in a compressed visual feature space to suppress overrepresented scene modes, and camera-conditioned fine-tuning freezes the encoder, injects multi-resolution camera prompts through small cross-attention adapters, and regresses a single per-image scale factor to convert relative depth to metric depth. With this recipe, DepthART-S
What carries the argument
BRDS (bias-resistant data sampling): density-aware stratified sampling in a visual feature space, which caps the number of samples drawn from dense, dominant scene modes and guarantees coverage of sparse long-tail modes, producing a 1.7M-image bias-resistant pool. CamFT (camera-conditioned fine-tuning): a frozen distilled encoder, pyramid camera-ray prompts injected through small cross-attention adapters, and a multi-query scale head that outputs one positive scale factor; metric depth is computed as the relative depth multiplied by that scale and a fixed maximum depth. The critical mechanism is that freezing the encoder preserves the transferred relative-depth structure, while the adapters
Load-bearing premise
The claim rests on the premise that metric depth can be obtained from the frozen relative-depth encoder by a single per-image scalar multiplied onto the relative depth with a fixed maximum, with no per-image offset or non-scalar correction.
What would settle it
Run DepthART, fine-tuned on KITTI, on a collection of handheld, drone, and telephoto images. Fit per-image optimal scale and shift to ground truth; if the optimal shift is systematically large and varies with scene content, the scale-only metric parameterization is falsified.
If this is right
- If correct, real-time on-device depth estimation can retain a large share of foundation-model generalization, closing the gap between research models and deployed systems.
- The recipe yields specific speedups: DepthART-S/B/L run at 347/298/191 FPS at 224² and 245/204/124 FPS at 448² on an RTX A6000 in strict FP32, with >15 FPS on a Jetson Nano.
- Zero-shot metric depth becomes practical without post-hoc scale-and-shift alignment, since metric predictions are directly usable as 3D point clouds.
- Density-aware sampling is more sample-efficient than random or k-means selection for distillation, improving zero-shot transfer even at fixed data budgets.
- Encoder-frozen, camera-conditioned fine-tuning substantially limits catastrophic forgetting compared to full fine-tuning, as measured by affine-invariant RMSE on out-of-distribution scenes.
Where Pith is reading between the lines
- The result implies that data diversity, not raw data volume, is the binding constraint for tiny models; the same density-aware sampling could improve distillation of other dense-prediction tasks (e.g., normals, segmentation) where small backbones overfit to dataset-specific cues.
- The freezing strategy suggests that catastrophic forgetting in tiny models can be largely avoided by keeping the feature extractor fixed; this principle may generalize to other metric-regression problems that build on pretrained representations.
- The scale-only metric parameterization is a testable bet; if scenes with strong offset variations (high viewpoint changes, telephoto crops) show systematic nonzero shift, an additional shift head would be needed—the paper's limitations acknowledge this.
- Combining BRDS with a larger teacher or multi-teacher pseudo-labeling could push tiny-model zero-shot accuracy further while preserving speed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DepthART, a family of compact monocular depth models (6M/11M/32M parameters) built from a TinyViM encoder and a DPT decoder, distilled from Depth Anything v2 Large. Two contributions are presented: (i) BRDS, a density-aware stratified sub-sampling scheme that reduces dataset-specific distribution bias when constructing a 1.7M-image distillation pool, and (ii) CamFT, an encoder-frozen, camera-conditioned fine-tuning protocol with camera adapters and a multi-query scale head that converts relative depth to metric depth using a single per-image scale factor. The paper reports zero-shot affine-invariant results on NYUD v2, KITTI, ETH3D, DIODE, and DDAD, metric fine-tuning results on NYUD v2 and KITTI with zero-shot transfer to several unseen datasets, and detailed latency/power measurements on RTX A6000 and Jetson Orin NX. The central claim is that DepthART consistently outperforms prior tiny MDE baselines in zero-shot generalization and metric accuracy while retaining real-time on-device efficiency.
Significance. If the empirical claims hold, this is a practically useful result: it shows that a substantial fraction of foundation-model monocular depth generalization can be transferred to models small enough for embedded deployment, and it provides a reproducible recipe (data rebalancing plus frozen-encoder calibration) that other lightweight depth systems could adopt. The paper is strong in breadth of evaluation: five zero-shot relative-depth benchmarks, four cross-dataset metric-transfer benchmarks, two fine-tuning domains, three model sizes, and careful latency/memory measurements. The ablations in Tables 3-5 are informative and the authors explicitly release code. No circular derivation is present; the gains are measured against external baselines. However, several load-bearing aspects need strengthening before the claims can be fully accepted: the scale-only metric mapping in Eq. (5) is tested only indirectly, the BRDS reduction factors are undisclosed, all results are single-run, and the zero-shot evaluation is intertwined with hyperparameter selection.
major comments (3)
- [Sec. 3.3, Eq. (5), Table 2] The central metric-transfer claim rests on the scale-only calibration d_metric = d_max * (s ⊙ d_rel), with no shift term. The paper itself states in Sec. 3.3 that shift correction is not universally unnecessary, and the zero-shot sets in Table 2 (iBims-1, SUN RGB-D, DIODE, DDAD, ETH3D) are exactly where per-scene shift varies with camera, indoor/outdoor content, and truncated depth range. A fixed d_max (10 m for NYUD-v2-trained, 80 m for KITTI-trained) also imposes a hard output range, which can truncate outdoor zero-shot depth. Fully fine-tuned baselines can absorb arbitrary scale and shift, so the comparison is not fully apples-to-apples. Please add a scale+shift control (or per-image affine alignment at test time) and/or report per-image scale-shift statistics on the zero-shot benchmarks to show that the claimed transfer does not depend on this restriction.
- [Sec. 4.3, Eq. (2)] The BRDS source-specific reduction factors γ are never disclosed. Eq. (2) defines the sampling only in terms of γ, and the text merely says that source-specific factors are applied to five large-scale datasets. These values determine the actual training set and are essential for reproducibility. In addition, BRDS hyperparameters (J=8, B=10), CamFT components (Tables 4-5), d_max, α, λ, and the top-10% truncation are selected using the same NYUD v2 and iBims-1 numbers later reported as zero-shot evidence. This makes the 'zero-shot' framing less clean than it appears. Please disclose the γ values, state which hyperparameters were fixed a priori and which were tuned, or present a held-out selection protocol.
- [Tables 1-5] All accuracy numbers are single-run with no seeds or error bars. Several comparisons are close enough that optimization noise could matter, e.g., Table 3 at the 0.5M budget, Table 4 (freeze-only vs. full fine-tuning), and Table 5 (adapter variants). Without repeated runs or interval estimates, the claims of 'consistently surpasses' and the module-level ablation conclusions are not statistically grounded. Add at least three seeds with mean±std for the main tables and the key ablations, or an equivalent uncertainty analysis.
minor comments (5)
- [Eq. (5)] s is a scalar, so the symbol ⊙ is misleading; use ordinary scalar multiplication or define broadcasting explicitly.
- [Eq. (2)] The notation U(D_c, m_c) is not defined. Specify that it denotes uniform sampling without replacement.
- [Table 2 caption] The caption refers to pink and cyan shading; if the paper is read or printed in grayscale, the distinction is hard to see. Please use letters or distinct markers in addition to color.
- [Fig. 3(b)] The 'scale-shift distribution' is described only by standard deviations. Define how the per-image scale and shift alignment parameters are estimated before reporting their statistics.
- [Abstract / Sec. 4.3] The manuscript says code is released on GitHub but gives only a project page URL. Include the repository URL in the text.
Circularity Check
No significant circularity: the core derivation is empirical, externally benchmarked, and does not reduce to its own inputs.
full rationale
The paper's derivation chain is not self-referential in a load-bearing way. Stage-1 distillation uses an external foundation teacher (Depth Anything v2) on a BRDS-selected multi-source corpus, and zero-shot affine-invariant depth is evaluated by the standard per-image least-squares scale/shift protocol on external datasets, so the claimed relative-depth generalization is not defined in terms of the method's own outputs. The metric calibration in Eq. (5), d_metric = d_max * (s ⊙ d_rel), is an explicit modeling choice rather than a fitted prediction being reported as a result: the scale head is trained with SiLog supervision against metric ground truth, and d_max is a disclosed per-dataset constant. The paper even acknowledges that 'shift correction is not universally unnecessary' (§3.3), which undercuts the universality of the scale-only assumption but does not make the claim circular. BRDS hyperparameters and CamFT design choices may have been selected using NYUD v2/KITTI/iBims-1 zero-shot numbers, which is a test-set selection concern and somewhat weakens the 'zero-shot' framing, but no equation or fitted constant is being renamed as a prediction. The only self-citations ([22], [40], [56]) are contextual or baseline references and do not carry the central construction. Therefore, no specific circular reduction can be exhibited, and the paper should be scored as free of constructional circularity.
Axiom & Free-Parameter Ledger
free parameters (7)
- BRDS source-specific reduction factors γ =
not disclosed (five values)
- PCA dimension J =
8
- Bins per dimension B =
10
- Scale-head query count Q =
8
- Max depth d_max =
10 (NYUD v2), 80 (KITTI)
- Distillation loss weight α and top-10% truncation =
α=0.5; top-10% highest-error pixels ignored
- Fine-tuning edge loss weight λ =
not given
axioms (6)
- domain assumption Depth Anything v2 Large teacher pseudo-depth is a reliable distillation target.
- domain assumption Stable Diffusion VAE + PCA binning captures image-level distribution bias relevant to depth learning.
- domain assumption A per-image global scale s and fixed d_max are sufficient to convert relative depth to metric depth.
- domain assumption Camera intrinsics are known at training and inference.
- domain assumption The evaluation benchmarks NYUD v2, KITTI, ETH3D, DIODE, and DDAD are not part of the 1.7M BRDS training pool.
- standard math Per-image least-squares scale-shift alignment is the accepted affine-invariant depth evaluation protocol.
read the original abstract
Recent geometric foundation models (e.g., Metric3D, Depth Anything and UniDepth) have substantially improved monocular depth estimation (MDE) in both cross-scene generalization and metric-scale prediction, yet these gains have not translated to tiny models. We bridge this gap with DepthART (Depth Anything Rethought for Tiny Models), which is a compact MDE model for on-device deployment across diverse scenes. We first identify two capacity-driven bottlenecks in tiny models: (i) overfitting to dataset-specific distribution bias and (ii) unstable metric adaptation under camera shift, where full fine-tuning easily damages transferable geometry. Accordingly, DepthART combines two simple but effective strategies: a bias-resistant data sampling scheme to reduce distribution bias under the same training budget, and a camera-conditioned fine-tuning protocol that freezes the distilled encoder and adjusts metric scale conditioned on intrinsics while better preserving cross-dataset generalization. Across datasets, DepthART consistently surpasses previous tiny baselines in both zero-shot generalization and metric accuracy (e.g., zero-shot $\delta_1$=0.964 for DepthART-S on NYUD v2), and in some cases approaches heavy models. We further provide a scalable model family, with DepthART-S reaching 347/245 FPS (strict FP32) on an RTX A6000 at $224^2/448^2$, 102 FPS (TF32) on a Orin NX 8GB, and over 15 FPS (FP32) on a Jetson Nano 4GB.
Figures
Reference graph
Works this paper leans on
-
[1]
Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka. 2021. Adabins: Depth estimation using adaptive bins. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2021
-
[2]
Shariq Farooq Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, and Matthias Müller
-
[3]
Lucy Chai, Richard Tucker, Zhengqi Li, Phillip Isola, and Noah Snavely. 2023. Persistent Nature: A Generative Model of Unbounded 3D Worlds. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2023
-
[4]
Jaehoon Cho, Dongbo Min, Youngjung Kim, and Kwanghoon Sohn. 2021. DIML/CVL RGB-D Dataset: 2M RGB-D Images of Natural Indoor and Outdoor Scenes.arXiv preprint arXiv:2110.11590(2021). arXiv:2110.11590
Pith/arXiv arXiv 2021
-
[5]
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus En- zweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. 2016. The cityscapes dataset for semantic urban scene understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2016
-
[6]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Ima- geNet: A large-scale hierarchical image database. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2009
-
[7]
David Eigen and Rob Fergus. 2015. Predicting depth, surface normals and se- mantic labels with a common multi-scale convolutional architecture. InIEEE international conference on computer vision
2015
-
[8]
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao. 2018. Deep ordinal regression network for monocular depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion
2018
-
[9]
Xiao Fu, Wei Yin, Mu Hu, Kaixuan Wang, Yuexin Ma, Ping Tan, Shaojie Shen, Dahua Lin, and Xiaoxiao Long. 2024. Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image. InECCV
2024
-
[10]
Ming Gui, Johannes Schusterbauer, Ulrich Prestel, Pingchuan Ma, Dmytro Ko- tovenko, Olga Grebenkova, Stefan Andreas Baumann, Vincent Tao Hu, and Björn Ommer. 2025. DepthFM: Fast Generative Monocular Depth Estimation with Flow Matching. InProceedings of the AAAI Conference on Artificial Intelligence
2025
-
[11]
Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Allan Raventos, and Adrien Gaidon
-
[12]
Jing He, Haodong Li, Mingzhi Sheng, and Ying-Cong Chen. 2025. Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model. arXiv:2512.01030 [cs.CV] https://arxiv.org/abs/2512.01030
Pith/arXiv arXiv 2025
-
[13]
Mu Hu, Wei Yin, Chi Zhang, Zhipeng Cai, Xiaoxiao Long, Hao Chen, Kaixuan Wang, Gang Yu, Chunhua Shen, and Shaojie Shen. 2024. Metric3D v2: A Versatile Monocular Geometric Foundation Model for Zero-Shot Metric Depth and Surface Normal Estimation.IEEE Transactions on Pattern Analysis and Machine Intelligence 46, 12 (2024), 10579–10596
2024
-
[14]
Yiwen Hua, Puneet Kohli, Pritish Uplavikar, Anand Ravi, Saravana Gunaseelan, Jason Orozco, and Edward Li. 2020. Holopix50k: A large-scale in-the-wild stereo image dataset.arXiv preprint arXiv:2003.11172(2020)
Pith/arXiv arXiv 2020
-
[15]
Xinyu Huang, Xinjing Cheng, Qichuan Geng, Binbin Cao, Dingfu Zhou, Peng Wang, Yuanqing Lin, and Ruigang Yang. 2018. The apolloscape dataset for autonomous driving. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition workshop
2018
-
[16]
Xueyang Kang, Zhengkang Xiang, Zezheng Zhang, and Kourosh Khoshelham
-
[17]
Fatemeh Baran Karimi, Amir Mehrpanah, and Reza Rawassizadeh. 2024. Light- Depth: A resource efficient depth estimation approach for dealing with ground truth sparsity via curriculum learning.Robotics and Autonomous Systems181 (2024), 104784
2024
-
[18]
Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Metzger, Rodrigo Caye Daudt, and Konrad Schindler. 2024. Repurposing diffusion-based image genera- tors for monocular depth estimation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2024
-
[19]
Krishnateja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Abir De, and Rishabh Iyer. 2021. GRAD-MATCH: Gradient Matching based Data Subset Selection for Efficient Deep Model Training. InProceedings of the International Conference on Machine Learning
2021
-
[20]
Tobias Koch, Lukas Liebel, Friedrich Fraundorfer, and Marco Korner. 2018. Eval- uation of cnn-based single-image depth estimation methods. InProceedings of the European Conference on Computer Vision Workshops
2018
-
[21]
Jaewook Lee, Filippo Aleotti, Diego Mazala, Guillermo Garcia-Hernando, Sara Vicente, Oliver James Johnston, Isabel Kraus-Liang, Jakub Powierza, Donghoon Shin, Jon E Froehlich, et al. 2025. Imaginatear: Ai-assisted in-situ authoring in augmented reality. InProceedings of the Annual ACM Symposium on User Interface Software and Technology
2025
-
[22]
Yihao Liu, Feng Xue, Anlong Ming, Mingshuai Zhao, Huadong Ma, and Nicu Sebe
-
[23]
Jiahuan Long and Xin Zhou. 2025. LMDepth: Lightweight Mamba-based Monocu- lar Depth Estimation for Real-World Deployment.arXiv preprint arXiv:2505.00980 (2025). arXiv:2505.00980
Pith/arXiv arXiv 2025
-
[24]
Xiaowen Ma, Zhenliang Ni, and Xinghao Chen. 2025. Tinyvim: Frequency decou- pling for tiny hybrid vision mamba. InProceedings of the IEEE/CVF International Conference on Computer Vision
2025
-
[25]
Nicholas Merrill, Patrick Geneva, Michael Paton, Zixin Li, Yulin Yang, and Guo- quan Huang. 2024. Fast and Robust Learned Single-View Depth-aided Monocular Visual-Inertial Initialization.The International Journal of Robotics Research43, 2 (2024), 237–257
2024
-
[26]
Baharan Mirzasoleiman, Jeff Bilmes, and Jure Leskovec. 2020. Coresets for Data- efficient Training of Machine Learning Models. InProceedings of the International Conference on Machine Learning
2020
-
[27]
Lorenzo Papa, Paolo Russo, and Irene Amerini. 2023. METER: A mobile vision transformer architecture for monocular depth estimation.IEEE Transactions on Circuits and Systems for Video Technology33, 10 (2023), 5882–5893
2023
-
[28]
Luigi Piccinelli, Christos Sakaridis, Mattia Segu, Yung-Hsu Yang, Siyuan Li, Wim Abbeloos, and Luc Van Gool. 2025. UniK3D: Universal Camera Monocular 3D Estimation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2025
-
[29]
Luigi Piccinelli, Christos Sakaridis, Yung-Hsu Yang, Mattia Segu, Siyuan Li, Wim Abbeloos, and Luc Van Gool. 2025. UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler.IEEE Transactions on Pattern Analysis and Machine Intelligence(2025), 1–14
2025
-
[30]
Luigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis, Mattia Segu, Siyuan Li, Luc Van Gool, and Fisher Yu. 2024. UniDepth: Universal Monocular Metric Depth Estimation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2024
-
[31]
René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. 2021. Vision Transformers for Dense Prediction. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 12179–12188. doi:10.1109/ICCV48922.2021.01196
arXiv 2021
-
[32]
René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun
-
[33]
Xuanchi Ren, Tianchang Shen, Jiahui Huang, Huan Ling, Yifan Lu, Merlin Nimier- David, Thomas Müller, Alexander Keller, Sanja Fidler, and Jun Gao. 2025. GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Con- trol. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2025
-
[34]
Zeyu Ren, Zeyu Zhang, Wukai Li, Qingxiang Liu, and Hao Tang. 2026. Any- Depth: Depth Estimation Made Easy.arXiv preprint arXiv:2601.02760(2026). arXiv:2601.02760
arXiv 2026
-
[35]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion
2022
-
[36]
Michael Rudolph, Youssef Dawoud, Ronja Güldenring, Lazaros Nalpantidis, and Vasileios Belagiannis. 2022. Lightweight monocular depth estimation through guided decoding. InProceedings of the IEEE International Conference on Robotics and Automation
2022
-
[37]
Thomas Schops, Johannes L Schonberger, Silvano Galliani, Torsten Sattler, Kon- rad Schindler, Marc Pollefeys, and Andreas Geiger. 2017. A multi-view stereo benchmark with high-resolution images and multi-camera videos. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2017
-
[38]
Ozan Sener and Silvio Savarese. 2018. Active Learning for Convolutional Neu- ral Networks: A Core-Set Approach. InInternational Conference on Learning Representations. DepthART: Scaling Foundation Monocular Depth to Tiny Models Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
2018
-
[39]
Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. 2019. Objects365: A large-scale, high-quality dataset for object detection. InProceedings of the IEEE/CVF International Conference on Com- puter Vision
2019
-
[40]
Fei Sheng, Feng Xue, Yicong Chang, Wenteng Liang, and Anlong Ming. 2022. Monocular depth distribution alignment with low computation. InProceedings of the IEEE International Conference on Robotics and Automation
2022
-
[41]
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. 2012. Indoor segmentation and support inference from rgbd images. InProceedings of the European Conference on Computer Vision
2012
-
[42]
Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao. 2015. Sun rgb-d: A rgb-d scene understanding benchmark suite. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2015
-
[43]
Wen Su, Haifeng Zhang, Jia Li, Wenzhen Yang, and Zengfu Wang. 2019. Monocu- lar depth estimation as regression of classification using piled residual networks. InProceedings of the ACM International Conference on Multimedia. 2161–2169
2019
-
[44]
Xiaohan Tu, Cheng Xu, Siping Liu, Renfa Li, Guoqi Xie, Jing Huang, and Lau- rence Tianruo Yang. 2021. Efficient Monocular Depth Estimation for Edge Devices in Internet of Things.IEEE Transactions on Industrial Informatics (TII)17, 4 (2021), 2821–2832
2021
-
[45]
Jonas Uhrig, Nick Schneider, Lukas Schneider, Uwe Franke, Thomas Brox, and Andreas Geiger. 2017. Sparsity Invariant CNNs. InProceedings of the International Conference on 3D Vision
2017
-
[46]
Igor Vasiljevic, Nick Kolkin, Shanyi Zhang, Ruotian Luo, Haochen Wang, Falcon Z Dai, Andrea F Daniele, Mohammadreza Mostajabi, Steven Basart, Matthew R Walter, et al. 2019. Diode: A dense indoor and outdoor depth dataset.arXiv preprint arXiv:1908.00463(2019). arXiv:1908.00463
Pith/arXiv arXiv 2019
-
[47]
Alexander Veicht, Paul-Edouard Sarlin, Philipp Lindenberger, and Marc Polle- feys. 2024. GeoCalib: Single-image Calibration with Geometric Optimization. In Proceedings of the European Conference on Computer Vision
2024
-
[48]
Ruicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang, Yu Deng, Xin Tong, and Jiaolong Yang. 2025. MoGe: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2025
-
[49]
Ruicheng Wang, Sicheng Xu, Yue Dong, Yu Deng, Jianfeng Xiang, Zelong Lv, Guangzhong Sun, Xin Tong, and Jiaolong Yang. 2025. MoGe-2: Accurate Monoc- ular Geometry with Metric Scale and Sharp Details. arXiv:2507.02546 [cs.CV] https://arxiv.org/abs/2507.02546
Pith/arXiv arXiv 2025
-
[50]
Tobias Weyand, Andre Araujo, Bingyi Cao, and Jack Sim. 2020. Google land- marks dataset v2-a large-scale benchmark for instance-level recognition and retrieval. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2020
-
[51]
Diana Wofk, Fangchang Ma, Tien-Ju Yang, Sertac Karaman, and Vivienne Sze
-
[52]
Zizhang Wu, Zhuozheng Li, Zhi-Gang Fan, Yunzhe Wu, Jian Pu, and Xianzhi Li. 2023. V2Depth: Monocular depth estimation via feature-Level virtual-view simulation and refinement. InProceedings of the ACM International Conference on Multimedia. 688–697
2023
-
[53]
Ke Xian, Chunhua Shen, Zhiguo Cao, Hao Lu, Yang Xiao, Ruibo Li, and Zhenbo Luo. 2018. Monocular relative depth perception with web stereo data supervi- sion. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2018
-
[54]
Ke Xian, Jianming Zhang, Oliver Wang, Long Mai, Zhe Lin, and Zhiguo Cao. 2020. Structure-Guided Ranking Loss for Single Image Depth Prediction. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2020
-
[55]
Guangkai Xu, Yongtao Ge, Mingyu Liu, Chengxiang Fan, Kangyang Xie, Zhiyue Zhao, Hao Chen, and Chunhua Shen. 2025. What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?. InInternational Conference on Learning Representations
2025
-
[56]
Feng Xue, Junfeng Cao, Yu Zhou, Fei Sheng, Yankai Wang, and Anlong Ming
-
[57]
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Heng- shuang Zhao. 2024. Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2024
-
[58]
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. 2024. Depth Anything V2. InAdvances in Neural Information Processing Systems, Vol. 37. 21875–21911
2024
-
[59]
Wei Yin, Yifan Liu, and Chunhua Shen. 2021. Virtual normal: Enforcing geometric constraints for accurate and robust depth prediction.TPAMI44, 10 (2021), 7282– 7295
2021
-
[60]
Wei Yin, Chi Zhang, Hao Chen, Zhipeng Cai, Gang Yu, Kaixuan Wang, Xiaozhi Chen, and Chunhua Shen. 2023. Metric3D: Towards Zero-shot Metric 3D Predic- tion from A Single Image. InProceedings of the IEEE/CVF International Conference on Computer Vision. 9009–9019. doi:10.1109/ICCV51070.2023.00830
arXiv 2023
-
[61]
Naoki Yokoyama, Sehoon Ha, Dhruv Batra, Jiuguang Wang, and Bernadette Bucher. 2024. Vision-Language Frontier Maps for Zero-Shot Semantic Navigation. InProceedings of the IEEE International Conference on Robotics and Automation
2024
-
[62]
Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell. 2020. Bdd100k: A diverse driving dataset for heterogeneous multitask learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2020
-
[63]
Fisher Yu, Yinda Zhang, Shuran Song, Ari Seff, and Jianxiong Xiao. 2015. LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop.arXiv preprint arXiv:1506.03365(2015)
Pith/arXiv arXiv 2015
-
[64]
Songsong Yu, Yifan Wang, Yunzhi Zhuge, Lijun Wang, and Huchuan Lu. 2024. DME: Unveiling the Bias for Better Generalized Monocular Depth Estimation. In Proceedings of the AAAI Conference on Artificial Intelligence
2024
-
[65]
Denis Zavadski, Johann-Friedrich Feiden, and Carsten Rother. 2024. ControlNet- XS: Rethinking the Control of Text-to-Image Diffusion Models as Feedback- Control Systems. InProceedings of the European Conference on Computer Vision
2024
-
[66]
Longjian Zeng, Zunjie Zhu, Rongfeng Lu, Ming Lu, Bolun Zheng, Chenggang Yan, and Anke Xue. 2025. DepthDark: robust monocular depth estimation for low-light environments. InProceedings of the ACM International Conference on Multimedia. 11239–11248
2025
-
[67]
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding conditional con- trol to text-to-image diffusion models. InProceedings of the IEEE/CVF International Conference on Computer Vision
2023
-
[68]
Wendong Zhang, Feng Gao, Bingbing Ni, Lingyu Duan, Yichao Yan, Jingwei Xu, and Xiaokang Yang. 2018. Depth structure preserving scene image generation. InProceedings of the ACM International Conference on Multimedia. 727–736
2018
-
[69]
Kecheng Zheng, Zheng-Jun Zha, Yang Cao, Xuejin Chen, and Feng Wu. 2018. La- net: Layout-aware dense network for monocular depth estimation. InProceedings of the ACM International Conference on Multimedia. 1381–1388
2018
-
[2019]
In Proceedings of the IEEE International Conference on Robotics and Automation
Fastdepth: Fast monocular depth estimation on embedded systems. In Proceedings of the IEEE International Conference on Robotics and Automation
-
[2020]
InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
3d packing for self-supervised monocular depth estimation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
-
[2021]
Boundary-induced and scene-aggregated network for monocular depth prediction.Pattern Recognition115 (2021), 107901
2021
-
[2022]
doi:10.1109/TPAMI.2020.3019967
Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero- Shot Cross-Dataset Transfer.IEEE Transactions on Pattern Analysis and Machine Intelligence44, 3 (2022), 1623–1637. doi:10.1109/TPAMI.2020.3019967
arXiv 2022
-
[2023]
arXiv:2302.12288 [cs.CV] https://arxiv.org/abs/2302.12288
ZoeDepth: Zero-Shot Transfer by Combining Relative and Metric Depth. arXiv:2302.12288 [cs.CV] https://arxiv.org/abs/2302.12288
-
[2024]
InProceedings of the ACM International Conference on Multimedia
SM4Depth: Seamless Monocular Metric Depth Estimation across Multiple Cameras and Scenes by One Model. InProceedings of the ACM International Conference on Multimedia
-
[2025]
InProceedings of the ACM International Conference on Multimedia
Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion. InProceedings of the ACM International Conference on Multimedia. 9375–9384
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.