REVIEW 4 major objections 4 minor 93 references
DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Benchmark probing shows human-like depth cues emerge in newer, larger vision models, including ones trained only on images with no depth supervision.
desk verdict A genuinely useful benchmark with a real validation confound: the correlation with depth is computed on the same frozen features on both sides, so the headline R² numbers are weaker than they look. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the DepthCues benchmark itself: six tasks, each operationalising one classic monocular depth cue from the human vision literature — horizon-line regression for elevation, shadow-object association for light-shadow, occlusion detection, vanishing-point estimation for perspective, 3D size comparison, and depth ordering on textured planes for texture-grad. The protocol that carries the argument is feature probing: for each frozen model, task-specific features are extracted by masked average pooling of object regions or by taking the full feature map, and a lightweight probe (an MLP for binary tasks, an attentive probe for the two regression tasks) is trained on them, with the best model layer chosen by validation performance. The benchmark's validity rests on its controls and correlations: human accuracy of about 95% shows the tasks are well posed, a coordinate-map baseline sets the floor, and DepthCues scores across the 20 models trace the same ranking as depth estimation accuracy while staying nearly uncorrelated with ImageNet classification, so the benchmark measures something geometric rather than generic visual quality.
What would settle it
Run the exact DepthCues probing protocol on two controls: features from a randomly initialized network, and features whose spatial layout is destroyed by global average pooling or patch shuffling. If either control solves elevation or perspective well above the coordinate-map trivial baseline, those tasks are being solved from image statistics rather than learned geometric understanding, and the emergence claim would not survive.
Extended reading notes
Core claim
The paper's central claim is that the information humans use for monocular depth — horizon position, shadow-object relations, occlusion, vanishing points, familiar object sizes, and texture compression — is present in the frozen features of recent large vision models even when those models were never trained on depth. It asserts that this emergence is a measurable phenomenon: probes placed on the features of DINOv2, Stable Diffusion, and other newer models solve the six DepthCues tasks well above trivial baselines, and the ranking of models on DepthCues tracks their ranking on actual depth estimation tasks, with $R^2 = 0.83$ on NYUv2 and $R^2 = 0.80$ on DIW. The paper also argues the causal direction runs both ways: injecting these cues through fine-tuning on DepthCues, with far sparser supervision than dense depth maps, still improves downstream depth estimation. Together these findings position DepthCues as a diagnostic for asking how and where depth perception develops in vision models.
Load-bearing premise
The load-bearing premise is that a probe trained on a model's frozen features solving a task means the model understands that depth cue, when probe success only demonstrates that some learnable signal, potentially a low-level statistical shortcut, exists in the features.
Editorial extensions
If this is right
- DepthCues can serve as a cheap proxy for depth estimation quality: probing a model on six cue tasks reveals its geometric understanding without requiring dense depth ground truth.
- The emergence result explains why self-supervised backbones such as DINOv2 transfer so well to downstream dense depth prediction: the geometric cues are already encoded in their features.
- Sparse cue-level supervision is a viable route to improving depth perception, since fine-tuning on DepthCues improved NYUv2 and DIW depth results for both DINOv2 and CLIP without dense depth labels.
- Multi-view pre-training leaves a specific signature: models trained across viewpoints, such as CroCo, LRM, and DUSt3R, dominate the local texture-gradient task, showing what each pre-training recipe contributes geometrically.
- No model masters all six cues, so the benchmark isolates which cue each pre-training objective fails to capture, giving a map for targeted improvement of geometric understanding.
Reading between the lines
- Because all six tasks share one probing regime, DepthCues could be re-run as a controlled experiment over pre-training objectives — freezing architecture and data scale while varying only the objective — to isolate which recipes produce geometric features; the authors note they cannot control these variables with public checkpoints.
- The strong correlation between DepthCues and depth estimation could be partly driven by shared low-level image statistics rather than shared geometric understanding; an adversarial test would perturb images (for example, removing texture or contrast) and check whether the two scores degrade together.
- The benchmark covers static-image cues only, so extending it to motion parallax and ego-motion cues would test whether the emergence generalises to temporal geometry, which the paper explicitly leaves out.
- The fine-tuning results hint at a curriculum for depth: cue-level objectives might serve as a warm start that makes dense depth fine-tuning faster or more data-efficient, an extension the paper does not test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces DepthCues, a benchmark suite of six tasks designed to probe whether large pre-trained vision models encode the monocular depth cues used by humans: elevation, light-shadow, occlusion, perspective, size, and texture gradient. The authors probe 20 frozen vision models with lightweight MLP or attentive probes, report accuracies per cue, and compare models by average performance. Their main findings are that recent self-supervised and geometry-estimation models score higher, that DepthCues scores correlate strongly with linear-probe depth estimation on NYUv2 and DIW (R^2 = 0.83 and 0.80), and that LoRA fine-tuning on DepthCues can improve downstream depth estimation when the new features are concatenated with the original features. The benchmark, baselines, human accuracy estimate, and code release are described in detail.
Significance. If the validation claim holds, DepthCues would be a useful and inexpensive proxy for probing monocular depth understanding in frozen vision models, complementary to dense depth evaluation. The benchmark is grounded in vision-science cues, includes six carefully constructed tasks with existing datasets plus a synthetic texture-gradient task, and ships with important controls: coordinate-map and end-to-end baselines, human accuracy, threshold validation, and layer selection. The breadth of 20 models spanning SSL, classification, VLM, generative, segmentation, and multi-view objectives is a strength. However, the headline claims—that depth cues emerge in more recent larger models and that DepthCues is validated by its correlation with depth estimation—rest on correlational and shared-protocol evidence that needs additional support before the benchmark can be adopted as a proxy for depth perception.
major comments (4)
- [5.1(v), Figs. A1-A2, Fig. 4] The central validation claim that DepthCues performance is highly correlated with depth estimation is based on two quantities obtained from the same frozen features: DepthCues uses layer-searched MLP/attentive probes (Sec. 4.2) while NYUv2 and DIW use a linear probe on patch tokens (Sec. 5.1). Any model-level property that makes features generally more probeable—feature conditioning, spatial resolution, feature dimensionality, or architecture—will tend to raise both sides of the correlation, so the reported R^2=0.83 and R^2=0.80 can overstate how specifically DepthCues tracks depth ability. The low ImageNet correlation does not remove this confound, because ImageNet linear probing is itself sensitive to class-token usage and object-centric features, as the authors note in Appendix A.1. Please re-validate DepthCues against depth metrics obtained from an independent protocol (e.g., published full-model depth accuracies or depth heads fine-tuned on depth data), or at minimum report partial correlations controlling for a measure of general probeability.
- [5.1(i), Fig. 1, Abstract] The claim that human-like depth cues 'emerge in more recent larger models' is inferred from a scatter plot against release date in which models also differ in architecture, pre-training objective, and dataset size (Table A6). This is a confounded trend, not a controlled emergence result; for example, DepthAnythingv2 is initialized from DINOv2 and trained with dense depth supervision, so its top score cannot be attributed to recency. The appendix acknowledges these confounds (Appendix A.1, Fig. A3), but the abstract and Sec. 5.1 state the conclusion without qualification. Please either restrict the claim to a descriptive statement about the evaluated checkpoints or provide matched comparisons that vary scale/recency while holding architecture and objective fixed.
- [4.2, Eqs. (1)-(2)] The benchmark's interpretational claim rests on the assumption that probe accuracy on frozen features measures the model's understanding of a depth cue. The coordinate-map and end-to-end baselines establish task-solvability floors, but they do not identify which feature dimensions the probe uses; a probe can solve a task from low-level artifacts (e.g., shadow boundaries, mask statistics, or texture statistics) without engaging any geometric representation. Without a control that removes or perturbs the cue-specific information, or an analysis of the probe's decision features, the statement that 'the better the model understands the task-specific cue' is stronger than what the protocol establishes. Please soften the wording or add such a control.
- [5.2, Table 2] The statement that fine-tuning on DepthCues 'resulted in improved depth perception' is not supported by the +DC rows alone: DINOv2+DC drops NYUv2 accuracy from 87.78 to 87.06, and CLIP+DC drops from 43.78 to 43.59; improvements appear only when the fine-tuned features are concatenated with the original features. The abstract and Sec. 1 claim that merely fine-tuning on DepthCues improves depth estimation. Please report the concatenation-dependent result as the actual finding, or provide a fine-tuning variant that improves the standalone features, and state the result at the level of precision the data support.
minor comments (4)
- [5.2] Because the size cue in DepthCues uses SUNRGBD images that overlap NYUv2, the NYUv2 numbers in Table 2 may be partially contaminated; the authors note this and point to DIW as fairer, but the main-text sentence should explicitly say that the NYUv2 improvements are subject to this caveat.
- [C.2, Figs. A15-A16] The threshold validation for the perspective and elevation accuracies is reported as a correlation with the raw error; please also report the absolute success rates at the chosen thresholds so readers can calibrate how strict the 0.2 and 0.1 thresholds are.
- [Throughout] Minor wording and naming inconsistencies: 'DepthAnyv2' vs. 'DepthAnythingv2', 'light-shadow' vs. 'light and shadow', and the informal phrase 'we evaluate our own performance' in Sec. 5.1. These do not affect the results but should be cleaned up.
- [A.1, Fig. A3] The reported Pearson r=0.69 for pre-training data size vs. DepthCues is computed after excluding the three language-supervised models; the text should state how the exclusion criterion was chosen and note that the remaining set still mixes objectives and architectures.
Circularity Check
No circular derivation: DepthCues labels come from external sources, depth validation uses held-out NYUv2/DIW, and the core claims do not reduce to fitted inputs or self-citation.
full rationale
I walked the paper's derivation chain. The DepthCues benchmark tasks are constructed from external datasets (HLW, SOBA, COCOA, NaturalScene, KITTI/SUNRGBD, and DTD textures), and the labels are not defined in terms of any depth-estimation score produced by the probed models. The validation claim (Sec. 5.1(v), Fig. 4) is an empirical Spearman correlation between DepthCues probe accuracy and linear-probe depth accuracy on held-out NYUv2 and DIW; neither side of the correlation is fitted to the other, and no parameter is optimized to force the R^2 values in Figs. A1-A2. The fine-tuning experiment (Sec. 5.2) is a genuine empirical test on external depth benchmarks, and the acknowledged SUNRGBD/NYUv2 overlap in the size task is explicitly mitigated by reporting DIW. The Sec. 4 assumption that probe accuracy reflects cue understanding is an interpretive assumption rather than a definitional reduction; the coordinate-map and end-to-end baselines are controls, not circular inputs. Self-citations by the authors (e.g., refs. 26, 27, 39, 47, 54) support related work or protocol choices and are not load-bearing for the benchmark's core validity. The limitation passages in Sec. 5.3 and Appendix A.1 openly acknowledge uncontrolled pre-training variables and probe implementation differences, which are validity caveats rather than circular steps. The shared-probe-protocol confound raised by skeptics is a possible threat to the correlation's interpretation, but it does not make any claimed quantity equal to its input by construction. Therefore, no significant circularity is present.
Assumptions & free parameters
free parameters (3)
- Size difference thresholds =
2.5 m^3 (KITTI), 0.4 m^3 (SUN-RGBD)
- Perspective success threshold =
0.2 normalized Euclidean distance
- Elevation success threshold =
0.1 normalized horizon error
assumptions (4)
- domain assumption Probing frozen features with a trained lightweight probe measures a model's understanding of a depth cue.
- domain assumption The six chosen cues (elevation, light-shadow, occlusion, perspective, size, texture-grad) are the relevant human monocular depth cues.
- domain assumption The synthetic Blender texture-gradient dataset isolates the texture-gradient cue.
- domain assumption The release-date trend is interpreted as emergence of depth cues rather than as a consequence of confounded pre-training scale, objective, or architecture (Fig. 1, Sec. 5.1(i)).
Cite this review
Pith. "Pith review of DepthCues: Evaluating Monocular Depth Perception in Large Vision Models." pith.science (2026). https://pith.science/paper/QM35DVSR
@misc{pith2026241117385,
author = {Pith},
title = {Pith review of: DepthCues: Evaluating Monocular Depth Perception in Large Vision Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/QM35DVSR}},
note = {Machine review of arXiv:2411.17385}
}
read the original abstract
Large-scale pre-trained vision models are becoming increasingly prevalent, offering expressive and generalizable visual representations that benefit various downstream tasks. Recent studies on the emergent properties of these models have revealed their high-level geometric understanding, in particular in the context of depth perception. However, it remains unclear how depth perception arises in these models without explicit depth supervision provided during pre-training. To investigate this, we examine whether the monocular depth cues, similar to those used by the human visual system, emerge in these models. We introduce a new benchmark, DepthCues, designed to evaluate depth cue understanding, and present findings across 20 diverse and representative pre-trained vision models. Our analysis shows that human-like depth cues emerge in more recent larger models. We also explore enhancing depth perception in large vision models by fine-tuning on DepthCues, and find that even without dense depth supervision, this improves depth estimation. To support further research, our benchmark and evaluation code will be made publicly available for studying depth perception in vision models.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Monocular depth es- timation using cues inspired by biological vision systems
Dylan Auty and Krystian Mikolajczyk. Monocular depth es- timation using cues inspired by biological vision systems. In ICPR, 2022. 3
2022
-
[2]
Language-Based Depth Hints for Monocular Depth Estimation
Dylan Auty and Krystian Mikolajczyk. Language- based depth hints for monocular depth estimation. arXiv:2403.15551, 2024. 3
work page Pith review arXiv 2024
-
[3]
Enhancing 2d represen- tation learning with a 3d prior
Mehmet Aygun, Prithviraj Dhar, Zhicheng Yan, Oisin Mac Aodha, and Rakesh Ranjan. Enhancing 2d represen- tation learning with a 3d prior. In CVPR Workshops, 2024. 2, 5
2024
-
[4]
Understanding Depth and Height Perception in Large Visual-Language Models
Shehreen Azad, Yash Jain, Rishit Garg, Yogesh S Rawat, and Vibhav Vineet. Geometer: Probing depth and height per- ception of large visual-language models. arXiv:2408.11748,
-
[5]
Revisiting feature prediction for learning visual rep- resentations from video
Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mido Assran, and Nicolas Ballas. Revisiting feature prediction for learning visual rep- resentations from video. TMLR, 2024. Featured Certifica- tion. 5, 15
2024
-
[6]
Adabins: Depth estimation using adaptive bins
Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka. Adabins: Depth estimation using adaptive bins. In CVPR,
-
[7]
Depth perception of surgeons in minimally invasive surgery.Surgi- cal innovation, 2016
Rositsa Bogdanova, Pierre Boulanger, and Bin Zheng. Depth perception of surgeons in minimally invasive surgery.Surgi- cal innovation, 2016. 1, 3
2016
-
[8]
Tenenbaum, and Alexei A
Tyler Bonnen, Stephanie Fu, Yutong Bai, Thomas O’Connell, Yoni Friedman, Nancy Kanwisher, Joshua B. Tenenbaum, and Alexei A. Efros. Evaluating multiview ob- ject consistency in humans and image models. In NeurIPS,
Show all 93 references
-
[9]
Intelligent image and video com- pression: communicating pictures
David Bull and Fan Zhang. Intelligent image and video com- pression: communicating pictures . Academic Press, 2021. 1, 3
2021
-
[10]
Mvsformer++: Revealing the devil in transformer’s details for multi-view stereo
Chenjie Cao, Xinlin Ren, and Yanwei Fu. Mvsformer++: Revealing the devil in transformer’s details for multi-view stereo. In ICLR, 2024. 1
2024
-
[11]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In ICCV, 2021. 5, 19, 21
2021
-
[12]
Single- image depth perception in the wild
Weifeng Chen, Zhao Fu, Dawei Yang, and Jia Deng. Single- image depth perception in the wild. NeurIPS, 2016. 5, 8, 20
2016
-
[13]
Cimpoi, S
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, , and A. Vedaldi. Describing textures in the wild. In CVPR, 2014. 3, 4
2014
-
[14]
Objaverse-xl: A universe of 10m+ 3d objects
Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram V oleti, Samir Yitzhafk Gadre, et al. Objaverse-xl: A universe of 10m+ 3d objects. NeurIPS, 2024. 21
2024
-
[15]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009. 5, 12, 20
2009
-
[16]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...
2021
-
[17]
Depth map prediction from a single image using a multi-scale deep net- work
David Eigen, Christian Puhrsch, and Rob Fergus. Depth map prediction from a single image using a multi-scale deep net- work. NeurIPS, 2014. 2
2014
-
[18]
Prob- ing the 3d awareness of visual foundation models
Mohamed El Banani, Amit Raj, Kevis-Kokitsi Maninis, Ab- hishek Kar, Yuanzhen Li, Michael Rubinstein, Deqing Sun, Leonidas Guibas, Justin Johnson, and Varun Jampani. Prob- ing the 3d awareness of visual foundation models. In CVPR,
-
[19]
Scalable pre- training of large autoregressive image models
Alaaeldin El-Nouby, Michal Klein, Shuangfei Zhai, Miguel Angel Bautista, Alexander Toshev, Vaishaal Shankar, Joshua M Susskind, and Armand Joulin. Scalable pre- training of large autoregressive image models. In ICML,
-
[20]
Deep ordinal regression net- work for monocular depth estimation
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Bat- manghelich, and Dacheng Tao. Deep ordinal regression net- work for monocular depth estimation. In CVPR, 2018. 2
2018
-
[21]
Dream- sim: Learning new dimensions of human visual similarity using synthetic data
Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. Dream- sim: Learning new dimensions of human visual similarity using synthetic data. NeurIPS, 2024. 3
2024
-
[22]
Unsupervised cnn for single view depth estimation: Geome- try to the rescue
Ravi Garg, Vijay Kumar Bg, Gustavo Carneiro, and Ian Reid. Unsupervised cnn for single view depth estimation: Geome- try to the rescue. In ECCV, 2016. 2
2016
-
[23]
Geobench: Benchmarking and analyzing monocular geome- try estimation models
Yongtao Ge, Guangkai Xu, Zhiyue Zhao, Libo Sun, Zheng Huang, Yanlong Sun, Hao Chen, and Chunhua Shen. Geobench: Benchmarking and analyzing monocular geome- try estimation models. arXiv:2406.12671, 2024. 1, 2, 15
2024 arXiv
-
[24]
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In CVPR, 2012. 2, 3, 4, 18
2012
-
[25]
Wichmann, and Wieland Brendel
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel. Imagenet-trained CNNs are biased towards texture; increas- ing shape bias improves accuracy and robustness. In ICLR,
-
[26]
Unsupervised monocular depth estimation with left- right consistency
Cl ´ement Godard, Oisin Mac Aodha, and Gabriel J Bros- tow. Unsupervised monocular depth estimation with left- right consistency. In CVPR, 2017. 2
2017
-
[27]
Digging into self-supervised monocular depth estimation
Cl ´ement Godard, Oisin Mac Aodha, Michael Firman, and Gabriel J Brostow. Digging into self-supervised monocular depth estimation. In CVPR, 2019. 2
2019
-
[28]
Depth from videos in the wild: Unsupervised monocular depth learning from unknown cameras
Ariel Gordon, Hanhan Li, Rico Jonschkowski, and Anelia Angelova. Depth from videos in the wild: Unsupervised monocular depth learning from unknown cameras. In ICCV,
-
[29]
Depthfm: Fast monocular depth estimation with flow match- ing
Ming Gui, Johannes S Fischer, Ulrich Prestel, Pingchuan Ma, Dmytro Kotovenko, Olga Grebenkova, Stefan An- dreas Baumann, Vincent Tao Hu, and Bj ¨orn Ommer. Depthfm: Fast monocular depth estimation with flow match- ing. arXiv:2403.13788, 2024. 1
2024 arXiv
-
[30]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR,
-
[31]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In CVPR, 2022. 5, 21
2022
-
[32]
The organization of behavior: A neu- ropsychological theory
Donald Olding Hebb. The organization of behavior: A neu- ropsychological theory. Psychology press, 2005. 2
2005
-
[33]
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv:1606.08415, 2016. 5, 20
2016 arXiv
-
[34]
Lrm: Large reconstruction model for single image to 3d
Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. Lrm: Large reconstruction model for single image to 3d. ICLR, 2024. 2, 5, 21
2024
-
[35]
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In ICLR, 2022. 8, 20
2022
-
[36]
Squeeze-and-excitation net- works
Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation net- works. In CVPR, 2018. 5, 20, 21
2018
-
[37]
Monodtr: Monocular 3d object detection with depth-aware transformer
Kuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, and Win- ston H Hsu. Monodtr: Monocular 3d object detection with depth-aware transformer. In CVPR, 2022. 1
2022
-
[38]
Receptive fields and functional architecture of monkey striate cortex.The Journal of Physiology, 1968
David H Hubel and Torsten N Wiesel. Receptive fields and functional architecture of monkey striate cortex.The Journal of Physiology, 1968. 2
1968
-
[39]
Brostow, and Jamie Watson
Sergio Izquierdo, Mohamed Sayed, Michael Firman, Guillermo Garcia-Hernando, Daniyar Turmukhambetov, Javier Civera, Oisin Mac Aodha, Gabriel J. Brostow, and Jamie Watson. MVSAnywhere: Zero shot multi-view stereo. In CVPR, 2025. 1
2025
-
[40]
The perception of depth
Michael Kalloniatis and Charles Luu. The perception of depth. Webvision-the organization of the retina and visual system, 2011. 1, 3
2011
-
[41]
Repurpos- ing diffusion-based image generators for monocular depth estimation
Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Met- zger, Rodrigo Caye Daudt, and Konrad Schindler. Repurpos- ing diffusion-based image generators for monocular depth estimation. In CVPR, 2024. 1, 2, 7
2024
-
[42]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 18, 20
2015
-
[43]
Segment any- thing
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment any- thing. In ICCV, 2023. 4, 5, 18, 21
2023
-
[44]
Temporally consistent horizon lines
Florian Kluger, Hanno Ackermann, Michael Ying Yang, and Bodo Rosenhahn. Temporally consistent horizon lines. In ICRA, 2020. 3
2020
-
[45]
On the viability of monocular depth pre-training for semantic segmentation
Dong Lao, Fengyu Yang, Daniel Wang, Hyoungseob Park, Samuel Lu, Alex Wong, and Stefano Soatto. On the viability of monocular depth pre-training for semantic segmentation. In ECCV, 2024. 1
2024
-
[46]
Perceptual depth indicator for s-3d con- tent based on binocular and monocular cues
Pierre Lebreton, Alexander Raake, Marcus Barkowsky, and Patrick Le Callet. Perceptual depth indicator for s-3d con- tent based on binocular and monocular cues. In Asilomar Conference on Signals, Systems and Computers, 2012. 1, 3
2012
-
[47]
Multi-task learning with 3d-aware regulariza- tion
Wei-Hong Li, Steven McDonagh, Ales Leonardis, and Hakan Bilen. Multi-task learning with 3d-aware regulariza- tion. In ICLR, 2024. 2
2024
-
[48]
Movideo: Motion-aware video generation with diffusion models
Jingyun Liang, Yuchen Fan, Kai Zhang, Radu Timofte, Luc Van Gool, and Rakesh Ranjan. Movideo: Motion-aware video generation with diffusion models. In ECCV, 2024. 1
2024
-
[49]
The 3d-pc: a benchmark for visual per- spective taking in humans and machines
Drew Linsley, Peisen Zhou, Alekh Karkada Ashok, Akash Nagaraj, Gaurav Gaonkar, Francis E Lewis, Zygmunt Pizlo, and Thomas Serre. The 3d-pc: a benchmark for visual per- spective taking in humans and machines. arXiv:2406.04138,
-
[50]
Vapid: A rapid vanishing point detector via learned optimizers
Shichen Liu, Yichao Zhou, and Yajie Zhao. Vapid: A rapid vanishing point detector via learned optimizers. In ICCV,
-
[51]
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feicht- enhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In CVPR, 2022. 5, 20, 21
2022
-
[52]
Ima- genet3d: Towards general-purpose object-level 3d under- standing
Wufei Ma, Guofeng Zhang, Qihao Liu, Guanning Zeng, Adam Kortylewski, Yaoyao Liu, and Alan Yuille. Ima- genet3d: Towards general-purpose object-level 3d under- standing. In NeurIPS, 2024. 15
2024
-
[53]
Lexicon3d: Probing visual foundation models for complex 3d scene understand- ing
Yunze Man, Shuhong Zheng, Zhipeng Bao, Martial Hebert, Liang-Yan Gui, and Yu-Xiong Wang. Lexicon3d: Probing visual foundation models for complex 3d scene understand- ing. In NeurIPS, 2024. 2
2024
-
[54]
Im- proving semantic correspondence with viewpoint-guided spherical maps
Octave Mariotti, Oisin Mac Aodha, and Hakan Bilen. Im- proving semantic correspondence with viewpoint-guided spherical maps. In CVPR, 2024. 8
2024
-
[55]
Indoor segmentation and support inference from rgbd images
Pushmeet Kohli Nathan Silberman, Derek Hoiem and Rob Fergus. Indoor segmentation and support inference from rgbd images. In ECCV, 2012. 2
2012
-
[56]
Approaching human 3d shape perception with neurally mappable models
Thomas P O’Connell, Tyler Bonnen, Yoni Friedman, Ayush Tewari, Josh B Tenenbaum, Vincent Sitzmann, and Nancy Kanwisher. Approaching human 3d shape perception with neurally mappable models. arXiv:2308.11300, 2023. 3
2023 arXiv
-
[57]
Maxime Oquab, Timoth ´ee Darcet, Th´eo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael ...
2024
-
[58]
The hippocampus as a cognitive map, 1978
J O’Keefe. The hippocampus as a cognitive map, 1978. 2
1978
-
[59]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In ICML, 2021. 2, 5, 21
2021
-
[60]
Large-scale, high-resolution comparison of the core visual object recogni- tion behavior of humans, monkeys, and state-of-the-art deep artificial neural networks
Rishi Rajalingham, Elias B Issa, Pouya Bashivan, Kohitij Kar, Kailyn Schmidt, and James J DiCarlo. Large-scale, high-resolution comparison of the core visual object recogni- tion behavior of humans, monkeys, and state-of-the-art deep artificial neural networks. Journal of Neur...
2018
-
[61]
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
Ren ´e Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. TPAMI, 2020. 2, 5, 20, 21
2020
-
[62]
Vi- sion transformers for dense prediction
Ren ´e Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vi- sion transformers for dense prediction. In ICCV, 2021. 12
2021
-
[63]
High-resolution image syn- 10 thesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- 10 thesis with latent diffusion models. In CVPR, 2022. 1, 2, 5, 21
2022
-
[64]
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. NeurIPS, 2022. 21
2022
-
[65]
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. Indoor segmentation and support inference from rgbd images. In ECCV, 2012. 5, 8, 12
2012
-
[66]
Sun rgb-d: A rgb-d scene understanding benchmark suite
Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao. Sun rgb-d: A rgb-d scene understanding benchmark suite. In CVPR, 2015. 3, 4, 18
2015
-
[67]
Deconstructing self-supervised monocular recon- struction: The design decisions that matter
Jaime Spencer, Chris Russell, Simon Hadfield, and Richard Bowden. Deconstructing self-supervised monocular recon- struction: The design decisions that matter. TMLR, 2022. 2
2022
-
[68]
An- alyzing results of depth estimation models with monocular criteria
Jonas Theiner, Nils Nommensen, Jim Rhotert, Matthias Springstein, Eric M ¨uller-Budack, and Ralph Ewerth. An- alyzing results of depth estimation models with monocular criteria. In CVPR Workshops, 2023. 2
2023
-
[69]
Transformer based line segment classifier with image context for real-time vanishing point detection in man- hattan world
Xin Tong, Xianghua Ying, Yongjie Shi, Ruibin Wang, and Jinfa Yang. Transformer based line segment classifier with image context for real-time vanishing point detection in man- hattan world. In CVPR, 2022. 4
2022
-
[70]
Deit iii: Revenge of the vit
Hugo Touvron, Matthieu Cord, and Herv ´e J ´egou. Deit iii: Revenge of the vit. In ECCV, 2022. 5, 20, 21
2022
-
[71]
Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision
Ruicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang, Yu Deng, Xin Tong, and Jiaolong Yang. Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision. arXiv preprint arXiv:2410.19115, 2024. 17
-
[72]
Dust3r: Geometric 3d vi- sion made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vi- sion made easy. In CVPR, 2024. 5, 21
2024
-
[73]
Instance shadow detection
Tianyu Wang, Xiaowei Hu, Qiong Wang, Pheng-Ann Heng, and Chi-Wing Fu. Instance shadow detection. In CVPR,
-
[74]
Croco: Self-supervised pre-training for 3d vision tasks by cross-view completion
Philippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Ro- main Br´egier, Yohann Cabon, Vaibhav Arora, Leonid Ants- feld, Boris Chidlovskii, Gabriela Csurka, and J ´erˆome Re- vaud. Croco: Self-supervised pre-training for 3d vision tasks by cross-view completion. NeurIPS, 2022. 2, 5, 21
2022
-
[75]
Pytorch image models
Ross Wightman. Pytorch image models. https : / / github . com / rwightman / pytorch - image - models, 2019. 20
2019
-
[76]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chau- mond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R ´emi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, ...
2020
-
[77]
Horizon lines in the wild
Scott Workman, Menghua Zhai, and Nathan Jacobs. Horizon lines in the wild. In BMVC, 2016. 3, 5, 19
2016
-
[78]
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Doll ´ar, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In CVPR, 2017. 5, 20, 21
2017
-
[79]
Depth anything: Unleashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. In CVPR, 2024. 2, 17
2024
-
[80]
Depth any- thing v2
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiao- gang Xu, Jiashi Feng, and Hengshuang Zhao. Depth any- thing v2. NeurIPS, 2024. 1, 2, 5, 7, 20, 21
2024
-
[81]
Improving 2d feature representations by 3d-aware fine-tuning
Yuanwen Yue, Anurag Das, Francis Engelmann, Siyu Tang, and Jan Eric Lenssen. Improving 2d feature representations by 3d-aware fine-tuning. In ECCV, 2024. 2, 5, 8
2024
-
[82]
Detect- ing vanishing points using global image context in a non- manhattan world
Menghua Zhai, Scott Workman, and Nathan Jacobs. Detect- ing vanishing points using global image context in a non- manhattan world. In CVPR, 2016. 4
2016
-
[83]
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In ICCV, 2023. 5, 21
2023
-
[84]
A general protocol to probe large vision models for 3d physical understanding
Guanqi Zhan, Chuanxia Zheng, Weidi Xie, and Andrew Zis- serman. A general protocol to probe large vision models for 3d physical understanding. In NeurIPS, 2024. 1, 2, 4, 5, 15
2024
-
[85]
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In ICCV, 2023. 1
2023
-
[86]
ConDense: Consistent 2d/3d pre- training for dense and sparse features from multi-view im- ages
Xiaoshuai Zhang, Zhicheng Wang, Howard Zhou, Soham Ghosh, Danushen Gnanapragasam, Varun Jampani, Hao Su, and Leonidas Guibas. ConDense: Consistent 2d/3d pre- training for dense and sparse features from multi-view im- ages. In ECCV, 2024. 2
2024
-
[87]
ibot: Image bert pre-training with online tokenizer
Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. ibot: Image bert pre-training with online tokenizer. ICLR, 2022. 5, 21
2022
-
[88]
Unsupervised learning of depth and ego-motion from video
Tinghui Zhou, Matthew Brown, Noah Snavely, and David G Lowe. Unsupervised learning of depth and ego-motion from video. In CVPR, 2017. 2
2017
-
[89]
Detect- ing dominant vanishing points in natural scenes with appli- cation to composition-sensitive image retrieval.Transactions on Multimedia, 2017
Zihan Zhou, Farshid Farhat, and James Z Wang. Detect- ing dominant vanishing points in natural scenes with appli- cation to composition-sensitive image retrieval.Transactions on Multimedia, 2017. 3, 4
2017
-
[90]
LLaV A-3D: A simple yet effective pathway to empowering LMMs with 3d-awareness
Chenming Zhu, Tai Wang, Wenwei Zhang, Jiangmiao Pang, and Xihui Liu. LLaV A-3D: A simple yet effective pathway to empowering LMMs with 3d-awareness. arXiv:2409.18125,
-
[91]
Semantic amodal segmentation
Yan Zhu, Yuandong Tian, Dimitris Metaxas, and Piotr Doll´ar. Semantic amodal segmentation. In CVPR, 2017. 3, 4, 15, 18 11 Appendix A. Additional Results on DepthCues A.1. Full Benchmark Results In the main paper, we reported the performance of 20 pre- trained vision models on ...
2017
-
[93]
patch size-14
are generally higher than our results. Possible rea- sons for this include the following factors. Firstly, while we train a simple linear layer to probe the patch tokens from the last layer of the vision models, a more complex non-linear probe similar to the DPT decoder [62] i...
-
[2024]
1, 2, 4, 5, 7, 12, 15, 19, 20
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.