REVIEW 3 major objections 5 minor 135 references
NTIRE 2025 Challenge on HR Depth from Images of Specular and Transparent Surfaces
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read In the NTIRE 2025 challenge, all submitted stereo and mono depth models outperformed their baselines on transparent and mirror surfaces.
desk verdict A solid NTIRE challenge report that gives a useful snapshot of stereo and mono depth methods on transparent/mirror surfaces, with one real caveat: the mono track's headline results rely on per-image ground-truth scale/shift alignment, so they measure affine-invariant relative shape, not metric depth. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the Booster benchmark evaluation protocol. It supplies 12-megapixel ($4112\times3008$) stereo and mono test images with ground-truth depth, splits every prediction into ToM (transparent-or-mirror), Other, and All pixel sets, ranks stereo entries by bad-2 (percentage of pixels whose disparity error exceeds 2 pixels), and ranks mono entries by $\delta<1.05$ (percentage of pixels whose predicted/ground-truth ratio is below 1.05) after a least-squares scale-and-shift alignment makes affine-invariant predictions metric. This protocol is what turns 'submitted methods are better' into a quantified, comparable claim.
What would settle it
Measure the winning stereo and mono models on a held-out set of transparent bottles and mirrors where the required depth is the physical surface, as in a robotic grasping task, using annotations collected independently of Booster. If the top-ranked challenge methods do not beat their baselines there with comparable margins, the reported gains are tied to Booster's labeling convention rather than to genuinely recovering transparent-surface depth.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that submissions to the NTIRE 2025 challenge outperform off-the-shelf baselines on transparent and reflective surfaces. In stereo, the best ToM bad-2 error falls from 59.64 percent (CREStereo baseline) to 42.47 percent (SRC-B), and the best all-pixel bad-2 error falls from 35.75 percent to 19.27 percent (weouibaguette), although CREStereo still wins on ordinary 'Other' pixels. In mono, every submitted method beats ZoeDepth on every pixel category; the winner, Lavreniuk, reaches 87.67 percent on the $\delta<1.05$ metric for ToM pixels versus 45.21 percent for the baseline, and the top three mono entries all exceed 85 percent, roughly 15 points above the previous challenge edition.
Load-bearing premise
The claim rests on treating Booster's ground-truth depth for glass and mirror regions as the correct target, even though the paper itself notes that depth is ambiguous for such surfaces, since it could mean the surface, the reflected scene, or the object behind the glass.
Editorial extensions
If this is right
- Stereo entries cut ToM bad-2 error by about 17 percentage points relative to CREStereo, and the best all-pixel stereo error drops by about 16 points.
- Monocular entries show that a fine-tuned depth foundation model or an inpainting pipeline can push the strict $\delta<1.05$ metric above 85 percent on transparent-or-mirror pixels, where the baseline sits at 45 percent.
- Because CREStereo still holds the best Other-pixel stereo score, gains on non-Lambertian surfaces do not yet transfer to ordinary scene regions in a single stereo model.
- The leading mono solution works by blending and inpainting away the glass or mirror content before depth estimation, which suggests the benchmark can be approached without explicitly modeling reflection physics.
Reading between the lines
- If the benchmark convention is what practitioners need for grasping or navigation, these results are directly actionable; if they need a different depth target (the reflected scene or the object behind glass), the ranking could reward models that fit Booster's labels rather than the underlying perception problem.
- The dominance of inpainting-based mono entries suggests a testable research direction: explicitly combining multi-layer depth (surface plus transmitted or reflected depth) with inpainting-based priors, rather than choosing one representation.
- A reader checking the manuscript will notice an internal inconsistency: Section 3 says five mono teams submitted, while the abstract and Table 2 list four; the 'all submitted methods' claim should be read as referring to the four tabulated mono entries and the four stereo entries.
- Because the stereo baseline still wins on Other pixels, a fair next experiment is an ablation that fine-tunes the same foundation model on Booster while measuring ToM and Other accuracy separately, to see whether the ToM gain trades away ordinary-region accuracy.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports the NTIRE 2025 challenge on high-resolution depth estimation for specular and transparent surfaces. The organizers use the Booster dataset (228 training stereo pairs and separate stereo/mono test splits) and define two tracks: stereo disparity estimation and monocular depth estimation. Four teams per track submitted final results. Stereo methods are evaluated with bad-tau, MAE, and RMSE on ToM/All/Other pixel sets; mono methods are evaluated with delta thresholds, Abs Rel, MAE, and RMSE after per-image least-squares scale/shift alignment (Eq. 1). The paper reports that all stereo submissions outperform CREStereo on ToM and All pixels, that all mono submissions outperform ZoeDepth, describes each method, and gives qualitative examples.
Significance. If taken at face value, the empirical finding is meaningful: fine-tuning depth foundation models and using mask-guided inpainting pipelines yields large gains over off-the-shelf CREStereo and ZoeDepth on 12Mpx non-Lambertian surfaces, including a roughly 15-point improvement in delta<1.05 on ToM regions over the previous edition. The report is transparent about the evaluation protocol and the baseline settings, and the team method descriptions are sufficiently detailed to reproduce the main design choices. The principal caveat is that the mono track's ranking is affine-invariant by construction (Eq. 1), so the headline mono claim is not about metric depth accuracy as the introduction's grasping motivation suggests. The stereo and mono tables do support the all-submitted-beat-baseline statements under the stated protocol.
major comments (3)
- [Section 3, Eq. (1)] The mono track's headline result ('all of the submitted methods consistently outperformed the ZoeDepth baseline', Section 4.2) is computed after per-image least-squares scale/shift alignment to the ground-truth depth map via Eq. (1). This makes the evaluation affine-invariant: a model can receive a perfect alpha and beta from the test-time oracle even if its raw output has incorrect metric scale and shift. Since the Introduction motivates the mono track with applications such as grasping glass objects, where no ground-truth oracle exists, the paper should report raw (unaligned) mono metrics or a fixed global calibration, and should explicitly state that the ranking reflects relative-depth accuracy rather than metric depth accuracy. Without these, the central mono claim is not verified in the form stated.
- [Table 1] In the CREStereo baseline row, the Other-pixel columns as printed read bad-2 = 8.34 and bad-4 = 13.93, which violates the monotonicity requirement bad-2 >= bad-4 >= bad-6 >= bad-8. This also conflicts with the Section 4.1 statement that 'the baseline still performs the best on Other pixels': under the natural column ordering the printed bad-4 would be invalid, while under the alternative ordering the baseline's Other bad-2 of 13.93 is worse than weouibaguette's 13.64. Please correct the table and re-check the sentence.
- [Section 4.2, Table 2] The statement that Lavreniuk achieves 'the best accuracy values on ToM, All and Other pixels' is not supported by Table 2 on every metric: colab has lower MAE on ToM (2.58 vs 2.71) and on All (2.97 vs 3.12), and higher delta<1.15 on ToM, All and Other. If 'accuracy' refers only to the delta<1.05 ranking metric used for the two official rankings, the sentence should say so explicitly; otherwise it overstates the results.
minor comments (5)
- [Abstract and Introduction] The Abstract states that 4 and 4 teams submitted in the final stage, while the Introduction says '5 and 4 teams, respectively, for the monocular and stereo tracks'; Section 4 says four teams per track. Please reconcile these numbers.
- [Table 1 caption] The caption says methods are ranked on 'delta < 1.05' for the stereo track, but Section 3 and Section 4.1 define the stereo ranking metric as bad-2. This appears to be a copy-paste error from the mono track caption.
- [Section 3, Datasets] The sentence 'with 38 and 26 for the two sets respectively' is unclear about whether 38 and 26 refer to scenes or pairs; please rephrase to avoid ambiguity with the earlier counts of 228 and 191 pairs.
- [Section 5.1.3] The text uses 'TOM regions' with uppercase O; the standard abbreviation used elsewhere in the paper is 'ToM'. Please make the capitalization consistent.
- [Section 2, Competitions/Challenges] There is a typo in 'the depth estimation task ... as been the object of several challenges'; it should be 'has been'.
Circularity Check
No circular derivation: the challenge report compares submitted models against fixed baselines on a public benchmark; the mono-track per-image scale/shift alignment is a disclosed, uniform evaluation convention, not a fitted prediction.
full rationale
This paper is a challenge report, not a derivation from first principles, so the standard circularity patterns (self-definitional quantities, uniqueness theorems imported from authors, ansatz smuggled via citation) do not apply. The central claims are empirical: all stereo submissions beat CREStereo on ToM and All pixels (Table 1), and all mono submissions beat ZoeDepth after the protocol described in Eq. (1). The mono protocol fits per-image scale and shift (alpha, beta) to ground-truth depth before computing error metrics; this makes the reported mono metrics affine-invariant rather than metric-scale tests, and the phrase 'recover metric depth' overstates what Eq. (1) does. However, this is a disclosed evaluation convention applied identically to every method, including the ZoeDepth baseline, so it does not force any particular ranking and is not a fitted input renamed as a prediction. The reliance on the organizers' Booster dataset and earlier editions is self-citation, but the dataset is a published benchmark with fixed annotations and the challenge explicitly reports results on that benchmark; no load-bearing argument reduces to an unverified self-citation. The Introduction's acknowledgment that 'the definition of depth itself becomes ambiguous in such circumstances' is a limitation of the benchmark target, not a circular step. No equation in the paper defines a claimed output in terms of the quantity it is supposed to predict, and no baseline comparison is constructed from the submitted methods' own outputs. Score 0.
Assumptions & free parameters
free parameters (1)
- per-image affine alignment (alpha, beta) for mono track =
LSE-fitted to ground truth per image via Eq. 1
assumptions (4)
- domain assumption Booster ground-truth depth provides the correct and unambiguous depth definition for transparent and mirror surfaces.
- domain assumption Participants did not use the public test-set labels for training or tuning.
- domain assumption The selected validation and test splits are representative of the Booster domain.
- standard math Standard evaluation metrics (bad-tau, delta, Abs Rel, MAE, RMSE) are appropriate for depth quality comparison.
Cite this review
Pith. "Pith review of NTIRE 2025 Challenge on HR Depth from Images of Specular and Transparent Surfaces." pith.science (2026). https://pith.science/paper/QVTMZTSS
@misc{pith2026250605815,
author = {Pith},
title = {Pith review of: NTIRE 2025 Challenge on HR Depth from Images of Specular and Transparent Surfaces},
year = {2026},
howpublished = {\url{https://pith.science/paper/QVTMZTSS}},
note = {Machine review of arXiv:2506.05815}
}
read the original abstract
This paper reports on the NTIRE 2025 challenge on HR Depth From images of Specular and Transparent surfaces, held in conjunction with the New Trends in Image Restoration and Enhancement (NTIRE) workshop at CVPR 2025. This challenge aims to advance the research on depth estimation, specifically to address two of the main open issues in the field: high-resolution and non-Lambertian surfaces. The challenge proposes two tracks on stereo and single-image depth estimation, attracting about 177 registered participants. In the final testing stage, 4 and 4 participating teams submitted their models and fact sheets for the two tracks.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Stable diffusion v1.5 model card, 2022
2022
-
[2]
Neural disparity refinement for arbitrary res- olution stereo
Filippo Aleotti, Fabio Tosi, Pierluigi Zama Ramirez, Mat- teo Poggi, Samuele Salti, Stefano Mattoccia, and Luigi Di Stefano. Neural disparity refinement for arbitrary res- olution stereo. InInternational Conference on 3D Vision,
-
[3]
Stereo anywhere: Robust zero-shot deep stereo matching even where either stereo or mono fail
Luca Bartolomei, Fabio Tosi, Matteo Poggi, and Stefano Mattoccia. Stereo anywhere: Robust zero-shot deep stereo matching even where either stereo or mono fail. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
2025
-
[4]
Zoedepth: Zero-shot trans- fer by combining relative and metric depth, 2023
Shariq Farooq Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, and Matthias M ¨uller. Zoedepth: Zero-shot trans- fer by combining relative and metric depth, 2023
2023
-
[5]
Depth pro: Sharp monocular metric depth in less than a second.arXiv preprint arXiv:2410.02073, 2024
Aleksei Bochkovskii, Ama ˜AG ¸ l Delaunoy, Hugo Germain, Marcel Santos, Yichao Zhou, Stephan R Richter, and Vladlen Koltun. Depth pro: Sharp monocular metric depth in less than a second.arXiv preprint arXiv:2410.02073, 2024
arXiv 2024
-
[6]
Virtual kitti 2.arXiv preprint arXiv:2001.10773, 2020
Yohann Cabon, Naila Murray, and Martin Humenberger. Virtual kitti 2.arXiv preprint arXiv:2001.10773, 2020
arXiv 2001
-
[7]
Pyramid stereo matching network
Jia-Ren Chang and Yong-Sheng Chen. Pyramid stereo matching network. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5410–5418, 2018
2018
-
[8]
Single-image depth perception in the wild
Weifeng Chen, Zhao Fu, Dawei Yang, and Jia Deng. Single-image depth perception in the wild. InProc. NeurIPS, 2016
2016
Show all 135 references
-
[9]
NTIRE 2025 challenge on image super-resolution (×4): Methods and results
Zheng Chen, Kai Liu, Jue Gong, Jingkai Wang, Lei Sun, Zongwei Wu, Radu Timofte, Yulun Zhang, et al. NTIRE 2025 challenge on image super-resolution (×4): Methods and results. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Work- shops, 2025
2025
-
[10]
NTIRE 2025 challenge on real-world face restoration: Methods and results
Zheng Chen, Jingkai Wang, Kai Liu, Jue Gong, Lei Sun, Zongwei Wu, Radu Timofte, Yulun Zhang, et al. NTIRE 2025 challenge on real-world face restoration: Methods and results. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Work- shops, 2025
2025
-
[11]
Monster: Marry monodepth to stereo unleashes power.arXiv preprint arXiv:2501.08643, 2025
Junda Cheng, Longliang Liu, Gangwei Xu, Xianqi Wang, Zhaoxing Zhang, Yong Deng, Jinliang Zang, Yurui Chen, Zhipeng Cai, and Xin Yang. Monster: Marry monodepth to stereo unleashes power.arXiv preprint arXiv:2501.08643, 2025
2025
-
[12]
Learn- ing depth with convolutional spatial propagation network
Xinjing Cheng, Peng Wang, and Ruigang Yang. Learn- ing depth with convolutional spatial propagation network. IEEE transactions on pattern analysis and machine intelli- gence, 42(10):2361–2379, 2019
2019
-
[13]
Hierarchical neural architecture search for deep stereo matching.Advances in Neural Information Pro- cessing Systems, 33, 2020
Xuelian Cheng, Yiran Zhong, Mehrtash Harandi, Yuchao Dai, Xiaojun Chang, Hongdong Li, Tom Drummond, and Zongyuan Ge. Hierarchical neural architecture search for deep stereo matching.Advances in Neural Information Pro- cessing Systems, 33, 2020
2020
-
[14]
Selfdeco: Self- supervised monocular depth completion in challenging in- door environments
Jaehoon Choi, Dongki Jung, Yonghan Lee, Deokhwa Kim, Dinesh Manocha, and Donghwan Lee. Selfdeco: Self- supervised monocular depth completion in challenging in- door environments. In2021 IEEE International Conference on Robotics and Automation (ICRA), pages 467–474. IEEE, 2021
2021
-
[15]
NTIRE 2025 challenge on raw image restoration and super-resolution
Marcos Conde, Radu Timofte, et al. NTIRE 2025 challenge on raw image restoration and super-resolution. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025
2025
-
[16]
Raw image reconstruc- tion from RGB on smartphones
Marcos Conde, Radu Timofte, et al. Raw image reconstruc- tion from RGB on smartphones. NTIRE 2025 challenge re- port. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR) Workshops, 2025
2025
-
[17]
Learning depth estimation for transparent and mirror sur- faces
Alex Costanzino, Pierluigi Zama Ramirez, Matteo Poggi, Fabio Tosi, Stefano Mattoccia, and Luigi Di Stefano. Learning depth estimation for transparent and mirror sur- faces. InThe IEEE International Conference on Computer Vision, 2023. ICCV
2023
-
[18]
Deeppruner: Learning efficient stereo matching via differentiable patchmatch
Shivam Duggal, Shenlong Wang, Wei-Chiu Ma, Rui Hu, and Raquel Urtasun. Deeppruner: Learning efficient stereo matching via differentiable patchmatch. InProceedings of the IEEE/CVF international conference on computer vi- sion, pages 4384–4393, 2019
2019
-
[19]
Depth map prediction from a single image using a multi-scale deep network
David Eigen, Christian Puhrsch, and Rob Fergus. Depth map prediction from a single image using a multi-scale deep network. InProc. NeurIPS, 2014
2014
-
[20]
NTIRE 2025 challenge on night photography rendering
Egor Ershov, Sergey Korchagin, Alexei Khalin, Artyom Panshin, Arseniy Terekhin, Ekaterina Zaychenkova, Georgiy Lobarev, Vsevolod Plokhotnyuk, Denis Abramov, Elisey Zhdanov, Sofia Dorogova, Yasin Mamedov, Nikola Banic, Georgii Perevozchikov, Radu Timofte, et al. NTIRE 2025 chal...
2025
-
[21]
Transcg: A large-scale real-world dataset for transparent object depth completion and a grasping baseline.IEEE Robotics and Automation Letters, 7(3):7383–7390, 2022
Hongjie Fang, Hao-Shu Fang, Sheng Xu, and Cewu Lu. Transcg: A large-scale real-world dataset for transparent object depth completion and a grasping baseline.IEEE Robotics and Automation Letters, 7(3):7383–7390, 2022
2022
-
[22]
Ge- owizard: Unleashing the diffusion priors for 3d geometry estimation from a single image
Xiao Fu, Wei Yin, Mu Hu, Kaixuan Wang, Yuexin Ma, Ping Tan, Shaojie Shen, Dahua Lin, and Xiaoxiao Long. Ge- owizard: Unleashing the diffusion priors for 3d geometry estimation from a single image. InECCV, 2024
2024
-
[23]
NTIRE 2025 challenge on cross-domain few-shot object detection: Methods and results
Yuqian Fu, Xingyu Qiu, Bin Ren Yanwei Fu, Radu Timofte, Nicu Sebe, Ming-Hsuan Yang, Luc Van Gool, et al. NTIRE 2025 challenge on cross-domain few-shot object detection: Methods and results. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CV...
2025
-
[24]
Dense depth for autonomous driving (DDAD) challenge (https://sites.google.com/ view/mono3d-workshop), 2021
Adrien Gaidon, Greg Shakhnarovich, Rares Ambrus, Vi- tor Guizilini, Igor Vasiljevic, Matthew Walter, Sudeep Pil- lai, and Nick Kolkin. Dense depth for autonomous driving (DDAD) challenge (https://sites.google.com/ view/mono3d-workshop), 2021
2021
-
[25]
Effi- cient large-scale stereo matching
Andreas Geiger, Martin Roser, and Raquel Urtasun. Effi- cient large-scale stereo matching. InAsian conference on computer vision, pages 25–38. Springer, 2010
2010
-
[26]
Unsupervised monocular depth estimation with left- right consistency
Cl ´ement Godard, Oisin Mac Aodha, and Gabriel J Bros- tow. Unsupervised monocular depth estimation with left- right consistency. InProc. CVPR, 2017
2017
-
[27]
Digging into self-supervised monocular depth estimation
Cl ´ement Godard, Oisin Mac Aodha, Michael Firman, and 9 Gabriel J Brostow. Digging into self-supervised monocular depth estimation. InProc. ICCV, 2019
2019
-
[28]
Forget about the lidar: Self-supervised depth estimators with med prob- ability volumes
Juan Luis GonzalezBello and Munchurl Kim. Forget about the lidar: Self-supervised depth estimators with med prob- ability volumes. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors,Advances in Neural In- formation Processing Systems, volume 33, pages ...
2020
-
[29]
Context-enhanced stereo transformer
Weiyu Guo, Zhaoshuo Li, Yongkui Yang, Zheng Wang, Russell H Taylor, Mathias Unberath, Alan Yuille, and Ying- wei Li. Context-enhanced stereo transformer. InEuropean Conference on Computer Vision, pages 263–279. Springer, 2022
2022
-
[30]
Learning monocular depth by distilling cross-domain stereo networks
Xiaoyang Guo, Hongsheng Li, Shuai Yi, Jimmy Ren, and Xiaogang Wang. Learning monocular depth by distilling cross-domain stereo networks. InProc. ECCV, 2018
2018
-
[31]
Group-wise correlation stereo network
Xiaoyang Guo, Kai Yang, Wukui Yang, Xiaogang Wang, and Hongsheng Li. Group-wise correlation stereo network. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3273–3282, 2019
2019
-
[32]
NTIRE 2025 challenge on text to image generation model qual- ity assessment
Shuhao Han, Haotian Fan, Fangyuan Kong, Wenjie Liao, Chunle Guo, Chongyi Li, Radu Timofte, et al. NTIRE 2025 challenge on text to image generation model qual- ity assessment. InProceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025
2025
-
[33]
Lotus: Diffusion-based visual foundation model for high-quality dense prediction.arXiv preprint arXiv:2409.18124, 2024
Jing He, Haodong Li, Wei Yin, Yixun Liang, Leheng Li, Kaiqiang Zhou, Hongbo Zhang, Bingbing Liu, and Ying- Cong Chen. Lotus: Diffusion-based visual foundation model for high-quality dense prediction.arXiv preprint arXiv:2409.18124, 2024
2024 arXiv
-
[34]
Depthcrafter: Generating consistent long depth sequences for open-world videos
Wenbo Hu, Xiangjun Gao, Xiaoyu Li, Sijie Zhao, Xi- aodong Cun, Yong Zhang, Long Quan, and Ying Shan. Depthcrafter: Generating consistent long depth sequences for open-world videos. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
2025
-
[35]
Fast and accurate single-image depth estimation on mobile devices, mobile ai 2021 challenge: Report
Andrey Ignatov, Grigory Malivenko, David Plowman, Samarth Shukla, and Radu Timofte. Fast and accurate single-image depth estimation on mobile devices, mobile ai 2021 challenge: Report. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) W...
2021
-
[36]
NTIRE 2025 challenge on video quality enhancement for video conferencing: Datasets, methods and results
Varun Jain, Zongwei Wu, Quan Zou, Louis Florentin, Henrik Turbell, Sandeep Siddhartha, Radu Timofte, et al. NTIRE 2025 challenge on video quality enhancement for video conferencing: Datasets, methods and results. InPro- ceedings of the IEEE/CVF Conference on Computer Vision an...
2025
-
[37]
Self- supervised relative depth learning for urban scene understanding
Huaizu Jiang, Gustav Larsson, Michael Maire Greg Shakhnarovich, and Erik Learned-Miller. Self- supervised relative depth learning for urban scene understanding. InProc. ECCV, 2018
2018
-
[38]
Defom- stereo: Depth foundation model based stereo matching
Hualie Jiang, Zhiqiang Lou, Laiyan Ding, Rui Xu, Minglang Tan, Wenjie Jiang, and Rui Huang. Defom- stereo: Depth foundation model based stereo matching. arXiv preprint arXiv:2501.09466, 2025
2025 arXiv
-
[39]
Uncertainty guided adaptive warping for robust and efficient stereo matching
Junpeng Jing, Jiankun Li, Pengfei Xiong, Jiangyu Liu, Shuaicheng Liu, Yichen Guo, Xin Deng, Mai Xu, Lai Jiang, and Leonid Sigal. Uncertainty guided adaptive warping for robust and efficient stereo matching. InProceedings of the IEEE/CVF International Conference on Computer Vis...
2023
-
[40]
Self-supervised monocular trained depth estimation using self-attention and discrete disparity volume
Adrian Johnston and Gustavo Carneiro. Self-supervised monocular trained depth estimation using self-attention and discrete disparity volume. InProc. CVPR, 2020
2020
-
[41]
On the importance of accurate geometry data for dense 3d vision tasks
HyunJun Jung, Patrick Ruhkamp, Guangyao Zhai, Nikolas Brasch, Yitong Li, Yannick Verdie, Jifei Song, Yiren Zhou, Anil Armagan, Slobodan Ilic, et al. On the importance of accurate geometry data for dense 3d vision tasks. InPro- ceedings of the IEEE/CVF Conference on Computer Vi...
2023
-
[42]
Housecat6d-a large-scale multi-modal category level 6d object perception dataset with household objects in realistic scenarios
HyunJun Jung, Shun-Cheng Wu, Patrick Ruhkamp, Guangyao Zhai, Hannah Schieber, Giulia Rizzoli, Pengyuan Wang, Hongcheng Zhao, Lorenzo Garattoni, Sven Meier, et al. Housecat6d-a large-scale multi-modal category level 6d object perception dataset with household objects in realist...
2024
-
[43]
Video depth without video models
Bingxin Ke, Dominik Narnhofer, Shengyu Huang, Lei Ke, Torben Peters, Katerina Fragkiadaki, Anton Obukhov, and Konrad Schindler. Video depth without video models. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), 2025
2025
-
[44]
Re- purposing diffusion-based image generators for monocular depth estimation
Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Met- zger, Rodrigo Caye Daudt, and Konrad Schindler. Re- purposing diffusion-based image generators for monocular depth estimation. InProceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[45]
End-to-end learning of geometry and context for deep stereo regression
Alex Kendall, Hayk Martirosyan, Saumitro Dasgupta, Pe- ter Henry, Ryan Kennedy, Abraham Bachrach, and Adam Bry. End-to-end learning of geometry and context for deep stereo regression. InThe IEEE International Conference on Computer Vision (ICCV), Oct 2017
2017
-
[46]
Stereonet: Guided hierarchical refinement for real-time edge-aware depth prediction
Sameh Khamis, Sean Fanello, Christoph Rhemann, Adarsh Kowdle, Julien Valentin, and Shahram Izadi. Stereonet: Guided hierarchical refinement for real-time edge-aware depth prediction. InProceedings of the European Confer- ence on Computer Vision (ECCV), pages 573–590, 2018
2018
-
[47]
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. InProceedings of the IEEE/CVF international conference on computer vision, pages 4015–4026, 2023
2023
-
[48]
Ai2-thor: An interactive 3d environment for visual ai.arXiv preprint arXiv:1712.05474, 2017
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli Vander- Bilt, Luca Weihs, Alvaro Herrasti, Matt Deitke, Kiana Ehsani, Daniel Gordon, Yuke Zhu, et al. Ai2-thor: An interactive 3d environment for visual ai.arXiv preprint arXiv:1712.05474, 2017
2017 arXiv
-
[49]
Alvarez, Yan Wang, Vincent Casser, Fisher Yu, Marco Pavone, Bo Li, Andreas Geiger, Peter Ondruska, Li Erran Li, Dragomir Angelov, John Leonard, and Luc Van Gool
Henrik Kretzschmar, Alex Liniger, Jose M. Alvarez, Yan Wang, Vincent Casser, Fisher Yu, Marco Pavone, Bo Li, Andreas Geiger, Peter Ondruska, Li Erran Li, Dragomir Angelov, John Leonard, and Luc Van Gool. Argov- erse stereo competition (https://cvpr2022.wad. vision/), 2021, 2022
2021
-
[50]
Deeper depth prediction with fully convolutional residual networks
Iro Laina, Christian Rupprecht, Vasileios Belagiannis, Fed- 10 erico Tombari, and Nassir Navab. Deeper depth prediction with fully convolutional residual networks. In2016 Fourth international conference on 3D vision (3DV), pages 239–
-
[51]
NTIRE 2025 challenge on efficient burst hdr and restoration: Datasets, methods, and results
Sangmin Lee, Eunpil Park, Angel Canelo, Hyunhee Park, Youngjo Kim, Hyungju Chun, Xin Jin, Chongyi Li, Chun- Le Guo, Radu Timofte, et al. NTIRE 2025 challenge on efficient burst hdr and restoration: Datasets, methods, and results. InProceedings of the IEEE/CVF Conference on Com...
2025
-
[52]
Practical stereo matching via cascaded recurrent net- work with adaptive correlation
Jiankun Li, Peisen Wang, Pengfei Xiong, Tao Cai, Ziwei Yan, Lei Yang, Jiangyu Liu, Haoqiang Fan, and Shuaicheng Liu. Practical stereo matching via cascaded recurrent net- work with adaptive correlation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...
2022
-
[53]
NTIRE 2025 challenge on day and night raindrop removal for dual-focused images: Methods and results
Xin Li, Yeying Jin, Xin Jin, Zongwei Wu, Bingchen Li, Yufei Wang, Wenhan Yang, Yu Li, Zhibo Chen, Bihan Wen, Robby Tan, Radu Timofte, et al. NTIRE 2025 challenge on day and night raindrop removal for dual-focused images: Methods and results. InProceedings of the IEEE/CVF Confe...
2025
-
[54]
NTIRE 2025 challenge on short-form ugc video quality assessment and enhancement: Kwaisr dataset and study
Xin Li, Xijun Wang, Bingchen Li, Kun Yuan, Yizhen Shao, Suhang Yao, Ming Sun, Chao Zhou, Radu Timofte, and Zhibo Chen. NTIRE 2025 challenge on short-form ugc video quality assessment and enhancement: Kwaisr dataset and study. InProceedings of the IEEE/CVF Conference on Compute...
2025
-
[55]
NTIRE 2025 challenge on short-form ugc video qual- ity assessment and enhancement: Methods and results
Xin Li, Kun Yuan, Bingchen Li, Fengbin Guan, Yizhen Shao, Zihao Yu, Xijun Wang, Yiting Lu, Wei Luo, Suhang Yao, Ming Sun, Chao Zhou, Zhibo Chen, Radu Timofte, et al. NTIRE 2025 challenge on short-form ugc video qual- ity assessment and enhancement: Methods and results. In Proc...
2025
-
[56]
Patch- fusion: An end-to-end tile-based framework for high- resolution monocular metric depth estimation
Zhenyu Li, Shariq Farooq Bhat, and Peter Wonka. Patch- fusion: An end-to-end tile-based framework for high- resolution monocular metric depth estimation. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[57]
Revisiting stereo depth estimation from a sequence- to-sequence perspective with transformers
Zhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy Ding, Francis X Creighton, Russell H Taylor, and Mathias Un- berath. Revisiting stereo depth estimation from a sequence- to-sequence perspective with transformers. InProceedings of the IEEE/CVF International Conference on Compute...
2021
-
[58]
NTIRE 2025 the 2nd restore any image model (RAIM) in the wild challenge
Jie Liang, Radu Timofte, Qiaosi Yi, Zhengqiang Zhang, Shuaizheng Liu, Lingchen Sun, Rongyuan Wu, Xindong Zhang, Hui Zeng, Lei Zhang, et al. NTIRE 2025 the 2nd restore any image model (RAIM) in the wild challenge. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion a...
2025
-
[59]
Delving into multi-illumination monocular depth esti- mation: A new dataset and method.IEEE Transactions on Multimedia, 2024
Yuan Liang, Zitian Zhang, Chuhua Xian, and Shengfeng He. Delving into multi-illumination monocular depth esti- mation: A new dataset and method.IEEE Transactions on Multimedia, 2024
2024
-
[60]
Learning for disparity estimation through feature constancy
Zhengfa Liang, Yiliu Feng, Yulan Guo, Hengzhu Liu, Wei Chen, Linbo Qiao, Li Zhou, and Jianfeng Zhang. Learning for disparity estimation through feature constancy. InPro- ceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018
2018
-
[61]
RAFT-Stereo: Multilevel Recurrent Field Transforms for Stereo Match- ing.arXiv preprint arXiv:2109.07547, 2021
Lahav Lipson, Zachary Teed, and Jia Deng. RAFT-Stereo: Multilevel Recurrent Field Transforms for Stereo Match- ing.arXiv preprint arXiv:2109.07547, 2021
2021 arXiv
-
[62]
NTIRE 2025 XGC quality assessment challenge: Methods and results
Xiaohong Liu, Xiongkuo Min, Qiang Hu, Xiaoyun Zhang, Jie Guo, et al. NTIRE 2025 XGC quality assessment challenge: Methods and results. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025
2025
-
[63]
NTIRE 2025 challenge on low light image enhancement: Methods and results
Xiaoning Liu, Zongwei Wu, Florin-Alexandru Vasluianu, Hailong Yan, Bin Ren, Yulun Zhang, Shuhang Gu, Le Zhang, Ce Zhu, Radu Timofte, et al. NTIRE 2025 challenge on low light image enhancement: Methods and results. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion ...
2025
-
[64]
Elfnet: Evidential local-global fusion for stereo matching
Jieming Lou, Weide Liu, Zhuo Chen, Fayao Liu, and Jun Cheng. Elfnet: Evidential local-global fusion for stereo matching. InProceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages 17784– 17793, 2023
2023
-
[65]
A large dataset to train convolutional networks for dispar- ity, optical flow, and scene flow estimation
Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for dispar- ity, optical flow, and scene flow estimation. InThe IEEE Conference on Computer Vision and Pattern Recogn...
2016
-
[66]
Depth-aware mirror seg- mentation
Haiyang Mei, Bo Dong, Wen Dong, Pieter Peers, Xin Yang, Qiang Zhang, and Xiaopeng Wei. Depth-aware mirror seg- mentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3044– 3053, 2021
2021
-
[67]
Object scene flow for autonomous vehicles
Moritz Menze and Andreas Geiger. Object scene flow for autonomous vehicles. InConference on Computer Vision and Pattern Recognition (CVPR), 2015
2015
-
[68]
Boosting monocular depth estima- tion models to high-resolution via content-adaptive multi- resolution merging
S Mahdi H Miangoleh, Sebastian Dille, Long Mai, Sylvain Paris, and Yagiz Aksoy. Boosting monocular depth estima- tion models to high-resolution via content-adaptive multi- resolution merging. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition...
2021
-
[69]
DeSouza, Tuan-Anh Yang, Minh-Quang Nguyen, Thien-Phuc Tran, Albert Luginov, 11 and Muhammad Shahzad
Anton Obukhov, Matteo Poggi, Fabio Tosi, Ripu- daman Singh Arora, Jaime Spencer, Chris Russell, Si- mon Hadfield, Richard Bowden, Shuaihang Wang, Zhenxin Ma, Weijie Chen, Baobei Xu, Fengyu Sun, Di Xie, Jiang Zhu, Mykola Lavreniuk, Haining Guan, Qun Wu, Yupei Zeng, Chao Lu, Hua...
2025
-
[70]
Ren, Chengxi Yang, and Qiong Yan
Jiahao Pang, Wenxiu Sun, Jimmy SJ. Ren, Chengxi Yang, and Qiong Yan. Cascade residual learning: A two-stage convolutional neural network for stereo matching. In The IEEE International Conference on Computer Vision (ICCV), Oct 2017
2017
-
[71]
On the uncertainty of self-supervised monocular depth estimation
Matteo Poggi, Filippo Aleotti, Fabio Tosi, and Stefano Mat- toccia. On the uncertainty of self-supervised monocular depth estimation. InProc. CVPR, 2020
2020
-
[72]
On the confidence of stereo matching in a deep-learning era: a quantitative evaluation.IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 2021
Matteo Poggi, Seungryong Kim, Fabio Tosi, Sunok Kim, Filippo Aleotti, Dongbo Min, Kwanghoon Sohn, and Ste- fano Mattoccia. On the confidence of stereo matching in a deep-learning era: a quantitative evaluation.IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 2021
2021
-
[73]
Federated online adaptation for deep stereo
Matteo Poggi and Fabio Tosi. Federated online adaptation for deep stereo. InCVPR, 2024
2024
-
[74]
On the synergies be- tween machine learning and binocular stereo for depth esti- mation from images: a survey.IEEE Transactions on Pat- tern Analysis and Machine Intelligence, 2021
Matteo Poggi, Fabio Tosi, Konstantinos Batsos, Philippos Mordohai, and Stefano Mattoccia. On the synergies be- tween machine learning and binocular stereo for depth esti- mation from images: a survey.IEEE Transactions on Pat- tern Analysis and Machine Intelligence, 2021
2021
-
[75]
Predicting sharp and accurate occlusion boundaries in monocular depth estimation using displacement fields
Michael Ramamonjisoa, Yuming Du, and Vincent Lep- etit. Predicting sharp and accurate occlusion boundaries in monocular depth estimation using displacement fields. In Proc. CVPR, 2020
2020
-
[76]
Ntire 2023 challenge on hr depth from images of specular and transparent surfaces
Pierluigi Zama Ramirez, Fabio Tosi, Luigi Di Stefano, Radu Timofte, Alex Costanzino, Matteo Poggi, Samuele Salti, Stefano Mattoccia, Jun Shi, Dafeng Zhang, et al. Ntire 2023 challenge on hr depth from images of specular and transparent surfaces. InProceedings of the IEEE/CVF C...
2023
-
[77]
Pierluigi Zama Ramirez, Fabio Tosi, Luigi Di Ste- fano, Radu Timofte, Alex Costanzino, Matteo Poggi, Samuele Salti, Stefano Mattoccia, Yangyang Zhang, Cailin Wu, Zhuangda He, Shuangshuang Yin, Jiaxu Dong, Yangchenxu Liu, Hao Jiang, Jun Shi, Yong A, Yixiang Jin, Dingzhe Li, Bin...
2024
-
[78]
Vi- sion transformers for dense prediction.ICCV, 2021
Ren ´e Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vi- sion transformers for dense prediction.ICCV, 2021
2021
-
[79]
Vi- sion transformers for dense prediction
Ren ´e Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vi- sion transformers for dense prediction. InProceedings of the IEEE/CVF International Conference on Computer Vi- sion (ICCV), pages 12179–12188, October 2021
2021
-
[80]
Towards robust monocu- lar depth estimation: Mixing datasets for zero-shot cross- dataset transfer.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(3), 2022
Ren ´e Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocu- lar depth estimation: Mixing datasets for zero-shot cross- dataset transfer.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(3), 2022
2022
-
[81]
The tenth NTIRE 2025 efficient super-resolution challenge report
Bin Ren, Hang Guo, Lei Sun, Zongwei Wu, Radu Tim- ofte, Yawei Li, et al. The tenth NTIRE 2025 efficient super-resolution challenge report. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025
2025
-
[82]
Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding
Mike Roberts, Jason Ramapuram, Anurag Ranjan, At- ulit Kumar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M Susskind. Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding. In Proceedings of the IEEE/CVF international conference o...
2021
-
[83]
NTIRE 2025 challenge on UGC video enhancement: Meth- ods and results
Nickolay Safonov, Alexey Bryntsev, Andrey Moskalenko, Dmitry Kulikov, Dmitriy Vatolin, Radu Timofte, et al. NTIRE 2025 challenge on UGC video enhancement: Meth- ods and results. InProceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) Worksh...
2025
-
[84]
Autodispnet: Improving disparity estimation with automl
Tonmoy Saikia, Yassine Marrakchi, Arber Zela, Frank Hut- ter, and Thomas Brox. Autodispnet: Improving disparity estimation with automl. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 1812– 1823, 2019
2019
-
[85]
Clear grasp: 3d shape estimation of transparent objects for ma- nipulation
Shreeyak Sajjan, Matthew Moore, Mike Pan, Ganesh Na- garaja, Johnny Lee, Andy Zeng, and Shuran Song. Clear grasp: 3d shape estimation of transparent objects for ma- nipulation. In2020 IEEE International Conference on Robotics and Automation (ICRA), pages 3634–3642. IEEE, 2020
2020
-
[86]
High-resolution stereo datasets with subpixel-accurate ground truth
Daniel Scharstein, Heiko Hirschm ¨uller, York Kitajima, Greg Krathwohl, Nera Ne ˇsi´c, Xi Wang, and Porter West- ling. High-resolution stereo datasets with subpixel-accurate ground truth. InGerman conference on pattern recognition, pages 31–42. Springer, 2014
2014
-
[87]
A multi-view stereo benchmark with high- resolution images and multi-camera videos
Thomas Schops, Johannes L Schonberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and An- dreas Geiger. A multi-view stereo benchmark with high- resolution images and multi-camera videos. InIEEE Con- ference on Computer Vision and Pattern Recognition,...
2017
-
[88]
Learning temporally consistent video depth from video diffusion priors
Jiahao Shao, Yuanbo Yang, Hongyu Zhou, Youmin Zhang, Yujun Shen, Vitor Guizilini, Yue Wang, Matteo Poggi, and Yiyi Liao. Learning temporally consistent video depth from video diffusion priors. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition ...
2025
-
[89]
Cfnet: Cas- cade and fused cost volume for robust stereo matching
Zhelun Shen, Yuchao Dai, and Zhibo Rao. Cfnet: Cas- cade and fused cost volume for robust stereo matching. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), pages 13906–13915, June 2021
2021
-
[90]
Asgrasp: Generalizable transparent object reconstruction and grasping from rgb-d active stereo cam- 12 era.arXiv preprint arXiv:2405.05648, 2024
Jun Shi, Yixiang Jin, Dingzhe Li, Haoyu Niu, Zhezhu Jin, He Wang, et al. Asgrasp: Generalizable transparent object reconstruction and grasping from rgb-d active stereo cam- 12 era.arXiv preprint arXiv:2405.05648, 2024
2024 arXiv
-
[91]
Edgestereo: A context integrated residual pyramid network for stereo matching
Xiao Song, Xu Zhao, Hanwen Hu, and Liangji Fang. Edgestereo: A context integrated residual pyramid network for stereo matching. InACCV, 2018
2018
-
[92]
Stella Qian, Chris Russell, Simon Had- field, Erich Graf, Wendy Adams, Andrew J
Jaime Spencer, C. Stella Qian, Chris Russell, Simon Had- field, Erich Graf, Wendy Adams, Andrew J. Schofield, James H. Elder, Richard Bowden, Heng Cong, Stefano Mattoccia, Matteo Poggi, Zeeshan Khan Suri, Yang Tang, Fabio Tosi, Hao Wang, Youmin Zhang, Yusheng Zhang, and Chaoqi...
2023
-
[93]
Stella Qian, Michaela Trescakova, Chris Russell, Simon Hadfield, Erich Graf, Wendy Adams, An- drew J
Jaime Spencer, C. Stella Qian, Michaela Trescakova, Chris Russell, Simon Hadfield, Erich Graf, Wendy Adams, An- drew J. Schofield, James Elder, Richard Bowden, Ali An- war, Hao Chen, Xiaozhi Chen, Kai Cheng, Yuchao Dai, Huynh Thai Hoa, Sadat Hossain, Jianmian Huang, Mo- han Ji...
2023
-
[94]
Jaime Spencer, Fabio Tosi, Matteo Poggi, Ripu- daman Singh Arora, Chris Russell, Simon Hadfield, Richard Bowden, GuangYuan Zhou, ZhengXin Li, Qiang Rao, YiPing Bao, Xiao Liu, Dohyeong Kim, Jinseong Kim, Myunghyun Kim, Mykola Lavreniuk, Rui Li, Qing Mao, Jiang Wu, Yu Zhu, Jinqi...
2024
-
[95]
NTIRE 2025 challenge on event-based image deblurring: Methods and results
Lei Sun, Andrea Alfarano, Peiqi Duan, Shaolin Su, Kaiwei Wang, Boxin Shi, Radu Timofte, Danda Pani Paudel, Luc Van Gool, et al. NTIRE 2025 challenge on event-based image deblurring: Methods and results. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...
2025
-
[96]
The tenth ntire 2025 image denoising challenge report
Lei Sun, Hang Guo, Bin Ren, Luc Van Gool, Radu Timo- fte, Yawei Li, et al. The tenth ntire 2025 image denoising challenge report. InProceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025
2025
-
[97]
Hitnet: Hierar- chical iterative tile refinement network for real-time stereo matching
Vladimir Tankovich, Christian Hane, Yinda Zhang, Adarsh Kowdle, Sean Fanello, and Sofien Bouaziz. Hitnet: Hierar- chical iterative tile refinement network for real-time stereo matching. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR),...
2021
-
[98]
Raft: Recurrent all-pairs field transforms for optical flow
Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. InEuropean conference on computer vision, pages 402–419. Springer, 2020
2020
-
[99]
Real-time self-adaptive deep stereo
Alessio Tonioni, Fabio Tosi, Matteo Poggi, Stefano Mat- toccia, and Luigi Di Stefano. Real-time self-adaptive deep stereo. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 195–204, 2019
2019
-
[100]
Learning monocular depth estimation infusing tradi- tional stereo knowledge
Fabio Tosi, Filippo Aleotti, Matteo Poggi, and Stefano Mat- toccia. Learning monocular depth estimation infusing tradi- tional stereo knowledge. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019
2019
-
[101]
Distilled semantics for comprehensive scene un- derstanding from videos
Fabio Tosi, Filippo Aleotti, Pierluigi Zama Ramirez, Mat- teo Poggi, Samuele Salti, Luigi Di Stefano, and Stefano Mattoccia. Distilled semantics for comprehensive scene un- derstanding from videos. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni...
2020
-
[102]
Neural Disparity Refinement .IEEE Transactions on Pattern Analysis & Machine Intelligence, 46(12):8900–8917, 2024
Fabio Tosi, Filippo Aleotti, Pierluigi Zama Ramirez, Mat- teo Poggi, Samuele Salti, Stefano Mattoccia, and Luigi Di Stefano. Neural Disparity Refinement .IEEE Transactions on Pattern Analysis & Machine Intelligence, 46(12):8900–8917, 2024
2024
-
[103]
A sur- vey on deep stereo matching in the twenties.International Journal of Computer Vision, pages 1–32, 2025
Fabio Tosi, Luca Bartolomei, and Matteo Poggi. A sur- vey on deep stereo matching in the twenties.International Journal of Computer Vision, pages 1–32, 2025
2025
-
[104]
Smd-nets: Stereo mixture density networks
Fabio Tosi, Yiyi Liao, Carolin Schmitt, and Andreas Geiger. Smd-nets: Stereo mixture density networks. InConfer- ence on Computer Vision and Pattern Recognition (CVPR), 2021
2021
-
[105]
Nerf-supervised deep stereo
Fabio Tosi, Alessio Tonioni, Daniele De Gregorio, and Mat- teo Poggi. Nerf-supervised deep stereo. InConference on Computer Vision and Pattern Recognition (CVPR), pages 855–866, June 2023
2023
-
[106]
Diffusion models for monocular depth estimation: Over- coming challenging conditions
Fabio Tosi, Pierluigi Zama Ramirez, and Matteo Poggi. Diffusion models for monocular depth estimation: Over- coming challenging conditions. InEuropean Conference on Computer Vision (ECCV), 2024
2024
-
[107]
NTIRE 2025 image shadow removal challenge report
Florin-Alexandru Vasluianu, Tim Seizinger, Zhuyun Zhou, Cailian Chen, Zongwei Wu, Radu Timofte, et al. NTIRE 2025 image shadow removal challenge report. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025
2025
-
[108]
NTIRE 2025 ambi- ent lighting normalization challenge
Florin-Alexandru Vasluianu, Tim Seizinger, Zhuyun Zhou, Zongwei Wu, Radu Timofte, et al. NTIRE 2025 ambi- ent lighting normalization challenge. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025
2025
-
[109]
CLIFFNet for monocular depth estimation with hierarchical embedding loss
Lijun Wang, Jianming Zhang, Yifan Wang, Huchuan Lu, and Xiang Ruan. CLIFFNet for monocular depth estimation with hierarchical embedding loss. InProc. ECCV, 2020
2020
-
[110]
Anytime stereo image depth estimation on mobile devices
Yan Wang, Zihang Lai, Gao Huang, Brian H Wang, Lau- rens Van Der Maaten, Mark Campbell, and Kilian Q Wein- berger. Anytime stereo image depth estimation on mobile devices. In2019 International Conference on Robotics and Automation (ICRA), pages 5893–5900, 2019
2019
-
[111]
NTIRE 2025 challenge on light 13 field image super-resolution: Methods and results
Yingqian Wang, Zhengyu Liang, Fengyuan Zhang, Lvli Tian, Longguang Wang, Juncheng Li, Jungang Yang, Radu Timofte, Yulan Guo, et al. NTIRE 2025 challenge on light 13 field image super-resolution: Methods and results. InPro- ceedings of the IEEE/CVF Conference on Computer Vision...
2025
-
[112]
Brostow, and Daniyar Turmukhambetov
Jamie Watson, Michael Firman, Gabriel J. Brostow, and Daniyar Turmukhambetov. Self-supervised monocular depth hints. InProc. ICCV, 2019
2019
-
[113]
Foundationstereo: Zero- shot stereo matching.arXiv, 2025
Bowen Wen, Matthew Trepte, Joseph Aribido, Jan Kautz, Orazio Gallo, and Stan Birchfield. Foundationstereo: Zero- shot stereo matching.arXiv, 2025
2025
-
[114]
Segmenting transparent ob- jects in the wild
Enze Xie, Wenjia Wang, Wenhai Wang, Mingyu Ding, Chunhua Shen, and Ping Luo. Segmenting transparent ob- jects in the wild. InComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIII 16, pages 696–711. Springer, 2020
2020
-
[116]
Iterative geometry encoding volume for stereo matching
Gangwei Xu, Xianqi Wang, Xiaohuan Ding, and Xin Yang. Iterative geometry encoding volume for stereo matching. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21919–21928, 2023
2023
-
[117]
Unifying flow, stereo and depth estimation.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023
Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao, and Andreas Geiger. Unifying flow, stereo and depth estimation.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023
2023
-
[118]
Hierarchical deep stereo matching on high- resolution images
Gengshan Yang, Joshua Manela, Michael Happold, and Deva Ramanan. Hierarchical deep stereo matching on high- resolution images. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 5515–5524, 2019
2019
-
[119]
Segstereo: Exploiting semantic infor- mation for disparity estimation
Guorun Yang, Hengshuang Zhao, Jianping Shi, Zhidong Deng, and Jiaya Jia. Segstereo: Exploiting semantic infor- mation for disparity estimation. InECCV, pages 636–651, 2018
2018
-
[120]
NTIRE 2025 challenge on single image reflection removal in the wild: Datasets, methods and results
Kangning Yang, Jie Cai, Ling Ouyang, Florin-Alexandru Vasluianu, Radu Timofte, Jiaming Ding, Huiming Sun, Lan Fu, Jinlong Li, Chiu Man Ho, Zibo Meng, et al. NTIRE 2025 challenge on single image reflection removal in the wild: Datasets, methods and results. InProceedings of the...
2025
-
[121]
Depth anything: Un- leashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Ji- ashi Feng, and Hengshuang Zhao. Depth anything: Un- leashing the power of large-scale unlabeled data. InCVPR, 2024
2024
-
[122]
Depth any- thing v2.Advances in Neural Information Processing Sys- tems, 37:21875–21911, 2024
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiao- gang Xu, Jiashi Feng, and Hengshuang Zhao. Depth any- thing v2.Advances in Neural Information Processing Sys- tems, 37:21875–21911, 2024
2024
-
[123]
Where is my mirror? In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8809–8818, 2019
Xin Yang, Haiyang Mei, Ke Xu, Xiaopeng Wei, Baocai Yin, and Rynson WH Lau. Where is my mirror? In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8809–8818, 2019
2019
-
[124]
Learning to recover 3d scene shape from a single image
Wei Yin, Jianming Zhang, Oliver Wang, Simon Niklaus, Long Mai, Simon Chen, and Chunhua Shen. Learning to recover 3d scene shape from a single image. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 204–213, 2021
2021
-
[125]
Hierarchi- cal discrete distribution decomposition for match density estimation
Zhichao Yin, Trevor Darrell, and Fisher Yu. Hierarchi- cal discrete distribution decomposition for match density estimation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6044– 6053, 2019
2019
-
[126]
Tricky 2024 challenge on monocu- lar depth from images of specular and transparent surfaces
Pierluigi Zama Ramirez, Alex Costanzino, Fabio Tosi, Mat- teo Poggi, Luigi Di Stefano, Jean-Baptiste Weibel, Do- minik Bauer, Doris Antensteiner, Markus Vincze, Jiaqi Li, Yachuan Huang, Junrui Zhang, Yiran Wang, Jinghong Zheng, Liao Shen, Zhiguo Cao, Ziyang Song, Zerong Wang, ...
2024
-
[127]
Booster: a benchmark for depth from im- ages of specular and transparent surfaces.arXiv preprint arXiv:2301.08245, 2023
Pierluigi Zama Ramirez, Alex Costanzino, Fabio Tosi, Mat- teo Poggi, Samuele Salti, Luigi Di Stefano, and Stefano Mattoccia. Booster: a benchmark for depth from im- ages of specular and transparent surfaces.arXiv preprint arXiv:2301.08245, 2023
2023 arXiv
-
[128]
Geometry meets semantics for semi-supervised monocular depth estima- tion
Pierluigi Zama Ramirez, Matteo Poggi, Fabio Tosi, Ste- fano Mattoccia, and Luigi Di Stefano. Geometry meets semantics for semi-supervised monocular depth estima- tion. InComputer Vision–ACCV 2018: 14th Asian Confer- ence on Computer Vision, Perth, Australia, December 2–6, 2018...
2018
-
[129]
NTIRE 2025 challenge on hr depth from images of specular and transparent surfaces
Pierluigi Zama Ramirez, Fabio Tosi, Luigi Di Stefano, Radu Timofte, Alex Costanzino, Matteo Poggi, Samuele Salti, Stefano Mattoccia, et al. NTIRE 2025 challenge on hr depth from images of specular and transparent surfaces. InProceedings of the IEEE/CVF Conference on Computer V...
2025
-
[130]
Open challenges in deep stereo: The booster dataset
Pierluigi Zama Ramirez, Fabio Tosi, Matteo Poggi, Samuele Salti, Stefano Mattoccia, and Luigi Di Stefano. Open challenges in deep stereo: The booster dataset. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), pages 21168–21178, June 2022
2022
-
[131]
Stereo matching by training a convolutional neural network to compare image patches.J
Jure Zbontar, Yann LeCun, et al. Stereo matching by training a convolutional neural network to compare image patches.J. Mach. Learn. Res., 17(1):2287–2318, 2016
2016
-
[132]
The robust vision challenge (http : / / www
Oliver Zendel, Angela Dai, Xavier Puig Fernandez, Andreas Geiger, Vladen Koltun, Peter Kontschieder, Adam Kortylewski, Tsung-Yi Lin, Torsten Sattler, Daniel Scharstein, Hendrik Schilling, Jonas Uhrig, and Jonas Wulff. The robust vision challenge (http : / / www . robustvision....
2018
-
[133]
Parameterized cost volume for stereo matching
Jiaxi Zeng, Chengtang Yao, Lidong Yu, Yuwei Wu, and Yunde Jia. Parameterized cost volume for stereo matching. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 18347–18357, October 2023
2023
-
[134]
GA-Net: Guided aggregation net for end- to-end stereo matching
Feihu Zhang, Victor Prisacariu, Ruigang Yang, and Philip HS Torr. GA-Net: Guided aggregation net for end- to-end stereo matching. InIEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), 2019. 14
2019
-
[135]
Gasmono: Geometry-aided self-supervised monocular depth estima- tion for indoor scenes
Chaoqiang Zhao, Matteo Poggi, Fabio Tosi, Lei Zhou, Qiyu Sun, Yang Tang, and Stefano Mattoccia. Gasmono: Geometry-aided self-supervised monocular depth estima- tion for indoor scenes. InProceedings of the IEEE/CVF In- ternational Conference on Computer Vision, pages 16209– 16220, 2023
2023
-
[136]
Unsupervised learning of stereo matching
Chao Zhou, Hong Zhang, Xiaoyong Shen, and Jiaya Jia. Unsupervised learning of stereo matching. InThe IEEE In- ternational Conference on Computer Vision (ICCV). IEEE, October 2017. 15
2017
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.