REVIEW 3 major objections 5 minor 73 references
This paper derives a closed-form geometric relationship between a drone's height, pitch, roll, and field of view and the ground-plane depth of every pixel, and shows that a monocular network which injects this relationship as a prior improv
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 07:26 UTC pith:ZQEWR6WH
load-bearing objection A well-engineered joint depth+pose model for UAV imagery with a useful new synthetic dataset, but the pose SOTA claim is undercut by an unclear baseline training protocol and a pitch-range contradiction. the 3 major comments →
DAPM: UAV Monocular Depth Estimation from Any Height, Pitch, Roll and FOV
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central discovery is a parameterization: a UAV image's perspective geometry is fully determined by height, pitch, roll, and field of view, and the ground-plane depth of any pixel is given by d = h·sqrt(tan²θx + tan²(φ+θy) + 1), where θx and θy are the pixel's angular offsets and φ is the pitch. Using this formula, the authors construct an Ideal Ground Depth map from estimated pose, feed it into the depth network as a geometric prior, and use it as a dense pose-supervision signal. They further show that coarse-to-fine per-pixel quantization bins, rather than globally shared depth bins, handle the large and shifting depth ranges of aerial scenes. On their new UAPD dataset, the full
What carries the argument
The Ideal Ground Depth (IGD) module is the central mechanism: a differentiable mapping from predicted pose parameters to a dense ground-plane depth image via the derived formula, which simultaneously supervises pose (by comparing IGD maps with ground truth) and enhances depth features (by concatenating the IGD map into the depth decoder). The Progressive Quantization Bins (PQB) module complements this with a coarse-to-fine classification scheme (16 then 64 bins per pixel, plus a binary sky/out-of-range head) that handles large depth scales. The UAPD dataset—42k simulator-rendered aerial images with continuously sampled height, pitch, roll, and FOV, plus depth and pose ground truth—provides t
Load-bearing premise
The load-bearing premise is that images and depth maps rendered by the simulator are an accurate proxy for real aerial drone imagery, so the pose-estimation gains measured on the synthetic dataset will transfer to real flight conditions.
What would settle it
Run the trained model on real drone footage with reference pose from an onboard RTK-GPS/IMU and reference depth from lidar: if the median roll or pitch error rises well above the synthetic-benchmark numbers (about 2–3 degrees), or if the ideal-ground-depth maps systematically misalign with lidar ground points, the simulator-to-real transfer claim fails.
If this is right
- A monocular drone can estimate both metric depth and its own camera pose (roll, pitch, FOV, height) without GPS/IMU, across a continuous range of flight attitudes and altitudes.
- The derived geometric formula creates a training signal that needs no extra labels: any image with a known ground plane can contribute pose supervision through the IGD loss.
- Per-pixel progressive quantization handles the huge dynamic range of aerial scenes more effectively than the globally shared bins used by prior depth estimators.
- The UAPD dataset offers a standardized benchmark for viewpoint robustness in UAV depth estimation, with pose annotations that existing aerial datasets lack.
- A standard convolutional backbone and real-time inference speed make the approach deployable on resource-constrained drone hardware.
Where Pith is reading between the lines
- The same ground-plane geometric prior could be turned into a self-supervised objective on real, unlabeled drone footage: predict pose, synthesize the IGD map, and enforce consistency with a monocular depth network, reducing the sim-to-real gap.
- The formula assumes a flat ground plane; on terrain with hills, buildings, or vegetation the prior becomes locally wrong. A testable extension is to fit the ground plane per image region or use the IGD residual as a terrain-non-flatness signal.
- Because the ideal ground depth encodes the horizon, the approach may double as a horizon and sky-ground segmentation prior, which could benefit other aerial tasks like object detection or visual odometry.
- The UAPD training distribution is bounded in height, pitch, roll, and FOV; extending it with more extreme attitudes or real images would test whether the geometric coupling holds outside the synthetic envelope.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DAPM, a joint monocular depth and camera-pose estimation framework for UAV imagery under continuously varying height, pitch, roll, and FOV. The authors derive a geometric relationship between camera pose and the ideal ground-plane depth distribution (Eq. 15), and use this relationship in an Ideal Ground Depth (IGD) module that provides dense pose supervision and geometric feature injection to the depth branch. A Progressive Quantization Bins (PQB) module is introduced for coarse-to-fine depth and pose estimation. The paper also presents UAPD, a new CARLA-based synthetic dataset of 42k images with continuous pose/FOV distributions. Experiments on UAPD report state-of-the-art depth and pose results, with additional results on Potsdam and extensive ablation studies.
Significance. If the claims hold, this is a useful contribution to UAV depth and pose estimation: it introduces a new public dataset with a wide continuous range of viewpoints, a clean geometric prior for aerial depth, and a joint training framework that demonstrably improves both tasks. The depth comparison uses a standardized toolbox, the ablations are thorough, and the authors promise to release code and data. The main risks are the unverified fairness of the pose baselines and the synthetic-only evaluation of the pose branch.
major comments (3)
- [§IV-D, Table VI] The pose comparison does not state whether DeepCalib, Perspective Fields, and GeoCalib were trained or fine-tuned on the UAPD training split. The text only says the experiments use the 'unified GeoCalib evaluation framework.' If these baselines are evaluated zero-shot with their original training, the reported gains (Roll 2.33° vs 3.25°, Pitch 2.36° vs 3.44°) may be artifacts of domain shift rather than evidence that the IGD depth-pose coupling improves pose estimation. Please clarify the training protocol and, if the baselines are zero-shot, retrain or fine-tune them on UAPD before claiming SOTA pose performance.
- [Tables I–II, Fig. 13] The pitch range is inconsistent across the manuscript. Table I lists Pitch as [-10,100] degrees, Table II lists [-100,10] degrees, and the Fig. 13 discussion describes a range from -100° to 0°. This ambiguity affects the pose normalization, the evaluation distribution, and the interpretation of pose ablation results. The correct range must be specified and used consistently throughout.
- [§III-B2, Eqs. (12)–(15)] The derivation of the ideal ground depth uses Hmax and Wmax in Eq. (12) without defining them. More importantly, the sign convention for the vertical image offset H_C (and hence θ_yI) is not stated. The expression y_w = h tan(φ + θ_yI) yields different depth patterns depending on whether H_C is positive in the upward or downward image direction. Since the IGD map is the core geometric prior and the basis for dense pose supervision, please define the coordinate frames and verify Eq. (15) with a concrete worked example.
minor comments (5)
- [§III-B2] The derivation assumes zero lens distortion, yet the UAPD dataset is collected with randomized lens distortion (Table II, 'Lens kcube' and 'Lens k'). Please discuss how this mismatch affects the validity of the IGD supervision and feature injection.
- [§III-B2, Eq. (1)] The notation P = [K T] is non-standard; typically the extrinsic is written as [R|t] or K[R|t]. Consider using a clearer convention.
- [§IV-D, Fig. 16] The text refers to 'UAPM' instead of 'DAPM' in the figure discussion. Please correct.
- [Table V] The DAPM(D) row reads '0.9440.315' — add a space before 0.315.
- [§III-C5] The loss-balancing weight ρ = 10 is introduced after Eq. (19) without explanation. A sentence justifying the value would improve reproducibility.
Circularity Check
No significant circularity: Eq. (15) is derived from pinhole geometry, the IGD loss is a reparameterized pose-supervision term using ground-truth pose, and no load-bearing claim reduces to its own input.
full rationale
The central derivation (Eq. 15) is obtained from the pinhole camera model: pixel offsets are converted to ray angles via focal length/FoV (Eqs. 10-12), projected onto a flat ground plane (Eq. 13), and the camera-to-ground distance is computed by Pythagoras (Eq. 14). No term in this chain is fitted to UAPD or defined in terms of the depth labels; the only inputs are assumed camera geometry (height, pitch, roll, fov) and the ideal-ground-plane model. The IGD loss likewise compares predicted pose and ground-truth pose after applying the same deterministic IGD transform; it is a reparameterized pose-supervision term, not a prediction of held-out depth. The depth-vs-pose loop (predicted pose -> IGD features -> depth refinement) is a network feedback path evaluated by holdout metrics and ablated in Tables VII-IX, so it does not make the claim true by construction. Self-citations in the introduction ([3],[5],[6],[13]) are background only and not load-bearing; there is no imported uniqueness theorem or ansatz-by-citation. The external Potsdam experiment provides an independent depth check. The two genuine concerns - the pose baseline training protocol in Table VI is unspecified (they may be zero-shot), and the pitch range is inconsistent between Table I (-10..100) and Table II (-100..10) - are reproducibility/fairness issues, not circularity. The conclusion's stated limitation (real-world data collection and complex terrain as future work) also confirms the evaluation is synthetic-only, but this does not enter the derivation as an input. No circular step can be exhibited.
Axiom & Free-Parameter Ledger
free parameters (5)
- Loss scaling ρ =
10
- Depth loss constants α, λ =
α=10, λ=0.85
- PQB bin counts =
16 and 64 (depth and pose)
- GT-pose mixing schedule =
Not specified (linear increase over epochs)
- Depth threshold for binary sky classification =
unspecified
axioms (4)
- domain assumption Ideal ground plane assumption
- domain assumption Pinhole camera with zero distortion, centered principal point, square pixels
- domain assumption Height, FoV, Pitch, Roll fully characterize imaging geometry (yaw and horizontal translation ignored)
- domain assumption CARLA-rendered depth maps are ground-truth metric depth
read the original abstract
Monocular depth estimation is a fundamental prerequisite for 3D reconstruction and autonomous navigation in Unmanned Aerial Vehicles (UAVs). In practical deployments, UAVs operate under highly dynamic camera poses characterized by continuous variations in height, pitch, roll, and field of view (FOV). Existing monocular depth estimation methods frequently fail to generalize across such diverse perspectives and the expansive scale of depth distributions inherent in aerial scenes. To address these challenges, we establish a quantitative representation of UAV viewing angles through rigorous theoretical analysis, deriving the geometric correspondence between viewing angles and view distances using the ground plane as a reference for observation. Building upon this, we propose Depth Estimation for Any Perspectives Model (DAPM), representing the first monocular framework specifically designed for UAV aerial imagery to jointly estimate camera pose and depth under continuously varying viewpoints. Specifically, we introduce an Ideal Ground Depth (IGD) module that leverages the derived geometric relationships between UAV perspectives and view distances to implement dense camera-pose supervision and enhance depth features. And we further develop a coarse-to-fine Progressive Quantization Bins (PQB) module. By incorporating progressive supervision and hierarchical quantization bins, the PQB module enables robust estimation in complex UAV aerial imagery. To evaluate the proposed framework, we present the UAV Any Perspectives Depth (UAPD) dataset, featuring comprehensive and continuous distributions of pose parameters. Experimental results on UAPD demonstrate that DAPM achieves state-of-the-art performance across both depth and camera-pose estimation metrics. The source code and datasets are available at: https://github.com/ThisIsLT/DAPM.
Figures
Reference graph
Works this paper leans on
-
[1]
Z. Wang, P. Cheng, M. Chen, P. a. Tian, Z. Wang, X. Li, X. Yang, and X. Sun, “Drones help drones: A collaborative framework for multi-drone object trajectory prediction and beyond,”arXiv preprint arXiv:2405.14674, 2024. 1
Pith/arXiv arXiv 2024
-
[2]
Ucdnet: Multi-uav collaborative 3d object detection network by reliable feature mapping,
P. Tian, Z. Wang, P. Cheng, Y . Wang, Z. Wang, L. Zhao, M. Yan, X. Yang, and X. Sun, “Ucdnet: Multi-uav collaborative 3d object detection network by reliable feature mapping,”IEEE Trans. Geosci. Remote Sens., 2024. 1
2024
-
[3]
Ringmo-aerial: An aerial remote sensing foundation model with affine transformation contrastive learning,
W. Diao, H. Yu, K. Kang, T. Ling, D. Liu, Y . Feng, H. Bi, L. Ren, X. Li, Y . Mao,et al., “Ringmo-aerial: An aerial remote sensing foundation model with affine transformation contrastive learning,”IEEE Trans. Pattern Anal. Mach. Intell., 2025. 1
2025
-
[4]
A progressive target- aware network for drone-based person detection using rgb-t images,
Z. He, B. Zhao, Y . Wu, Y . Jiang, and Q. Zhao, “A progressive target- aware network for drone-based person detection using rgb-t images,” Remote Sens., vol. 17, no. 19, p. 3361, 2025. 1
2025
-
[5]
H. Hu, P. Wang, Y . Feng, K. Wei, W. Yin, W. Diao, M. Wang, H. Bi, K. Kang, T. Ling,et al., “Ringmo-agent: A unified remote sensing foundation model for multi-platform and multi-modal reasoning,”arXiv preprint arXiv:2507.20776, 2025. 1
arXiv 2025
-
[6]
Supermot: Decoupling motion and fusing temporal pyramid features for uav multi-object tracking,
L. Ren, W. Yin, W. Diao, K. Fu, and X. Sun, “Supermot: Decoupling motion and fusing temporal pyramid features for uav multi-object tracking,”IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens., 2025. 1
2025
-
[7]
Uav-on: A benchmark for open-world object goal navigation with aerial agents,
J. Xiao, Y . Sun, Y . Shao, B. Gan, R. Liu, Y . Wu, W. Guan, and X. Deng, “Uav-on: A benchmark for open-world object goal navigation with aerial agents,” inACM MM, pp. 13023–13029, 2025. 1
2025
-
[8]
Explainable ai and monocular vision for enhanced uav navigation in smart cities: prospects and challenges,
S. Javaid, M. A. Khan, H. Fahim, B. He, and N. Saeed, “Explainable ai and monocular vision for enhanced uav navigation in smart cities: prospects and challenges,”Front. Sustain. Cities, vol. 7, p. 1561404,
-
[9]
An improved deep q-learning approach for navigation of an autonomous uav agent in 3d obstacle-cluttered environment,
G. Farid, M. Bilal, L. Zhang, A. Alharbi, I. Ahmed, and M. Azhar, “An improved deep q-learning approach for navigation of an autonomous uav agent in 3d obstacle-cluttered environment,”Drones, vol. 9, no. 8, p. 518, 2025. 1
2025
-
[10]
Single drone-based 3d reconstruction approach to improve public engagement in conservation of heritage buildings: A case of hakka tulou,
Q. Li, G. Yang, C. Gao, Y . Huang, J. Zhang, D. Huang, B. Zhao, X. Chen, and B. M. Chen, “Single drone-based 3d reconstruction approach to improve public engagement in conservation of heritage buildings: A case of hakka tulou,”J. Build. Eng., vol. 87, p. 108954,
-
[11]
A review on viewpoints and path planning for uav-based 3-d reconstruction,
M. Maboudi, M. Homaei, S. Song, S. Malihi, M. Saadatseresht, and M. Gerke, “A review on viewpoints and path planning for uav-based 3-d reconstruction,”IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens., vol. 16, pp. 5026–5048, 2023. 1
2023
-
[12]
Super-resolution techniques in photogrammetric 3d reconstruction from close-range uav imagery,
A. Panagiotopoulou, L. Grammatikopoulos, A. El Saer, E. Petsa, E. Charou, L. Ragia, and G. Karras, “Super-resolution techniques in photogrammetric 3d reconstruction from close-range uav imagery,” Heritage, vol. 6, no. 3, pp. 2701–2715, 2023. 1
2023
-
[13]
F. Yao, Y . Yue, Y . Liu, X. Sun, and K. Fu, “Aeroverse: Uav-agent benchmark suite for simulating, pre-training, finetuning, and evaluating aerospace embodied world models,”arXiv preprint arXiv:2408.15511,
-
[14]
Uavs meet agentic ai: A multidomain survey of autonomous aerial intelligence and agentic uavs,
R. Sapkota, K. I. Roumeliotis, and M. Karkee, “Uavs meet agentic ai: A multidomain survey of autonomous aerial intelligence and agentic uavs,” arXiv preprint arXiv:2506.08045, 2025. 1
Pith/arXiv arXiv 2025
-
[15]
Bedi: A compre- hensive benchmark for evaluating embodied agents on uavs,
M. Guo, M. Wu, J. He, S. Li, H. Li, and C. Tao, “Bedi: A compre- hensive benchmark for evaluating embodied agents on uavs,”ISPRS J. Photogramm. Remote Sens., vol. 232, pp. 910–936, 2026. 1
2026
-
[16]
Depthmaster: Taming diffusion models for monocular depth estimation,
Z. Song, Z. Wang, B. Li, H. Zhang, R. Zhu, L. Liu, P.-T. Jiang, and T. Zhang, “Depthmaster: Taming diffusion models for monocular depth estimation,”arXiv preprint arXiv:2501.02576, 2025. 1
Pith/arXiv arXiv 2025
-
[17]
Depthfm: Fast generative monocular depth estimation with flow matching,
M. Gui, J. Schusterbauer, U. Prestel, P. Ma, D. Kotovenko, O. Grebenkova, S. A. Baumann, V . T. Hu, and B. Ommer, “Depthfm: Fast generative monocular depth estimation with flow matching,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, pp. 3203–3211, 2025. 1
2025
-
[18]
Adabins: Depth estimation using adaptive bins,
S. F. Bhat, I. Alhashim, and P. Wonka, “Adabins: Depth estimation using adaptive bins,” inCVPR, pp. 4009–4018, 2021. 1, 3, 12, 14
2021
-
[19]
Binsformer: Revisiting adaptive bins for monocular depth estimation,
Z. Li, X. Wang, X. Liu, and J. Jiang, “Binsformer: Revisiting adaptive bins for monocular depth estimation,”IEEE Trans. Image Process.,
-
[20]
Perspective fields for single image camera calibration,
L. Jin, J. Zhang, Y . Hold-Geoffroy, O. Wang, K. Blackburn-Matzen, M. Sticha, and D. F. Fouhey, “Perspective fields for single image camera calibration,” inCVPR, pp. 17307–17316, 2023. 2, 3, 14, 15
2023
-
[21]
Geocalib: Learning single-image calibration with geometric optimization,
A. Veicht, P.-E. Sarlin, P. Lindenberger, and M. Pollefeys, “Geocalib: Learning single-image calibration with geometric optimization,” in ECCV, pp. 1–20, Springer, 2024. 2, 3, 4, 11, 14, 15
2024
-
[22]
Vision meets robotics: The kitti dataset,
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets robotics: The kitti dataset,”Int. J. Robot. Res., vol. 32, no. 11, pp. 1231–1237,
-
[23]
Indoor segmentation and support inference from rgbd images,
N. Silberman, D. Hoiem, P. Kohli, and R. Fergus, “Indoor segmentation and support inference from rgbd images,” inECCV, pp. 746–760, 2012. 2, 3
2012
-
[24]
Height aware understanding of remote sensing images based on cross-task interaction,
Y . Feng, X. Sun, W. Diao, J. Li, R. Niu, X. Gao, and K. Fu, “Height aware understanding of remote sensing images based on cross-task interaction,”ISPRS J. Photogramm. Remote Sens., vol. 195, pp. 233– 249, 2023. 2, 3
2023
-
[25]
Monocular depth estimation with improved long-range accuracy for uav environment perception,
V .-C. Miclea and S. Nedevschi, “Monocular depth estimation with improved long-range accuracy for uav environment perception,”IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–15, 2021. 2, 3
2021
-
[26]
A novel recurrent encoder-decoder structure for large- scale multi-view stereo reconstruction from an open aerial dataset,
J. Liu and S. Ji, “A novel recurrent encoder-decoder structure for large- scale multi-view stereo reconstruction from an open aerial dataset,” in CVPR, pp. 6050–6059, 2020. 2, 3
2020
-
[27]
Syndrone-multi- modal uav dataset for urban scenarios,
G. Rizzoli, F. Barbato, M. Caligiuri, and P. Zanuttigh, “Syndrone-multi- modal uav dataset for urban scenarios,” inICCV, pp. 2210–2220, 2023. 2, 3, 10
2023
-
[28]
Depth map prediction from a single image using a multi-scale deep network,
D. Eigen, C. Puhrsch, and R. Fergus, “Depth map prediction from a single image using a multi-scale deep network,”NeurIPS, vol. 27, 2014. 3, 9
2014
-
[29]
Cliffnet for monoc- ular depth estimation with hierarchical embedding loss,
L. Wang, J. Zhang, Y . Wang, H. Lu, and X. Ruan, “Cliffnet for monoc- ular depth estimation with hierarchical embedding loss,” inECCV, pp. 316–331, Springer, 2020. 3
2020
-
[30]
Vision transformers for dense prediction,
R. Ranftl, A. Bochkovskiy, and V . Koltun, “Vision transformers for dense prediction,” inICCV, pp. 12179–12188, 2021. 3, 12
2021
-
[31]
Self-supervised monocular trained depth estimation using self-attention and discrete disparity volume,
A. Johnston and G. Carneiro, “Self-supervised monocular trained depth estimation using self-attention and discrete disparity volume,” inCVPR, pp. 4756–4765, 2020. 3
2020
-
[32]
Depth anything v2,
L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything v2,”NeurIPS, vol. 37, pp. 21875–21911, 2024. 3
2024
-
[33]
Unidepthv2: Universal monocular metric depth estimation made simpler,
L. Piccinelli, C. Sakaridis, Y .-H. Yang, M. Segu, S. Li, W. Abbeloos, and L. Van Gool, “Unidepthv2: Universal monocular metric depth estimation made simpler,”arXiv preprint arXiv:2502.20110, 2025. 3
Pith/arXiv arXiv 2025
-
[34]
Vggt: Visual geometry grounded transformer,
J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny, “Vggt: Visual geometry grounded transformer,” inCVPR, pp. 5294– 5306, 2025. 3
2025
-
[35]
A hierarchical deformable deep neural network and an aerial image benchmark dataset for surface multiview stereo reconstruction,
J. Li, X. Huang, Y . Feng, Z. Ji, S. Zhang, and D. Wen, “A hierarchical deformable deep neural network and an aerial image benchmark dataset for surface multiview stereo reconstruction,”IEEE Trans. Geosci. Re- mote Sens., vol. 61, pp. 1–12, 2023. 3
2023
-
[36]
Self-supervised monocular depth estimation from oblique uav videos,
L. Madhuanand, F. Nex, and M. Y . Yang, “Self-supervised monocular depth estimation from oblique uav videos,”ISPRS J. Photogramm. Remote Sens., vol. 176, pp. 1–14, 2021. 3
2021
-
[37]
Pix2pix-based monocular depth estimation for drones with optical flow on airsim,
T. Shimada, H. Nishikawa, X. Kong, and H. Tomiyama, “Pix2pix-based monocular depth estimation for drones with optical flow on airsim,” Sensors, vol. 22, no. 6, p. 2097, 2022. 3
2097
-
[38]
Scene-aware refinement network for un- supervised monocular depth estimation in ultra-low altitude oblique photography of uav,
K. Yu, H. Li, L. Xing, T. Wen, D. Fu, Y . Yang, C. Zhou, R. Chang, S. Zhao, L. Xing,et al., “Scene-aware refinement network for un- supervised monocular depth estimation in ultra-low altitude oblique photography of uav,”ISPRS J. Photogramm. Remote Sens., vol. 205, pp. 284–300, 2023. 3
2023
-
[39]
Posenet: A convolutional network for real-time 6-dof camera relocalization,
A. Kendall, M. c. Grimes, and R. Cipolla, “Posenet: A convolutional network for real-time 6-dof camera relocalization,” inICCV, pp. 2938– 2946, 2015. 3, 4
2015
-
[40]
Visual camera re-localization from rgb and rgb-d images using dsac,
E. Brachmann and C. Rother, “Visual camera re-localization from rgb and rgb-d images using dsac,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 9, pp. 5847–5865, 2021. 3
2021
-
[41]
Hierarchical scene coordinate classification and regression for visual localization,
X. Li, S. Wang, Y . Zhao, J. Verbeek, and J. Kannala, “Hierarchical scene coordinate classification and regression for visual localization,” in CVPR, pp. 11983–11992, 2020. 3
2020
-
[42]
Mid-air: A multi-modal dataset for extremely low altitude drone flights,
M. Fonder and M. Van Droogenbroeck, “Mid-air: A multi-modal dataset for extremely low altitude drone flights,” inCVPRW, pp. 0–0, 2019. 3, 10
2019
-
[43]
Wilduav: Monocular uav dataset for depth estimation tasks,
H. Florea, V .-C. Miclea, and S. Nedevschi, “Wilduav: Monocular uav dataset for depth estimation tasks,” inICCP, pp. 291–298, IEEE, 2021. 3
2021
-
[44]
Uemm-air: Make unmanned aerial vehicles perform more multi-modal tasks,
L. Yao, F. Liu, S. Xu, C. Zhang, X. Ma, J. Jiang, Z. Wang, S. Di, and J. Zhou, “Uemm-air: Make unmanned aerial vehicles perform more multi-modal tasks,”arXiv preprint arXiv:2406.06230, 2024. 3
Pith/arXiv arXiv 2024
-
[45]
Ddos: the drone depth and obstacle segmentation dataset,
B. Kolbeinsson and K. Mikolajczyk, “Ddos: the drone depth and obstacle segmentation dataset,” inCVPR, pp. 7328–7337, 2024. 3, 10
2024
-
[46]
Undemon: Unsu- pervised deep network for depth and ego-motion estimation,
V . M. Babu, K. Das, A. Majumdar, and S. Kumar, “Undemon: Unsu- pervised deep network for depth and ego-motion estimation,” inIROS, pp. 1082–1088, 2018. 3 JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 18
2018
-
[47]
Unsupervised learning of depth, optical flow and pose with occlusion from 3d geometry,
G. Wang, C. Zhang, H. Wang, J. Wang, Y . Wang, and X. Wang, “Unsupervised learning of depth, optical flow and pose with occlusion from 3d geometry,”IEEE Trans. Intell. Transp. Syst., vol. 23, no. 1, pp. 308–320, 2020. 3
2020
-
[48]
Unsupervised monocular depth and camera pose estimation with multiple masks and geometric consistency constraints,
X. Zhang, B. Zhao, J. Yao, and G. Wu, “Unsupervised monocular depth and camera pose estimation with multiple masks and geometric consistency constraints,”Sensors, vol. 23, no. 11, p. 5329, 2023. 3
2023
-
[49]
Joint unsupervised learning of depth, pose, ground normal vector and ground segmentation by a monocular camera sensor,
L. Xiong, Y . Wen, Y . Huang, J. Zhao, and W. Tian, “Joint unsupervised learning of depth, pose, ground normal vector and ground segmentation by a monocular camera sensor,”Sensors, vol. 20, no. 13, p. 3737, 2020. 3
2020
-
[50]
Self- supervised monocular depth and ego-motion estimation in endoscopy: Appearance flow to the rescue,
S. Shao, Z. Pei, W. Chen, W. Zhu, X. Wu, D. Sun, and B. Zhang, “Self- supervised monocular depth and ego-motion estimation in endoscopy: Appearance flow to the rescue,”Med. Image Anal., vol. 77, p. 102338,
-
[51]
En- hanced self-supervised monocular depth estimation with self-attention and joint depth-pose loss for laparoscopic images,
W. Li, Y . Hayashi, M. Oda, T. Kitasaka, K. Misawa, and K. Mori, “En- hanced self-supervised monocular depth estimation with self-attention and joint depth-pose loss for laparoscopic images,”Int. J. Comput. Assist. Radiol. Surg., pp. 1–11, 2025. 3
2025
-
[52]
Relative pose from deep learned depth and a single affine correspondence,
I. Eichhardt and D. Barath, “Relative pose from deep learned depth and a single affine correspondence,” inECCV, pp. 627–644, 2020. 3
2020
-
[53]
Fixing the scale and shift in monocular depth for camera pose estimation,
Y . Ding, V . V´avra, V . Kocur, J. Yang, T. Sattler, and Z. Kukelova, “Fixing the scale and shift in monocular depth for camera pose estimation,”arXiv preprint arXiv:2501.12345, 2025. 3
Pith/arXiv arXiv 2025
-
[54]
A self-supervised monocular depth estimation approach based on uav aerial images,
Y . Zhang, Q. Yu, K. H. Low, and C. Lv, “A self-supervised monocular depth estimation approach based on uav aerial images,” inDASC, pp. 1– 8, 2022. 3
2022
-
[55]
Self-supervised monocular depth estimation using global and local mixed multi-scale feature enhancement network for low-altitude uav remote sensing,
R. Chang, K. Yu, and Y . Yang, “Self-supervised monocular depth estimation using global and local mixed multi-scale feature enhancement network for low-altitude uav remote sensing,”Remote Sens., vol. 15, no. 13, p. 3275, 2023. 3
2023
-
[56]
Fast and high- quality monocular depth estimation with optical flow for autonomous drones,
T. Shimada, H. Nishikawa, X. Kong, and H. Tomiyama, “Fast and high- quality monocular depth estimation with optical flow for autonomous drones,”Drones, vol. 7, no. 2, p. 134, 2023. 3
2023
-
[57]
Tandepth: Leveraging global dems for metric monocular depth estimation in uavs,
H. Florea and S. Nedevschi, “Tandepth: Leveraging global dems for metric monocular depth estimation in uavs,”IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens., 2025. 3
2025
-
[58]
Skyscenes: A synthetic dataset for aerial scene understanding,
S. Khose, A. Pal, A. Agarwal, Deepanshi, J. Hoffman, and P. Chattopad- hyay, “Skyscenes: A synthetic dataset for aerial scene understanding,” inECCV, pp. 19–35, 2024. 10
2024
-
[59]
Uemm-air: Make unmanned aerial vehicles perform more multi-modal tasks,
L. Yao, F. Liu, S. Xu, C. Zhang, X. Ma, J. Jiang, Z. Wang, S. Di, and J. Zhou, “Uemm-air: Make unmanned aerial vehicles perform more multi-modal tasks,” 2025. 10
2025
-
[60]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” inCVPR, pp. 770–778, 2016. 11
2016
-
[61]
From big to small: Multi-scale local planar guidance for monocular depth estimation,
J. H. Lee, M.-K. Han, D. W. Ko, and I. H. Suh, “From big to small: Multi-scale local planar guidance for monocular depth estimation,”arXiv preprint arXiv:1907.10326, 2019. 12
Pith/arXiv arXiv 1907
-
[62]
Depthformer: Exploiting long- range correlation and local information for accurate monocular depth estimation,
Z. Li, Z. Chen, X. Liu, and J. Jiang, “Depthformer: Exploiting long- range correlation and local information for accurate monocular depth estimation,”Mach. Intell. Res., vol. 20, no. 6, pp. 837–854, 2023. 12
2023
-
[63]
Neural window fully- connected crfs for monocular depth estimation,
W. Yuan, X. Gu, Z. Dai, S. Zhu, and P. Tan, “Neural window fully- connected crfs for monocular depth estimation,” inCVPR, pp. 3916– 3925, 2022. 12
2022
-
[64]
Simipu: Simple 2d image and 3d point cloud unsupervised pre-training for spatial-aware visual representations,
Z. Li, Z. Chen, A. Li, L. Fang, Q. Jiang, X. Liu, J. Jiang, B. Zhou, and H. Zhao, “Simipu: Simple 2d image and 3d point cloud unsupervised pre-training for spatial-aware visual representations,” inAAAI, vol. 36, pp. 1500–1508, 2022. 12
2022
-
[65]
Monocular depth estimation toolbox
Z. Li, “Monocular depth estimation toolbox.” https://github.com/ zhyever/Monocular-Depth-Estimation-Toolbox, 2022. 12
2022
-
[66]
On regression losses for deep depth estimation,
M. Carvalho, B. Le Saux, P. Trouv ´e-Peloux, A. Almansa, and F. Cham- pagnat, “On regression losses for deep depth estimation,” inICIP, pp. 2915–2919, 2018. 14
2018
-
[67]
Multi-path fusion network for high-resolution height estimation from a single orthophoto,
Y . Zhang and X. Chen, “Multi-path fusion network for high-resolution height estimation from a single orthophoto,” inICMEW, pp. 186–191,
-
[68]
Height estimation from single aerial images using a deep convolutional encoder-decoder network,
H. A. Amirkolaee and H. Arefi, “Height estimation from single aerial images using a deep convolutional encoder-decoder network,”ISPRS J. Photogramm. Remote Sens., vol. 149, pp. 50–66, 2019. 14
2019
-
[69]
Boundary-aware multitask learning for remote sensing imagery,
Y . Wang, W. Ding, R. Zhang, and H. Li, “Boundary-aware multitask learning for remote sensing imagery,”IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens., vol. 14, pp. 951–963, 2020. 14
2020
-
[70]
Inverted pyramid multi-task transformer for dense scene understanding,
H. Ye and D. Xu, “Inverted pyramid multi-task transformer for dense scene understanding,” inECCV, pp. 514–530, 2022. 14
2022
-
[71]
Taskexpert: Dynamically assembling multi-task representations with memorial mixture-of-experts,
H. Ye and D. Xu, “Taskexpert: Dynamically assembling multi-task representations with memorial mixture-of-experts,” inICCV, pp. 21828– 21837, 2023. 14
2023
-
[72]
Knowledge- guided multi-task network for remote sensing imagery,
M. Li, G. Wang, T. Li, Y . Yang, W. Li, X. Liu, and Y . Liu, “Knowledge- guided multi-task network for remote sensing imagery,”Remote Sensing, vol. 17, no. 3, p. 496, 2025. 14
2025
-
[73]
Deepcalib: A deep learning approach for automatic intrinsic calibration of wide field- of-view cameras,
O. Bogdan, V . Eckstein, F. Rameau, and J.-C. Bazin, “Deepcalib: A deep learning approach for automatic intrinsic calibration of wide field- of-view cameras,” inCVMP, pp. 1–10, 2018. 14, 15
2018
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.