{"id":"bce8d2a7-8910-4d09-9ca3-7112d7574e66","arxiv_id":"2501.11841","paper_version":4,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of monocular metric depth estimation methods, datasets, and open challenges, with no new experimental results.","lead":"This preprint surveys monocular metric depth estimation, a computer vision task that predicts absolute distance from a single image. A generalist would read it for a compact map of current models, datasets, and open problems, though it contains no new experiments.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Tables 2 and 3 contain internal attribution and classification errors, so the survey's reliability claim is unsupported unless the data are corrected.","rationale":"The reader's weakest assumption—that the secondary data in Tables 2 and 3 are accurate—is exactly where the load-bearing concern lies. My stress-test found internal evidence supporting this concern rather than merely absence of external verification: the ZeroDepth row's citation contradicts Section 5, and Table 3 labels for A2D2, Diode, and Objaverse are demonstrably inconsistent with the datasets' own definitions. These are not matters of interpretation; they are factual errors in the survey's central comparative machinery. Since the paper's contribution is to provide a reliable structured reference, these errors directly undermine that contribution as written. The errors appear correctable, so a conditional verdict is appropriate: the survey would be acceptable once the tables are fixed and the numerical rows are verified against their cited sources. I therefore agree with the reader's identification of the weakest assumption, but go further by showing that the assumption is not merely unverified—it is violated by the preprint's own internal inconsistencies. The 'critical gap' claim is also asserted without a systematic comparison to the existing surveys cited in Section 1, but that is secondary to the data-integrity issue because a novel positioning cannot rescue an inaccurate reference work.","tokens_in":13326,"tokens_out":5030,"duration_ms":51679,"concrete_test":"Verify each numerical row of Table 2 against the original DepthPro paper (Bochkovskii et al., 2024, arXiv:2410.02073); specifically check whether the ZeroDepth row corresponds to Guizilini et al. 2023 and whether its values match. For Table 3, check the A2D2, Diode, and Objaverse entries against their official documentation and dataset releases to confirm scene type and supervision labels.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The survey's central claim to be a reliable structured reference depends on the accuracy of its two comparative tables, and that premise is already violated internally. In Table 2, the row 'ZeroDepth' is attributed to Bhat et al. 2023, but Section 5 identifies ZeroDepth with Guizilini et al. 2023, while Bhat et al. 2023 is ZoeDepth, which also appears separately in the same table. This conflation means the numerical row may belong to a different model, so the comparison is not trustworthy without checking the source. Table 3 contains similar classification errors: A2D2 is listed as both indoor and outdoor although it is an outdoor autonomous driving dataset; Diode is listed as indoor-only although it contains outdoor scenes; and Objaverse is labeled as providing relative depth despite being a repository of 3D object models without image-depth pairs. Because the 'comprehensive review' argument rests on these tables, the paper cannot currently support its headline contribution; any comparative conclusions drawn from Tables 2 and 3 are unreliable until the entries are verified and corrected.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews monocular metric depth estimation (MMDE), covering the task formulation, classical and deep-learning methods, zero-shot relative-depth precursors such as MiDaS, recent metric-depth architectures (ZoeDepth, Metric3D, UniDepth, Depth Anything, Depth Pro, and others), loss and training strategies, sensor-assisted approaches, and a large dataset overview. The paper argues that existing surveys are either outdated or domain-specific, and that it fills a critical gap by providing a comprehensive, structured reference for MMDE, with comparative performance tables and a dataset taxonomy as central evidence.","tokens_in":13374,"tokens_out":3676,"duration_ms":39014,"significance":"If its comparative data were reliable, the survey would be a useful entry point to the MMDE literature: it is broadly scoped, recent, and organized around practically relevant axes such as zero-shot generalization, boundary preservation, patch-based inference, and generative modeling. The paper also deserves credit for explicitly noting when important claims rest on unavailable artifacts, as in the case of DMD, and for naming the absence of standardized benchmarks as a field-level problem. However, the central claim of being a reliable structured reference depends on the accuracy of the two large comparative tables, and both contain attribution and classification errors that currently undermine that claim. These issues are local and correctable, so the contribution can be salvaged with careful verification.","major_comments":[{"comment":"The row labeled 'ZeroDepth' is cited as (Bhat et al., 2023), but Section 5 and the reference list identify ZeroDepth with Guizilini et al. (2023), while Bhat et al. (2023) is ZoeDepth, which already appears as a separate row in the same table. This conflation means the numerical values in the ZeroDepth row may belong to a different model, so the cross-model comparison in Table 2 is not trustworthy as printed. The citation and the numbers need to be verified against the original ZeroDepth and ZoeDepth papers.","section":"Table 2 (Section 6.4)"},{"comment":"The table caption states that all results are 'reported by DepthPro (Bochkovskii et al., 2024)'. Relying on a single model paper as the sole source for every entry makes the comparison unverifiable and potentially biased, since DepthPro is itself one of the compared methods. The survey should either reproduce results from each model's original source, cite the benchmark leaderboards, or explicitly qualify the table as a secondary re-reporting with all attendant caveats. Without this, the quantitative claims in Section 6.4 cannot be independently checked.","section":"Table 2 caption / Section 6.4"},{"comment":"Several dataset classifications in Table 3 are incorrect or internally inconsistent. A2D2 is listed as 'Indoor Outdoor' and described as 'including both indoor and outdoor scenes', but A2D2 is an outdoor autonomous-driving dataset. Diode is listed with the cell 'Indoor Indoor' and described as indoor-only, although the Diode benchmark contains both indoor and outdoor scenes. Objaverse is labeled as providing relative depth supervision, but it is a repository of 3D object models without image–depth pairs, so it cannot support the depth-estimation roles claimed for it. These errors directly affect the survey's comprehensiveness claim and the reported count of metric-depth datasets, so every row of Table 3 should be checked against the primary dataset documentation.","section":"Table 3"}],"minor_comments":[{"comment":"The mathematical notation is malformed: 'D := (R)H×W' and 'I := (R)H×W ×3' should use superscripts, and 'pixel ii,j' should read 'pixel i_{i,j}'.","section":"Section 2, first paragraph"},{"comment":"MiDaS is attributed to (Birkl et al., 2023), but the MiDaS method originates with Ranftl et al.; the cited Birkl et al. paper is the MiDaS v3.1 model zoo. Please cite the original MiDaS paper alongside or instead of the model-zoo paper.","section":"Section 4"},{"comment":"The 'Indoor Outdoor' column for Diode contains the duplicated and contradictory entry 'Indoor Indoor'; this appears to be a formatting error that should be corrected to the dataset's actual scene composition.","section":"Table 3, Diode row"},{"comment":"The label 'MiDAS v3.1' should be 'MiDaS v3.1', and the caption would benefit from stating which hardware and settings were used for the timing and memory measurements.","section":"Figure 2"},{"comment":"The 'Output' column uses 'open' and 'close' to describe code release; 'close' should be 'closed' or 'closed-source' for clarity.","section":"Table 1"},{"comment":"The reference entry reads 'arxiv 2022. arXiv preprint arXiv:2203.01502'; the duplicated 'arxiv 2022' should be removed.","section":"Reference list, BinsFormer"}],"recommendation":"major_revision","confidential_remarks":"The survey has no self-cited prior work, so there is no circularity concern; the main risk is purely one of secondary-data accuracy. The errors in Tables 2 and 3 are concrete and verifiable, so a revision that corrects them and documents the provenance of every table entry would make the paper acceptable. I would also encourage the editor to ask the author to state the survey's inclusion criteria and search strategy, since the current 'comprehensive review' claim is otherwise hard to assess."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a straightforward survey of monocular metric depth estimation (MMDE). It organizes the recent zero-shot methods, datasets, and challenges into a clear narrative, and for someone new to the area it does give a reasonable map of the main threads: MiDaS to ZoeDepth, Depth Anything, Metric3D, patch-based and generative approaches, plus the standard indoor/outdoor/synthetic dataset split. The writing is plain and the structure is logical. I would not call it a groundbreaking contribution, but it is a competent entry point.\n\nThe soft spots are not minor, though. The stress-test note is right: Tables 2 and 3, which carry the survey's weight, have real problems. Table 2 attributes the ZeroDepth row to Bhat et al. 2023, but Section 5 and the references identify ZeroDepth as Guizilini et al. 2023, while Bhat et al. 2023 is ZoeDepth, also in the same table. That conflation throws the row's numerical values into doubt. Table 3 mislabels A2D2 as both indoor and outdoor (it is an outdoor driving dataset), Diode as indoor-only (it includes outdoor scenes), and Objaverse as providing relative depth (it is a 3D object repository without image-depth pairs). These are not trivial; they are exactly the kind of errors that make readers distrust a survey's secondary data. Additionally, Table 2 appears to take all zero-shot numbers from a single model paper (DepthPro) without independent cross-checking, which makes the comparison fragile. The claim of filling a 'critical gap' is also asserted rather than demonstrated: the paper cites several 2022–2024 surveys but never systematically states how it differs from them.\n\nThe paper does have one honest caveat: it flags that DMD has no public release, which shows the author is not hiding uncertainty. But the central reliability claim is unsupported as written.\n\nWho is this for? A research group looking for a quick bibliography and a sketch of the area might skim it, but they should not trust the tables without going back to the primary sources. I would not cite it in its current form. A serious referee could push the author to verify every entry in Tables 2 and 3 and to situate the survey against prior ones; the topic is important enough that a corrected version would be worth publishing. So I would send it to review, but with a clear expectation of major revision.","headline":"A readable but unreliable survey: the anchoring tables have citation and classification errors that undercut its own comprehensiveness claim.","tokens_in":13952,"tokens_out":1879,"would_cite":false,"duration_ms":20816,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Monocular metric depth estimation has become a field of its own, and this survey organizes its datasets, models, and open problems into one reference.","keywords":["monocular depth estimation","metric depth","zero-shot generalization","depth datasets","adaptive binning","diffusion models","benchmark comparison","survey"],"falsifier":"Re-run the eight zero-shot benchmarks in Table 2 with the official checkpoints of Depth Anything, Depth Anything V2, Metric3D, Metric3D v2, PatchFusion, UniDepth, ZeroDepth, ZoeDepth, and Depth Pro using the original evaluation protocols; if the reproduced numbers diverge substantially from Table 2, the survey's comparative evaluation is not trustworthy. As a lighter check, verify whether BDD100K and Mapillary Vistas provide metric ground truth; if they do, Table 3's 'Relative' labels are wrong and the claim that 32 of 38 datasets are metric needs revision.","tokens_in":13001,"feed_emoji":"📏","tokens_out":8136,"duration_ms":76891,"temperature":0.7,"pith_summary":"Monocular metric depth estimation (MMDE) is the task of predicting depth in physical units from a single image, and this survey claims the task has become a distinct research area with its own trajectory and its own open problems. The paper's central assertion is that progress in MMDE is driven by two things working together: datasets that supply true metric labels, and a sequence of methodological advances such as adaptive binning, camera-aware normalization, zero-shot transfer, patch-based inference, and generative refinement. It argues that earlier depth surveys missed this coalescing, either by predating the zero-shot metric era or by focusing on relative depth and specialized domains. If the survey is right, a reader can rely on it as a single entry point for choosing datasets, comparing models, and locating the field's unsolved problems.","feed_headline":"Metric depth from one image gets a full survey","feed_subtitle":"Datasets, zero-shot models, and open challenges, organized into one reference for absolute-scale depth.","key_machinery":"The central organizing object is the distinction between relative depth and metric depth, operationalized through scale-invariant versus scale-aware training and evaluation. The survey's load-bearing mechanisms are its comparison tables: Table 1's timeline of MMDE methods, Table 2's zero-shot benchmark numbers reproduced from Depth Pro (Bochkovskii et al., 2024) across eight datasets, and Table 3's 38-dataset taxonomy. These tables let the survey argue that dataset choice and metric-supervision type shape model behavior, and that no standardized protocol yet exists for fair cross-model comparison.","core_discovery":"On the paper's own terms, the discovery is a synthesis: the field has moved from ordinal, scale-invariant predictions to models that output metric depth, and the decisive ingredients are dataset scale and diversity together with architectural choices such as adaptive depth bins, camera-intrinsic normalization, patch-based fusion, and diffusion-based refinement. The evidence is organized in three tables: a timeline of key MMDE methods, a zero-shot benchmark comparison across eight datasets using numbers reported by Depth Pro (Bochkovskii et al., 2024), and a taxonomy of 38 datasets labeled by scene type, real or synthetic origin, modality, and whether the labels are metric or relative. The survey also establishes that zero-shot generalization is the leading open problem and that no standardized evaluation protocol yet exists for comparing models fairly.","pith_inferences":["A reader who wants to compare models should treat Table 2 as a transcription of Depth Pro's evaluation rather than an independent benchmark: the survey's own Section 6.4 says the absence of standardized protocols hinders fair comparison, so the numbers are best read as indicative.","The classification of BDD100K and Mapillary Vistas as relative-depth datasets in Table 3 is worth revisiting, because dataset releases change and that label determines the paper's count of 32 metric datasets.","A natural next step, not taken by the survey, would be a meta-analysis correlating dataset properties such as synthetic versus real origin, indoor versus outdoor scenes, and LiDAR density with the zero-shot scores in Table 2 to identify which dataset characteristics actually drive transfer.","The paper's own remark that DMD is closed-source suggests a reproducibility criterion for future surveys: separate released models from paper-only results when summarizing the state of the art."],"forward_implications":["A practitioner selecting a model can use the survey's comparison to choose between single-inference speed, patch-based detail, and generative fidelity, since the survey shows these are the three current trade-off regimes.","Dataset choice is load-bearing: with 32 of 38 datasets providing metric depth, supervised MMDE training has enough raw material, but the six relative-depth datasets still matter for scale-agnostic pretraining.","The survey's timeline implies that the next advance will likely combine large-scale metric supervision with an architecture that removes camera-intrinsic dependence, since that is the direction shared by ZoeDepth, UniDepth, and DAC.","Because generative models currently produce mostly relative depth, the survey leaves open the possibility that metric diffusion models will be the next frontier once inference cost and release of checkpoints are addressed."],"supporting_citations":[{"why":"Reports the eight-benchmark zero-shot numbers that Table 2 reproduces as the survey's main comparative evidence.","marker":"(Bochkovskii et al., 2024)"},{"why":"Anchors the timeline as the breakthrough that extended MiDaS-style relative depth to metric depth with adaptive binning and scene routing.","marker":"(Bhat et al., 2023)"},{"why":"Provides the camera-intrinsics-aware metric prediction approach that the survey credits as an early MMDE method.","marker":"(Yin et al., 2023)"},{"why":"Supplies the camera-intrinsics-free zero-shot metric model used in the survey's comparison and outlook.","marker":"(Piccinelli et al., 2024)"},{"why":"Supports the claim that large-scale semi-supervised data improves cross-domain generalization.","marker":"(Yang et al., 2024a)"},{"why":"The only reported generative model for metric depth; its closed source is the basis for the survey's open-challenge claim.","marker":"(Saxena et al., 2023)"},{"why":"Represents patch-based multi-resolution inference and the accuracy-latency trade-off the survey discusses.","marker":"(Li et al., 2024a)"},{"why":"Extends metric depth to fisheye and 360-degree cameras, supporting the generalization discussion.","marker":"(Guo et al., 2025)"},{"why":"The canonical outdoor driving dataset that anchors Table 3's metric-supervision claims.","marker":"(Geiger et al., 2013)"},{"why":"The canonical indoor benchmark used across Table 3 and the model comparisons.","marker":"(Silberman et al., 2012)"}],"fun_headline_variants":["Absolute depth from a single photo: the survey","Metric depth in one shot: a survey of methods","From pixels to meters: survey on metric depth","Zero-shot metric depth: survey and open challenges","One image, true scale: the metric depth survey"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the benchmark numbers in Table 2 and the dataset labels in Table 3 are accurate as reported, even though the survey did not independently verify them.","fun_headline_variants_meta":{"raw":{"variants":["Absolute depth from a single photo: the survey","Metric depth in one shot: a survey of methods","From pixels to meters: survey on metric depth","Zero-shot metric depth: survey and open challenges","One image, true scale: the metric depth survey"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1343,"prompt_tokens":905,"completion_tokens":438,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":365}},"tokens_in":521,"tokens_out":438,"duration_ms":4809,"temperature":1.0,"reasoning_tokens":365,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:47:29.849846+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the eight zero-shot benchmarks in Table 2 with the official checkpoints of Depth Anything, Depth Anything V2, Metric3D, Metric3D v2, PatchFusion, UniDepth, ZeroDepth, ZoeDepth, and Depth Pro using the original evaluation protocols; if the reproduced numbers diverge substantially from Table 2, the survey's comparative evaluation is not trustworthy. As a lighter check, verify whether BDD100K and Mapillary Vistas provide metric ground truth; if they do, Table 3's 'Relative' labels are wrong and the claim that 32 of 38 datasets are metric needs revision.","supporting_citations":[],"review_version":1}