{"id":"d496a8e7-1c3e-419a-b056-d3eb7417bbe6","arxiv_id":"2504.17787","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The fourth MDEC winner, HRI, achieved a 3D F-score of 23.05% on SYNS-Patches, a small improvement over the previous best of 22.58%, under a new two-degree-of-freedom alignment protocol.","lead":"This paper reports the results of the fourth Monocular Depth Estimation Challenge, where teams submitted models that estimate depth from single images and were tested on a diverse benchmark called SYNS-Patches. The winning team raised the top 3D reconstruction score from 22.58% to 23.05%, and the organizers changed the evaluation to allow affine-invariant depth predictions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 0.47-point F-Score gain over PICO-MR is only shown under the new 2-DOF LSE protocol and carries no uncertainty estimate; the headline comparison may be a protocol artifact.","rationale":"The reader's CONDITIONAL verdict is appropriate. The reader's weakest assumption (LSE alignment validity) is indeed the central risk, but I would sharpen it: the more actionable weakness is the absence of any sensitivity analysis connecting the headline gain to the alignment protocol and to sampling noise. The paper is internally consistent and does not overclaim -- it states that PICO-MR was re-evaluated under the new protocol and that most leading methods are affine-invariant. However, because the main quantitative result is a 0.47-point gap with no error bars, and because the 2-DOF LSE protocol is precisely what affine-invariant methods are designed to exploit, the claim 'current best zero-shot result' should be conditioned on that protocol. A scale-only re-evaluation or a bootstrap over images would settle whether the ranking is stable. This does not change the reader's CONDITIONAL verdict; it identifies the specific experiment that would upgrade or overturn it.","tokens_in":20658,"tokens_out":5921,"duration_ms":59459,"concrete_test":"Recompute the leaderboard using the same submitted predictions but with 1-DOF scale-only least-squares alignment instead of the 2-DOF scale-and-shift alignment, and bootstrap 95% confidence intervals over per-image F-Scores for HRI and PICO-MR on the test split. If HRI's F-Score no longer exceeds PICO-MR under scale-only alignment, or if the bootstrap intervals overlap under the original LSE protocol, the headline 22.58-to-23.05 improvement is a protocol or noise artifact rather than an established gain.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that HRI improved the 3D F-Score from 22.58% to 23.05% (abstract, Table 1). This claim stands only if (i) the two numbers are evaluated under exactly the same protocol and (ii) the 2-DOF least-squares alignment used in Section 3 is a faithful way to compare affine-invariant, disparity, and metric predictions. Condition (i) is stated but not demonstrated with per-image results; PICO-MR is re-evaluated under the new protocol, but the paper does not report how much the re-evaluation changed its score or whether the same test predictions and upsampling were used. Condition (ii) is the more serious issue: HRI is an affine-invariant predictor (marked 'A' in Table 1), so it receives an oracle per-image scale-and-shift correction, while PICO-MR (not marked 'A') is also allowed a shift that a metric or disparity model should not need. Under a scale-only alignment -- the natural protocol for metric/disparity predictions and closer to the median scaling used in earlier editions -- the ordering could change. The margin is also small: 0.47 F-Score points with no confidence intervals, and Table 1 contains tied/duplicate rows (ranks 11-12, 15-16), suggesting per-image variability is not negligible. The headline 'improvement' is therefore not yet robust to the choice of alignment and to sampling noise.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports the results of the fourth Monocular Depth Estimation Challenge (MDEC), held with CVPR 2025, on the SYNS-Patches benchmark. Relative to previous editions, the evaluation protocol was changed from median scaling to a two-degree-of-freedom least-squares (LSE) alignment in order to admit metric, disparity, scale-invariant, and affine-invariant predictions, and the baselines were updated to Depth Anything v2 and Marigold. Twenty-four submissions are reported as outperforming the baselines on at least one metric, and ten teams additionally provided written method reports. The headline result is that team HRI achieves a reconstruction F-score of 23.05%, which the abstract presents as an improvement over the previous edition's best result of 22.58% (PICO-MR, re-evaluated under the new protocol). The paper contains a full results table, qualitative comparisons, and short technical descriptions of each reported method.","tokens_in":20903,"tokens_out":12418,"duration_ms":105630,"significance":"If the results hold, the paper provides the community with (i) a current snapshot of zero-shot monocular depth estimation on SYNS-Patches, (ii) an evaluation protocol aligned with current affine-invariant evaluation practice, and (iii) evidence that fine-tuned foundation models with affine-invariant output dominate the leaderboard. The organizational effort is substantial and worth acknowledging: 24 submissions, 10 method reports, updated and publicly described baselines, a public starter-pack repository, and a re-evaluation of the previous winner under the new protocol. The leaderboard is an external empirical measurement rather than a derived claim, and the flagged duplicate rows make the table more honest than many challenge reports. The main weakness is that the headline improvement is small (0.47 F-score points) and is demonstrated only under a single alignment protocol with no uncertainty estimates, so the paper's central numerical claim is less robust than its framing suggests.","major_comments":[{"comment":"There is a direct inconsistency between the two descriptions of the evaluation protocol. Section 3 states that all submissions were aligned using least squares with two degrees of freedom (scale and shift), whereas Section 5 states that predictions were aligned 'according to median depth scaling or least-squares alignment, as requested by the participants.' If participants could select their alignment, Table 1 is not computed under a single protocol and the headline improvement from 22.58% to 23.05% is not well-defined; if the Section 5 sentence is carried over from the previous editions' protocol, it must be removed or corrected.","section":"§5 vs. §3 (Evaluation)"},{"comment":"The claimed improvement from 22.58% to 23.05% requires the two numbers to be evaluated under the same protocol, but the manuscript is ambiguous about which protocol produced the 22.58% figure: the abstract calls it 'the previous edition's best result' (third edition, median scaling), while Table 1 reports PICO-MR at 22.58 under the new LSE protocol. The authors should state explicitly whether these two 22.58% values coincide exactly, by rounding, or only coincidentally, and should report PICO-MR's score under both the old and the new protocols so that the reader can see how much the re-evaluation changed its result.","section":"Abstract and §5.1, Table 1"},{"comment":"The rank ordering between HRI (affine-invariant, marked 'A' in Table 1) and PICO-MR (not marked affine-invariant) depends on the choice of a two-degree-of-freedom LSE alignment, because affine-invariant methods receive an oracle per-image scale-and-shift correction while metric or disparity predictions are thereby also allowed a free shift that they should not need. Since the margin is only 0.47 F-score points, the authors should either provide a sensitivity analysis of the top entries under scale-only alignment and under the previous median-scaling protocol, or argue explicitly why the 2-DOF protocol is the only principled way to compare all three prediction types; as it stands, the paper does not rule out a protocol artifact.","section":"§3 Evaluation; Table 1"},{"comment":"The headline gap of 0.47 F-score points is presented without any uncertainty estimate, and Table 1 itself contains tied or near-identical rows (ranks 11–12, 13, and 15–16), indicating that per-image variability is not negligible. The authors should add bootstrapped confidence intervals or per-image standard deviations for at least the top three entries, and should temper the precision of the 'raising it from 22.58% to 23.05%' claim accordingly.","section":"§5.1, Table 1"}],"minor_comments":[{"comment":"The description of the new alignment procedure is too terse to reproduce; please specify whether the LSE fit is computed per image on depth or inverse-depth values, whether it is weighted, and whether upsampling precedes or follows the fit, and point to the exact evaluation script in the starter pack.","section":"§3, Evaluation"},{"comment":"The paper reports only submissions that beat the baselines on at least one metric; for transparency it should state the total number of test-phase submissions received and how many were excluded, so that the reported '24 submissions' is interpretable.","section":"§3 and §5"},{"comment":"The flagged near-duplicate rows (ranks 11–13 and 15–16, the latter identical to Marigold) are retained in the count; the text says '24 teams,' but the number of distinct methods is smaller and should be stated explicitly.","section":"Table 1"},{"comment":"Minor editorial issues include 'alongside to' (Section 1), inconsistent spelling of 'DepthAnything' versus 'Depth Anything', broken superscripts in the delta-metric column headers of Table 1, and the space in the URL 'toshas/mdec benchmark' (footnote 2).","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"Nothing in the paper suggests fabrication; the concern is entirely about protocol robustness and the precision of the headline claim. The dual role of the organizing team as authors of previous challenge editions and of one of the baselines (Marigold) is normal for this genre and is disclosed by authorship and citation; I would only ask that the final version disclose the submission-selection criterion and the number of excluded submissions. The paper fits the workshop-proceedings venue well, and the requested changes are additive rather than structural."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things up front. First, this is a genuine benchmark update: the fourth MDEC introduces a 2-DOF least-squares alignment protocol, re-evaluates PICO-MR under that same protocol, adds Depth Anything v2 and Marigold as baselines, and reports a clean per-team metric table. Second, the headline improvement is exactly as small as it looks: 23.05 vs 22.58 F-score, a 0.47-point gain with no error bars, no significance test, and a protocol that changed relative to the third edition. The paper itself admits the gain is marginal, which I appreciate.\n\nWhat is actually new: the LSE alignment protocol is a reasonable response to the field moving toward affine-invariant and disparity predictions, and the authors are transparent about how it works. The paper does not oversell HRI's win; it notes most leading methods are fine-tuned foundation models and that the benchmark is approaching saturation. The per-team write-ups are concise and useful, and the comparison to PICO-MR is internally consistent because PICO-MR was re-evaluated under the new protocol.\n\nSoft spots, in proportion. The biggest one is the protocol change: the abstract frames 23.05 as surpassing the previous edition's best, but the previous number was produced under median scaling. The re-evaluated PICO-MR still shows 22.58, which is lucky (or a coincidence), but the paper never reports how much the re-evaluation changed the score or whether identical test predictions and upsampling were used. That leaves the door open to the stress-test concern: HRI is affine-invariant and benefits from a per-image scale-and-shift correction, while a metric or disparity method should not need the shift. Under a scale-only alignment the ordering could change. The margin is too small to shrug off. Second, Table 1 contains duplicate rows (ranks 11-12 and 15-16) that are flagged but not removed; that does not affect the top of the leaderboard, but it makes the table look sloppy and the per-image variability is clearly not negligible. Third, there is no uncertainty estimate anywhere, which is common for challenge reports but matters here because the headline claim is a 0.47-point difference.\n\nMy read: the paper deserves peer review as a workshop challenge report. The numbers appear honest, the protocol is stated, and the authors are appropriately cautious. I would send it out with a request for error bars or a per-image variance estimate, and for the authors to state explicitly what the re-evaluated PICO-MR score was under median scaling versus LSE. Without that, the headline should be read as provisional.","headline":"A competent challenge report that is honest about marginal gains; the headline 0.47-point F-score improvement is real under the stated protocol but fragile because the protocol changed and no uncertainty is reported.","tokens_in":823,"tokens_out":877,"would_cite":false,"duration_ms":26722,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The fourth Monocular Depth Estimation Challenge reports a new best zero-shot result on SYNS-Patches: a 23.05% reconstruction F-Score, up from 22.58% in the previous edition, under a revised least-squares alignment protocol that puts…","keywords":["monocular depth estimation","zero-shot generalization","SYNS-Patches benchmark","least-squares alignment","affine-invariant depth","reconstruction F-Score","depth foundation models","benchmark saturation"],"falsifier":"Recompute the Table 1 ranking after fixing each prediction to a single global scale learned from held-out images rather than per-image least-squares scale and shift; if the winning method changes or its F-Score drops below the previous edition's result, the reported improvement is an artifact of the alignment protocol.","tokens_in":20471,"feed_emoji":"📐","tokens_out":4819,"duration_ms":43453,"temperature":0.7,"pith_summary":"The paper is the report of the fourth Monocular Depth Estimation Challenge, a zero-shot generalization benchmark on the SYNS-Patches dataset. It claims that after switching the evaluation to least-squares scale-and-shift alignment, the winning method improved the reconstruction F-Score from 22.58% to 23.05%, the current best zero-shot result on this benchmark. It also documents that most competitive submissions rely on affine-invariant predictions built on top of public depth foundation models, and that the benchmark is close to saturation.","feed_headline":"Depth-model winner beats prior zero-shot best by 0.47 points","feed_subtitle":"A scale-and-shift alignment rule lets affine-invariant models compete; the field is nearing its ceiling on SYNS-Patches.","key_machinery":"The central object is the evaluation protocol: each prediction is bilinearly upsampled, disparity maps are inverted, and predictions are aligned to ground truth by a least-squares fit of an affine scale and shift before computing the pointcloud-based reconstruction F-Score used for ranking. This two-degree-of-freedom alignment is what makes it possible to compare metric, disparity, and affine-invariant predictions on equal footing on the SYNS-Patches benchmark.","core_discovery":"On its own terms, the paper establishes a new best zero-shot monocular depth result on SYNS-Patches: the top submission reaches a reconstruction F-Score of 23.05%, ahead of the previous edition's winner re-evaluated at 22.58% under the same least-squares alignment. The paper also shows that the leading approaches almost all predict affine-invariant depth and are fine-tuned from public foundation models, and that only one reporting team outperformed the previous champion. It concludes that despite sharper qualitative predictions, the benchmark is nearing saturation, with web-scale foundation-model priors now dominating.","pith_inferences":["An implication the organizers leave implicit: because the two-degree-of-freedom alignment absorbs a global scale and shift per image, a model that predicts only relative order can score as well as a metric model on the F-Score, so a metric-only track would likely change which method wins.","The gap between first and second place is 0.10 F-Score points, smaller than the cross-edition gain of 0.47, so a testable extension would be to report confidence intervals or per-scene variance before declaring a new state of the art.","The near-duplicate rows among anonymous entries suggest participants can re-submit minor variants; a future edition could detect and merge such duplicates before ranking."],"forward_implications":["Under this protocol, affine-invariant predictions are not a handicap: most top-ranked methods chose them, and the winner is one of them.","Zero-shot performance on SYNS-Patches appears close to saturation: the F-Score gain over the previous edition is 0.47 points, while qualitative sharpness improved more.","Web-scale foundation models such as Depth Anything v2 and Marigold are now the effective starting point for competitive depth estimation; custom architectures that avoid them only beat baselines on a single metric.","Further progress will likely require harder settings such as non-Lambertian surfaces, metric depth recovery, and alternative 3D representations rather than more data or larger foundation models."],"supporting_citations":[{"why":"Defines the SYNS-Patches benchmark images and dense LiDAR ground truth that all submissions are evaluated on.","marker":"[1, 96]"},{"why":"Supplies the third-edition winning result re-evaluated under the new alignment protocol, the baseline the winner must beat.","marker":"[99]"},{"why":"Marigold is one of the two new off-the-shelf baselines, the diffusion-based method most teams need to outperform.","marker":"[49]"},{"why":"Depth Anything v2 is the other new baseline and the foundation model most top teams fine-tuned.","marker":"[119]"},{"why":"Supplies the pointcloud-based reconstruction F-Score used as the leaderboard ranking metric.","marker":"[67]"},{"why":"Metric3D v2 provides the pre-trained ViT-G encoder weights that the winning submission initialized from.","marker":"[41]"},{"why":"One of the references behind the least-squares alignment standard adopted in this edition's evaluation.","marker":"[81, 82]"}],"fun_headline_variants":["Depth challenge winner tops prior zero-shot F-Score by 0.47","Zero-shot depth best rises to 23.05% F-Score in challenge","Monocular depth benchmark: new winner 23.05%, field near ceiling","Fourth MDEC winner: 23.05% F-Score, up 0.47 from prior"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire ranking assumes that least-squares scale-and-shift alignment is a fair common evaluation for metric, disparity, and affine-invariant depth predictions; if that alignment hides true metric errors, the F-Score comparison is an artifact of the protocol.","fun_headline_variants_meta":{"raw":{"variants":["Depth challenge winner tops prior zero-shot F-Score by 0.47","Zero-shot depth best rises to 23.05% F-Score in challenge","Monocular depth benchmark: new winner 23.05%, field near ceiling","Fourth MDEC winner: 23.05% F-Score, up 0.47 from prior"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000145,"raw_usage":{"total_tokens":1110,"prompt_tokens":809,"completion_tokens":301,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":425,"completion_tokens_details":{"reasoning_tokens":211}},"tokens_in":425,"tokens_out":301,"duration_ms":3335,"temperature":1.0,"reasoning_tokens":211,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:30:59.748198+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the Table 1 ranking after fixing each prediction to a single global scale learned from held-out images rather than per-image least-squares scale and shift; if the winning method changes or its F-Score drops below the previous edition's result, the reported improvement is an artifact of the alignment protocol.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the third-edition winning result re-evaluated under the new alignment protocol, the baseline the winner must beat."},{"cited_title":"Depth anything v2","cited_arxiv_id":null,"evidence_quote":"Depth Anything v2 is the other new baseline and the foundation model most top teams fine-tuned."}],"review_version":1}