{"id":"38a155b5-db0f-435f-98cb-9debbcbb8ecb","arxiv_id":"2607.04101","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"MRAC gates sparse anchors via Theil–Sen + MAD consistency with a frozen foundation's relative depth, repairing multipath outliers that collapse residual-on-CFA and blind VI-Depth while winning 84% of same-backbone cells.","lead":"A parameter-free inference wrapper called MRAC filters corrupted sparse depth anchors by checking consistency with a frozen monocular foundation's relative depth before calibration. It closes a multipath failure mode in deployed methods like VI-Depth and cuts KITTI multipath error 3.2\times with no retraining.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The strongest claim rests on two independent pillars: (1) same-backbone isolation of the selection mechanism (Tables I–III, 54/64 wins) and (2) the multipath blind-spot diagnosis of VI-Depth (Table II, Fig. 5, K=150 re-eval in Sec. IX). Both are tightly controlled; the backbone confound is confined to clean cells and is explicitly not claimed. The foundation-consistency premise is the natural soft spot, yet the paper already measures its failure modes (gate diagnostics, breakdown signature at 40% dropout, spatial-propagation Fig. S1) and shows that even imperfect gating (P/R 0.60/0.64 on the hardest cell) still yields the 3.2\times gap. No further assumption is required for the reported numbers to hold, and the method is parameter-free and immediately deployable. Therefore the reader's ACCEPT / high-confidence / low-correctness-risk judgment stands; the concrete oracle-gate test is only a confirmatory diagnostic, not a necessary condition for acceptance.","tokens_in":43330,"tokens_out":575,"duration_ms":5085,"concrete_test":"Re-run the eight headline cells (4 datasets × {near, dropout} at 25%) after replacing the Theil–Sen+MAD gate with an oracle that retains only true inliers (using the injected-outlier labels); if MRAC AbsRel improves by more than ~0.02–0.03 on KITTI-near relative to the reported 0.151, the residual gate misses are material; otherwise the foundation-consistency premise is already near-optimal for the claimed regime.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is well-supported by the same-backbone controls (MRAC vs vanilla/B' on DAv2, 84% strict wins across 64 cells) and the multipath-specific results against VI-Depth (all 12 corrupted near cells + all 16 KITTI cells, 3.2\times AbsRel reduction). The reader's weakest assumption—that legitimate anchors obey a single global affine with d_rel—is the method's premise, but the paper already quantifies its practical limits: gate P/R (mean 0.75/0.83, KITTI-near 0.60/0.64), Theil–Sen breakdown past ~29–40%, and clean-cell regressions of at most a few AbsRel points. These are disclosed and do not overturn the empirical win rates or the structural diagnosis of VI-Depth's missing present-but-wrong rejection step. No internal inconsistency or untested load-bearing gap remains that would reverse the ACCEPT.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies sparse-anchor metric calibration of frozen monocular depth foundations under sensor outliers that are present with wrong values (multipath, mixed pixels, dropout, uniform), not merely missing. It shows that residual-on-CFA collapses under such corruption and that VI-Depth, while robust to dropout, has a structural multipath blind spot and falls behind an unprotected baseline on three of four datasets when anchors are present-but-wrong. The proposed MRAC is a parameter-free inference-time wrapper: a Theil–Sen fit plus MAD gate on foundation consistency selects inliers, then a single residual-on-CFA head call produces metric depth. On a 320-cell benchmark with same-backbone controls, MRAC wins 84% of same-backbone cells across four outlier families and, against VI-Depth, all twelve corrupted multipath cells and all sixteen KITTI cells, cutting KITTI multipath AbsRel from 0.489 to 0.151 at ~50 µs CPU overhead and no retraining, while serving K∈[5,200] from one checkpoint.","tokens_in":43524,"tokens_out":1590,"duration_ms":33709,"significance":"If the results hold, the paper makes a clear, practically useful contribution: it isolates a real failure mode of the strongest publicly deployed sparse-anchor calibrator, supplies a systematic outlier-robustness benchmark that the field lacked, and offers a drop-in, K-agnostic, zero-parameter fix with negligible latency. Strengths include same-backbone same-architecture controls (vanilla and B′), three-seed headline bars, explicit Theil–Sen breakdown analysis, gate precision/recall diagnostics, a K=150 re-evaluation of VI-Depth at its training budget, and honest disclosure of clean-cell cost and gate misses. The work is more diagnostic and engineering than algorithmically novel (Theil–Sen+MAD is classical), but the controlled evaluation and structural diagnosis of present-but-wrong vs missing anchors are valuable for metric depth deployment from frozen foundations.","major_comments":[{"comment":"Abstract and §V-B headline the 3.2× KITTI multipath AbsRel reduction (0.489→0.151) against VI-Depth. On the same cell, unprotected vanilla already reaches 0.170 and B′ 0.164 (Table II), so most of the 3.2× gap is VI-Depth’s collapse relative to a competent residual-on-CFA baseline; MRAC’s same-backbone incremental gain is modest (~0.170→0.151). The structural blind-spot claim is well supported, but the abstract/intro should more carefully separate (i) VI-Depth worse than unprotected baseline from (ii) MRAC’s incremental same-backbone gain, so the wrapper’s contribution is not overstated relative to simply avoiding VI-Depth on multipath.","section":"Abstract, §V-B, Table II"},{"comment":"The entire 320-cell robustness grid (§IV-E, §V) injects synthetic outliers (uniform, near, dropout, mixed-pixel). The motivation and title center on real ToF multipath and related sensor faults, yet no experiment uses real multipath-corrupted range measurements. Synthetic families are sensor-grounded and the structural diagnosis of VI-Depth does not require real data, but transfer risk (outlier magnitude, spatial correlation, depth-range dependence) remains untested. A short real-sensor case study, or at least an explicit discussion of which multipath statistics the synthetic near model does and does not capture, is needed for the applied claim.","section":"§III-D, §IV-E, §V"},{"comment":"On the headline KITTI near/25% cell, MAD-gate precision/recall is only 0.60/0.64 (§IX-B, Fig. 5, Table SII): ~40% of clean anchors are falsely rejected and ~36% of multipath outliers pass. The residual head R_θ is trained on clean anchors with K∈{5,10,25,50} and no outlier injection (§III-F). False rejection therefore changes effective K and spatial support relative to training, which may explain residual losses (e.g., KITTI near/40% where vanilla 0.275 edges MRAC 0.299; §V-F). The paper should analyze residual-head behavior under gated (sparser/biased) anchors and whether false rejects, not only missed outliers, drive the remaining same-backbone gaps.","section":"§III-E–F, §V-F, §IX-B, Fig. 5"}],"minor_comments":[{"comment":"Fig. 3 averages AbsRel over four datasets; shaded bands are ±1 std across datasets, which can mask per-dataset inversions (e.g., DIODE clean wins for VI-Depth). Consider per-dataset curves in the supplement or a note that the average is for trend only.","section":"Fig. 3, §V-C"},{"comment":"Clean AbsRel for residual at K=50 is reported as 0.099 in the K-sweep (§VI-B) and 0.104 in the outlier harness (Table IV); the footnote attributes this to independent anchor draws. State the harness seed protocol once in §IV-E so readers do not treat the two as contradictory.","section":"§VI-B, Table IV"},{"comment":"VI-Depth’s strongest published backbone (dpt-beit-large-512) was unreachable (§IV-B, §IX-B). A one-sentence note on whether the public swin2 config is the recommended deployment setting would help readers interpret the cross-backbone comparison.","section":"§IV-B, §IX-B"},{"comment":"κ sensitivity is swept and flat (§III-E, Table SI); consider stating in the main text that κ=1.5 is marginally better on most cells so practitioners know the default is slightly conservative for precision.","section":"§III-E, Supplement Table SI"},{"comment":"Related work on robust regression (§II-C) is appropriate; a brief pointer to robust depth-completion or multipath-compensation literature beyond the cited ToF surveys would better situate the sensor-outlier families.","section":"§II-C, §III-D"},{"comment":"Algorithm 1 is clear; explicitly note that CFA is recomputed on M inside the single R_θ call (line 8) so readers do not assume the original a_cfa, b_cfa are reused.","section":"Algorithm 1, §III-E"}],"recommendation":"minor_revision","confidential_remarks":"Solid TPAMI-level empirical methods paper: careful controls, honest limitations, and a real deployment-relevant failure mode. Novelty is primarily diagnostic and systems-level rather than a new estimator; that is acceptable given the evaluation depth. The three major comments are framing and validation gaps, not correctness failures—fixable without new theory. I would not block on real multipath data if the authors add a clear transfer discussion, but it would strengthen the title claim. Fit for TPAMI is good."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is that VI-Depth really does have a structural multipath blind spot—robust to missing anchors, worse than an unprotected residual-on-CFA baseline when anchors are present but wrong—and they measure it cleanly. MRAC is a parameter-free Theil–Sen + MAD gate on foundation consistency before one residual-head call; ~50 µs, no new params, one checkpoint for K in [5,200].\n\nWhat is new is the diagnosis plus the 320-cell outlier grid with same-backbone same-architecture controls (vanilla and B′). Residual-on-CFA and Theil–Sen/MAD are prior art; packaging them as a K-agnostic inference wrapper that closes the present-but-wrong gap is the contribution. The 84% same-backbone win rate and the 3.2× KITTI multipath AbsRel drop (0.489→0.151) are backed by three-seed bars on the headline cells, gate P/R, and a K=150 re-check that the blind spot is not a budget artifact. They also show RANSAC fails on correlated dropout, which is the right diagnostic.\n\nSoft spots are real but disclosed and proportional. Backbone confound on clean DIODE/SUN cells is acknowledged; the load-bearing claim is the same-backbone comparison. Foundation consistency assumes a global affine with d_rel—gate P/R is only 0.60/0.64 on KITTI near, and Theil–Sen breaks past ~29–40%. Clean accuracy is at worst mildly conservative. No code release is a practical minus for a systems paper. None of that overturns the central empirical result.\n\nMath and citation pattern look fine: closed-form CFA substrate is correctly attributed, losses are ablated, related work covers foundations, sparse calibration, and robust regression without padding. This is for people shipping sparse-anchor metric depth from frozen foundations. I would bring it to reading group if we have depth people in the room, cite the multipath diagnosis and the wrapper if I work in this stack, and send it to peer review. Accept with the usual request for code and a bit more on when the affine premise fails.","headline":"Solid diagnosis of a real VI-Depth failure mode plus a cheap, well-controlled fix; immediately usable systems work, not a paradigm shift.","tokens_in":44166,"tokens_out":548,"would_cite":true,"duration_ms":9330,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Sparse-anchor metric depth methods collapse when range sensors return present-but-wrong values; a parameter-free foundation-consistency gate closes the multipath blind spot and cuts KITTI multipath error 3.2× with no retraining.","keywords":["monocular depth estimation","metric depth","sparse depth","robust estimation","depth foundation models","sensor outliers","calibration"],"falsifier":"On a domain where the foundation’s relative ordering is systematically biased, measure whether MRAC’s MAD gate rejects most clean anchors or accepts multipath ones and underperforms the unprotected residual head on multipath cells; that outcome would refute the foundation-consistency premise.","tokens_in":44233,"feed_emoji":"📏","tokens_out":964,"duration_ms":18884,"temperature":0.7,"pith_summary":"Monocular depth foundations give strong relative scene geometry but no absolute scale. A handful of sparse metric anchors from a range sensor can recover that scale, yet real sensors often return depths that are present with the wrong value—time-of-flight multipath, mixed pixels, max-range dropouts—not merely missing. The paper shows that the field’s residual-on-CFA calibration recipe collapses under these outliers, and that the strongest public method is prepared only for missing anchors, so it falls behind an unprotected baseline on three of four datasets when anchors are present but wrong. The authors introduce MRAC, an inference-time wrapper that tests each anchor against the foundation’s own relative-depth ordering with a Theil–Sen fit and a median-absolute-deviation gate, then feeds only the consistent anchors into a single calibration pass. On a 320-cell benchmark it wins 84% of same-backbone cells across four outlier families and all multipath-corrupted cells against the prior method, at roughly 50 microseconds of CPU cost, no new parameters, and one checkpoint for anchor budgets from 5 to 200.","feed_headline":"Wrong-valued anchors blind depth calibrators; free gate cuts error 3.2×","feed_subtitle":"A 50-microsecond foundation-consistency check rejects multipath and mixed-pixel returns before calibration.","key_machinery":"Multipath-Robust Anchor Calibration (MRAC): a parameter-free inference-time wrapper that estimates the affine link between the foundation’s relative depths and the sparse anchors with Theil–Sen, gates anchors by a median-absolute-deviation residual test, and passes only the survivors to one call of the residual-on-CFA calibration head.","core_discovery":"The strongest publicly deployed sparse-anchor calibrator has a structural multipath blind spot: its training prepares it for missing anchors but supplies no mechanism that can reject anchors present with the wrong value. Under multipath-style corruption its accuracy falls behind even an unprotected baseline on three of four datasets. Gating anchors by consistency with the frozen foundation’s relative-depth field—via a Theil–Sen affine fit and a MAD residual test—before a single residual-on-CFA head call restores accuracy across four sensor-grounded outlier families without retraining or added parameters.","pith_inferences":["Any anchor-conditioned metric head that never consults the foundation’s relative ordering will inherit the same present-but-wrong blind spot.","The same consistency gate can wrap other frozen geometry foundations without retraining those foundations or the gate.","Sensor-fusion stacks that treat ToF or LiDAR returns as trustworthy sparse cues will systematically degrade under multipath unless a similar consistency check is added."],"forward_implications":["Robustness to missing anchors and robustness to wrong-valued anchors are distinct problems; training that only drops anchors does not protect against multipath.","A single shared residual head plus a ~50 µs CPU gate can serve anchor budgets from 5 to 200 without per-budget checkpoints.","The foundation’s relative-depth geometry is a cheap, sufficient witness for rejecting multipath, dropout, mixed-pixel, and uniform sensor outliers.","Existing residual-on-CFA pipelines can be made multipath-resistant as a drop-in wrapper with no architectural change and zero retraining."],"fun_headline_variants":["Wrong anchors present but corrupted blind sparse depth calibrators","50μs foundation gate rejects multipath, cuts KITTI AbsRel 3.2× free","VI-Depth multipath blind spot: falls behind baseline on wrong values","Theil-Sen MAD consistency filter restores metric calibration zero-train","Parameter-free MRAC wins 84% cells across four multipath outlier families"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"Legitimate anchors obey a single global affine relationship with the foundation’s relative-depth field, so residuals cleanly separate clean measurements from present-but-wrong outliers of every family.","fun_headline_variants_meta":{"raw":{"variants":["Wrong anchors present but corrupted blind sparse depth calibrators","50μs foundation gate rejects multipath, cuts KITTI AbsRel 3.2× free","VI-Depth multipath blind spot: falls behind baseline on wrong values","Theil-Sen MAD consistency filter restores metric calibration zero-train","Parameter-free MRAC wins 84% cells across four multipath outlier families"]},"model":"grok-4.5","effort":"low","cost_usd":0.00641,"raw_usage":{"total_tokens":1739,"prompt_tokens":910,"num_sources_used":0,"completion_tokens":82,"cost_in_usd_ticks":64100000,"prompt_tokens_details":{"text_tokens":910,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":747,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":910,"tokens_out":82,"duration_ms":7055,"temperature":1.0,"reasoning_tokens":747,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T21:40:40.480027+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a domain where the foundation’s relative ordering is systematically biased, measure whether MRAC’s MAD gate rejects most clean anchors or accepts multipath ones and underperforms the unprotected residual head on multipath cells; that outcome would refute the foundation-consistency premise.","supporting_citations":[],"review_version":1}