{"id":"27d4198e-9e2a-45ea-a62d-64234ad40e60","arxiv_id":"2411.19278","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"OMNI-DC combines a multi-resolution depth integrator, a Laplacian loss, and scale-normalized synthetic training to achieve state-of-the-art zero-shot depth completion on seven datasets.","lead":"Depth completion fills in dense depth from an RGB image plus a handful of known depths; most models fail when the sensor or scene changes. This paper introduces OMNI-DC, a model trained on synthetic data that predicts dense depth across seven real-world datasets without retraining, cutting error by up to 43% on one benchmark.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Potential train/test image overlap between BlendedMVS and ETH3D may undermine the ETH3D zero-shot result, the source of the headline 43% claim.","rationale":"The reader's weakest assumption focuses on whether five synthetic datasets with simulated sparse patterns are representative enough for zero-shot transfer. That is a legitimate probabilistic concern, but I see a sharper, checkable threat to the same claim: BlendedMVS is not only a synthetic-rendering source with simulated patterns; it is a real-image multi-view stereo dataset whose composition is known to include ETH3D imagery. The paper's evaluation on ETH3D-SfM uses the ETH3D dataset, and no disjointness check between BlendedMVS training images and ETH3D evaluation images is reported. If the exact images overlap, the ETH3D-SfM result is not zero-shot, which removes the primary basis for the headline '43% error reduction' and weakens the 'directly applied to view synthesis' demonstration, since the 3DGS application is also on ETH3D scenes. This is more decisive than the Laplacian-loss novelty concern or the arbitrary 5.0 outdoor RMSE scaling in Tab. 2: those affect peripheral claims or presentation, whereas the overlap concern targets the central generalization result. I am not asserting that the overlap exists; I am asserting that the paper neither establishes disjointness nor acknowledges the need for it. The proposed hash and file-list comparison is a direct, low-cost way to settle the question. If the check comes back clean, the central claim retains its empirical support and the remaining issues are secondary. If it does not come back clean, the zero-shot claim requires substantial revision. Therefore I recommend keeping the verdict conditional, now conditioned specifically on the BlendedMVS-ETH3D disjointness check.","tokens_in":30301,"tokens_out":6941,"duration_ms":63931,"concrete_test":"Compare image-level identifiers or perceptual/photometric hashes between the BlendedMVS training file list and the 13 ETH3D scenes (454 images) used for evaluation in Appendix J. Also compare scene-level names in the released training and evaluation lists. If any exact or near-duplicate image appears in both, retrain OMNI-DC on BlendedMVS with those overlapping scenes removed and re-run the ETH3D-SfM rows of Table 3. If the RMSE and the 43% margin persist, the concern is cleared; if they change materially, the zero-shot claim must be re-quantified and the ETH3D results reported without the overlapping data.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is zero-shot state-of-the-art, with the 43% reduction anchored on the ETH3D-SfM outdoor result (Sec. 4.5, Table 3: RMSE 1.069 vs 1.883 for Marigold). For this to be valid, no ETH3D evaluation content may appear in training. The training mix in Table 1 includes BlendedMVS, a multi-view stereo dataset whose source imagery is partially derived from ETH3D scenes. If BlendedMVS contains any of the same images or scenes used for the ETH3D-SfM evaluation in Appendix J (13 scenes, 454 images), then that result is not zero-shot, and the strongest quantitative support for the generalization claim collapses. The paper does not report a check that BlendedMVS and the ETH3D evaluation split are disjoint. This is a concrete correctness risk distinct from the reader's representativeness concern: if the overlap exists, the loss is not a matter of distribution shift but of direct train/test contamination. Because the 43% claim, the 'zero-shot to real SfM' narrative, and the view-synthesis application all rely on this held-out status, the overlap check is load-bearing.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents OMNI-DC, a depth completion model aimed at zero-shot generalization across datasets and sparse depth sensor patterns. The main components are a multi-resolution variant of the Differentiable Depth Integrator (DDI), a Laplacian-based probabilistic loss, a log-depth scale normalization scheme with a claimed scale-equivariance guarantee, and a training pipeline that mixes five synthetic datasets with synthetically generated SfM, LiDAR, and noise patterns. The method is evaluated on seven real-world benchmarks (KITTI, NYUv2, VOID, ETH3D, iBims, ARKitScenes, DIODE) and is reported to outperform zero-shot baselines consistently, including a 43% RMSE reduction on ETH3D-SfM outdoor compared with Marigold. A downstream application to 3D Gaussian Splatting view synthesis is also presented.","tokens_in":30514,"tokens_out":11074,"duration_ms":93883,"significance":"If the claims hold, this is a substantively useful contribution: a single model that handles a wide range of densities, noise levels, and sensor types, with a parameter-free integration layer, a scale-equivariance property, and released code/checkpoints. The synthetic-only training recipe is interesting and the view-synthesis application is a practical demonstration. The main caveat is that the headline zero-shot result on ETH3D-SfM depends on the training data being truly held out, a property the manuscript does not verify. The scale-equivariance derivation and the careful ablations on held-out validation splits are strengths, as is the breadth of the evaluation.","major_comments":[{"comment":"The training set includes BlendedMVS (115K images) and the evaluation includes ETH3D-SfM (13 scenes, 454 images). BlendedMVS is publicly known to incorporate scene content derived from ETH3D, and the manuscript does not report any check that the ETH3D evaluation scenes are disjoint from the BlendedMVS training scenes. Since the headline result (Sec. 4.5, Table 3: RMSE 1.069 vs 1.883 for Marigold, a 43% reduction) and the claim of zero-shot generalization to real SfM both rest on ETH3D being unseen, this is a load-bearing issue. Please provide a scene-level and image-level disjointness check, such as listing the scene identifiers in BlendedMVS and the ETH3D evaluation scenes, or otherwise rule out overlap.","section":"Sec. 3.6, Table 1; Sec. 4.1, Appendix J"},{"comment":"The ETH3D-SfM evaluation projects COLMAP sparse points into image space, but the manuscript does not state how the arbitrary scale of the COLMAP reconstruction is converted to metric units before computing RMSE. If the ground-truth LiDAR is used to align the COLMAP scale, that should be stated explicitly, because it affects what the zero-shot claim covers. Please specify the exact scale-alignment procedure and any use of ground-truth data for scale recovery.","section":"Sec. 4.1, Appendix J"},{"comment":"Aggregating indoor and outdoor results by dividing outdoor RMSE by an arbitrary factor of 5.0 can change the ranking of methods and is not a standard metric. The per-dataset breakdown is relegated to Appendix N, which makes the main aggregated table difficult to interpret. Please either report the per-dataset numbers in the main table or use a scale-invariant aggregation such as REL, with the choice of the scaling factor justified.","section":"Sec. 4.6, Table 2"}],"minor_comments":[{"comment":"The statement that this work is \"the first to apply probability-based losses to depth estimation or depth completion\" is too strong. Heteroscedastic Laplacian and Gaussian losses have been used in prior uncertainty-aware depth regression and depth estimation works. Please temper the novelty claim accordingly.","section":"Sec. 3.4"},{"comment":"No error bars or multi-seed results are reported for the main benchmark tables. Given that several reported margins are modest, please report run-to-run variance or results from multiple seeds for the main comparisons.","section":"Tables 2-5"},{"comment":"The error-accumulation analysis assumes i.i.d. Gaussian gradient noise and a single known pixel, and the statement that multiresolution integration reduces the number of integration steps from n to n/2^{R-1} is heuristic. The ablation in Table 4 supports the design empirically, but the theoretical motivation should be qualified as an illustrative model.","section":"Sec. 3.2, Fig. 3"},{"comment":"The loss weights 0.5 and 2.0 in the final loss are not ablated. Please include a small sensitivity study of these weights or explain how they were selected.","section":"Eq. (10)"},{"comment":"The grouping \"Trained on KITTI/NYU\" includes methods that are evaluated zero-shot on ETH3D, which may confuse readers. Please clarify in the caption that these methods are trained on the in-domain datasets and tested on ETH3D without fine-tuning.","section":"Table 3"},{"comment":"The novel-view-synthesis results are reported on a single random split of 1/8 of the views. Please report variance across multiple splits or at least state whether the reported numbers are stable.","section":"Sec. 4.8"}],"recommendation":"major_revision","confidential_remarks":"The most serious issue is the potential train/test overlap between BlendedMVS and ETH3D. This should be resolved before acceptance, as the 43% headline claim depends on ETH3D being held out. The COLMAP scale-alignment protocol for ETH3D-SfM also needs clarification. The 'first to apply Laplacian loss' claim should be corrected; it is an overstatement given prior uncertainty-aware depth regression works. The scale-equivariance derivation and the breadth of the evaluation are strong, and the paper is likely to be a solid contribution once these issues are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"OMNI-DC is a serious, well-engineered system. The real contributions are the multi-resolution DDI, which fixes a genuine weakness in OGNI-DC on very sparse input; the log-space median normalization with provable scale equivariance; and the synthetic sparse-pattern training recipe. The scale-equivariance derivation in Appendix L is real, and the ablations use held-out validation. Seven datasets with consistently strong numbers plus a useful 3DGS demo give the paper its weight. The limitations section is honest about failure cases. That deserves credit.\n\nThe soft spots are real but localized. The claim of being 'first to apply probability-based loss to depth estimation' is not defensible; aleatoric and Laplacian losses have appeared in depth and dense prediction before. Soften or drop it. The Tab 2 aggregation divides outdoor RMSE by an arbitrary 5.0; the appendix reports separated numbers, so the information is there, but the main table's scaling is unjustified. No error bars or multi-seed runs, though the headline margins are large enough that this is a minor issue for the main comparisons.\n\nThe bigger concern is the stress-test note about train/test overlap. The training mix includes BlendedMVS, whose source imagery is partly derived from ETH3D scenes. The paper reports zero-shot results on ETH3D-SfM, anchors the 43% claim there, and uses ETH3D for the view-synthesis demo, but it never checks disjointness between BlendedMVS content and the ETH3D evaluation split. That is a load-bearing risk. If overlap exists, the strong zero-shot claim collapses. The authors need to report a scene-name or image-level overlap check.\n\nIf that check comes back clean, I'd trust the empirical body. As it stands, the paper deserves peer review, not desk rejection. The engineering is sound and the multi-res DDI plus scale normalization are useful beyond the zero-shot claim. A referee should ask for the overlap check, the loss-prior fix, and a justification or removal of the 5.0 scaling. This is for depth-completion practitioners and anyone using sparse depth in 3DGS pipelines. I'd cite it once the ETH3D overlap is ruled out.","headline":"Strong zero-shot depth completion with genuine contributions, but the headline ETH3D result needs a train/test overlap check before the 43% claim is trusted.","tokens_in":31131,"tokens_out":2895,"would_cite":true,"duration_ms":24941,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single depth-completion model, trained only on synthetic data, transfers zero-shot to real LiDAR, SfM, and VIO sparse depth and beats per-dataset-trained rivals on several benchmarks.","keywords":["depth completion","zero-shot generalization","multi-resolution integration","Laplacian loss","scale equivariance","synthetic training data","LiDAR simulation","novel view synthesis"],"falsifier":"Take a real sparse depth map from an unseen sensor (for instance mmWave radar on the ZJU-4DRadarCam benchmark) and compare OMNI-DC's zero-shot RMSE against the in-domain-trained RadarCam-Depth model; the paper's appendix already shows a two-fold gap. If that gap persists or widens when realistic radar patterns are added to the training simulation, the claim that the model is sensor-agnostic is falsified. A more direct test: on a controlled input with known sparse points, measure the RMSE of the predicted depth versus distance to the nearest known pixel; the Multi-resolution DDI predicts the growth slows with the number of resolutions, so a linear-in-distance growth matching the single-resolution bound would contradict the paper's mechanism.","tokens_in":30051,"feed_emoji":"📐","tokens_out":13109,"duration_ms":102162,"temperature":0.7,"pith_summary":"This paper tries to establish that depth completion — predicting a dense depth map from an RGB image plus a few known depth points — can be made to work reliably across many sensors, scene types, and sparsity levels without retraining, using a single model. The authors claim their multi-resolution depth integrator, uncertainty-aware Laplacian loss, and synthetic-only training recipe with simulated LiDAR and SfM patterns let one model outperform all zero-shot baselines on seven real-world datasets, with up to a 43 percent error reduction, and even beat models trained directly on KITTI on one metric. A reader should care because today's depth completion models typically fail on unseen sparse patterns and must be retrained per dataset, which blocks practical use in robotics, reconstruction, and view synthesis.","feed_headline":"Zero-shot depth model beats per-dataset rivals on 7 benchmarks","feed_subtitle":"Trained only on synthetic data, it transfers zero-shot to real LiDAR, SfM, and VIO inputs and improves 3DGS rendering.","key_machinery":"The Multi-resolution Depth Integrator is a parameter-free layer that estimates the dense depth map as the solution of a linear least-squares problem, balancing a sparse-depth fidelity term (with a learned confidence weight) against predicted depth-gradient constraints computed at several resolutions: full resolution and successive average-pooled half-resolutions. Down-sampling the optimization target before taking finite differences makes the effective integration step from a known pixel to a far pixel shrink by a factor of $2^{R-1}$, which the paper's 1D analysis shows reduces accumulated variance from $n\\sigma^2$ toward $n\\sigma^2/2^{R-1}$. The same least-squares formulation, with sparse depths expressed in log space and the network input normalized by the median log-depth, yields the paper's proven guarantee that the output scales linearly with any multiplicative change in input depth.","core_discovery":"The paper's central claim is that the poor generalization of depth completion to new datasets and unseen sparse depth patterns is not an intrinsic limit of the task but a fixable design problem. The authors diagnose the failure mode by analyzing a simplified 1D version of the optimization-guided neural iteration: with an i.i.d. Gaussian error on predicted depth gradients $\\hat{G}_i = G^{\\mathrm{gt}}_i + n_i$, integrating from a single known pixel at position 0 yields $\\hat{D}_n \\sim N(D_0 + \\sum_i G^{\\mathrm{gt}}_i,\\; n\\,\\sigma^2)$, so uncertainty grows linearly with distance to the nearest known pixel. Because real sparse depth from SfM or active sensors often leaves large holes, this error accumulation makes single-resolution integration unusable on extremely sparse inputs. The paper's remedy is a Multi-resolution Depth Integrator that solves a least-squares problem enforcing gradient constraints at several down-sampled resolutions, reducing the integration distance and the accumulated variance. Combined with a per-pixel Laplacian loss that models depth ambiguity, a log-depth scale normalization that gives guaranteed scale equivariance, and training on five synthetic datasets with simulated SIFT-based, LiDAR-line, and noise-corrupted sparse patterns, the resulting model generalizes zero-shot to LiDAR, SfM, VIO, and consumer depth across seven datasets, with the largest reported gain a 43 percent RMSE reduction on the outdoor ETH3D-SfM split and a zero-shot KITTI MAE that beats all in-domain-trained baselines. The same depth prior, injected as a loss term into 3D Gaussian Splatting training, lifts rendering PSNR from 15.64 to 20.38 on ETH3D.","pith_inferences":["An implication the authors leave implicit is that the multi-resolution gradient-integration idea should transfer to other integrating dense prediction tasks — such as surface-normal or optical-flow estimation from sparse constraints — where single-resolution propagation suffers the same variance growth with distance.","If the pattern simulator is the key to zero-shot transfer, then extending it to cover radar, structured-light holes, and object-removal gaps should extend the model's coverage; the paper's own radar experiments identify exactly this as the current weak spot.","The paper's finding that mixing real NYU data into training hurts performance — because real ground-truth depth is blurry — suggests that label sharpness, not domain realism, drives the success, a hypothesis the paper does not directly test.","A practical extension would be to let the model fall back to relative (monocular) depth when the sparse input is empty or nearly empty, since the current architecture explicitly does not handle the zero-sparse-point case; a smooth interpolation between the two regimes would broaden the model's applicability to monocular depth estimation."],"forward_implications":["A single OMNI-DC checkpoint can replace per-dataset-trained depth completion systems on any new sensor or scene whose sparse pattern resembles the simulated LiDAR, SfM, and noise patterns used in training.","The model's zero-shot KITTI results (MAE 0.597 on 8-line LiDAR, better than every in-domain-trained baseline) indicate that broad synthetic training with realistic pattern simulation can beat in-domain training, not merely approach it.","Because the output is provably scale-equivariant, users can feed SfM reconstructions with unknown metric scale directly into the model and get consistent relative depth without estimating a scale factor.","OMNI-DC's dense depth can serve as a geometry prior for 3D Gaussian Splatting, improving novel view synthesis on sparse-view scenes (PSNR 20.38 vs 15.64 for vanilla 3DGS).","The multi-resolution integrator's robustness to extremely sparse inputs (0.03% density, REL=0.034 in the synthetic-pattern suite) should carry over to downstream tasks where measurements are scarce, such as VIO with few tracked points."],"supporting_citations":[{"why":"The prior method this work builds on: its single-resolution Differentiable Depth Integrator is analyzed, found to accumulate integration error on sparse inputs, and extended to multiple resolutions.","marker":"[74]"},{"why":"Supplies the CompletionFormer backbone network used to extract features and predict multi-resolution depth gradients.","marker":"[73]"},{"why":"Source of the gradient-matching loss and the dataset-mixing philosophy the paper adapts to the depth completion setting.","marker":"[41]"},{"why":"ETH3D with real COLMAP SfM points is the testbed for the paper's headline 43 percent error reduction on the outdoor split.","marker":"[47]"},{"why":"KITTI provides the LiDAR benchmark where zero-shot OMNI-DC beats all in-domain-trained baselines on MAE for 8-line input.","marker":"[55]"},{"why":"The closest zero-shot baseline, G2-MonoDepth, used in nearly every comparison table as the method to beat.","marker":"[57]"},{"why":"3D Gaussian Splatting is the downstream view-synthesis framework used to demonstrate the model's practical value.","marker":"[22]"},{"why":"Provides the depth-loss regularization recipe (DN-Splatter) used to inject OMNI-DC's depth into 3DGS training.","marker":"[54]"}],"fun_headline_variants":["Depth completion that zero-shots to any sensor","Synthetic-only training, real-world depth wins","43% error cut: depth model that generalizes","OMNI-DC: one depth model, all sparse inputs","From synthetic to LiDAR, SfM, VIO: zero-shot depth"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the five synthetic training datasets and the simulated sparse patterns (SIFT keypoints, randomized LiDAR lines, and injected noise) are representative enough of real sensors and scenes that a model trained on them alone transfers zero-shot to real LiDAR, SfM, VIO, and consumer-depth data.","fun_headline_variants_meta":{"raw":{"variants":["Depth completion that zero-shots to any sensor","Synthetic-only training, real-world depth wins","43% error cut: depth model that generalizes","OMNI-DC: one depth model, all sparse inputs","From synthetic to LiDAR, SfM, VIO: zero-shot depth"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00055,"raw_usage":{"total_tokens":2684,"prompt_tokens":1062,"completion_tokens":1622,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":678,"completion_tokens_details":{"reasoning_tokens":1540}},"tokens_in":678,"tokens_out":1622,"duration_ms":11311,"temperature":1.0,"reasoning_tokens":1540,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:20:25.235861+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a real sparse depth map from an unseen sensor (for instance mmWave radar on the ZJU-4DRadarCam benchmark) and compare OMNI-DC's zero-shot RMSE against the in-domain-trained RadarCam-Depth model; the paper's appendix already shows a two-fold gap. If that gap persists or widens when realistic radar patterns are added to the training simulation, the claim that the model is sensor-agnostic is falsified. A more direct test: on a controlled input with known sparse points, measure the RMSE of the predicted depth versus distance to the nearest known pixel; the Multi-resolution DDI predicts the growth slows with the number of resolutions, so a linear-in-distance growth matching the single-resolution bound would contradict the paper's mechanism.","supporting_citations":[{"cited_title":"generalizable","cited_arxiv_id":null,"evidence_quote":"The prior method this work builds on: its single-resolution Differentiable Depth Integrator is analyzed, found to accumulate integration error on sparse inputs, and extended to multiple resolutions."},{"cited_title":"Completionformer: Depth completion with convolutions and vision transform- ers","cited_arxiv_id":null,"evidence_quote":"Supplies the CompletionFormer backbone network used to extract features and predict multi-resolution depth gradients."},{"cited_title":"Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer","cited_arxiv_id":null,"evidence_quote":"Source of the gradient-matching loss and the dataset-mixing philosophy the paper adapts to the depth completion setting."},{"cited_title":"Sch¨onberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and An- dreas Geiger","cited_arxiv_id":null,"evidence_quote":"ETH3D with real COLMAP SfM points is the testbed for the paper's headline 43 percent error reduction on the outdoor split."},{"cited_title":"Sparsity invariant cnns","cited_arxiv_id":null,"evidence_quote":"KITTI provides the LiDAR benchmark where zero-shot OMNI-DC beats all in-domain-trained baselines on MAE for 8-line input."},{"cited_title":"G2- monodepth: A general framework of generalized depth in- ference from monocular rgb+ x data","cited_arxiv_id":null,"evidence_quote":"The closest zero-shot baseline, G2-MonoDepth, used in nearly every comparison table as the method to beat."},{"cited_title":"Dn-splatter: Depth and normal priors for gaussian splatting and meshing","cited_arxiv_id":null,"evidence_quote":"Provides the depth-loss regularization recipe (DN-Splatter) used to inject OMNI-DC's depth into 3DGS training."}],"review_version":1}