{"id":"874ad04d-8db3-4ad2-8d80-0bd9acf48010","arxiv_id":"1908.06257","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"OmniMVS learns omnidirectional depth from multi-view fisheye images end-to-end and reports lower error than prior omnidirectional and stitched conventional stereo methods on synthetic benchmarks, with qualitative real-world demos.","lead":"This paper presents OmniMVS, a neural network that estimates full-surround depth from four fisheye cameras on a rig by sweeping learned features over concentric spheres and regularizing with a 3D encoder-decoder. It also introduces two large synthetic datasets for training and testing omnidirectional stereo algorithms. A generalist would read it as a step toward all-around depth sensing for robots and vehicles.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Synthetic-to-real transfer is the load-bearing premise: Sec. 5.3 provides only qualitative real-world point clouds with no ground-truth depth, so the central 'viable and better replacement' claim is not quantitatively supported.","rationale":"Good-faith reading: the paper's main technical contribution is a credible end-to-end spherical-sweeping architecture plus large synthetic datasets. The synthetic evaluation is carefully designed with multiple baselines, fine-tuned variants, and held-out test frames. I do not find an internal inconsistency in the spherical sweeping formulation or the network architecture. The most load-bearing concern is the unsupported leap from synthetic scores to real-world viability. The abstract claims excellent results in real-world environments, and Sec. 5.3 asserts 'clean and detailed reconstructions' from real data, but no ground-truth metric is reported. Since the conclusion frames OmniMVS as a replacement for multi-stage omnidirectional stereo in physical environments, the lack of any quantitative real-world validation leaves the central practical claim unverified. Visually plausible point clouds can conceal systematic depth bias, scale drift, or distorted surfaces. The proposed LiDAR/reference evaluation would settle this concern directly. If the transfer holds, the conditional can be lifted; if it does not, the claims must be scoped to synthetic data. This matches the reader's weakest_assumption, so the CONDITIONAL verdict stands unchanged.","tokens_in":14063,"tokens_out":22658,"duration_ms":231949,"concrete_test":"Acquire real-world ground truth for the same four-camera rig used in Sec. 5.3 (e.g., a co-calibrated LiDAR or structured-light scan) for a set of scenes; run OmniMVS-ft and the two strongest baselines (SweepNet+SGM and DispNet-CSS) under the same settings; then compute the Sec. 5.2 metric, Eq. 3, on the same H=160 equatorial crop and validity mask used for Table 3. If OmniMVS-ft's real-world MAE is not below the baselines by a margin comparable to the synthetic margins, or is far above its synthetic MAE, the sim-to-real transfer claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The synthetic benchmark in Sec. 5.2/Table 3 is well structured, but the paper's broader claim is that OmniMVS works 'in real-world environments' and is therefore a viable replacement for multi-stage omnidirectional stereo. That real-world evidence is exclusively qualitative: Sec. 5.3 shows point clouds and inverse-depth maps with no ground-truth depth, LiDAR, SfM reference, or any metric. The network is trained on Blender-rendered OmniThings and fine-tuned on synthetic Sunny/OmniHouse, then applied to real fisheye images without retraining. Sim-to-real gaps in lens distortion, radiometry, noise, and scene statistics could produce visually plausible but geometrically biased depths. Because the central practical claim is about a real-world replacement, the unmeasured synthetic-to-real transfer is load-bearing. Table 3's synthetic numbers cannot resolve this.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes OmniMVS, an end-to-end CNN for omnidirectional depth estimation from four wide-FOV fisheye cameras on a rig. After extracting 2D unary features, the network warps the feature maps onto concentric spheres at sampled inverse depths using calibrated camera parameters, forms a concatenated 4D cost volume, regularizes it with a 3D encoder-decoder, and regresses inverse depth with softargmin. The authors also introduce two Blender-rendered synthetic datasets, OmniThings and OmniHouse, totaling roughly 12.8K depth maps and 51K fisheye images, and evaluate the method on these plus the existing Sunny/Cloudy/Sunset datasets against spherical-sweeping baselines and stitched conventional stereo methods. The paper claims state-of-the-art results on all synthetic test sets and demonstrates real-world reconstructions qualitatively.","tokens_in":14221,"tokens_out":7368,"duration_ms":75600,"significance":"If the real-world claims are substantiated, this is a useful contribution: it is one of the first end-to-end learned systems for omnidirectional multi-view stereo, it introduces a large publicly described synthetic benchmark, and the architecture is described in sufficient detail to be credible. The synthetic comparison in Table 3 is internally consistent, uses held-out test frames, and includes several baselines, which are real strengths. The significance is currently tempered by the absence of any quantitative real-world evaluation and by the fact that all tabulated metrics are computed on a cropped equatorial band rather than the full sphere.","major_comments":[{"comment":"The real-world claim is supported only by qualitative point clouds and inverse-depth visualizations. There is no ground-truth depth, LiDAR reference, SfM point cloud, or any metric for the real fisheye sequences, while the abstract states that the method 'generates excellent results in both synthetic and real-world environments' and the conclusion says it 'successfully reconstructs accurate omnidirectional depth.' Because the network is trained only on Blender-rendered data, the synthetic-to-real transfer is load-bearing. Please add a quantitative real-world evaluation (for example, comparison with LiDAR or a structure-from-motion reference) or explicitly restrict the real-world claims to qualitative demonstrations.","section":"5.3, Abstract, Conclusion"},{"comment":"All quantitative metrics are computed on a crop with H = 160, i.e., phi in [-pi/4, pi/4], excluding the top and bottom 45 degrees, while the paper repeatedly claims omnidirectional depth estimation. The reported 'best on all five datasets' results therefore hold only for this equatorial band. In addition, the depth sampling parameters d0, dN-1, and the actual depth range are never stated, and the text says the inverse radius d_n is swept from 0 to d_max, which makes Eq. (1) undefined at d_n = 0. Please specify the full evaluation protocol, including the depth range and the handling of d_n, and discuss or evaluate the pole regions if the omnidirectional claim is retained.","section":"5.1, Eq. (3), Table 3"},{"comment":"The headline OmniMVS-ft results are obtained after fine-tuning on 'Sunny and OmniHouse' and then evaluated on Sunny and OmniHouse test frames. Even if the test frames themselves are held out, the paper does not state the exact fine-tuning split, and for OmniHouse the training and test frames are rendered from the same 451 house models, so the fine-tuned numbers may reflect scene-level familiarity rather than generalization. Please report the precise train/test split used for fine-tuning, present the OmniMVS-only results as the primary cross-dataset numbers, and show that fine-tuning does not rely on test-scene geometry.","section":"Table 3, Table 4, Sec. 5.1"}],"minor_comments":[{"comment":"The abstract says the datasets consist of 11K ground-truth depth maps and 45K fisheye images, but Table 2 sums to 12,800 scenes and 51,200 fisheye images if test frames are included. Please clarify whether the abstract refers to training frames only.","section":"Abstract, Table 2"},{"comment":"The sentence 'our end-to-end networks perform better in all datasets and metrics' is not true for OmniMVS-ft on OmniThings, where the MAE is 3.52 compared with 2.40 for OmniMVS. Please qualify the statement or discuss the fine-tuning trade-off.","section":"Sec. 5.2"},{"comment":"The claim that textureless walls are straight and small objects are reconstructed accurately is based on visual inspection. If quantitative evaluation is not added, please soften the wording from 'accurate' to 'qualitatively plausible.'","section":"Sec. 5.3 and Fig. 7"},{"comment":"There are several typos: 'exsiting' near Table 2, 'OmniThngs' in Sec. 5.2, 'tranined' in the supplementary caption of Fig. 1, and 'texureless' in the caption of Fig. 7.","section":"Typographical"},{"comment":"The MAE and RMS values are percentages of the inverse-depth index range, not metric depth errors. Please state this explicitly in the main text so that readers do not interpret the numbers as meters.","section":"Eq. (3)"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the synthetic benchmark study is careful and the datasets are a genuine community resource, so the paper is worth a revision round. The main missing element is a quantitative real-world evaluation or a proportionate narrowing of the real-world claims; the fine-tuning split also needs to be clarified before the Table 3 numbers can be taken at face value."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, well-executed paper. It is the first end-to-end network for wide-baseline omnidirectional multi-view stereo, and the two synthetic datasets it contributes (OmniThings, OmniHouse) are a real resource for a subfield that had only SweepNet's small driving set. The synthetic benchmark evidence is convincing; the real-world evidence is suggestive but not quantified, and the 'omnidirectional' label is a bit strong given the evaluation crops to a 90° band around the equator.\n\nWhat is genuinely new: spherical sweeping applied directly to deep feature maps in a rig coordinate system, followed by a 3D encoder-decoder and softargmin regression. SweepNet stops at learned matching costs plus SGM; conventional end-to-end stereo networks assume rectified pinhole pairs. The architecture is described in enough detail to be credible, and the training details (random permutation of the view order, random rig rotations) show sensible care.\n\nThe quantitative evaluation is properly structured: five test datasets, comparison against both spherical-sweeping baselines (ZNCC+SGM, MC-CNN+SGM, SweepNet+SGM) and stitched conventional stereo (PSMNet, DispNet-CSS), with published pretrained weights and fine-tuned versions. On OmniThings, OmniMVS drops MAE from 4.06 (DispNet-CSS) to 2.40; on the Sunny set, OmniMVS-ft gets 0.79 MAE versus 1.31 for SweepNet+SGM. These numbers are internally consistent and use held-out test frames. The fine-tuned results are obtained after fine-tuning on the same dataset families, which is a minor caveat, but OmniMVS without fine-tuning is already competitive or better on OmniThings and only slightly behind on OmniHouse, so the headline claim does not rest on that.\n\nThe soft spots are real but not fatal. Section 5.3 shows only qualitative point clouds and inverse-depth maps for real scenes; there is no LiDAR, SfM reference, or any metric. The paper's abstract says 'excellent results in both synthetic and real-world environments,' but the real-world part is not substantiated quantitatively. That is a gap in confidence, not a flaw in the synthetic evaluation. The polar cropping (evaluation limited to −π/4 ≤ φ ≤ π/4) also undercuts the 'omnidirectional' claim; the network might work near the poles, but we are not shown that. Finally, no link to code or data is given, which hurts reproducibility.\n\nWho is this for? Anyone working on multi-view stereo, omnidirectional depth, or synthetic-to-real training. The datasets alone are worth having, and the network is a strong baseline for future work. I would send it to peer review. The main revisions I would ask for: either add a quantitative real-world evaluation (sparse LiDAR or a few measured distances would suffice) or tone down the real-world claim, and make the code and datasets publicly available.","headline":"Solid end-to-end omnidirectional stereo paper with strong synthetic benchmarks and two useful datasets; the real-world evidence is qualitative and the polar regions are cropped out, so the 'omnidirectional' claim is slightly oversold.","tokens_in":14759,"tokens_out":2541,"would_cite":true,"duration_ms":26509,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An end-to-end network using spherical sweeping estimates a full 360-degree depth map from four fisheye cameras, outperforming all compared stereo pipelines on five test sets.","keywords":["omnidirectional depth estimation","spherical sweeping","multi-view stereo","fisheye cameras","end-to-end deep learning","synthetic datasets","inverse depth","3D encoder-decoder"],"falsifier":"Run the OmniMVS-ft network on a real four-fisheye rig in a room with a textureless wall and a reflective floor, and compare the predicted inverse depth against LiDAR or structured-light ground truth; if the mean absolute error is no better than the SweepNet+SGM baseline, the paper's claim that learned global context resolves real-world textureless and reflective surfaces collapses.","tokens_in":13832,"feed_emoji":"🌐","tokens_out":10380,"duration_ms":94596,"temperature":0.7,"pith_summary":"The paper tries to establish that full surround depth can be obtained from four fisheye cameras by one end-to-end trained network, replacing the usual pipeline of computed matching costs followed by hand-designed smoothing such as semi-global matching. The proposed OmniMVS network extracts learned features from each fisheye image, warps them onto concentric spheres at many candidate depths, and lets a 3D encoder-decoder regularize the resulting volume before regressing a complete inverse depth map. The authors report that this single network outperforms prior omnidirectional methods and stitched conventional stereo on all five test datasets, with large gains on thin objects and on large textureless or reflective surfaces. They also contribute two large synthetic datasets, OmniThings and OmniHouse, with roughly 11K ground-truth depth maps and 45K fisheye images, and argue that networks trained on these scenes transfer to real fisheye imagery.","feed_headline":"Learned spherical sweeping beats stitched stereo on every test set","feed_subtitle":"The single end-to-end network replaces hand-built cost aggregation and wins on all five omnidirectional datasets.","key_machinery":"The load-bearing object is the spherical feature volume built by calibration-based inverse-depth sweeping. After a shared 2D CNN extracts unary feature maps from each grayscale fisheye image, the feature maps are warped onto concentric spheres at a sequence of inverse depths using each camera's intrinsics and extrinsics, so that every sphere corresponds to one hypothesized distance. The warped maps from all cameras are concatenated and fused with 3D convolutions into a 4D volume indexed by spherical direction, depth, and channel, and a 3D encoder-decoder with skip connections produces a regularized cost volume from which softargmin yields the final inverse depth. This mechanism is what carries the argument: it makes the geometry differentiable, merges evidence from all views before any decision, and lets global context resolve occlusion and multiple-true-match ambiguity.","core_discovery":"The central claim is that spherical sweeping, the omnidirectional analogue of plane sweeping, can be made fully differentiable and trained end-to-end for a wide-baseline multi-camera rig, and that the learned network resolves problems that defeat the previous generation of omnidirectional stereo. Earlier pipelines such as SweepNet compute a matching-cost volume from warped spherical images and then smooth it with semi-global matching; in a global sweep a single viewing ray can pass through several objects, producing multiple true matches that SGM cannot handle. OmniMVS instead warps deep feature maps from all cameras onto concentric spheres indexed by inverse depth, concatenates and fuses them into a 4D volume, and runs a 3D encoder-decoder that uses global context to regularize the cost before softargmin regression. On the five test datasets (Sunny, Cloudy, Sunset, OmniThings, and OmniHouse) the authors report the lowest mean absolute error among all compared methods, with OmniMVS-ft reaching MAE 0.79 on Sunny versus 1.31 for SweepNet+SGM. The paper's claim, in short, is that end-to-end learned spherical sweeping is not only feasible but strictly better than the multi-stage pipeline it replaces.","pith_inferences":["The paper does not quantify the sim-to-real gap; collecting LiDAR or structured-light ground truth on the same rig and comparing OmniMVS-ft's error against those measurements would turn the qualitative real-world results into a testable metric.","Because warping uses calibrated intrinsics and extrinsics, the same network design should transfer to other multi-camera layouts, different numbers of cameras, and non-omnidirectional wide-FOV rigs; the paper only demonstrates one four-camera layout.","The inverse-depth parameterization concentrates depth samples near the rig, so downstream users who care about metric accuracy at long range would need to account for the non-uniform depth resolution rather than treating the error numbers as uniform in meters."],"forward_implications":["An omnidirectional rig with four calibrated fisheye cameras can produce a full spherical depth map in about one second per frame on a single GPU, with no separate cost aggregation or stitching stage.","Networks trained on random synthetic scenes (OmniThings) transfer to indoor, outdoor, and real-image settings, and fine-tuning on OmniHouse and Sunny further improves textureless and reflective surfaces.","Conventional stereo networks applied by stitching four rectified pairs are a strictly weaker route to omnidirectional depth than a single network that reasons over all views at once.","Ablating or replacing the 3D encoder-decoder should cause the multiple-true-match failures seen in SGM-based pipelines to reappear, confirming that global-context regularization is the source of the gain."],"supporting_citations":[{"why":"Supplies the previous spherical-sweeping pipeline and the Sunny, Cloudy, and Sunset datasets that OmniMVS is compared against.","marker":"[30]"},{"why":"Contributes the 3D encoder-decoder cost-volume architecture and softargmin regression that OmniMVS adapts for omnidirectional input.","marker":"[14]"},{"why":"Provides the state-of-the-art conventional stereo network used as a stitching baseline.","marker":"[4]"},{"why":"Supplies the synthetic-data recipe and the DispNet baseline, and its object-generation approach inspires OmniThings.","marker":"[16]"},{"why":"Gives DispNet-CSS, the strongest stitching baseline in the comparison.","marker":"[12]"},{"why":"Is the semi-global matching regularizer that the multi-stage baselines rely on and that the end-to-end network replaces.","marker":"[10]"},{"why":"Provides MC-CNN patch matching, a baseline cost computation method for spherical sweeping.","marker":"[32]"},{"why":"Supplies the indoor scene models from which OmniHouse is rendered.","marker":"[27]"}],"fun_headline_variants":["End-to-end spherical sweeping beats all previously published stereo","Learned sweeping replaces hand-built aggregation, wins every test","OmniMVS: one network to rule all omnidirectional depth tasks","Spherical sweep made end-to-end, tops prior art on all sets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that the network works in real environments rests on the untested premise that training on synthetic rendered scenes transfers to real fisheye images; the paper shows qualitative point clouds but never measures real-world ground-truth depth.","fun_headline_variants_meta":{"raw":{"variants":["End-to-end spherical sweeping beats all previously published stereo","Learned sweeping replaces hand-built aggregation, wins every test","OmniMVS: one network to rule all omnidirectional depth tasks","Spherical sweep made end-to-end, tops prior art on all sets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000337,"raw_usage":{"total_tokens":1878,"prompt_tokens":972,"completion_tokens":906,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":588,"completion_tokens_details":{"reasoning_tokens":834}},"tokens_in":588,"tokens_out":906,"duration_ms":9247,"temperature":1.0,"reasoning_tokens":834,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:51:34.852478+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the OmniMVS-ft network on a real four-fisheye rig in a room with a textureless wall and a reflective floor, and compare the predicted inverse depth against LiDAR or structured-light ground truth; if the mean absolute error is no better than the SweepNet+SGM baseline, the paper's claim that learned global context resolves real-world textureless and reflective surfaces collapses.","supporting_citations":[{"cited_title":"Improved wide-angle, fisheye and omnidirectional camera calibration","cited_arxiv_id":null,"evidence_quote":"Supplies the previous spherical-sweeping pipeline and the Sunny, Cloudy, and Sunset datasets that OmniMVS is compared against."},{"cited_title":"Occlusions, motion and depth boundaries with a generic network for disparity, optical flow or scene flow estimation","cited_arxiv_id":null,"evidence_quote":"Contributes the 3D encoder-decoder cost-volume architecture and softargmin regression that OmniMVS adapts for omnidirectional input."},{"cited_title":"Fast approximate energy minimization via graph cuts","cited_arxiv_id":null,"evidence_quote":"Provides the state-of-the-art conventional stereo network used as a stitching baseline."},{"cited_title":"End-to-end learning of geometry and context for deep stereo regression","cited_arxiv_id":null,"evidence_quote":"Supplies the synthetic-data recipe and the DispNet baseline, and its object-generation approach inspires OmniThings."},{"cited_title":"Stereo processing by semiglobal matching and mutual information","cited_arxiv_id":null,"evidence_quote":"Gives DispNet-CSS, the strongest stitching baseline in the comparison."},{"cited_title":"Displets: Resolving stereo ambiguities using object knowledge","cited_arxiv_id":null,"evidence_quote":"Is the semi-global matching regularizer that the multi-stage baselines rely on and that the end-to-end network replaces."},{"cited_title":"SweepNet: Wide-baseline Omnidirectional Depth Estimation","cited_arxiv_id":"1902.10904","evidence_quote":"Provides MC-CNN patch matching, a baseline cost computation method for spherical sweeping."},{"cited_title":"Omnidirectional 3d reconstruction in augmented manhattan worlds","cited_arxiv_id":null,"evidence_quote":"Supplies the indoor scene models from which OmniHouse is rendered."}],"review_version":1}