{"id":"ea39320d-98e5-4cd1-b392-62a441689770","arxiv_id":"1908.09474","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A systematic evaluation shows that learned descriptors, GMS pruning, and a coarse-to-fine RANSAC plus LMedS combination improve fundamental matrix estimation over the SIFT plus RANSAC baseline.","lead":"This paper benchmarks modern feature matchers, correspondence pruners, and robust estimators for two-view geometry estimation on four public datasets. It finds that learned descriptors, GMS pruning, and a coarse-to-fine RANSAC plus LMedS combination improve results over the classic SIFT plus RANSAC pipeline.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'significantly outperforms' claim for CF-RSC is not supported on TUM and KITTI: Table 5 differences over LMedS are 0.1 and 0.5 pp, within sampling error for 1000 pairs, and no confidence intervals or significance tests are reported.","rationale":"The reader's formal weakest assumption is the SIFT-based pair selection and COLMAP ground truth on T&T/CPC. That is a real threat to the relative ranking of local features on those datasets. However, for the proposed CF-RSC claim, all estimators in Table 5 consume the same SIFT matches, so the SIFT-selection bias does not directly explain away CF-RSC's advantage on T&T/CPC. The more direct threat to the central 'significantly outperforms' claim is statistical: the TUM and KITTI rows are effectively ties, and the paper provides no uncertainty quantification despite using a random subset of pairs, which invites resampling analysis. The proposed matching systems in Table 4 show consistent gains over the baseline and are competitive with CODE, and the CF-RSC gains on wide-baseline datasets are plausible because GC-RANSAC is known to find good inlier sets and LMedS is a strong final fitter. The correct remedy is therefore not rejection but conditioning the strongest claim on statistical significance and on the regime. This is consistent with the reader's CONDITIONAL verdict; my concern refines it rather than overturning it.","tokens_in":11426,"tokens_out":5094,"duration_ms":51920,"concrete_test":"Report paired bootstrap confidence intervals for the Table 5 %Recall differences (CF-RSC minus LMedS, and CF-RSC minus GC-RSC) using per-pair success indicators on the same 1000 pairs per dataset, with 10,000 resamples. Alternatively, run a McNemar test on per-pair accuracy. If the TUM and KITTI 95% CIs include zero, revise the claim to state that CF-RSC significantly outperforms alternatives only on wide-baseline datasets (T&T, CPC) and is comparable on short-baseline ones.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central, load-bearing claim is that the proposed Coarse-to-Fine RANSAC 'significantly outperforms other alternatives' (Sec. 5.3, Table 5). The evidence on two of the four datasets does not support that wording. On TUM, CF-RSC reaches 69.30% recall vs 69.20% for LMedS; on KITTI, 92.30% vs 91.80%. With N=1000 randomly chosen pairs, a binomial standard error for TUM is about 1.5 percentage points, so a 0.1 pp difference is negligible; KITTI's 0.5 pp difference is also within one standard error. No confidence intervals, repeated runs, or paired tests are provided anywhere in the paper. The large gains on T&T and CPC (roughly 7-10 pp over the next best estimator) are credible, but 'significantly outperforms' is stated globally, without restricting it to wide-baseline data or demonstrating statistical significance. This is not a fatal flaw: the proposed method may genuinely help on wide-baseline geometry, and on short-baseline datasets it is essentially on par with LMedS rather than worse. But the overclaim should be conditioned before the headline result is accepted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a comparative evaluation of local feature descriptors, correspondence pruning methods, and robust estimators for two-view fundamental matrix estimation. The evaluation is conducted on four datasets (TUM, KITTI, Tanks and Temples, and a Community Photo Collection) using a uniform pipeline with a SIFT baseline. The authors propose three matching systems that combine recent features with GMS pruning and LMedS fitting, and a Coarse-to-Fine RANSAC (CF-RSC) that uses GC-RANSAC for outlier pruning followed by LMedS model fitting. They report that CF-RSC significantly outperforms alternative estimators, and that the proposed matching systems perform on par with the much slower CODE system.","tokens_in":11701,"tokens_out":4979,"duration_ms":45245,"significance":"The paper addresses a timely and practical question: whether recent advances in local features, correspondence pruning, and robust estimation translate into improvements in a standard matching pipeline. The scale and breadth of the evaluation, the public release of the evaluation pipeline, and the careful definition of normalized metrics are genuine strengths. If the results are robust, the paper offers a useful practical guide and a simple, effective estimator. However, the headline claim of significant improvement for CF-RSC is currently under-supported, and the benchmark construction for two of the four datasets has a potential SIFT-related circularity. These issues need to be resolved before the paper's conclusions can be fully accepted.","major_comments":[{"comment":"The statement that CF-RSC 'significantly outperforms other alternatives' is not justified by the reported numbers on the short-baseline datasets. On TUM, CF-RSC achieves 69.30% recall versus 69.20% for LMedS (a 0.1 percentage point difference); on KITTI the difference is 92.30% versus 91.80% (0.5 percentage points). With 1000 randomly selected test pairs per dataset, the approximate binomial standard error at these recall levels is about 1.5 and 0.9 percentage points, respectively, so both differences are within sampling error. No confidence intervals, variance over repeated runs, or paired significance tests are reported anywhere in the paper. The claim should be restricted to the wide-baseline datasets (T&T and CPC) where the gains are 7–16 percentage points, or it should be accompanied by a proper statistical analysis.","section":"Section 5.3, Table 5"},{"comment":"For the T&T and CPC datasets, the test pairs are selected based on the number of SIFT inliers (more than 20), and the ground-truth fundamental matrices are derived from COLMAP reconstructions, which themselves rely on SIFT matching. This introduces a potential circularity that can systematically favor SIFT-based features and methods operating on SIFT correspondences, and it may inflate the apparent improvements of learned descriptors and pruning methods on these datasets. The authors should explicitly acknowledge this limitation and, ideally, validate the main conclusions on a subset of pairs that are selected independently of SIFT or on data with sensor ground truth. The TUM and KITTI results provide some independent support, but the large differences on T&T and CPC, which drive much of the paper's practical advice, are affected by this concern.","section":"Section 4 (Image Pairs Construction and Ground Truth)"},{"comment":"The paper does not report any measure of uncertainty for the %Recall and %Inlier numbers. Given that many comparisons in Table 2 are within 1–2 percentage points, the authors should either provide error bars, confidence intervals, or significance tests, or explicitly caution the reader against interpreting small differences as meaningful. This is particularly important for the feature comparison, where the ranking of methods on TUM and KITTI could change with a different random split.","section":"General (Tables 2, 4, 5)"}],"minor_comments":[{"comment":"The runtime table mixes measurements from two different machines (L and W), so runtimes of methods evaluated on different machines are not directly comparable. Please state this explicitly in the caption or, better, report runtimes of all methods on a single machine for the headline comparisons.","section":"Table 3"},{"comment":"The choice of the %Recall threshold (0.05) and the inlier threshold parameter α (0.003) is arbitrary. The authors note that recall curves can be used; providing recall curves for the key comparisons (e.g., CF-RSC vs. LMedS) would strengthen the evaluation.","section":"Section 3.1"},{"comment":"The abbreviation 'GC-RSC' is used in the tables but the method is introduced as GC-RANSAC in the text. Define the abbreviation in the table captions or use 'GC-RANSAC' throughout for clarity.","section":"Tables 2 and 5"},{"comment":"The claim that learned descriptors that perform better on standard benchmarks do not necessarily improve matching quality cites Balntas et al. [4]; consider also citing the HPatches benchmark [5] here, as it is already in the bibliography and directly supports this point.","section":"Section 1 (Introduction)"},{"comment":"There is a typo in the reference [22] title: 'calibrarion' should be 'calibration'. Also, 'Armanguèet al.' in Section 2 should be 'Armangué et al.' with proper accent and spacing.","section":"References and Section 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid benchmark contribution, but the headline claim about CF-RSC is oversold and the SIFT-circularity in the benchmark construction is a real concern that the authors should address. I would be comfortable with acceptance after a major revision that either provides statistical support for the claim or carefully conditions it, and that discusses the circularity and its implications for the feature ranking. I also suggest the editor ask the authors to release the exact test splits to make the benchmark reproducible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a solid, broad benchmark: the first joint evaluation of local features, correspondence pruning, and robust estimators in the classic SIFT+RANSAC+8-point pipeline, on four datasets with sensor ground truth for TUM/KITTI. Second, the headline result—the proposed Coarse-to-Fine RANSAC 'significantly outperforms' its alternatives—is overstated on short-baseline data. On TUM it beats plain LMedS by 0.1 percentage points; on KITTI by 0.5. For 1000-pair tests those differences are sampling noise, and the paper gives no confidence intervals or significance tests. The claim holds on wide-baseline T&T and CPC (7–10 pp), so the method is genuinely useful there, but the global wording should be conditioned.\n\nThe paper does real work. The normalized SGD metric is a sensible way to compare across image resolutions. The three proposed matching systems—RootSIFT-PCA+GMS+LMedS, HardNet+++GMS, and HesAffNet+HardNet+++GMS—reach CODE-level accuracy while being orders of magnitude faster, which practitioners will find useful. The evaluation is carefully organized, and the authors are upfront about their choices. The NSGD contribution is modest, but the ablation-style comparison is the main value.\n\nSoft spots, in order of seriousness. The significance overclaim is the biggest. The T&T/CPC ground truth is derived from COLMAP, which uses SIFT matching, and the test pairs are selected by SIFT inlier counts. That can bias the feature ranking on those datasets in favor of SIFT-like methods. But the TUM and KITTI results use sensor ground truth, so the estimator and pruning conclusions have independent support. A minor issue is that runtime comparisons run on different machines (Linux server for features, Windows laptop for pruning), which makes cross-method runtime claims shaky. None of these are fatal; the benchmark is still informative.\n\nWho should read this: anyone building a two-view matching pipeline or choosing features and estimators for SfM/SLAM. It deserves a serious referee; the overclaim needs to be fixed and confidence intervals added, but the experimental core is valuable.","headline":"Useful benchmark and practical matching systems, but the CF-RSC 'significantly outperforms' claim only holds on wide-baseline data.","tokens_in":12256,"tokens_out":2714,"would_cite":true,"duration_ms":25806,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The classic SIFT+RANSAC pipeline can be upgraded modularly to estimate fundamental matrices more accurately, and a two-stage Coarse-to-Fine RANSAC does best on all tested datasets.","keywords":["fundamental matrix estimation","feature matching","correspondence pruning","RANSAC","local features","evaluation benchmark","epipolar geometry","Coarse-to-Fine RANSAC"],"falsifier":"Select pairs for the same four datasets using a learned matcher's own inlier criterion rather than SIFT's, and obtain ground-truth cameras from known sensor poses or an independent reconstruction; if the recall ranking of HardNet++, GMS, LMedS, and CF-RSC changes substantially, the benchmark's SIFT-centric construction is the cause.","tokens_in":1522,"feed_emoji":"🎯","tokens_out":1473,"duration_ms":65672,"temperature":0.7,"pith_summary":"This paper evaluates whether swapping modern components into the classic two-view matching pipeline—SIFT features, ratio test, RANSAC with the 8-point algorithm—improves fundamental matrix estimation. Using four large-scale datasets and resolution-normalized metrics, it reports that learned descriptors (HardNet++, RootSIFT-PCA), correspondence pruning (GMS, LPM), and the LMedS estimator each outperform their classic counterparts in the right settings. The authors then combine the winners into three practical matching systems that reach the accuracy of the expensive CODE matcher at far lower runtime. They also propose Coarse-to-Fine RANSAC, which uses graph-cut RANSAC only to prune outliers and LMedS to fit the model, and report that it achieves the highest recall on every dataset, well above the classic baseline.","feed_headline":"Coarse-to-Fine RANSAC wins all four matching benchmarks","feed_subtitle":"Swapping in learned features, GMS pruning, then LMedS after GC-RANSAC raises recall on every dataset.","key_machinery":"The evaluation is carried by three standardized metrics: the normalized symmetric geometric distance (NSGD), which divides the symmetric epipolar distance by the image diagonal so errors are comparable across resolutions; %Recall, the fraction of estimates with NSGD below 0.05; and %Inlier, the ratio of matches lying within a normalized epipolar threshold. The proposed method's engine is a two-stage estimator: graph-cut RANSAC first supplies a clean inlier set, then LMedS fits the fundamental matrix. The practical matching systems pair DoG or HesAffNet detectors with RootSIFT-PCA or HardNet++ descriptors, prune with GMS, and finish with LMedS fitting.","core_discovery":"The central finding is that the long-standing SIFT + ratio test + RANSAC/8-point recipe is not the ceiling for fundamental matrix estimation: each pipeline stage can be swapped for a newer alternative and yield measurable gains, and the gains compound. On the evaluation's normalized metric, HardNet++ and RootSIFT-PCA consistently beat SIFT descriptors, GMS and LPM raise inlier rates while keeping enough matches, and LMedS fits better than RANSAC when inlier rates are high. Most importantly, GC-RANSAC and USAC, despite their poor standalone fitting recall on short-baseline data, are excellent outlier pruners; feeding their cleaned correspondences into LMedS gives Coarse-to-Fine RANSAC, which tops all four datasets (for instance 90.7% recall on T&T versus 70.0% baseline and 60.9% on CPC versus 29.2%).","pith_inferences":["The benchmark constructs pairs using SIFT inliers and structure-from-motion geometry; if SIFT systematically misses pairs that learned features would match, the ranking could shift when pair selection is matcher-agnostic.","CF-RSC's separation of outlier removal from model fitting might generalize to essential matrix and homography estimation, where a similar coarse/fine behavior could be tested.","The normalized error metric makes cross-dataset comparison possible; reporting recall curves instead of a single threshold would show whether the ranking is threshold-dependent."],"forward_implications":["A practitioner can choose components per scenario: DoG+HardNet++ with GMS and LMedS for general scenes, and HesAffNet+HardNet++ for wide-baseline scenes.","The two-stage idea of aggressive pruning followed by robust fitting can be dropped into any geometric estimation task that currently uses a RANSAC-family estimator.","The reported gains imply that benchmark scores of local features alone do not predict end-to-end geometry quality; the full matching-and-estimation pipeline must be evaluated.","Because CODE's accuracy is matched at several orders of magnitude lower cost, the proposed systems are practical for real-time or large-scale structure-from-motion and SLAM.","The evaluation protocol itself, with NSGD and %Recall, lets different datasets and image resolutions be compared on a single scale."],"supporting_citations":[{"why":"Supplies the SIFT detector/descriptor and the ratio test that define the classic baseline and generate initial correspondences for pair selection.","marker":"[24]"},{"why":"Defines RANSAC, the baseline robust estimator against which all other estimators are compared.","marker":"[14]"},{"why":"Provides the 8-point algorithm used to compute the fundamental matrix from correspondences in the baseline.","marker":"[18]"},{"why":"Defines the symmetric geometric distance used as the basis for the normalized evaluation metric (NSGD).","marker":"[48]"},{"why":"Provides the structure-from-motion pipeline used to derive ground-truth cameras for the datasets without known poses.","marker":"[37]"},{"why":"Supplies the LMedS estimator used both as a strong baseline and as the fitting stage of Coarse-to-Fine RANSAC.","marker":"[36]"},{"why":"Provides the graph-cut RANSAC used as the coarse outlier-pruning stage of CF-RSC.","marker":"[6]"},{"why":"Supplies the CODE matcher, the state-of-the-art pruning baseline whose accuracy the proposed systems match at lower cost.","marker":"[23]"},{"why":"Supplies GMS, the correspondence-pruning method adopted in the proposed matching systems.","marker":"[7]"}],"fun_headline_variants":["Coarse-to-Fine: prune with GC-RANSAC, fit with LMedS","Swap each SIFT stage for new methods, gains compound","Each modern swap beats SIFT baseline; combined wins all","Prune with GC-RANSAC, fit with LMedS: new top pipeline"],"cache_read_input_tokens":14336,"weakest_assumption_plain":"The result rests on the benchmark's way of choosing which image pairs count as matchable and how ground-truth cameras are obtained: if SIFT-based pair selection and structure-from-motion reconstruction bias the test set toward SIFT-like features, the reported ranking of matchers and systems would be partly an artifact of the benchmark construction rather than intrinsic quality.","fun_headline_variants_meta":{"raw":{"variants":["Coarse-to-Fine: prune with GC-RANSAC, fit with LMedS","Swap each SIFT stage for new methods, gains compound","Each modern swap beats SIFT baseline; combined wins all","Prune with GC-RANSAC, fit with LMedS: new top pipeline"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000644,"raw_usage":{"total_tokens":2953,"prompt_tokens":928,"completion_tokens":2025,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":1945}},"tokens_in":544,"tokens_out":2025,"duration_ms":15537,"temperature":1.0,"reasoning_tokens":1945,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:10:28.852610+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Select pairs for the same four datasets using a learned matcher's own inlier criterion rather than SIFT's, and obtain ground-truth cameras from known sensor poses or an independent reconstruction; if the recall ranking of HardNet++, GMS, LMedS, and CF-RSC changes substantially, the benchmark's SIFT-centric construction is the cause.","supporting_citations":[{"cited_title":"Distinctive image features from scale-invariant keypoints.International Journal on Computer Vision (IJCV), 60(2):91–110, 2004","cited_arxiv_id":null,"evidence_quote":"Supplies the SIFT detector/descriptor and the ratio test that define the classic baseline and generate initial correspondences for pair selection."},{"cited_title":"Random sample consensus: a paradigm for model ﬁtting with applications to image analysis and automated cartography","cited_arxiv_id":null,"evidence_quote":"Defines RANSAC, the baseline robust estimator against which all other estimators are compared."},{"cited_title":"In defense of the eight-point algorithm","cited_arxiv_id":null,"evidence_quote":"Provides the 8-point algorithm used to compute the fundamental matrix from correspondences in the baseline."},{"cited_title":"Determining the epipolar geometry and its uncertainty: A review","cited_arxiv_id":null,"evidence_quote":"Defines the symmetric geometric distance used as the basis for the normalized evaluation metric (NSGD)."},{"cited_title":"Structure-from-motion revisited","cited_arxiv_id":null,"evidence_quote":"Provides the structure-from-motion pipeline used to derive ground-truth cameras for the datasets without known poses."},{"cited_title":"Robust regression and outlier detection , volume 589","cited_arxiv_id":null,"evidence_quote":"Supplies the LMedS estimator used both as a strong baseline and as the fitting stage of Coarse-to-Fine RANSAC."},{"cited_title":"Graph-Cut RANSAC","cited_arxiv_id":null,"evidence_quote":"Provides the graph-cut RANSAC used as the coarse outlier-pruning stage of CF-RSC."},{"cited_title":"CODE: Coherence based decision boundaries for feature cor- respondence","cited_arxiv_id":null,"evidence_quote":"Supplies the CODE matcher, the state-of-the-art pruning baseline whose accuracy the proposed systems match at lower cost."},{"cited_title":"GMS: Grid-based motion statistics for fast, ultra-robust feature correspondence","cited_arxiv_id":null,"evidence_quote":"Supplies GMS, the correspondence-pruning method adopted in the proposed matching systems."}],"review_version":1}