{"id":"56989849-39a0-4915-a6a0-a64bbab4651b","arxiv_id":"2608.08685","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A 9-parameter calibrated Laplace mixture that models coarse-assignment failures as a heavy tail improves downstream geometric estimation when used as soft weights in a posterior refit.","lead":"This paper proposes a lightweight post-hoc uncertainty model for semi-dense image matching that combines local refinement noise with coarse-assignment failure tails, and uses it to reweight correspondences in a final geometric refit. It shows consistent accuracy gains in homography estimation and visual localization across six matchers and five robust estimators, with only 9 calibration parameters.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CoRe's posterior weights inherit the initial geometry's error (Eq. 8); Supp B.1 shows instability above 30 px, so the claimed consistency needs a scope caveat.","rationale":"I read the paper as claiming that modeling coarse-assignment failures, not just local confidence, yields downstream geometric gains through CoRe. The experiments are extensive and the ablation studies support the mixture model as a reasonable fit. The weakest point is not the mixture form itself but the transfer of a model calibrated on matching residuals to residuals under theta0: Eq. 8 introduces geometry error into the input to Eq. 10, and Supp B.1 confirms a degradation band. This does not refute the paper, but it makes the headline claim conditional on initial-geometry quality. The reader already identified this exact issue and assigned CONDITIONAL, so I recommend no change to the verdict. The absence of error bars and the 1 px drops on advanced estimators are real but secondary; the theta0 dependence is the load-bearing concern.","tokens_in":18032,"tokens_out":7057,"duration_ms":87068,"concrete_test":"Re-run the HPatches evaluation with controlled perturbations of the initial homography: take the RANSAC output H0, compose it with known transforms simulating initial ACEs of 0, 5, 15, 30, 60, and 120 px, and run CoRe from each perturbed H0. Report AUC@3px and the fraction of improved pairs with bootstrap confidence intervals. If gains vanish or become negative for ACE above 30 px, the 'consistently improves' claim must be scoped to initial-geometry quality; if gains persist, the concern is refuted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"CoRe's central mechanism (Eq. 10) replaces the matching residual r_i = yhat_i - y_i with the reprojection residual r^(0)_i = yhat_i - Pi_theta0(x_i) under the initial robust estimate (Eq. 8). The fine/coarse Laplace likelihoods in Eq. 10 are calibrated to matching residuals, not to residuals contaminated by geometry error. When theta0 is inaccurate, correct correspondences can have large r^(0), so their coarse-success posterior is suppressed and the refit can downweight exactly the matches needed to correct theta0. The paper's own Supp B.1 (Table 6) shows this: for initial ACE e0 < 1 px, CoRe is neutral-to-negative (47.9% improved, median delta e = -0.004 px), and for e0 > 30 px behavior becomes unstable, with only 58.3-66.7% improved in small samples. Supp Fig. 8 validates posterior quality in aggregate but does not condition on e0, which is exactly the regime where theta0 contamination is largest. Because the headline claim is 'consistently improves' across six matchers and five estimators, and the aggregate tables may be dominated by easy and moderate pairs, the failure regime is a genuine scope limitation rather than a purely cosmetic caveat. The central claim is therefore conditional on the initial estimate lying in a favorable error band.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses uncertainty quantification for coarse-to-fine semi-dense image matchers. It proposes a post-hoc calibrated two-component Laplace mixture model with nine learnable parameters: one component captures local refinement noise and the other captures the heavier tail of coarse-assignment failures. A learned gate combines coarse and fine uncertainty cues into a mixture likelihood. The authors then introduce CoRe (Coarse-success posterior Refit), a final geometric refitting step that uses the posterior probability of coarse-assignment success, computed under an initial robust estimate, as soft weights over all correspondences. The method is evaluated on homography estimation (HPatches, MTV) and visual localization (Aachen v1.1) across six pretrained matchers and five robust estimators, with parameters calibrated on MegaDepth and transferred zero-shot. The main claim is that CoRe consistently improves downstream geometric accuracy with minimal computational overhead.","tokens_in":18335,"tokens_out":7301,"duration_ms":74197,"significance":"If the claims hold, the contribution is practically valuable: it requires no retraining, adds only nine calibration parameters, applies to multiple off-the-shelf matchers and robust estimators, and is accompanied by released code and a broad experimental study. The paper also contains useful ablations comparing different weighting strategies, robustness across estimators, and a supplementary sensitivity analysis that is more informative than is typical. The external evaluation on datasets not used for calibration (HPatches, MTV, Aachen) and the explicit zero-shot transfer are strengths, as is the careful distinction between fine-only uncertainty and the proposed mixture. However, the central 'consistent improvement' claim is stronger than the evidence supports, particularly with respect to dependence on the initial geometric estimate and to several reported decreases in Table 2 and Table 4. The method is defensible as a lightweight way to improve geometry in the moderate-error regime, but the paper needs to scope the claim and address the failure regimes it already documents.","major_comments":[{"comment":"The claim that CoRe 'consistently improves' geometric estimation is not supported over the full range of initial geometry quality. The coarse-success posterior in Eq. (10) is evaluated on the reprojection residual r^(0)_i = yhat_i − Π_{θ0}(x_i), so when the initial θ0 is inaccurate, the residual is dominated by geometry error rather than by matching error; correct correspondences can then be down-weighted, exactly when they are needed to correct θ0. The paper's own Table 6 shows that for e0 < 1 px, CoRe improves only 47.9% of pairs with a median change of −0.004 px, and for e0 > 30 px the improvement frequency is 58.3–100% in bins containing only 3–12 pairs, i.e., unstable. Figure 7 confirms the same pattern across matchers. The authors should report the distribution of initial errors e0 in the evaluation datasets, condition the headline results on e0 or on difficulty, and either add a safeguard that retains the initial model when the posterior is unreliable or revise the universal improvement claim.","section":"§3.3 (Eqs. 8–10); Supp B.1 (Table 6, Fig. 7)"},{"comment":"The abstract's statement that the method 'consistently improves downstream geometric accuracy across various pretrained-only matchers and robust estimators' is contradicted by several reported cells. In Table 2, GIM-LoFTR Night at (0.25 m, 2°) decreases from 68.6 to 67.0, EfficientLoFTR Night decreases from 73.3 to 72.8, and CoMatch Day decreases from 85.2 to 85.1. In Table 4, the 1 px AUC decreases for LO-RANSAC, PROSAC, and GC-RANSAC. The main text correctly says 'in most settings' for localization, but the abstract and contribution (iii) should be reworded to reflect improvements at moderate thresholds and difficulty levels rather than universal consistency, or the authors should provide error bars and statistical tests supporting the aggregate claim.","section":"Abstract and Tables 1–4"},{"comment":"Equation (3) collapses the marginalization over all incorrect coarse cells into a single coarse-error component p_c_i(r). This is a strong approximation: the residual distribution after a wrong coarse assignment depends on which coarse cell was selected, so the two-component mixture is not an exact generative model. Consequently, the weight ω_i in Eq. (10) should be described as a calibrated heuristic rather than an exact posterior probability of coarse-assignment success. Because CoRe's usefulness depends on this posterior tracking true coarse success, the current validation in Supp Fig. 8 is only aggregate and should be augmented by conditioning on the initial error e0; otherwise the failure regime identified above also undermines the posterior interpretation.","section":"§3.2, Eq. (3)"}],"minor_comments":[{"comment":"The corner error value '372445508502019.8 -> 296482922247990.4 px' appears to be a typo or numerical overflow; please correct it.","section":"Fig. 3"},{"comment":"Calibration uses ground-truth correspondences y_i on MegaDepth, but the paper does not state how these correspondences are computed (e.g., depth and pose warping, pseudo-ground-truth, or manual labeling); this information is needed for reproducibility.","section":"Algorithm 1 and Sec. 4.1"},{"comment":"The square-root notation over the learned multipliers a and b is confusing; please clarify whether the learnable parameters are the spatial scales themselves or their squares.","section":"Eq. (5)"},{"comment":"The mapping from the main-text notation Θ = {a, b, w, β} to the nine parameters {a_x, a_y, b_x, b_y, k_s, k_m, t_x, t_y, t_m} should appear in the main paper, since the main text never defines the dimensions of w and β.","section":"Supp. A and Sec. 3.2"},{"comment":"The baseline weighting strategies 'mconf refit' and 'Raw fine std' should be defined in the caption or in the text, because their exact construction is not otherwise specified.","section":"Table 3"},{"comment":"There are minor text issues: 'solely only captures' is redundant, and the code URL in the abstract contains an unwanted line break ('Probabilistic\\n-matching').","section":"Sec. 3.1 and Abstract"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern raised in the reader's report does land: the supplementary's own Table 6 and Fig. 7 show that CoRe's behavior is neutral for very easy pairs and unstable for very hard pairs, which directly qualifies the headline 'consistent improvement' claim. The paper is otherwise well designed, with broad external evaluation and a calibration protocol that is not circular. I recommend major revision rather than rejection because the central mechanism is simple and defensible once properly scoped, and the required changes are to claims and to the presentation of the sensitivity analysis rather than to the core experiments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: the paper proposes a post-hoc, 9-parameter calibrated two-component Laplace mixture for semi-dense matcher uncertainty, and a coarse-success posterior refit (CoRe) that weights correspondences for geometric estimation. The core claim—that this improves homography and pose estimation across six frozen matchers and five robust estimators—holds in the aggregate, with gains that are often large and compute overhead that is genuinely small.\n\nWhat is new: nobody else has offered a calibration stage that explicitly models coarse-assignment failures as a separate tail and then uses the posterior of coarse success as soft weights in a refit. The closest work, SURE, only models local uncertainty. The experiments are broad and the calibration is clean: fitted on MegaDepth, transferred zero-shot to HPatches, MTV, and Aachen. The ablation against uniform, matching-confidence, and binary weighting is convincing, and the support analysis shows the posterior tracks real coarse accuracy.\n\nThe soft spot is the one the stress-test identified. CoRe computes residuals under the initial geometry theta0 (Eq. 8). When theta0 is poor, the residuals are dominated by geometry error, so the posterior no longer measures coarse-assignment success. The paper's own Table 6 shows exactly this: for e0 < 1px, only 47.9% of pairs improve; for e0 > 30px, behavior is unstable. That means the abstract's 'consistently improves' is too strong. The accurate claim is 'improves reliably when the initial estimate is in a moderate-error band, roughly 1-30px, and is neutral or unstable outside it.' This is a real scope limitation and it belongs in the main text, not only the supplement. That said, the aggregate tables are not cherry-picked—most real pairs fall in the favorable band—so the method itself stands.\n\nMinor issues: no error bars on the main tables; one qualitative MTV case in Fig. 3 reports a corner error of 3.7e14 px, which looks like a degenerate homography or a reporting bug; no comparison against SURE; and Eq. (3) assumes all coarse failures collapse into one component, which is a practical approximation but deserves one sentence of justification.\n\nWho this is for: anyone building geometric estimation pipelines with semi-dense matchers, and anyone working on uncertainty calibration for correspondence. It deserves a serious referee. I would accept it for review and ask for the scope claim to be fixed, the broken example to be repaired, and a small sensitivity experiment in the main paper.","headline":"A useful plug-in for semi-dense matching uncertainty, with a real scope caveat the paper itself buries in the supplement.","tokens_in":18865,"tokens_out":3178,"would_cite":true,"duration_ms":31458,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that semi-dense matching uncertainty is not just local confidence: a 9-parameter post-hoc Laplace mixture that models coarse-assignment failures improves homography and pose estimation across six pretrained matchers.","keywords":["semi-dense matching","uncertainty estimation","coarse-to-fine matching","Laplace mixture model","geometric refit","robust estimation","homography estimation","visual localization"],"falsifier":"Take an image pair with a known homography, deliberately perturb the initial model by an increasing amount, and check whether the CoRe posterior still ranks matches by whether their coarse guess was correct; if the ranking degrades to chance as the perturbation grows past roughly 30 px, the weights are encoding geometry error rather than matching error, contradicting the central claim in that regime.","tokens_in":17839,"feed_emoji":"🎯","tokens_out":8304,"duration_ms":80466,"temperature":0.7,"pith_summary":"Coarse-to-fine semi-dense matchers produce two very different kinds of error: small local refinement noise and occasional catastrophic failures where the coarse search picks the wrong region entirely. The paper argues that standard uncertainty estimates, which only look at local refinement confidence, miss the second kind and are therefore overconfident on hard pairs. It proposes a post-hoc, calibrated two-component Laplace mixture, fit with only 9 parameters and no retraining, that captures both error regimes, and a refit step called CoRe that uses the posterior probability of coarse-assignment success as soft weights in the final geometric estimate. Across six pretrained matchers and five robust estimators, this consistently improves homography and pose accuracy at modest computational cost, especially on moderately hard pairs.","feed_headline":"Refitting with coarse-failure odds lifts matching accuracy","feed_subtitle":"A 9-parameter mixture turns rejected matches into soft weights that sharpen six matchers.","key_machinery":"The load-bearing object is the calibrated two-component Laplace mixture $\\hat p_i(r) = (1-\\alpha_i)\\mathrm{Lap}(r|0,s_i^f) + \\alpha_i\\mathrm{Lap}(r|0,s_i^c)$, where $s_i^f$ and $s_i^c$ are per-axis scales calibrated from the matcher's fine- and coarse-level outputs and $\\alpha_i = \\sigma(w^\\top\\phi_i + \\beta)$ is a sigmoid gate over coarse cues. Bayes' rule turns this likelihood, evaluated at the residual under an initial robust estimate $\\theta_0$, into the coarse-success posterior $\\omega_i$ used as a soft correspondence weight in one weighted geometric refit $\\theta^\\star = \\arg\\min_\\theta \\sum_i \\omega_i \\rho(x_i,\\hat y_i;\\theta)$. The mixture is what lets the method express both a sharp inlier peak and a heavy failure tail; the posterior refit is what converts that model into better geometry.","core_discovery":"The paper's central claim is that the overall error of a semi-dense coarse-to-fine matcher is a two-source mixture, and that current uncertainty estimates capture only one source. Concretely, it models the residual $\\hat y_i - y_i$ as $(1-\\alpha_i)\\mathrm{Lap}(r|0,s_i^f) + \\alpha_i\\mathrm{Lap}(r|0,s_i^c)$: the fine component reflects local refinement noise when the coarse assignment is correct, and the coarse component absorbs the heavy tail of matches where the coarse search failed entirely. The paper then derives, via Bayes' rule, the posterior probability that the coarse assignment succeeded given the residual under an initial robust model, and uses that posterior as a soft weight in a single final geometric refit (CoRe). It reports that this refit improves homography AUC and pose accuracy across six pretrained matchers and five robust estimators, with parameters calibrated once on MegaDepth and transferred zero-shot.","pith_inferences":["Beyond the paper's experiments, the same coarse-success posterior could be used inside iterative estimation loops: re-estimate the model, recompute residuals, recompute weights, and refit again, rather than only performing a single final refit.","Because the mixture is fit post-hoc from cues the matcher already outputs, the posterior could also serve as a soft inlier prior for training robust estimators end-to-end, which the paper does not explore.","The axis-factorized Laplace treats horizontal and vertical residuals independently; a correlated error model might behave differently on slanted or perspective-distorted surfaces, a testable variation not covered in the paper."],"forward_implications":["Any coarse-to-fine matcher can improve downstream homography and pose accuracy by adding the calibrated mixture and CoRe, with no retraining and only tens of milliseconds of overhead.","Soft weighting with the coarse-success posterior beats hard inlier/outlier masking at 3px and looser thresholds, so standard RANSAC final steps discard useful residual information.","Fine-only uncertainty is bounded by the refinement window and cannot express large errors, so modeling the coarse-failure tail is necessary for calibrated uncertainty on hard pairs.","The 9 calibration parameters transfer zero-shot across datasets and resolutions, making calibration a one-time lightweight step.","Gains concentrate on moderately hard pairs with initial error between 1 and 30 pixels, while near-perfect fits may lose a little at the strict 1px threshold."],"supporting_citations":[{"why":"Defines the coarse-to-fine dual-softmax matching paradigm that the paper's error model is built on and is one of the evaluated matchers.","marker":"[37]"},{"why":"EfficientLoFTR is the matcher used for the main ablations, including refit weighting strategies and robust-estimator comparisons.","marker":"[44]"},{"why":"EDM is one of the six evaluated matchers and supplies an axis-wise fine refinement uncertainty used as a cue in the mixture.","marker":"[21]"},{"why":"CoMatch is one of the six evaluated matchers, providing two-stage correlation refinement cues for the calibrated model.","marker":"[23]"},{"why":"RANSAC is the initial robust estimator whose hard inlier threshold and discarded residuals CoRe replaces with a soft-weighted refit.","marker":"[11]"},{"why":"HPatches is the main homography benchmark used to measure AUC improvements across matchers and estimators.","marker":"[2]"},{"why":"MegaDepth supplies the held-out calibration pairs and the resolution-transfer evaluation set.","marker":"[22]"},{"why":"MAGSAC++ is one of the advanced robust estimators CoRe is applied to, showing gains across estimator choices.","marker":"[5]"},{"why":"LO-RANSAC is one of the advanced robust estimators CoRe is applied to, demonstrating compatibility with locally optimized solvers.","marker":"[8]"}],"fun_headline_variants":["9 parameters fix coarse-failure blind spots in matching","Two-source uncertainty model sharpens six matchers","CoRe refit: 9 parameters turn coarse failures into soft weights","Post-hoc 9-parameter mixture reveals hidden coarse-failure tail","Improve matcher accuracy by weighting with coarse-success odds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The refit assumes that residuals under the initial geometry estimate are close to the true matching errors, so when the initial geometry is wrong by tens of pixels the posterior weights stop measuring matching success; the paper's own supplementary analysis shows unstable behavior above roughly 30 px.","fun_headline_variants_meta":{"raw":{"variants":["9 parameters fix coarse-failure blind spots in matching","Two-source uncertainty model sharpens six matchers","CoRe refit: 9 parameters turn coarse failures into soft weights","Post-hoc 9-parameter mixture reveals hidden coarse-failure tail","Improve matcher accuracy by weighting with coarse-success odds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000775,"raw_usage":{"total_tokens":3410,"prompt_tokens":910,"completion_tokens":2500,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":2417}},"tokens_in":526,"tokens_out":2500,"duration_ms":17275,"temperature":1.0,"reasoning_tokens":2417,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:27:01.335719+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take an image pair with a known homography, deliberately perturb the initial model by an increasing amount, and check whether the CoRe posterior still ranks matches by whether their coarse guess was correct; if the ranking degrades to chance as the perturbation grows past roughly 30 px, the weights are encoding geometry error rather than matching error, contradicting the central claim in that regime.","supporting_citations":[{"cited_title":"LoFTR: Detector-free local feature matching with transformers.CVPR, 2021","cited_arxiv_id":null,"evidence_quote":"Defines the coarse-to-fine dual-softmax matching paradigm that the paper's error model is built on and is one of the evaluated matchers."},{"cited_title":"Efficient LoFTR: Semi-dense local feature matching with sparse-like speed","cited_arxiv_id":null,"evidence_quote":"EfficientLoFTR is the matcher used for the main ablations, including refit weighting strategies and robust-estimator comparisons."},{"cited_title":"EDM: efficient deep fea- ture matching","cited_arxiv_id":null,"evidence_quote":"EDM is one of the six evaluated matchers and supplies an axis-wise fine refinement uncertainty used as a cue in the mixture."},{"cited_title":"Comatch: Dynamic covisibility-aware transformer for bilateral subpixel-level semi-dense image matching","cited_arxiv_id":null,"evidence_quote":"CoMatch is one of the six evaluated matchers, providing two-stage correlation refinement cues for the calibrated model."},{"cited_title":"Fischler and Robert C","cited_arxiv_id":null,"evidence_quote":"RANSAC is the initial robust estimator whose hard inlier threshold and discarded residuals CoRe replaces with a soft-weighted refit."},{"cited_title":"Hpatches: A benchmark and evaluation of handcrafted and learned local descriptors","cited_arxiv_id":null,"evidence_quote":"HPatches is the main homography benchmark used to measure AUC improvements across matchers and estimators."},{"cited_title":"Megadepth: Learning single- view depth prediction from internet photos","cited_arxiv_id":null,"evidence_quote":"MegaDepth supplies the held-out calibration pairs and the resolution-transfer evaluation set."},{"cited_title":"Magsac++, a fast, reliable and accurate robust estima- tor","cited_arxiv_id":null,"evidence_quote":"MAGSAC++ is one of the advanced robust estimators CoRe is applied to, showing gains across estimator choices."},{"cited_title":"Locally opti- mized RANSAC","cited_arxiv_id":null,"evidence_quote":"LO-RANSAC is one of the advanced robust estimators CoRe is applied to, demonstrating compatibility with locally optimized solvers."}],"review_version":1}