{"id":"1c58851f-9da6-4cdd-9e14-c9aa355b9dc5","arxiv_id":"2502.01002","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The MultiResSAR dataset (10,850 pairs, 0.16-10 m, 4 satellites) shows all 16 tested SAR-optical registration algorithms fail on sub-meter imagery; best SR: RIFT 66.51%, XoFTR 40.58%.","lead":"This paper introduces MultiResSAR, a new public dataset of 10,850 radar and optical satellite image pairs spanning resolutions from 16 centimeters to 10 meters, and benchmarks 16 registration algorithms on it. The key finding is that no algorithm registers sub-meter pairs reliably, with the best success rates at 66.51% for RIFT and 40.58% for XoFTR.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MultiResSAR's ground truth is seeded by an unnamed automatic registration method, so the Table 7 rankings and the resolution-degradation conclusion may be circular; an independent control-point re-evaluation is needed.","rationale":"The reader's weakest assumption - that automatic registration may bias the control points and therefore skew the comparative evaluation - is the most load-bearing concern. I agree with it. The paper's headline contribution is the MultiResSAR benchmark and the observed algorithm rankings; all of those conclusions pass through the ground-truth generation step. If the reference points inherit the behavior of an unnamed automatic registration method, the rankings in Table 7 become circular for any tested algorithm that shares that method's assumptions. The concern is concrete and testable, not a matter of taste. I am not raising a different objection because the resolution-degradation claim, while currently lacking per-resolution tables in the text, is a reporting gap that can be fixed from the existing figures; the ground-truth bias, by contrast, affects the validity of every number in the benchmark. The paper does have independent support: the dataset is released, code links for all 16 methods are provided, and the qualitative figures are useful. However, the evaluation code is not released, and the automatic registration method is undisclosed, which makes the key assumption impossible to audit. The reader's CONDITIONAL verdict is appropriate: the dataset is valuable, but the benchmark conclusions should not be treated as definitive until independent control-point validation is performed or the ground-truth generation is fully disclosed and shown to be unbiased. Therefore the verdict should remain unchanged.","tokens_in":38078,"tokens_out":4877,"duration_ms":56608,"concrete_test":"Select a stratified random sample of at least 200 MultiResSAR pairs (50 per resolution class). Have two independent remote-sensing operators manually select control points from the raw, pre-alignment image products, without seeing any automatic registration output, and compute inter-operator agreement as a ground-truth noise floor. Re-run all 16 methods against this independent control-point set and compare SR and RMSE rankings to Table 7. If the ordering (RIFT and XoFTR at the top, sub-meter pairs mostly failing) reproduces within roughly 5% success rate, the circular-ground-truth concern is not material. If rankings or the resolution trend shift, the current ground truth is biased and the benchmark must be re-released with independent control points and the automatic method disclosed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing concern is the ground-truth generation described in Section 4.3. For every image pair, control points are selected 'based on the automatic registration results,' then refined by professionals to within one pixel. The automatic method is never named, and the workflow means the reference transform is a human-filtered version of an automatic alignment. If that automatic method is one of the 16 benchmarked algorithms - or is closely related to RIFT, HOWP, or another phase-consistency method - then the top success rates in Table 7 are partially self-fulfilling rather than independent measurements. Even if the automatic method is not in the benchmark, its biases can still be baked into the ground truth, skewing RMSE and success-rate comparisons across all resolutions. The one-pixel accuracy claim is especially hard to credit for the 0.16 m Umbra sub-meter pairs, where SAR speckle and geometric layover make sub-pixel manual tie-point identification implausible without independent geodetic checkpoints. Because the central claims - no algorithm reaches 100% success, performance degrades with resolution, RIFT best among traditional and XoFTR best among deep methods - all depend on this benchmark, the unnamed automatic seeding is a single point of failure for the paper's empirical conclusions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript combines a survey of SAR-optical image registration methods with the introduction of a new benchmark dataset, MultiResSAR, containing 10,850 image pairs from four SAR satellites (Sentinel-1, HT1-A, GF-3, Umbra) with resolutions from 10 m to 0.16 m and scenes covering urban, rural, plain, hill, mountain, and water areas. Sixteen registration algorithms are evaluated with four metrics: success rate (SR), number of correct matches (NCM), RMSE, and matching time (TM). The reported headline findings are that no algorithm achieves 100% success across resolutions and scenes, performance degrades as resolution increases with nearly all methods failing on sub-meter pairs, RIFT is the best traditional method (66.51% SR), and XoFTR is the best deep-learning method (40.58% SR). The paper concludes with future research directions including noise suppression, 3D geometric information fusion, cross-view transformation modeling, and deep-learning optimization.","tokens_in":38254,"tokens_out":5117,"duration_ms":50667,"significance":"The main contribution is a public, multi-source, multi-resolution, multi-scene SAR-optical registration dataset with sub-meter Umbra data, together with a broad comparison of 16 existing methods. If the ground truth is reliable, this is a genuinely useful resource for the remote-sensing registration community: it provides a stress test for both traditional feature-based methods and deep-learning matchers, the dataset and code are made available, and the empirical claims are falsifiable. The survey portion is comprehensive in coverage, though largely descriptive. The benchmark claims are conditional on the integrity of the ground-truth construction, and the manuscript currently does not provide enough independent validation to fully support the comparative rankings and the resolution-degradation conclusion.","major_comments":[{"comment":"The ground-truth construction is algorithm-in-the-loop: control points are selected \"based on the automatic registration results\" using an unnamed automatic registration method, then manually refined to within one pixel. Since the same benchmark is later used to rank 16 algorithms, including phase-consistency and self-similarity methods (RIFT, HOWP, ASS, MOSS), the rankings in Table 7 may be partly self-fulfilling if the seeding method shares the same feature and transformation assumptions as any of the benchmarked methods. Please (i) identify the automatic registration method used for seeding, (ii) state explicitly whether it is one of the 16 methods in Table 6 or a different external method, and (iii) provide an independent validation of a random subset of ground-truth control points, for example manually selected tie points without algorithm seeding or photogrammetric checkpoints, with reported residuals. Without this, the central claims about relative algorithm performance cannot be fully assessed.","section":"Section 4.3"},{"comment":"The claim that registration performance \"deteriorates as resolution increases\" and that \"almost all matches fail in sub-meter resolution image pairs\" is not supported by any resolution-stratified numerical result. Table 7 reports only aggregate SR, RMSE, NCM, and TM across the entire dataset, while Fig. 10 shows qualitative examples for the 850 Umbra pairs but no per-resolution success-rate table. Please add a breakdown by resolution class (10 m, 3 m, 1 m, 0.16 m) or by SAR source, including pair counts and the four metrics. This is load-bearing because the resolution-degradation conclusion is one of the paper's primary empirical findings.","section":"Section 5 / Table 7"},{"comment":"The evaluation metrics lack uncertainty quantification, and the success criterion is not fully pinned down. Equation (1) is typeset in a corrupted form, making the indicator function and summation ambiguous. Beyond reformatting, the paper should state whether each algorithm was run once or multiple times, and report variance or confidence intervals for SR and RMSE, especially for stochastic deep-learning matchers and for RANSAC-based methods. As presented, small differences such as XoFTR's 40.58% versus RoMa's 35.26% cannot be distinguished from run-to-run variability, and the ranking in Table 7 should not be treated as exact without such information.","section":"Section 5 / Eq. (1)-(2)"}],"minor_comments":[{"comment":"The equation for SR is garbled; please rewrite it with a clear indicator function I(p_i), the threshold N_min, and a summation over image pairs.","section":"Section 5, Eq. (1)"},{"comment":"The XoFTR reference (Ö, T., Köksal, A., et al., 2024) has an incomplete author name; it should read Tuzcuoglu, O., Köksal, A., et al.","section":"References"},{"comment":"The HT1-A satellite appears in Table 5 but is not introduced in Section 1 or Section 4; please add a brief description of its band, resolution, and operating characteristics.","section":"Section 4.4 / Table 5"},{"comment":"The ultra-high-resolution example figures would benefit from scale bars and explicit chip sizes or geographic extents, since the text claims 0.16 m resolution but the figures do not show any scale information.","section":"Figures 4 and 10"},{"comment":"The text states that \"only the RoMa algorithm achieved correct registration results in the four image groups,\" but the relationship between these four groups and the 850-pair Umbra subset is unclear; please clarify the denominator and report the corresponding success counts.","section":"Section 5, Fig. 10"}],"recommendation":"major_revision","confidential_remarks":"The paper's contribution is primarily a dataset and benchmark, so the ground-truth generation protocol is the single most important methodological element. The unnamed automatic seeding of control points is a specific, fixable concern rather than a reason to reject, but the authors should be required to disclose the seeding method, verify independence from the benchmarked algorithms, and provide per-resolution tables with uncertainty information before the empirical claims can be accepted. The survey content is useful but does not by itself carry the paper's headline conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the MultiResSAR dataset is the real contribution: no prior public dataset spans 0.16 m to 10 m SAR-optical pairs from four satellites, and the paper puts 10,850 pairs plus a 16-algorithm benchmark online. Second, the empirical claims—no algorithm reaches 100% success, and sub-meter registration essentially fails—are probably true in direction, but the ground-truth protocol makes the precise numbers and rankings untrustworthy as published.\n\nWhat the paper does well: the dataset construction is careful in coverage (four sensors, six scene types, global distribution), the experimental setup uses author-released code with recommended settings, and the results, even if biased, are likely to be reproduced by others. The review portion is a competent survey, not novel, but useful as a point of entry. The benchmark result that RIFT leads traditional methods and XoFTR leads deep methods is consistent with prior literature, which adds credibility.\n\nThe soft spots are real. Section 4.3 says control points are selected 'based on the automatic registration results' and then refined by professionals. The automatic method is never named. If it is RIFT or a close relative—and RIFT is the top scorer in Table 7—the rankings become partially self-fulfilling. Even if it is not in the benchmark, the biases of that method are baked into the ground truth. This is a single point of failure for the evaluation. Also, there is no per-resolution breakdown of SR/RMSE, no error bars, and the sub-meter manual verification to one pixel is hard to credit given SAR speckle and layover. These are fixable: name the seeding method, re-verify a random sample with independent tie points, release the evaluation code, and provide results by resolution bin.\n\nWho gets value from this: researchers working on SAR-optical registration, especially those needing a multi-resolution benchmark with sub-meter pairs. For that audience the dataset is worth engaging with now, but the numbers should be treated as preliminary until the ground-truth chain is clarified.\n\nMy recommendation: accept for peer review with a request for major revision. The dataset and the empirical direction are valuable; the ground-truth protocol and the lack of per-resolution reporting need to be fixed before the rankings are cited as definitive.","headline":"MultiResSAR is a genuinely useful dataset, but the benchmark's ground truth is seeded by an unnamed automatic method, so the headline rankings should not be taken as final until that is addressed.","tokens_in":38854,"tokens_out":2302,"would_cite":true,"duration_ms":23279,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"On a new 10,850-pair benchmark spanning 0.16 to 10 meter resolution, none of 16 registration algorithms succeeds across all resolutions and scenes, and almost all fail on sub-meter SAR-optical pairs.","keywords":["SAR-optical image registration","MultiResSAR dataset","multi-resolution benchmark","sub-meter SAR","image registration evaluation","remote sensing data fusion","RIFT","XoFTR"],"falsifier":"Independently re-annotate a random stratified sample of MultiResSAR pairs with fresh control points chosen by different operators, or validate them against geodetic ground control, then recompute all 16 methods' success rates; if RIFT's lead over XoFTR, or the near-total failure on 0.16-meter pairs, changes materially under that independent ground truth, the paper's central empirical claim is an artifact of its annotation process.","tokens_in":37843,"feed_emoji":"🛰️","tokens_out":4563,"duration_ms":43592,"temperature":0.7,"pith_summary":"This paper argues that the hard open problem in SAR-optical image registration is no longer modality difference alone but resolution, especially sub-meter data. The authors build and release MultiResSAR, a public dataset of 10,850 multi-source, multi-resolution, multi-scene SAR-optical pairs, and benchmark 16 state-of-the-art algorithms on it. They find that no algorithm achieves 100% success, that performance drops sharply as resolution increases, and that nearly all matching fails on 0.16-meter Umbra imagery. On this benchmark the best traditional method, RIFT, reaches 66.51% success, while the best deep learning method, XoFTR, reaches 40.58%. The paper concludes that future progress depends on noise suppression, 3D geometric fusion, cross-view transformation modeling, and deep learning optimization rather than incremental descriptor tweaks.","feed_headline":"Best SAR-optical matcher succeeds on only two-thirds of pairs","feed_subtitle":"A 10,850-pair benchmark from 0.16 to 10 meters shows performance collapsing at sub-meter resolution; deep learning tops out at 40.6 percent.","key_machinery":"The load-bearing object is the MultiResSAR dataset: 10,850 SAR-optical pairs built from four SAR satellites (Sentinel-1 at 10 meters, HT1-A at 3 meters, GF-3 at 1 meter, and Umbra at 0.16 meters) with optical images from Google Earth and six scene types including urban, rural, plains, hills, mountains, and water. The argument runs through this dataset because its resolution spread and source diversity are what expose the resolution-dependent collapse. The evaluation uses four metrics: Success Rate (a pair counts as successful when it yields at least 20 correct matches, with a root-mean-square error at most 10 pixels), Number of Correct Matches, RMSE, and matching time; ground truth is produced by automatic registration followed by manual visual inspection in which professionals select control points and keep error within one pixel.","core_discovery":"The central result is an empirical measurement: registration accuracy on SAR-optical image pairs degrades with spatial resolution, and no current method generalizes across sources, scenes, and resolutions. On the full MultiResSAR dataset, RIFT has the highest success rate at 66.51%, with HOWP at 52.63% and ASS at 39.34% among traditional methods; among deep learning methods XoFTR leads at 40.58%, followed by XFeat at 36.29% and RoMa at 35.26%. Most algorithms fall below 50%. On the 850 ultra-high-resolution Umbra pairs at 0.16 meters, almost every method's matches fail, and only RoMa produces correct registration results in four image groups, with few and unevenly distributed points. The authors present MultiResSAR as the first public benchmark combining multiple satellites, resolutions from 0.16 to 10 meters, and diverse scenes, positioned to fill the gap left by existing datasets that are single-resolution or lack accurate ground truth.","pith_inferences":["If sub-meter failure stems from speckle noise entangled with fine structure and 3D layover effects, then approaches that jointly despeckle and register, or that explicitly model SAR imaging geometry, may succeed where descriptor-based and transformer matchers fail; this is a testable extension the paper does not run.","The benchmark suggests that conclusions drawn from 10-meter-only datasets such as SEN1-2 may not transfer to modern high-resolution satellites; resolution should be treated as a covariate in future dataset design.","Because the ground truth was built from automatic registration plus manual inspection, the reported rankings are only as trustworthy as that one-pixel error claim; an independent geodetic validation of a subset would materially strengthen the benchmark's conclusions.","Combining phase-congruency features (the strength of RIFT) with learned refinement and sub-pixel matching (the strength of RoMa and XoFTR) is a natural next architecture to test on the sub-meter subset."],"forward_implications":["A public multi-resolution benchmark now exists on which future SAR-optical registration claims can be tested fairly, including sub-meter data that was previously missing.","Resolution is a first-order driver of failure: methods that appear strong on 10-meter or 1-meter data cannot be assumed to work on sub-meter imagery.","The best traditional method (RIFT at 66.51%) outperforms the best deep learning method (XoFTR at 40.58%) on this benchmark, so the deep learning advantage seen on same-modality matching does not automatically transfer to SAR-optical registration.","The stated research agenda follows directly: suppress speckle noise, fuse 3D geometric information, model cross-view transformations, and optimize deep learning architectures for high-resolution multi-modal data."],"supporting_citations":[{"why":"Provides RIFT, the best-performing traditional method on MultiResSAR and the phase-consistency baseline the benchmark's top traditional result rests on.","marker":"Li et al. 2020"},{"why":"Provides XoFTR, the best-performing deep learning method on the benchmark, setting the deep learning reference point.","marker":"Tuzcuoğlu et al. 2024"},{"why":"Provides RoMa, the only method that registered any of the ultra-high-resolution sub-meter image pairs.","marker":"Edstedt et al. 2024"},{"why":"SARptical, a prior SAR-optical dataset with high-resolution imagery but no multi-resolution or sub-meter coverage, used to frame the dataset gap.","marker":"Wang and Zhu 2018"},{"why":"SEN1-2, a 10-meter-resolution benchmark dataset, used to show that existing large-scale data lacks resolution variation.","marker":"Schmitt et al. 2018"},{"why":"QXS-SAROPT, a 1-meter Gaofen-3 and Google Earth dataset, used to illustrate that prior public data does not reach sub-meter resolution.","marker":"Huang et al. 2021"},{"why":"OSEval, an existing sub-meter optical-SAR evaluation set, used as the closest prior evaluation platform and a point of comparison for MultiResSAR.","marker":"Xiang et al. 2023"},{"why":"SOPatch, a large-scale patch dataset and deep descriptor, used to represent the state of deep feature learning for SAR-optical matching before the MultiResSAR benchmark.","marker":"Xu et al. 2023"}],"fun_headline_variants":["SAR-optical registration: no method tops 66.5% success","Sub-meter SAR-optical pairs stump all matchers","New benchmark shows resolution kills SAR-optical alignment","Deep learning lags traditional on SAR-optical registration","MultiResSAR: 10k pairs, best success only 66.5%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"All benchmark conclusions depend on the MultiResSAR ground truth being unbiased and accurate to within one pixel; if the automatic registration used to seed control points systematically favors certain algorithms, the success-rate rankings and the sub-meter failure finding would be skewed.","fun_headline_variants_meta":{"raw":{"variants":["SAR-optical registration: no method tops 66.5% success","Sub-meter SAR-optical pairs stump all matchers","New benchmark shows resolution kills SAR-optical alignment","Deep learning lags traditional on SAR-optical registration","MultiResSAR: 10k pairs, best success only 66.5%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000213,"raw_usage":{"total_tokens":1455,"prompt_tokens":1011,"completion_tokens":444,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":627,"completion_tokens_details":{"reasoning_tokens":372}},"tokens_in":627,"tokens_out":444,"duration_ms":4192,"temperature":1.0,"reasoning_tokens":372,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T16:53:57.511589+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independently re-annotate a random stratified sample of MultiResSAR pairs with fresh control points chosen by different operators, or validate them against geodetic ground control, then recompute all 16 methods' success rates; if RIFT's lead over XoFTR, or the near-total failure on 0.16-meter pairs, changes materially under that independent ground truth, the paper's central empirical claim is an artifact of its annotation process.","supporting_citations":[],"review_version":1}