REVIEW 3 major objections 5 minor 59 references
LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper claims that dense matching in global-scale remote sensing should be reformulated as localization followed by registration, and that LoRetta, built on this reformulation, reaches 83.3% AUC on a new global benchmark.
desk verdict A substantial new benchmark and a clean two-stage matcher, with a real but addressable evaluation-circularity concern that deserves referee time rather than desk rejection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the matchability-weighted affine prior. From coarse soft-assignment correspondences and a predicted matchability map, the model solves $A^\star = \arg\min_{A\in\mathbb{R}^{2\times3}} \sum_{i\in\Omega_A} m_i \lVert q_i - A\tilde{p}_i \rVert_2^2$, a least-squares fit over correspondences weighted by predicted matchability. That single affine transform is what separates reliable overlap from non-overlap and turns global search into local refinement: it warps the reference image and coarse matchability into the sensed frame, initializes the residual displacement at zero, and defines where the dense registration branch is trusted. The registration branch then predicts only the residual field at progressively finer scales, and the final warp is the affine map composed with the residual. The same matchability concept doubles as supervision: pseudo-matchability labels are generated by template matching with geometric verification, thresholded to a binary verified-matchability indicator, and the evaluation metrics are computed only on pixels the pseudo-label marks as fully matchable.
What would settle it
Recompute the LEVIR-GM AUC and PCK comparisons on a human-annotated matchability mask: have annotators mark confident correspondence regions on a sample of the same test pairs, then score LoRetta and the strongest prior dense matcher only on those confirmed pixels. If the 1.6-point AUC and 6.5/8.2-point PCK advantages shrink or disappear, the headline result depends on the pseudo-matchability mask rather than on intrinsic matching quality; a second check would be to train the strongest baseline with the same matchability-weighted losses and thresholded mask and see whether the gap persists.
Extended reading notes
Core claim
The central claim is that the localization-and-registration decomposition, not a more powerful direct flow regressor, is what unlocks dense matching under global, multi-temporal, cross-resolution conditions. LoRetta predicts coarse correspondences and a matchability map, fits a global affine prior by matchability-weighted least squares, warps the reference image into the sensed frame, and then estimates multi-scale residual displacement fields and a final matchability map. The paper shows this design is not decorative: removing the affine localization drops AUC from 83.3% to 38.2%, removing the localization guidance drops it to 73.4%, and replacing matchability-weighted sampling with uniform sampling drops it to 80.9%. The claim is that these numbers, together with the land-cover-wise and downstream localization results, establish matchability-aware affine localization as the correct inductive bias for this regime.
Load-bearing premise
The whole comparison is scored only on pixels that the paper's automatic template-matching procedure labels as reliably matchable, and LoRetta was trained on those same labels; if that procedure is not a fair ground truth, the reported gains over other methods may come from fitting the label rather than from better matching.
Editorial extensions
If this is right
- If the claim holds, global-scale dense matching becomes practical for time-sensitive pipelines: LoRetta reports 64.8 ms per 512x512 pair, roughly 1.9 times faster than the strongest prior dense matcher at 124.1 ms.
- The LEVIR-GM benchmark would give the community a common testbed where sparse, semi-dense, and dense matchers are scored as dense registration systems through the same random-sample-consensus plus thin-plate-spline protocol, making comparisons meaningful across output densities.
- Matchability labels would allow other matchers to be trained to refuse non-overlapping, cloudy, or changed pixels rather than being forced to produce a correspondence everywhere.
- LoRetta's downstream transfer results imply that a single aligner trained on satellite pairs can serve as a reusable geometric component in astronaut-to-satellite and UAV-to-satellite localization, raising the astronaut localization success rate from 93.0% to 97.5%.
- The ablation results imply that skipping the affine localization stage is catastrophic (AUC drops from 83.3% to 38.2%), so existing direct-regression matchers may need an explicit global geometry stage rather than more capacity.
Reading between the lines
- My inference: the same affine-localization then residual-registration decomposition should transfer to cross-modal pairs such as optical and synthetic-aperture radar imagery, where global offsets and local appearance mismatch are even more extreme; the paper's geometric argument is not tied to optical-optical sensors.
- My inference: because the pseudo-matchability labels are produced by classical template-matching agreement, the benchmark may systematically mark scenes with strong seasonal appearance change but stable geometry as unmatchable, which would make the evaluation conservative for all methods; a human audit on such scenes is a natural extension.
- My inference: the matchability-weighted affine fitter could be reused as a cheap initialization for other alignment problems, such as video stabilization or multi-view satellite reconstruction, where a global similarity or affine estimate is available before local refinement.
- My inference: since LoRetta's gains concentrate at 1-2 pixel thresholds while 5-10 pixel thresholds are near saturation, the practical relevance of the model depends on the target task's required accuracy, with change detection and high-precision cartography benefiting most.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper addresses global-scale dense image matching for multi-temporal remote sensing. The authors propose LoRetta, a two-stage architecture that first estimates a matchability-weighted affine prior from coarse correspondences and then refines dense residual displacements in the affine-aligned frame. They also introduce LEVIR-GM, a large optical-optical matching dataset with 103K aligned and 827K augmented pairs, spanning six continents, resolutions from 0.5 to 1024 m, and synthetic warp targets with pseudo-matchability labels derived from NMI/NCC template matching. On a held-out LEVIR-GM test split, LoRetta reports 83.3% AUC, outperforming RoMa v2 by 1.6 AUC points and by 6.5/8.2 PCK points at 1/2 px, with 47.8% lower latency. Additional experiments cover land-cover-wise robustness, component ablations, and astronaut/UAV-to-satellite geolocalization.
Significance. If the reported results hold, the paper makes two useful contributions: a large, diverse training and evaluation resource for remote sensing dense matching, and a strong two-stage baseline with a controlled comparison to RoMa v2 using the same frozen DINOv3 backbone. The design is transparent: losses, thresholds, augmentation ranges, and evaluation protocol are specified, and the ablations in Table VI isolate the contributions of affine localization, dense registration, localization guidance, and matchability sampling. The main risk is that the evaluation mask and LoRetta's supervision share the same noisy pseudo-matchability labels, so the headline gains may partly reflect training to the evaluation subset rather than genuinely better alignment. This is addressable with additional experiments, but it is load-bearing for the central claim.
major comments (3)
- [Sec. V-A, Eq. (24); Sec. III-E, Eqs. (13)-(19)] The evaluation mask Omega = {p | \bar M_{1/1}(p) = 1} in Eq. (24) is produced by the same NMI/NCC pseudo-matchability procedure (Eq. 13) that defines LoRetta's geometric supervision set Omega_l (Eq. 16) and matchability targets (Eq. 19). LoRetta is therefore explicitly trained to concentrate its warp regression on exactly the pixels used for scoring, while the baselines are fine-tuned without this mask. This is not direct circularity because the endpoint error is computed against the known augmentation warp \bar W_{1/1}, not against the mask; however, if the thresholds in Eq. (13) label genuinely difficult pixels (low texture, strong change, parallax) as unmatchable, those pixels are excluded from the score for all methods, and LoRetta has been trained to ignore them. The paper gives no independent validation of the pseudo-labels (human annotation or alternative geometric consistency) and no sensitivity analysis for S_u >= 0.05 or the peak-agreement radius of 2. I request that the authors (i) report the main comparison on all pixels with a valid warp target, (ii) repeat scoring on a mask built from an independent criterion such as forward-backward flow consistency, and (iii) vary the Eq. (13) thresholds to show the ranking is stable. Without these, the reported 1.6 AUC and 6.5/8.2 PCK gains over RoMa v2 are not fully established.
- [Sec. V-B, Table III; Sec. V-F, Table VI] The main results and ablations are single-run numbers on a single test split. The headline advantage over RoMa v2 is 1.6 AUC points (83.3 vs 81.7), which is small relative to typical training-seed variation for dense matchers. Please report at least three training seeds with mean and standard deviation for LoRetta and RoMa v2, or provide paired bootstrap confidence intervals over test pairs. If the best checkpoint is selected on the validation split, the selection procedure should be described so the reader can assess optimism in the reported test numbers.
- [Sec. V-A, baseline fine-tuning] The learned baselines are initialized from GIM or official weights and fine-tuned on LEVIR-GM, but the paper does not state whether their training losses are masked to matchable pixels. If baselines are trained with dense losses on all pixels while LoRetta's geometric loss is restricted to Omega_l, then part of the observed gain may come from the training mask rather than from the localization-and-registration architecture. Please clarify the baseline training objectives and, ideally, fine-tune a RoMa v2 variant with a matchability-weighted loss equivalent to Eq. (17) to isolate the architectural contribution.
minor comments (5)
- [Sec. V-A, Eq. (26)] The displayed AUC formula appears to have a typesetting artifact ('1 tK−t 1 K−1X'); please correct the formula so the normalization by t_K - t_1 is unambiguous.
- [Sec. IV-A] The construction of the aligned base layer (how the 103K multi-temporal pairs were co-registered and quality-checked) is not described; this is important for understanding the synthetic warp targets and the validity of the pseudo-matchability labels.
- [Abstract and Sec. V-E] The text contains several instances of 'UA V-to-satellite' that should read 'UAV-to-satellite'.
- [Sec. V-E, Table V] Please clarify whether the baseline numbers from EarthMatch were produced with the same RANSAC/TPS evaluation protocol or with their original scoring, since only the LoRetta row is re-evaluated here; a direct re-run of all methods under the same protocol would make the comparison cleaner.
- [Sec. III-C, Eq. (3)] The affine fit uses hard thresholding on m_i > tau_A; the paper does not discuss how gradients flow through this selection during training. A brief note on the differentiable approximation or straight-through estimator would improve reproducibility.
Circularity Check
LEVIR-GM evaluation mask is defined by the same pseudo-matchability labels used to supervise LoRetta, making part of the reported gain a product of the evaluation protocol.
-
fitted input called prediction
[Section III-E (Eqs. 16, 17, 19) and Section V-A (Eqs. 24-26)]
"matchability prediction is supervised by the full soft target ¯Mℓ using class-balanced binary cross-entropy (BCE) ... For geometric regression, we supervise the fully matchable locations Ωℓ ={p| ¯Mℓ(p) = 1} ... Following the matchability supervision defined in Section III-E, the metrics are computed only on fully matchable pixels: Ω ={p| ¯M1/1(p) = 1}."
The same pseudo-matchability map M̄, generated by NMI/NCC template matching in Eq. 13, is used in three roles: (i) as supervision for LoRetta's matchability head via BCE (Eq. 19), (ii) as the mask Ωℓ restricting LoRetta's geometric regression (Eqs. 16-17), and (iii) as the evaluation mask Ω on which PCK and AUC are computed (Eqs. 24-26). Thus the benchmark's scoreable pixels are exactly the pixels LoRetta was trained to identify and align, while the baselines are evaluated on the same mask without receiving equivalent matchability supervision.
full rationale
The paper's central claim is an empirical benchmark result, not a derived equation, so there is no equation-level circularity. The circularity concern is localized to the evaluation protocol: the mask defining scoreable pixels (Eq. 24) is identical in construction to the pseudo-matchability target used to supervise LoRetta (Eqs. 16, 17, 19). Because LoRetta is explicitly trained to reproduce these pseudo-labels and to regress warps only within the fully matchable set, its AUC and PCK numbers are measured on a set that its own training objective is designed to match, while baselines are scored on the same mask without that supervision. This does not make the warp accuracy itself tautological, but it means a learned matchability bias could account for part of the reported gains. No load-bearing self-citation or imported-uniqueness pattern is present; the paper's ablations and downstream tests provide additional independent evidence. The 4.0 score reflects one significant protocol-level circularity rather than a fully forced derivation.
Assumptions & free parameters
free parameters (7)
- Affine fit matchability threshold tau_A =
0.3
- Pseudo-matchability thresholds =
S>=0.05; NMI/NCC peak agreement <=2 px; exclude boundary
- Robust loss scale c =
0.1
- Loss weights lambda_mat, lambda_cls, lambda_aff =
0.01, 1e-4, 0.05
- Scale weights alpha_l =
0.1, 0.1, 0.1, 0.2, 0.5
- Local correlation radii =
5, 3, 2
- Augmentation perturbation ranges =
scale 1.0-4.0; rotation +-pi; shear +-pi/6; local centers 8/16/24/32; amplitudes 32/24/16/8
assumptions (5)
- domain assumption The mapping of a locally planar ground surface between two satellite observations can be modeled by a single affine transformation under weak-perspective imaging.
- domain assumption The NMI/NCC hybrid peak verification (Eq. 13) yields correct pseudo-matchability labels.
- domain assumption Frozen DINOv3 features provide a sufficient representation for global-scale remote sensing matching.
- domain assumption The held-out LEVIR-GM test split, generated by the same augmentation pipeline as training, is a representative measure of global-scale matching performance.
- standard math Standard mathematical operations (least squares, softmax, bilinear sampling, TPS fitting) are valid and correctly implemented.
Cite this review
Pith. "Pith review of LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching." pith.science (2026). https://pith.science/paper/NLLMKEGL
@misc{pith2026260804106,
author = {Pith},
title = {Pith review of: LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching},
year = {2026},
howpublished = {\url{https://pith.science/paper/NLLMKEGL}},
note = {Machine review of arXiv:2608.04106}
}
read the original abstract
Dense image matching establishes pixel-wise correspondences and underpins broad applications in computer vision and photogrammetry. However, extending dense matching to global-scale remote sensing remains challenging because image pairs may differ in acquisition time, season, viewpoint, spatial resolution, and land-cover state. The resulting large geometric offsets, partial overlap, and intrinsically unmatchable regions make direct dense correspondence prediction unreliable and inefficient. We thus reformulate dense matching as localization-and-registration: first localizing the matchable overlap and affine geometry, then refining dense residuals within the aligned frame. Based on this formulation, we propose LoRetta, a foundation model coupling matchability-aware affine localization with guided dense registration. We also introduce LEVIR-GM, a global-scale multi-temporal optical matching benchmark with dataset-native matchability labels (103K aligned, 827K augmented pairs, six continents, five years, 0.5-1024 m resolution). We further establish a unified evaluation protocol for sparse, semi-dense, and dense matchers. On LEVIR-GM, LoRetta achieves an area under the curve (AUC) of 83.3%, outperforming the strongest baseline RoMa v2 by 1.6 points, with larger percentage of correct keypoints (PCK) gains of 6.5 and 8.2 points at 1 and 2 pixels, while reducing inference latency by 47.8%. Astronaut-to-satellite and unmanned aerial vehicle (UAV)-to-satellite geolocalization experiments further demonstrate its transferability as a reusable geometric aligner.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
Local feature matching using deep learning: A survey,
S. Xu, S. Chen, R. Xu, C. Wang, P. Lu, and L. Guo, “Local feature matching using deep learning: A survey,”Information Fusion, vol. 107, p. 102344, 2024
work page 2024
-
[2]
Deep learning in remote sensing image matching: A survey,
L. Li, L. Han, Y . Ye, Y . Xiang, and T. Zhang, “Deep learning in remote sensing image matching: A survey,”ISPRS Journal of Photogrammetry and Remote Sensing, vol. 225, pp. 88–112, 2025
2025
-
[3]
PDC-Net+: En- hanced probabilistic dense correspondence network,
P. Truong, M. Danelljan, R. Timofte, and L. V . Gool, “PDC-Net+: En- hanced probabilistic dense correspondence network,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 8, pp. 10 247– 10 266, 2023
work page 2023
-
[4]
DKM: Dense kernelized feature matching for geometry estimation,
J. Edstedt, I. Athanasiadis, M. Wadenb ¨ack, and M. Felsberg, “DKM: Dense kernelized feature matching for geometry estimation,” inPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 17 765–17 775
work page 2023
-
[5]
RoMa: Robust dense feature matching,
J. Edstedt, Q. Sun, G. B ¨okman, M. Wadenb ¨ack, and M. Fels- berg, “RoMa: Robust dense feature matching,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 19 790–19 800
work page 2024
-
[6]
RoMa v2: Harder better faster denser feature matching,
J. Edstedt, D. Nordstr ¨om, Y . Zhang, G. B¨okman, J. Astermark, V . Lars- son, A. Heyden, F. Kahl, M. Wadenb ¨ack, and M. Felsberg, “RoMa v2: Harder better faster denser feature matching,” arXiv preprint arXiv:2511.15706, 2025
arXiv 2025
-
[7]
SAR-optical feature matching: A large-scale patch dataset and a deep local descriptor,
W. Xu, X. Yuan, Q. Hu, and J. Li, “SAR-optical feature matching: A large-scale patch dataset and a deep local descriptor,”International Journal of Applied Earth Observation and Geoinformation, vol. 122, p. 103433, 2023
work page 2023
-
[8]
The SEN1-2 dataset for deep learning in SAR-optical data fusion,
M. Schmitt, L. H. Hughes, and X. X. Zhu, “The SEN1-2 dataset for deep learning in SAR-optical data fusion,”ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. IV-1, pp. 141–146, 2018
work page 2018
Show all 59 references
-
[9]
3MOS: A multi-source, multi-resolution, and multi-scene optical-SAR dataset with insights for multi-modal image matching,
Y . Ye, X. Teng, H. Yang, S. Chen, Y . Sun, Y . Bian, T. Tan, Z. Li, and Q. Yu, “3MOS: A multi-source, multi-resolution, and multi-scene optical-SAR dataset with insights for multi-modal image matching,” Visual Intelligence, vol. 3, no. 1, pp. 1–27, 2025
2025
-
[10]
SOMA- 1M: A large-scale SAR-optical multi-resolution alignment dataset for multi-task remote sensing,
P. Wu, Y . Yao, Y . Wan, W. Zhang, R. Zhao, J. Li, and Y . Zhang, “SOMA- 1M: A large-scale SAR-optical multi-resolution alignment dataset for multi-task remote sensing,” arXiv preprint arXiv:2602.05480, 2026
2026
-
[11]
LightGlue: Local feature matching at light speed,
P. Lindenberger, P.-E. Sarlin, and M. Pollefeys, “LightGlue: Local feature matching at light speed,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 17 581–17 592
2023
-
[12]
LoFTR: Detector- free local feature matching with transformers,
J. Sun, Z. Shen, Y . Wang, H. Bao, and X. Zhou, “LoFTR: Detector- free local feature matching with transformers,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 8918–8927
2021
-
[13]
Distinctive image features from scale-invariant keypoints,
D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” International Journal of Computer Vision, vol. 60, no. 2, pp. 91–110, 2004
2004
-
[14]
SuperPoint: Self- supervised interest point detection and description,
D. DeTone, T. Malisiewicz, and A. Rabinovich, “SuperPoint: Self- supervised interest point detection and description,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Jun. 2018, pp. 224–236
2018
-
[15]
SuperGlue: Learning feature matching with graph neural networks,
P.-E. Sarlin, D. DeTone, T. Malisiewicz, and A. Rabinovich, “SuperGlue: Learning feature matching with graph neural networks,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2020, pp. 4937–4946
2020
-
[16]
ASpanFormer: Detector-free image matching with adaptive span transformer,
H. Chen, Z. Luo, L. Zhou, Y . Tian, M. Zhen, T. Fang, D. McKinnon, Y . Tsin, and L. Quan, “ASpanFormer: Detector-free image matching with adaptive span transformer,” inComputer Vision - ECCV 2022, 2022, pp. 20–36
2022
-
[17]
Efficient LoFTR: Semi- dense local feature matching with sparse-like speed,
Y . Wang, X. He, S. Peng, D. Tan, and X. Zhou, “Efficient LoFTR: Semi- dense local feature matching with sparse-like speed,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 21 666–21 675
2024
-
[18]
Raising the ceiling: Conflict-free local feature matching with dynamic view switching,
X. Lu and S. Du, “Raising the ceiling: Conflict-free local feature matching with dynamic view switching,” inComputer Vision - ECCV 2024, 2024, pp. 256–273
2024
-
[19]
Toward free-form local feature matching,
X. Lu, S. Du, Y . Yan, X. Lu, and T. Ikenaga, “Toward free-form local feature matching,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 2, pp. 1470–1484, 2026
2026
-
[20]
Learning accurate dense correspondences and when to trust them,
P. Truong, M. Danelljan, L. V . Gool, and R. Timofte, “Learning accurate dense correspondences and when to trust them,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 5710–5720
2021
-
[21]
DINOv2: Learning robust visual features without supervision,
M. Oquab, T. Darcet, T. Moutakanni, H. V . V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . Huang, S.- W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. J ´egou, J. Mairal, P. Labat...
2024
-
[22]
Sim ´eoni, H
O. Sim ´eoni, H. V . V o, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V . Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa, F. Massa, D. Haz- iza, L. Wehrstedt, J. Wang, T. Darcet, T. Moutakanni, L. Sentana, C. Roberts, A. Vedaldi, J. Tolan, J. Brandt, C. Couprie, J. Mairal,...
2025 arXiv
-
[23]
UFM: A simple path towards unified dense correspondence with flow,
Y . Zhang, N. Keetha, C. Lyu, B. Jhamb, Y . Chen, Y . Qiu, J. Karhade, S. Jha, Y . Hu, D. Ramanan, S. Scherer, and W. Wang, “UFM: A simple path towards unified dense correspondence with flow,” inAdvances in Neural Information Processing Systems, vol. 38. Curran Associates, Inc...
2025
-
[24]
Repeatability is not enough: Learning affine regions via discriminability,
D. Mishkin, F. Radenovic, and J. Matas, “Repeatability is not enough: Learning affine regions via discriminability,” inComputer Vision - ECCV 2018, 2018, pp. 287–304
2018
-
[25]
Structured epipolar matcher for local feature matching,
J. Chang, J. Yu, and T. Zhang, “Structured epipolar matcher for local feature matching,” inProceedings of the IEEE/CVF Conference on ARXIV PREPRINT 17 Computer Vision and Pattern Recognition Workshops, 2023, pp. 6177– 6186
2023
-
[26]
MESA: Effective matching redundancy reduction by semantic area segmentation,
Y . Zhang, S. Shen, and X. Zhao, “MESA: Effective matching redundancy reduction by semantic area segmentation,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 4, pp. 4454–4472, 2026
2026
-
[27]
GIM: Learning generalizable image matcher from internet videos,
X. Shen, Z. Cai, W. Yin, M. M ¨uller, Z. Li, K. Wang, X. Chen, and C. Wang, “GIM: Learning generalizable image matcher from internet videos,” inInternational Conference on Learning Representations, 2024. [Online]. Available: https://openreview.net/forum?id=NYN1b8GRGS
2024
-
[28]
MINIMA: Modality invariant image matching,
J. Ren, X. Jiang, Z. Li, D. Liang, X. Zhou, and X. Bai, “MINIMA: Modality invariant image matching,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 23 059–23 068
2025
-
[29]
MatchAnything: Universal cross-modality image matching with large- scale pre-training,
X. He, H. Yu, S. Peng, D. Tan, Z. Shen, H. Bao, and X. Zhou, “MatchAnything: Universal cross-modality image matching with large- scale pre-training,” arXiv preprint arXiv:2501.07556, 2025
2025 arXiv
-
[30]
A deep learning framework for remote sensing image registration,
S. Wang, D. Quan, X. Liang, M. Ning, Y . Guo, and L. Jiao, “A deep learning framework for remote sensing image registration,”ISPRS Journal of Photogrammetry and Remote Sensing, vol. 145, pp. 148–164, 2018
2018
-
[31]
Optical and SAR image matching using pixelwise deep dense features,
H. Zhang, L. Lei, W. Ni, T. Tang, J. Wu, D. Xiang, and G. Kuang, “Optical and SAR image matching using pixelwise deep dense features,” IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1–5, 2022
2022
-
[32]
Remote sensing image registration using convolutional neural network features,
F. Ye, Y . Su, H. Xiao, X. Zhao, and W. Min, “Remote sensing image registration using convolutional neural network features,”IEEE Geoscience and Remote Sensing Letters, vol. 15, no. 2, pp. 232–236, 2018
2018
-
[33]
Multimodal remote sensing image matching via learning features and attention mechanism,
Y . Zhang, C. Lan, H. Zhang, G. Ma, and H. Li, “Multimodal remote sensing image matching via learning features and attention mechanism,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1– 20, 2024
2024
-
[34]
A two-stream symmetric network with bidirectional ensemble for aerial image matching,
J.-H. Park, W.-J. Nam, and S.-W. Lee, “A two-stream symmetric network with bidirectional ensemble for aerial image matching,”Remote Sensing, vol. 12, no. 3, p. 465, 2020
2020
-
[35]
Precise aerial image matching based on deep homography estimation,
M.-S. Oh, Y .-J. Lee, and S.-W. Lee, “Precise aerial image matching based on deep homography estimation,” arXiv preprint arXiv:2107.08768, 2021
2021 arXiv
-
[36]
Multimodal image fusion framework for end-to-end remote sensing image registration,
L. Li, L. Han, M. Ding, and H. Cao, “Multimodal image fusion framework for end-to-end remote sensing image registration,”IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–14, 2023
2023
-
[37]
Remote sensing image registration based upon extensive convolutional architecture with transfer learning and network pruning,
H.-H. Chang, “Remote sensing image registration based upon extensive convolutional architecture with transfer learning and network pruning,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1– 16, 2023
2023
-
[38]
Unsupervised image registration for video SAR,
X. Huang, J. Ding, and Q. Guo, “Unsupervised image registration for video SAR,”IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 14, pp. 1075–1083, 2021
2021
-
[39]
Unsupervised multistep deformable registration of remote sensing imagery based on deep learning,
M. Papadomanolaki, S. Christodoulidis, K. Karantzalos, and M. Vakalopoulou, “Unsupervised multistep deformable registration of remote sensing imagery based on deep learning,”Remote Sensing, vol. 13, no. 7, p. 1294, 2021
2021
-
[40]
A multiscale framework with unsupervised learning for remote sensing image regis- tration,
Y . Ye, T. Tang, B. Zhu, C. Yang, B. Li, and S. Hao, “A multiscale framework with unsupervised learning for remote sensing image regis- tration,”IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–15, 2022
2022
-
[41]
MID: A novel mountainous remote sensing imagery registration dataset assessed by a coarse-to-fine unsupervised cascading network,
R. Feng, X. Li, J. Bai, and Y . Ye, “MID: A novel mountainous remote sensing imagery registration dataset assessed by a coarse-to-fine unsupervised cascading network,”Remote Sensing, vol. 14, no. 17, p. 4178, 2022
2022
-
[42]
OSFlowNet: Optical and SAR image dense registration using a robust deep optical flow framework,
H. Zhang, L. Lei, W. Ni, X. Yang, T. Tang, K. Cheng, D. Xiang, and G. Kuang, “OSFlowNet: Optical and SAR image dense registration using a robust deep optical flow framework,”IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 16, pp. 1269– 1294, 2023
2023
-
[43]
OS3Flow: Optical and SAR image registration using symmetry-guided semi-dense optical flow,
Z. Sun, S. Zhi, K. Huo, X. Liu, W. Jiang, and Y . Liu, “OS3Flow: Optical and SAR image registration using symmetry-guided semi-dense optical flow,”IEEE Geoscience and Remote Sensing Letters, vol. 21, pp. 1–5, 2024
2024
-
[44]
Very deep convolutional networks for large-scale image recognition,
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” inInternational Conference on Learning Representations, 2015
2015
-
[45]
An overlap invariant entropy measure of 3d medical image alignment,
C. Studholme, D. L. G. Hill, and D. J. Hawkes, “An overlap invariant entropy measure of 3d medical image alignment,”Pattern Recognition, vol. 32, no. 1, pp. 71–86, 1999
1999
-
[46]
Fast template matching,
J. P. Lewis, “Fast template matching,” inProceedings of Vision Interface 1995, 1995, pp. 120–123
1995
-
[47]
A general and adaptive robust loss function,
J. T. Barron, “A general and adaptive robust loss function,” inPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 4326–4334
2019
-
[48]
A deep learning semantic template matching framework for remote sensing image registration,
L. Li, L. Han, M. Ding, H. Cao, and H. Hu, “A deep learning semantic template matching framework for remote sensing image registration,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 181, pp. 205–217, 2021
2021
-
[49]
The QXS-SAROPT dataset for deep learning in SAR- optical data fusion,
M. Huang, Y . Xu, L. Qian, W. Shi, Y . Zhang, W. Bao, N. Wang, X. Liu, and X. Xiang, “The QXS-SAROPT dataset for deep learning in SAR- optical data fusion,” arXiv preprint arXiv:2103.08259, 2021
2021 arXiv
-
[50]
The SARptical dataset for joint analysis of SAR and optical image in dense urban area,
Y . Wang and X. X. Zhu, “The SARptical dataset for joint analysis of SAR and optical image in dense urban area,” inProceedings of the IEEE International Geoscience and Remote Sensing Symposium, 2018, pp. 6840–6843
2018
-
[51]
A global-to- local algorithm for high-resolution optical and SAR image registration,
Y . Xiang, X. Wang, F. Wang, H. You, X. Qiu, and K. Fu, “A global-to- local algorithm for high-resolution optical and SAR image registration,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1– 20, 2023
2023
-
[52]
Automatic registration of optical and SAR images via improved phase congruency model,
Y . Xiang, R. Tao, F. Wang, H. You, and B. Han, “Automatic registration of optical and SAR images via improved phase congruency model,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 13, pp. 5847–5861, 2020
2020
-
[53]
SpaceNet 6: Multi-sensor all weather mapping dataset,
J. Shermeyer, D. Hogan, J. Brown, A. V . Etten, N. Weir, F. Paci- fici, R. Hansch, A. Bastidas, S. Soenen, T. Bacastow, and R. Lewis, “SpaceNet 6: Multi-sensor all weather mapping dataset,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion W...
2020
-
[54]
Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography,
M. A. Fischler and R. C. Bolles, “Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography,”Communications of the ACM, vol. 24, no. 6, pp. 381–395, 1981
1981
-
[55]
Principal warps: Thin-plate splines and the decom- position of deformations,
F. L. Bookstein, “Principal warps: Thin-plate splines and the decom- position of deformations,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 11, no. 6, pp. 567–585, 1989
1989
-
[56]
Decoupled weight decay regularization,
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations, 2019. [Online]. Available: https://openreview.net/forum?id=Bkg6RiCqY7
2019
-
[57]
EarthMatch: Iterative coregistration for fine-grained localization of astronaut photography,
G. Berton, G. Goletto, G. Trivigno, A. Stoken, B. Caputo, and C. Ma- sone, “EarthMatch: Iterative coregistration for fine-grained localization of astronaut photography,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Jun. 2024, p...
2024
-
[58]
Find my astronaut photo: Automated local- ization and georectification of astronaut photography,
A. Stoken and K. Fisher, “Find my astronaut photo: Automated local- ization and georectification of astronaut photography,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Jun. 2023, pp. 6196–6205
2023
-
[59]
UA VLoc-M3 UA V visual localization dataset,
Z. Nie, J. Huang, Y . Li, and K. Ren, “UA VLoc-M3 UA V visual localization dataset,” Science Data Bank, Version 2, Mar. 2026. [Online]. Available: https://doi.org/10.57760/sciencedb.29772
2026 doi
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.