Pith. sign in

REVIEW 3 major objections 5 minor 194 references

Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A unified benchmark of image matchers finds accuracy rises with matching density, while semi-dense methods win at localization.

desk verdict Useful survey and taxonomy, but the ScanNet benchmark confounds training domain with method family, so the central dense-vs-sparse hierarchy claim is overstated. read the letter →

arxiv 2608.11093 v1 pith:CL6B47CP submitted 2026-08-11 cs.LG cs.CV

classification cs.LGcs.CV
keywords cross-viewfeaturematchinglocalvisionfoundationmodelssparsematchersemi-densedensevisuallocalizationrobustestimation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey tries to give the field of cross-view feature matching—finding corresponding points across images taken from very different viewpoints—one shared map. It proposes a six-axis taxonomy covering feature extraction, single-type matchers, multi-type matchers, vision-foundation-model methods, training strategy, and robust estimation, then runs representative methods through a unified benchmarking protocol on standard datasets. Its central finding is a consistent performance hierarchy: dense methods reach the highest accuracy but cost the most, sparse methods stay efficient yet are capped by keypoint detector quality, and semi-dense methods give the best overall visual-localization performance. The paper matters because it translates a fragmented literature into a single ordering that researchers can use to choose methods and target gaps.

What carries the argument

The load-bearing object is the six-axis taxonomy, which treats cross-view feature matching not as a single algorithm family but as a layered design space spanning feature extraction, sparse/semi-dense/dense matching, multi-type feature fusion, vision-foundation-model use, training strategy, and robust estimation. The matching-paradigm split (sparse, semi-dense, dense) is the axis that does the explanatory work in the benchmark. The second mechanism is the unified evaluation protocol: one detector for sparse methods, one resizing rule for semi-dense methods, per-paper settings for dense methods, shared RANSAC thresholds, and the standard hierarchical localization pipeline, which together allow the paper to read performance differences as paradigm effects rather than implementation noise.

What would settle it

Retrain every sparse and semi-dense method on ScanNet's own training split, the same data dense methods see, and re-run relative pose estimation. If the dense advantage at AUC@20° shrinks below the gap reported in Table I, or if a sparse method overtakes a dense one, the claimed paradigm hierarchy reflects training conditions rather than matching density.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that the apparent chaos of cross-view matching methods resolves into a structured design space plus a reproducible performance ordering. The proposed taxonomy organizes dozens of methods along six orthogonal dimensions and, within matching paradigms, by motivation: accuracy, efficiency, generalization, and distribution. Under the survey's consistent evaluation—SuperPoint for all sparse methods, fixed resizing for semi-dense methods, each dense method at its native settings, RANSAC for pose and homography, and the standard HLoc pipeline for localization—three results stand out. First, on relative pose estimation (MegaDepth, ScanNet) and homography estimation (HPatches), accuracy rises with matching density: dense methods such as RoMa and RoMa v2 lead, cascaded semi-dense methods such as IMAmatch approach them, and sparse methods lag most on weakly textured indoor scenes. Second, sparse methods' ceiling is set by the detector: even strong attention-based matchers collapse on ScanNet because they inherit limited keypoints. Third, in visual localization (Aachen Day-Night, InLoc), semi-dense methods are the best overall, with dense methods losing ground at night and sparse methods falling below 80% indoors.

Load-bearing premise

The comparison is fair enough to rank paradigms; that holds only if results obtained under different training-data conditions are directly comparable, since dense methods are trained separately on each dataset, some semi-dense ScanNet numbers come from ScanNet-trained models (marked by underlining), and parameter counts are missing for several methods without public code.

Editorial extensions

If this is right

  • Keypoint detector quality is the binding constraint on sparse matchers: on weakly textured indoor scenes their pose-estimation AUC drops sharply no matter how strong the attention module is.
  • Accuracy scales with matching density, so applications that need pose quality in low-texture environments should prefer dense or cascaded semi-dense models, accepting their larger parameter counts.
  • Semi-dense matching is currently the best default for visual localization, combining indoor robustness with outdoor day/night reliability, while dense methods are more vulnerable to low-light noise.
  • Foundation models are best used as coarse semantic or geometric guides inside specialized matching networks, not as wholesale replacements, because their features are too coarse for precise localization.
  • The survey's listed open problems—uncertainty-aware matching, explicit geometric reasoning at inference, non-rigid scenes, and quadratic-cost global matching—define the next generation of correspondence models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because dense methods were trained separately on each dataset while sparse and most semi-dense methods used a single MegaDepth-trained model, part of the dense advantage on ScanNet may reflect extra in-domain training; re-running the benchmark with all methods trained on identical data would isolate algorithmic merit from training conditions.
  • The strong result of cascaded semi-dense methods like IMAmatch suggests the sparse/dense dichotomy is dissolving, and hybrid detector-free pipelines with adaptive density will keep closing the gap to dense methods at lower cost.
  • If the detector bottleneck is real, replacing SuperPoint with a stronger detector or a detector-free sparse matcher could lift sparse methods on indoor benchmarks more than any further attention redesign.
  • The nighttime localization results hint that dense methods overfit to day-lit appearance, so testing dense matchers with illumination augmentation or multi-hypothesis matching could recover their localization performance.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper is a survey of cross-view feature matching that proposes a six-part taxonomy (feature extraction, sparse/semi-dense/dense matchers, multi-type feature matchers, vision-foundation-model methods, training strategies, and robust estimation), reviews recent methods within that taxonomy, and presents a benchmark on MegaDepth, ScanNet, HPatches, Aachen Day-Night, and InLoc. The authors claim that the benchmark reveals a consistent performance hierarchy: dense methods achieve the highest pose and homography accuracy at higher computational cost, sparse methods are efficient but detector-limited, and semi-dense methods perform best overall in visual localization.

Significance. The survey's taxonomic organization is useful and its literature coverage is broad; the review of learning paradigms, VFM-based methods, and robust estimation provides a reasonable entry point for the field. The authors are also transparent in Table I about some training and code-availability caveats, which is a strength. However, the central empirical claims are currently not established because the benchmark does not control for training domain and dataset version across the compared methods. If the authors either add controlled comparisons or substantially weaken the hierarchy claims, the survey would be a valuable reference. The paper contains no derivations or machine-checked proofs; its value rests on the accuracy of the taxonomy and the credibility of the benchmark.

major comments (3)
  1. [Section III.C, Table I] The claimed dense > semi-dense > sparse hierarchy for relative pose is confounded by training domain on ScanNet. The table's note states that dense models are trained separately on each dataset, while sparse and most semi-dense methods use a single MegaDepth-trained model, and several semi-dense ScanNet results (underlined) come from ScanNet-trained models. RoMa v2's AUC@20 of 73.8 versus LightGlue's 51.8 is therefore not a controlled comparison of matching density; the same table shows EcoMatcher (ScanNet-trained) at 64.0 versus ELoFTR (MegaDepth-trained) at 53.6, indicating that training data alone can move AUC by roughly 10 points. Please add ScanNet-trained sparse/semi-dense baselines or MegaDepth-trained dense evaluations, or explicitly restrict the hierarchy claim to settings with matched training data.
  2. [Section III.E, Table III] The visual localization comparison mixes dataset versions: the settings paragraph says sparse methods use Aachen Day-Night v1.0 while semi-dense and dense methods use v1.1, yet Table III reports a single Aachen Day-Night column group. Because v1.0 and v1.1 have different reference sets and query sets, the sparse versus semi-dense/dense comparisons on this benchmark are not mutually comparable. Please use one dataset version for all methods, or report results separately by version and avoid cross-category claims when versions differ.
  3. [Section III, Table I note] The abstract and Section III describe a 'unified experimental benchmarking' under 'consistent protocols', but the Table I note admits that several semi-dense rows are literature-reported values from non-public code with unknown parameter counts, and dense methods keep their original image settings. This means the benchmark is not a single controlled implementation. Please state explicitly which rows were re-run by the authors, provide evaluation code and logs, and replace 'unified' and 'consistent' with a more precise description that lists the harmonized settings and the admitted exceptions.
minor comments (5)
  1. [Section III.D, Table II] The statement that dense methods 'attain the highest overall performance' is too strong: IMAmatch, a semi-dense method, achieves AUC@3px=72.4 versus RoMa's 72.2 and is within 0.2 AUC at the other thresholds. Please soften the wording to 'comparable or slightly higher'.
  2. [Section III.C, Table I] The underlining used to mark ScanNet-trained semi-dense results is not visible in the provided text; ensure the final PDF renders the underlining clearly and add an explicit legend in the table caption.
  3. [Section III.C, Table I] Parameter counts are missing for SEM and EcoMatcher; the efficiency discussion in Section III.C should avoid drawing parameter-cost conclusions from incomplete data.
  4. [Section II] The taxonomy labels CasMTR and IMAmatch as semi-dense methods that bridge to dense matching, while DKM and RoMa are classified as dense; please state the classification criterion explicitly so that coarse-to-fine architecture alone does not become an ambiguous discriminator.
  5. [Section III.E] The evaluation uses Aachen Day-Night v1.0 for sparse methods and v1.1 for semi-dense and dense methods; even if the final version reports each separately, the table should include a version column so readers can see the protocol difference immediately.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the benchmark hierarchy is an empirical claim with an admitted training-domain confound, not an equation-level reduction or a self-citation-forced conclusion.

full rationale

The paper contains no derivation chain of the kind the circularity checklist targets: Section II is a descriptive taxonomy, and Section III is an empirical benchmark. No theorem is derived, no quantity is defined in terms of another, and no fitted parameter is renamed as a prediction. The central empirical claim in Section III.C is that "the performance of two-view pose estimation methods systematically improves with increasing matching density: dense methods consistently outperform semi-dense approaches, which in turn generally surpass sparse methods." This is an empirical generalization from Table I, not a prediction obtained from fitted inputs. The most serious validity threat is explicitly disclosed by the paper itself in the Table I caption: "FOR DENSE METHODS, MODELS ARE TRAINED SEPARATELY ON EACH RESPECTIVE DATASET" and "THEIR PUBLISHED SCANNET RESULTS WERE OBTAINED USING MODELS TRAINED ON SCANNET." Because sparse and most semi-dense methods are evaluated with a single MegaDepth-trained model while dense methods are trained in-domain on ScanNet, the ScanNet columns confound matching density with training domain. That weakness undermines the strength of the hierarchy claim and should be recorded as a correctness risk, but it is not circular: the numbers are externally measured and the protocol difference is admitted rather than concealed. The author self-citations (JamMa, NCTR, DNC-Net, SAM, ParaFormer, MSFormer, FineFormer, ContextMatcher, RCM, ProMa, and others) are bibliographic entries for methods surveyed or benchmarked; none is invoked as a uniqueness theorem, ansatz authority, or load-bearing justification that would make the taxonomy or benchmark conclusions equivalent to a self-cited assumption. Consequently, no circular step meets the quote-and-reduction bar, and the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper introduces no free parameters or new entities; it surveys existing methods. Its central assumptions concern the validity of benchmark datasets and evaluation protocols.

assumptions (2)
  • domain assumption Ground-truth poses and depth maps in the benchmark datasets (MegaDepth, ScanNet, HPatches, Aachen Day-Night, InLoc) are accurate enough for reliable evaluation.
    Section III-A describes the datasets and trusts their standard reconstructions.
  • domain assumption Standard robust estimation (RANSAC) with fixed thresholds is an appropriate common post-processing step for all evaluated matchers.
    Section III-C describes RANSAC thresholds for each dataset; different thresholds are used for different datasets.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives." pith.science (2026). https://pith.science/paper/CL6B47CP

@misc{pith2026260811093,
  author       = {Pith},
  title        = {Pith review of: Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CL6B47CP}},
  note         = {Machine review of arXiv:2608.11093}
}
read the original abstract

Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field has evolved from task-specific models toward increasingly unified and generalizable correspondence models, with recent progress further driven by the emergence of vision foundation models (VFMs). Despite these advances, existing studies remain highly diverse in their problem formulations, model architectures, training paradigms, and evaluation protocols, making it difficult to obtain a unified understanding of the field. In this survey, we present a unified review of cross-view feature matching. We first introduce a structured taxonomy covering feature extraction, single-type feature matcher, multi-type feature matcher, VFMs based methods, training strategy and robust estimation, providing a coherent framework for analysis and comparison. We further examine recent advances, distilling key design principles and highlighting the shift toward unified and generalizable correspondence models. We also provide a unified experimental benchmarking of representative state-of-the-art methods under consistent protocols, enabling fair and comprehensive performance comparisons. In addition, we discuss open challenges and future directions, including efficiency, robustness under extreme conditions, and cross-domain generalization. This survey aims to provide a comprehensive and structured reference for understanding the evolution, current landscape, and future development of cross-view feature matching in the era of vision foundation models.

Figures

Figures reproduced from arXiv: 2608.11093 by the authors.

Figure 1
Figure 1. Taxonomy of cross-view feature matching methods. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Taxonomy and timeline of sparse matching methods categorized by their motivation (Section II-B). [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Taxonomy and timeline of semi-dense matching methods categorized by their motivation (Section II-C). [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Taxonomy and timeline of dense matching methods categorized by their motivation (Section II-D). [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Taxonomy and timeline of multi-type feature matchers (Section II-E). [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Timeline of VFM based image matching approaches (Section II-F). [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Timeline of typical training strategies (Section II-G). [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Taxonomy and timeline of robust estimation methods (Section II-H). [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 9
Figure 9. Figure 9: Visualization of typical cross-view image matching results from sparse, semi-dense, and dense matchers. [PITH_FULL_IMAGE:figures/full_fig_p012_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

194 extracted references · 70 canonical work pages

  1. [1]

    Deep non-rigid structure from motion with missing data,

    C. Kong and S. Lucey, “Deep non-rigid structure from motion with missing data,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 43, no. 12, pp. 4365–4377, 2021

  2. [2]

    InLoc: Indoor visual localization with dense matching and view synthesis,

    H. Taira, M. Okutomi, T. Sattler, M. Cimpoi, M. Pollefeys, J. Sivic, T. Pajdla, and A. Torii, “InLoc: Indoor visual localization with dense matching and view synthesis,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 43, no. 4, pp. 1293–1307, 2021

  3. [3]

    Linear RGB-D SLAM for structured environments,

    K. Joo, P. Kim, M. Hebert, I. S. Kweon, and H. J. Kim, “Linear RGB-D SLAM for structured environments,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 11, pp. 8403–8419, 2022

  4. [4]

    SuperGlue: Learning feature matching with graph neural networks,

    P. Sarlin, D. DeTone, T. Malisiewicz, and A. Rabinovich, “SuperGlue: Learning feature matching with graph neural networks,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020, pp. 4938–4947

  5. [5]

    LoFTR: Detector-free local feature matching with transformers,

    J. Sun, Z. Shen, Y . Wang, H. Bao, and X. Zhou, “LoFTR: Detector-free local feature matching with transformers,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2021, pp. 8922–8931

  6. [6]

    Jamma: Ultra-lightweight local feature matching with joint mamba,

    X. Lu and S. Du, “Jamma: Ultra-lightweight local feature matching with joint mamba,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025, pp. 14 934–14 943

  7. [7]

    MambaGlue: Fast and robust local feature matching with mamba,

    K. Ryoo, H. Lim, and H. Myung, “MambaGlue: Fast and robust local feature matching with mamba,” inProc. IEEE Int. Conf. Robot. Autom. (ICRA), 2025, pp. 5758–5765

  8. [8]

    LightGlue: Local feature matching at light speed,

    P. Lindenberger, P.-E. Sarlin, and M. Pollefeys, “LightGlue: Local feature matching at light speed,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023, pp. 17 627–17 638

Show all 194 references
  1. [9]

    OmniGlue: Generalizable feature matching with foundation model guidance,

    H. Jiang, A. Karpur, B. Cao, Q. Huang, and A. Araujo, “OmniGlue: Generalizable feature matching with foundation model guidance,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 19 865–19 875

  2. [10]

    Roma: Robust dense feature matching,

    J. Edstedt, Q. Sun, G. B ¨okman, M. Wadenb ¨ack, and M. Felsberg, “Roma: Robust dense feature matching,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024, pp. 19 790–19 800

  3. [11]

    Grounding image matching in 3D with MASt3R,

    V . Leroy, Y . Cabon, and J. Revaud, “Grounding image matching in 3D with MASt3R,” inProc. Eur. Conf. Comput. Vis. (ECCV), 2024, pp. 71–91

  4. [12]

    VGGT: Visual geometry grounded transformer,

    J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny, “VGGT: Visual geometry grounded transformer,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025, pp. 5294–5306. 17

  5. [13]

    Image matching from handcrafted to deep features: A survey,

    J. Ma, X. Jiang, A. Fan, J. Jiang, and J. Yan, “Image matching from handcrafted to deep features: A survey,”Int. J. Comput. Vis., vol. 129, no. 1, pp. 23–79, 2021

  6. [14]

    A survey of feature matching methods,

    Q. Huang, X. Guo, Y . Wang, H. Sun, and L. Yang, “A survey of feature matching methods,”IET Image Process., vol. 18, no. 6, pp. 1385–1410, 2024

  7. [15]

    A review of multimodal image matching: Methods and applications,

    X. Jiang, J. Ma, G. Xiao, Z. Shao, and X. Guo, “A review of multimodal image matching: Methods and applications,”Inf. Fusion, vol. 73, pp. 22–71, 2021

  8. [16]

    A combined corner and edge detector,

    C. Harris, M. Stephenset al., “A combined corner and edge detector,” inProc. Alvey Vis. Conf. (AVC), 1988, pp. 147–151

  9. [17]

    Distinctive image features from scale-invariant key- points,

    D. G. Lowe, “Distinctive image features from scale-invariant key- points,”Int. J. Comput. Vis., vol. 60, no. 2, pp. 91–110, 2004

  10. [18]

    SURF: Speeded up robust features,

    H. Bay, T. Tuytelaars, and L. Van G., “SURF: Speeded up robust features,” inProc. Eur. Conf. Comput. Vis. (ECCV), 2006, pp. 404– 417

  11. [19]

    Robust wide-baseline stereo from maximally stable extremal regions,

    J. Matas, O. Chum, M. Urban, and T. Pajdla, “Robust wide-baseline stereo from maximally stable extremal regions,”Image Vis. Comput., vol. 22, no. 10, pp. 761–767, 2004

  12. [20]

    Machine learning for high-speed corner detection,

    E. Rosten and T. Drummond, “Machine learning for high-speed corner detection,” inProc. Eur. Conf. Comput. Vis. (ECCV), 2006, pp. 430– 443

  13. [21]

    Adaptive and generic corner detection based on the accelerated segment test,

    E. Mair, G. D. Hager, D. Burschka, M. Suppa, and G. Hirzinger, “Adaptive and generic corner detection based on the accelerated segment test,” inProc. Eur. Conf. Comput. Vis. (ECCV), 2010, pp. 183–196

  14. [22]

    KAZE features,

    P. F. Alcantarilla, A. Bartoli, and A. J. Davison, “KAZE features,” in Proc. Eur. Conf. Comput. Vis. (ECCV), 2012, pp. 214–227

  15. [23]

    Fast explicit diffusion for accelerated features in nonlinear scale spaces,

    P. F. Alcantarilla, J. Nuevo, and A. Bartoli, “Fast explicit diffusion for accelerated features in nonlinear scale spaces,” inProc. Br. Mach. Vis. Conf. (BMVC), 2013

  16. [24]

    BRISK: Binary robust invariant scalable keypoints,

    S. Leutenegger, M. Chli, and R. Y . Siegwart, “BRISK: Binary robust invariant scalable keypoints,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2011, pp. 2548–2555

  17. [25]

    LIFT: Learned invariant feature transform,

    K. M. Yi, E. Trulls, V . Lepetit, and P. Fua, “LIFT: Learned invariant feature transform,” inProc. Eur. Conf. Comput. Vis. (ECCV), 2016, pp. 467–483

  18. [26]

    TILDE: A temporally invariant learned detector,

    Y . Verdie, K. M. Yi, P. Fua, and V . Lepetit, “TILDE: A temporally invariant learned detector,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2015, pp. 5279–5288

  19. [27]

    SuperPoint: Self- supervised interest point detection and description,

    D. DeTone, T. Malisiewicz, and A. Rabinovich, “SuperPoint: Self- supervised interest point detection and description,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), 2018, pp. 224–236

  20. [28]

    D2-Net: A trainable CNN for joint description and detection of local features,

    M. Dusmanu, I. Rocco, T. Pajdla, M. Pollefeys, J. Sivic, A. Torii, and T. Sattler, “D2-Net: A trainable CNN for joint description and detection of local features,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019, pp. 8084–8093

  21. [29]

    R2D2: Reliable and repeatable detector and descriptor,

    J. Revaud, C. De Souza, M. Humenberger, and P. Weinzaepfel, “R2D2: Reliable and repeatable detector and descriptor,” inProc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2019, pp. 12 414–12 424

  22. [30]

    Key.Net: Keypoint detection by handcrafted and learned CNN filters,

    A. B. Laguna, E. Riba, D. Ponsa, and K. Mikolajczyk, “Key.Net: Keypoint detection by handcrafted and learned CNN filters,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2019, pp. 5835–5843

  23. [31]

    Key.Net: Keypoint detection by handcrafted and learned CNN filters revisited,

    A. Barroso-Laguna and K. Mikolajczyk, “Key.Net: Keypoint detection by handcrafted and learned CNN filters revisited,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 1, pp. 698–711, 2023

  24. [32]

    ALIKE: Accurate and lightweight keypoint detection and descriptor extraction,

    X. Zhao, X. Wu, J. Miao, W. Chen, P. C. Y . Chen, and Z. Li, “ALIKE: Accurate and lightweight keypoint detection and descriptor extraction,” IEEE Trans. Multimedia, vol. 25, pp. 3101–3112, 2023

  25. [33]

    Quad- Networks: Unsupervised learning to rank for interest point detection,

    N. Savinov, A. Seki, L. Ladick ´y, T. Sattler, and M. Pollefeys, “Quad- Networks: Unsupervised learning to rank for interest point detection,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2017, pp. 3929–3937

  26. [34]

    Disk: Learning local fea- tures with policy gradient,

    M. Tyszkiewicz, P. Fua, and E. Trulls, “Disk: Learning local fea- tures with policy gradient,” inProc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 33, 2020, pp. 14 254–14 265

  27. [35]

    Dedode v2: Analyzing and improving the dedode keypoint detector,

    J. Edstedt, G. B ¨okman, and Z. Zhao, “Dedode v2: Analyzing and improving the dedode keypoint detector,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024, pp. 4245–4253

  28. [36]

    DaD: Distilled reinforcement learning for diverse keypoint detection,

    J. Edstedt, G. B ¨okman, M. Wadenb ¨ack, and M. Felsberg, “DaD: Distilled reinforcement learning for diverse keypoint detection,”arXiv preprint arXiv:2503.07347, 2025

  29. [37]

    PCA-SIFT: a more distinctive representation for local image descriptors,

    Y . Ke and R. Sukthankar, “PCA-SIFT: a more distinctive representation for local image descriptors,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2004

  30. [38]

    A performance evaluation of local descriptors,

    K. Mikolajczyk and C. Schmid, “A performance evaluation of local descriptors,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 27, no. 10, pp. 1615–1630, 2005

  31. [39]

    BRIEF: Binary robust independent elementary features,

    M. Calonder, V . Lepetit, C. Strecha, and P. Fua, “BRIEF: Binary robust independent elementary features,” inProc. Eur. Conf. Comput. Vis. (ECCV), 2010, pp. 778–792

  32. [40]

    FREAK: Fast retina key- point,

    A. Alahi, R. Ortiz, and P. Vandergheynst, “FREAK: Fast retina key- point,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2012, pp. 510–517

  33. [41]

    Local intensity order pattern for feature description,

    Z. Wang, B. Fan, and F. Wu, “Local intensity order pattern for feature description,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2011, pp. 603–610

  34. [42]

    Rethinking the sGLOH descriptor,

    F. Bellavia and C. Colombo, “Rethinking the sGLOH descriptor,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 40, no. 4, pp. 931–944, 2018

  35. [43]

    Discriminative learning of deep convolutional feature point descriptors,

    E. Simo-Serra, E. Trulls, L. Ferraz, I. Kokkinos, P. Fua, and F. Moreno- Noguer, “Discriminative learning of deep convolutional feature point descriptors,” inProc. IEEE Int. Conf. Comput. Vis. (ICCV), 2015, pp. 118–126

  36. [44]

    L2-Net: Deep learning of discrimina- tive patch descriptor in Euclidean space,

    Y . Tian, B. Fan, and F. Wu, “L2-Net: Deep learning of discrimina- tive patch descriptor in Euclidean space,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2017, pp. 6128–6136

  37. [45]

    Working hard to know your neighbor’s margins: local descriptor learning loss,

    A. Mishchuk, D. Mishkin, F. Radenovi ´c, and J. Matas, “Working hard to know your neighbor’s margins: local descriptor learning loss,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2017, pp. 4829–4840

  38. [46]

    Learning local image descriptors with deep siamese and triplet convolutional networks by minimizing global loss functions,

    B. Vijay Kumar, G. Carneiro, and I. Reid, “Learning local image descriptors with deep siamese and triplet convolutional networks by minimizing global loss functions,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2016, pp. 5385–5394

  39. [47]

    GeoDesc: Learning local descriptors by integrating geometry constraints,

    Z. Luo, T. Shen, L. Zhou, S. Zhu, R. Zhang, Y . Yao, T. Fang, and L. Quan, “GeoDesc: Learning local descriptors by integrating geometry constraints,” inProc. Eur. Conf. Comput. Vis. (ECCV), 2018, p. 170¨C185

  40. [48]

    Local descriptors optimized for average precision,

    K. He, Y . Lu, and S. Sclaroff, “Local descriptors optimized for average precision,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018, pp. 596–605

  41. [49]

    ContextDesc: Local descriptor augmentation with cross- modality context,

    Z. Luo, T. Shen, L. Zhou, J. Zhang, Y . Yao, S. Li, T. Fang, and L. Quan, “ContextDesc: Local descriptor augmentation with cross- modality context,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019, pp. 2522–2531

  42. [50]

    HyNet: learning local descriptor with hybrid similarity measure and triplet loss,

    Y . Tian, A. Barroso-Laguna, T. Ng, V . Balntas, and K. Mikolajczyk, “HyNet: learning local descriptor with hybrid similarity measure and triplet loss,” inProc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2020, pp. 7401–7412

  43. [51]

    SOSNet: Second order similarity regularization for local descriptor learning,

    Y . Tian, X. Yu, B. Fan, F. Wu, H. Heijnen, and V . Balntas, “SOSNet: Second order similarity regularization for local descriptor learning,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019, pp. 11 008–11 017

  44. [52]

    Zippypoint: Fast interest point detection, description, and matching through mixed precision discretization,

    M. Kanakis, S. Maurer, M. Spallanzani, A. Chhatkuli, and L. Van Gool, “Zippypoint: Fast interest point detection, description, and matching through mixed precision discretization,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023, pp. 6114–6123

  45. [53]

    Xfeat: Accelerated features for lightweight image matching,

    G. Potje, F. Cadar, A. Araujo, R. Martins, and E. R. Nascimento, “Xfeat: Accelerated features for lightweight image matching,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024, pp. 2682–2691

  46. [54]

    NCTR: Neighborhood consensus transformer for feature matching,

    X. Lu and S. Du, “NCTR: Neighborhood consensus transformer for feature matching,” inProc. IEEE Int. Conf. Image Process. (ICIP), 2022, pp. 2726–2730

  47. [55]

    DNC-Net: Dual-neighbourhood consensus net- work for feature matching,

    Q. Jiang and S. Du, “DNC-Net: Dual-neighbourhood consensus net- work for feature matching,” inProc. IEEE Int. Conf. Image Process. (ICIP), 2023, pp. 610–614

  48. [56]

    IMP: Iterative matching and pose estimation with adaptive pooling,

    F. Xue, I. Budvytis, and R. Cipolla, “IMP: Iterative matching and pose estimation with adaptive pooling,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023, pp. 21 317–21 326

  49. [57]

    Scene-aware feature matching,

    X. Lu, Y . Yan, T. Wei, and S. Du, “Scene-aware feature matching,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023, pp. 3704–3713

  50. [58]

    S2H- GNN: Learning soft to hard feature matching with sparsified graph neural network,

    T. Xie, R. Li, Z. Jiang, K. Dai, K. Wang, L. Zhao, and S. Li, “S2H- GNN: Learning soft to hard feature matching with sparsified graph neural network,” inProc. IEEE Int. Conf. Real-Time Comput. Robot. (RCAR), 2023, pp. 756–761

  51. [59]

    HTMatch: An efficient hybrid transformer based graph neural network for local feature match- ing,

    Y . Cai, L. Li, D. Wang, X. Li, and X. Liu, “HTMatch: An efficient hybrid transformer based graph neural network for local feature match- ing,”Signal Process., vol. 204, p. 108859, 2023

  52. [60]

    DiffGlue: Diffusion-aided image feature match- ing,

    S. Zhang and J. Ma, “DiffGlue: Diffusion-aided image feature match- ing,” inProc. ACM Int. Conf. Multimedia (ACM MM), 2024, pp. 8451– 8460. 18

  53. [61]

    ResMatch: Residual attention learning for feature matching,

    Y . Deng, K. Zhang, S. Zhang, Y . Li, and J. Ma, “ResMatch: Residual attention learning for feature matching,” inProc. AAAI Conf. Artif. Intell. (AAAI), 2024, pp. 1501–1509

  54. [62]

    Matching while perceiving: Enhance image feature matching with applicable semantic amalgama- tion,

    S. Zhang, Z. Zhu, Z. Li, T. Lu, and J. Ma, “Matching while perceiving: Enhance image feature matching with applicable semantic amalgama- tion,” inProc. AAAI Conf. Artif. Intell. (AAAI), vol. 39, no. 10, 2025, pp. 10 094–10 102

  55. [63]

    SemMatcher: Semantic- aware feature matching with neighborhood consensus,

    Q. Jiang, X. Lu, D. Liang, and S. Du, “SemMatcher: Semantic- aware feature matching with neighborhood consensus,”J. Vis. Commun. Image Represent., p. 104611, 2025

  56. [64]

    CoMatcher: Multi-view collaborative feature matching,

    J. Zhang, Z. Xia, M. Dong, S. Shen, L. Yue, and X. Zheng, “CoMatcher: Multi-view collaborative feature matching,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025, pp. 21 970–21 980

  57. [65]

    AAPMatcher: Adaptive attention pruning matcher for accurate local feature matching,

    X. Fan, S. Liu, S. Liu, L. Zhao, and R. Li, “AAPMatcher: Adaptive attention pruning matcher for accurate local feature matching,”Neural Netw., vol. 188, p. 107403, 2025

  58. [66]

    Learning to match features with seeded graph matching network,

    H. Chen, Z. Luo, J. Zhang, L. Zhou, X. Bai, Z. Hu, C.-L. Tai, and L. Quan, “Learning to match features with seeded graph matching network,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021, pp. 6301–6310

  59. [67]

    Improving sparse graph attention for feature matching by informative keypoints exploration,

    X. Jiang, S. Zhang, X.-P. Zhang, and J. Ma, “Improving sparse graph attention for feature matching by informative keypoints exploration,” Comput. Vis. Image Underst., vol. 235, p. 103803, 2023

  60. [68]

    ClusterGNN: Cluster-based coarse-to-fine graph neural network for efficient feature matching,

    Y . Shi, J.-X. Cai, Y . Shavit, T.-J. Mu, W. Feng, and K. Zhang, “ClusterGNN: Cluster-based coarse-to-fine graph neural network for efficient feature matching,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022, pp. 12 507–12 516

  61. [69]

    Learning to match features with discriminative sparse graph neural network,

    Y . Shi, J.-X. Cai, M. Fan, W. Feng, and K. Zhang, “Learning to match features with discriminative sparse graph neural network,”Pattern Recognit., vol. 156, p. 110784, 2024

  62. [70]

    ParaFormer: Parallel attention transformer for efficient feature matching,

    X. Lu, Y . Yan, B. Kang, and S. Du, “ParaFormer: Parallel attention transformer for efficient feature matching,” inProc. AAAI Conf. Artif. Intell. (AAAI), 2023, pp. 1853–1860

  63. [71]

    AMatFormer: Efficient feature matching via anchor matching transformer,

    B. Jiang, S. Luo, X. Wang, C. Li, and J. Tang, “AMatFormer: Efficient feature matching via anchor matching transformer,”IEEE Trans. Multimedia, vol. 26, pp. 1504–1515, 2023

  64. [72]

    LightSGM: Local feature matching with lightweight seeded,

    S. Feng, H. Qian, W. Wanget al., “LightSGM: Local feature matching with lightweight seeded,”J. King Saud Univ. - Comput. Inf. Sci., vol. 36, no. 6, p. 102095, 2024

  65. [73]

    FilterGNN: Image feature matching with cascaded outlier filters and linear attention,

    J.-X. Cai, T.-J. Mu, and Y .-K. Lai, “FilterGNN: Image feature matching with cascaded outlier filters and linear attention,”Comput. Vis. Media, vol. 10, no. 5, pp. 873–884, 2024

  66. [74]

    Learning feature matching via matchable keypoint- assisted graph neural network,

    Z. Li and J. Ma, “Learning feature matching via matchable keypoint- assisted graph neural network,”IEEE Trans. Image Process., 2024

  67. [75]

    Ada- Matcher: A deep detector-based local feature matcher with adaptive weight sharing,

    F. Zheng, C. Cao, Z. Zhang, T. Sun, J. Zhang, and L. Zhao, “Ada- Matcher: A deep detector-based local feature matcher with adaptive weight sharing,”Knowl.-Based Syst., vol. 316, p. 113350, 2025

  68. [76]

    LFM-3D: Learnable feature matching across wide baselines using 3D signals,

    A. Karpur, G. Perrotta, R. Martin-Brualla, H. Zhou, and A. Araujo, “LFM-3D: Learnable feature matching across wide baselines using 3D signals,” inProc. Int. Conf. 3D Vis. (3DV), 2024, pp. 11–20

  69. [77]

    MapGlue: Multimodal remote sensing image matching,

    P. Wu, Y . Yao, W. Zhang, D. Wei, Y . Wan, Y . Li, and Y . Zhang, “MapGlue: Multimodal remote sensing image matching,”arXiv preprint arXiv:2503.16185, 2025

  70. [78]

    Patch2Pix: Epipolar-guided pixel-level correspondences,

    Q. Zhou, T. Sattler, and L. Leal-Taixe, “Patch2Pix: Epipolar-guided pixel-level correspondences,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2021, pp. 4669–4678

  71. [79]

    Match- Former: Interleaving attention in transformers for feature matching,

    Q. Wang, J. Zhang, K. Yang, K. Peng, and R. Stiefelhagen, “Match- Former: Interleaving attention in transformers for feature matching,” in Proc. Asian Conf. Comput. Vis. (ACCV), 2022, pp. 2746–2762

  72. [80]

    ASpanFormer: Detector-free image matching with adaptive span transformer,

    H. Chen, Z. Luo, L. Zhou, Y . Tian, M. Zhen, T. Fang, D. Mckinnon, Y . Tsin, and L. Quan, “ASpanFormer: Detector-free image matching with adaptive span transformer,” inProc. Eur. Conf. Comput. Vis. (ECCV), 2022, pp. 20–36

  73. [81]

    3DG-STFM: 3D geometric guided student-teacher feature matching,

    R. Mao, C. Bai, Y . An, F. Zhu, and C. Lu, “3DG-STFM: 3D geometric guided student-teacher feature matching,” inProc. Eur. Conf. Comput. Vis. (ECCV), 2022, pp. 125–142

  74. [82]

    MSFormer: Multi-scale trans- former with neighborhood consensus for feature matching,

    D. Li, Y . Yan, D. Liang, and S. Du, “MSFormer: Multi-scale trans- former with neighborhood consensus for feature matching,” inProc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP), 2023, pp. 1–5

  75. [83]

    TopicFM: Robust and interpretable topic-assisted feature matching,

    K. T. Giang, S. Song, and S. Jo, “TopicFM: Robust and interpretable topic-assisted feature matching,” inProc. AAAI Conf. Artif. Intell. (AAAI), vol. 37, no. 2, 2023, pp. 2447–2455

  76. [84]

    TopicFM+: Boosting accuracy and efficiency of topic-assisted feature matching,

    Giang, Khang Truong and Song, Soohwan and Jo, Sungho, “TopicFM+: Boosting accuracy and efficiency of topic-assisted feature matching,” IEEE Trans. Image Process., vol. 33, pp. 6016–6028, 2024

  77. [85]

    Adaptive spot- guided transformer for consistent local feature matching,

    J. Yu, J. Chang, J. He, T. Zhang, J. Yu, and F. Wu, “Adaptive spot- guided transformer for consistent local feature matching,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023, pp. 21 898–21 908

  78. [86]

    Structured epipolar matcher for local feature matching,

    J. Chang, J. Yu, and T. Zhang, “Structured epipolar matcher for local feature matching,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023, pp. 6177–6186

  79. [87]

    Affine-based deformable attention and selective fusion for semi-dense matching,

    H. Chen, Z. Luo, Y . Tian, X. Bai, Z. Wang, L. Zhou, M. Zhen, T. Fang, D. McKinnon, Y . Tsinet al., “Affine-based deformable attention and selective fusion for semi-dense matching,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024, pp. 4254–4263

  80. [88]

    Beyond global cues: Unveiling the power of fine details in image matching,

    D. Li and S. Du, “Beyond global cues: Unveiling the power of fine details in image matching,” inProc. IEEE Int. Conf. Multimedia Expo (ICME), 2024, pp. 1–6

  81. [89]

    PRISM: Progressive dependency maximization for scale-invariant image matching,

    X. Cai, Y . Wang, L. Luo, M. Wang, D. Li, J. Xu, W. Gu, and R. Ai, “PRISM: Progressive dependency maximization for scale-invariant image matching,” inProc. ACM Int. Conf. Multimedia (ACM MM), 2024, pp. 5250–5259

  82. [90]

    ContextMatcher: Detector-free feature matching with cross-modality context,

    D. Li and S. Du, “ContextMatcher: Detector-free feature matching with cross-modality context,”IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 9, pp. 7922–7934, 2024

  83. [91]

    DeepMatcher: A deep transformer-based network for robust and accurate local feature matching,

    T. Xie, K. Dai, K. Wang, R. Li, and L. Zhao, “DeepMatcher: A deep transformer-based network for robust and accurate local feature matching,”Expert Syst. Appl., vol. 237, p. 121361, 2024

  84. [92]

    FMAP: Learning robust and accurate local feature matching with anchor points,

    K. Dai, T. Xie, K. Wang, Z. Jiang, R. Li, and L. Zhao, “FMAP: Learning robust and accurate local feature matching with anchor points,”Expert Syst. Appl., vol. 236, p. 121328, 2024

  85. [93]

    OAMatcher: An overlapping areas- based network with label credibility for robust and accurate feature matching,

    Dai, Kun and Xie, Tao and Wang, Ke and Jiang, Zhiqiang and Li, Ruifeng and Zhao, Lijun, “OAMatcher: An overlapping areas- based network with label credibility for robust and accurate feature matching,”Pattern Recognit., vol. 147, p. 110094, 2024

  86. [94]

    DSAP: Dynamic sparse attention perception matcher for accurate local feature matching,

    K. Dai, K. Wang, T. Xie, T. Sun, J. Zhang, Q. Kong, Z. Jiang, R. Li, L. Zhao, and M. Omar, “DSAP: Dynamic sparse attention perception matcher for accurate local feature matching,”IEEE Trans. Instrum. Meas., vol. 73, pp. 1–16, 2024

  87. [95]

    CorMatcher: A corners-guided graph neural network for local feature matching,

    H. Luo, T. Xie, A. Wang, K. Dai, C. Cao, and L. Zhao, “CorMatcher: A corners-guided graph neural network for local feature matching,” Expert Syst. Appl., vol. 258, p. 125190, 2024

  88. [96]

    MR-Matcher: a multirouting transformer-based network for accurate local feature matching,

    Z. Jiang, K. Wang, Q. Kong, K. Dai, T. Xie, Z. Qin, R. Li, P. Perner, and L. Zhao, “MR-Matcher: a multirouting transformer-based network for accurate local feature matching,”IEEE Trans. Instrum. Meas., vol. 73, pp. 1–15, 2024

  89. [97]

    FMRT: Learning accurate feature matching with reconciliatory transformer,

    L. Wang, X. Zhang, Z. Jiang, K. Dai, T. Xie, L. Yang, W. Yu, Y . Shen, B. Xu, and J. Li, “FMRT: Learning accurate feature matching with reconciliatory transformer,”IEEE Trans. Autom. Sci. Eng., 2025

  90. [98]

    WinMRSI: Feature matching with window attention for multi-modal remote sensing image,

    Y . Di, Y . Liao, Y . Liu, H. Zhou, K. Zhu, M. Lu, Q. Duan, and J. Liu, “WinMRSI: Feature matching with window attention for multi-modal remote sensing image,”IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens., 2025

  91. [99]

    LGFCTR: Local and global feature convolu- tional transformer for image matching,

    W. Zhong and J. Jiang, “LGFCTR: Local and global feature convolu- tional transformer for image matching,”Expert Syst. Appl., vol. 270, p. 126393, 2025

  92. [100]

    CoMatch: Dynamic covisibility-aware transformer for bilateral subpixel-level semi-dense image matching,

    Z. Li, Y . Lu, L. Tang, S. Zhang, and J. Ma, “CoMatch: Dynamic covisibility-aware transformer for bilateral subpixel-level semi-dense image matching,”arXiv preprint arXiv:2503.23925, 2025

  93. [101]

    Mind the gap: Aligning vision foundation models to image feature matching,

    Y . Liu, J. Fu, Y . Wu, K. Wu, P. Li, J. Wu, S. Zhou, and J. Xin, “Mind the gap: Aligning vision foundation models to image feature matching,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025, pp. 20 313– 20 323

  94. [102]

    EcoMatcher: Efficient clustering oriented matcher for detector-free image matching,

    P. Chen, L. Yu, Y . Wan, Y . Zhang, J. Wang, L. Zhong, J. Chen, and M. Yang, “EcoMatcher: Efficient clustering oriented matcher for detector-free image matching,” inProc. Eur. Conf. Comput. Vis. (ECCV). Springer, 2024, pp. 344–360

  95. [103]

    Efficient LoFTR: Semi-dense local feature matching with sparse-like speed,

    Y . Wang, X. He, S. Peng, D. Tan, and X. Zhou, “Efficient LoFTR: Semi-dense local feature matching with sparse-like speed,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024, pp. 21 666–21 675

  96. [104]

    MAIM: a mixer MLP architecture for image matching,

    Z. Shen, B. Kong, and X. Dong, “MAIM: a mixer MLP architecture for image matching,”Vis. Comput., vol. 40, no. 3, pp. 1327–1337, 2024

  97. [105]

    EDM: Efficient deep feature matching,

    X. Li, T. Rao, and C. Pan, “EDM: Efficient deep feature matching,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), October 2025, pp. 26 198–26 208

  98. [106]

    CasP: Improving semi-dense feature match- ing pipeline leveraging cascaded correspondence priors for guidance,

    P. Chen, L. Yu, Y . Wan, Y . Pei, X. Liu, Y . Yao, Y . Zhang, L. Ru, L. Zhong, J. Chenet al., “CasP: Improving semi-dense feature match- ing pipeline leveraging cascaded correspondence priors for guidance,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025, pp. 28 063– 28 072. 19

  99. [107]

    VD-Matcher: A very deep local feature matcher with weight recycling and keypoint detection,

    K. Dai, Z. Zhou, Z. Jiang, Q. Sun, T. Xie, H. Gao, T. An, R. Li, and L. Zhao, “VD-Matcher: A very deep local feature matcher with weight recycling and keypoint detection,”IEEE Trans. Circuits Syst. Video Technol., vol. 35, no. 11, pp. 11 416–11 431, 2025

  100. [108]

    LiteSAM: Lightweight and robust feature matching for satellite and aerial imagery,

    B. Wang, S. Wang, Y . Han, L. Xu, and D. Ye, “LiteSAM: Lightweight and robust feature matching for satellite and aerial imagery,”Remote Sens., vol. 17, no. 19, p. 3349, 2025

  101. [109]

    Progressive matchability guided hybrid interaction for latency-balanced local feature matching,

    X. Lu, S. Du, G. Xiao, and G. Wen, “Progressive matchability guided hybrid interaction for latency-balanced local feature matching,”IEEE Trans. Multimedia, 2026

  102. [110]

    PATS: Patch area transportation with subdivision for local feature matching,

    J. Ni, Y . Li, Z. Huang, H. Li, H. Bao, Z. Cui, and G. Zhang, “PATS: Patch area transportation with subdivision for local feature matching,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023, pp. 17 776–17 786

  103. [111]

    Adaptive assignment for geometry aware local feature matching,

    D. Huang, Y . Chen, Y . Liu, J. Liu, S. Xu, W. Wu, Y . Ding, F. Tang, and C. Wang, “Adaptive assignment for geometry aware local feature matching,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023, pp. 5425–5434

  104. [112]

    Raising the ceiling: conflict-free local feature matching with dynamic view switching,

    X. Lu and S. Du, “Raising the ceiling: conflict-free local feature matching with dynamic view switching,” inProc. Eur. Conf. Comput. Vis. (ECCV). Springer, 2024, pp. 256–273

  105. [113]

    Improving transformer-based image matching by cascaded capturing spatially informative keypoints,

    C. Cao and Y . Fu, “Improving transformer-based image matching by cascaded capturing spatially informative keypoints,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023, pp. 12 129–12 139

  106. [114]

    HomoMatcher: Dense feature matching results with semi-dense efficiency by homography estimation,

    X. Wang, L. Yu, Y . Zhang, J. Lao, L. Ru, L. Zhong, J. Chen, Y . Zhang, and M. Yang, “HomoMatcher: Dense feature matching results with semi-dense efficiency by homography estimation,”arXiv preprint arXiv:2411.06700, 2024

  107. [115]

    Semi-dense feature matching with increased matching amount,

    Y . Di, Y . Liao, Y . Liu, M. Lu, Q. Duan, and J. Liu, “Semi-dense feature matching with increased matching amount,”Vis. Comput., vol. 41, no. 13, pp. 11 407–11 426, 2025

  108. [116]

    Towards free-form local feature matching,

    X. Lu, S. Du, Y . Yan, X. Lu, and T. Ikenaga, “Towards free-form local feature matching,”IEEE Trans. Pattern Anal. Mach. Intell., 2025

  109. [117]

    Convolutional neural network architecture for geometric matching,

    I. Rocco, R. Arandjelovic, and J. Sivic, “Convolutional neural network architecture for geometric matching,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2017, pp. 6148–6157

  110. [118]

    DGC-Net: Dense geometric correspondence network,

    I. Melekhov, A. Tiulpin, T. Sattler, M. Pollefeys, E. Rahtu, and J. Kannala, “DGC-Net: Dense geometric correspondence network,” in Proc. IEEE Winter Conf. Appl. Comput. Vis. (WACV), 2019, pp. 1034– 1042

  111. [119]

    GLU-Net: Global-local universal network for dense flow and correspondences,

    P. Truong, M. Danelljan, and R. Timofte, “GLU-Net: Global-local universal network for dense flow and correspondences,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020, pp. 6258–6268

  112. [120]

    GOCor: Bringing globally optimized correspondence volumes into your neural network,

    P. Truong, M. Danelljan, L. Van Gool, and R. Timofte, “GOCor: Bringing globally optimized correspondence volumes into your neural network,” inProc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2020, pp. 14 278–14 290

  113. [121]

    Learning accurate dense correspondences and when to trust them,

    Truong, Prune and Danelljan, Martin and Van Gool, Luc and Timofte, Radu, “Learning accurate dense correspondences and when to trust them,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2021, pp. 5714–5724

  114. [122]

    DKM: Dense kernelized feature matching for geometry estimation,

    J. Edstedt, I. Athanasiadis, M. Wadenb ¨ack, and M. Felsberg, “DKM: Dense kernelized feature matching for geometry estimation,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023, pp. 17 765–17 775

  115. [123]

    RGM: A robust gen- eralizable matching model,

    S. Zhang, X. Sun, H. Chen, B. Li, and C. Shen, “RGM: A robust gen- eralizable matching model,”arXiv preprint arXiv:2310.11755, 2023

  116. [124]

    PanMatch: Unleashing the potential of large vision models for unified matching models,

    Y . Zhang, L. Wang, K. Li, Y . Zhang, Y . Wang, L. Lin, and Y . Guo, “PanMatch: Unleashing the potential of large vision models for unified matching models,”arXiv preprint arXiv:2507.08400, 2025

  117. [125]

    Handling multiple hypotheses in coarse-to-fine dense image matching,

    M. Vilain, R. Giraud, Y . Berthoumieu, and G. Bourmaud, “Handling multiple hypotheses in coarse-to-fine dense image matching,” inProc. IEEE Int. Conf. Image Process. (ICIP), 2025, pp. 1265–1270

  118. [126]

    UFM: A simple path towards unified dense correspondence with flow,

    Y . Zhang, N. Keetha, C. Lyu, B. Jhamb, Y . Chen, Y . Qiu, J. Karhade, S. Jha, Y . Hu, D. Ramananet al., “UFM: A simple path towards unified dense correspondence with flow,”arXiv preprint arXiv:2506.09278, 2025

  119. [127]

    Roma v2: Harder better faster denser feature matching,

    J. Edstedt, D. Nordstr ¨om, Y . Zhang, G. B ¨okman, J. Astermark, V . Larsson, A. Heyden, F. Kahl, M. Wadenb ¨ack, and M. Felsberg, “Roma v2: Harder better faster denser feature matching,”arXiv preprint arXiv:2511.15706, 2025

  120. [128]

    Adapting dense matching for homography estimation with grid-based acceleration,

    K. Zhang, Y . Deng, J. Ma, and P. Favaro, “Adapting dense matching for homography estimation with grid-based acceleration,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025, pp. 6294–6303

  121. [129]

    ArgMatch: Adaptive refinement gathering for efficient dense matching,

    Y . Deng, K. Zhang, L. Tang, J. Yang, and J. Ma, “ArgMatch: Adaptive refinement gathering for efficient dense matching,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025, pp. 27 369–27 379

  122. [130]

    HDPL: a hybrid descriptor for points and lines based on graph neural networks,

    Z. Guo, H. Lu, Q. Yu, R. Guo, J. Xiao, and H. Yu, “HDPL: a hybrid descriptor for points and lines based on graph neural networks,”Ind. Robot, vol. 48, no. 5, pp. 737–744, 2021

  123. [131]

    GlueStick: Robust image matching by sticking points and lines together,

    R. Pautrat, I. Su ´arez, Y . Yu, M. Pollefeys, and V . Larsson, “GlueStick: Robust image matching by sticking points and lines together,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023, pp. 9706–9716

  124. [132]

    Searching from area to point: A hierarchical framework for semantic-geometric combined feature matching,

    Y . Zhang and X. Zhao, “Searching from area to point: A hierarchical framework for semantic-geometric combined feature matching,”arXiv preprint arXiv:2305.00194, 2023

  125. [133]

    MESA: Matching everything by segmenting anything,

    Zhang, Yesheng and Zhao, Xu, “MESA: Matching everything by segmenting anything,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024, pp. 20 217–20 226

  126. [134]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Loet al., “Segment anything,” inProc. Int. Conf. Comput. Vis. (ICCV), 2023, pp. 4015–4026

  127. [135]

    SGAD: Semantic and geometric-aware descriptor for local feature matching,

    X. Liu, C. Wang, G. Shi, X. Zhang, Q. Miao, and M. Fan, “SGAD: Semantic and geometric-aware descriptor for local feature matching,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025, pp. 27 095– 27 104

  128. [136]

    Single-image depth prediction makes feature matching easier,

    C. Toft, D. Turmukhambetov, T. Sattler, F. Kahl, and G. J. Brostow, “Single-image depth prediction makes feature matching easier,” in Proc. Eur. Conf. Comput. Vis. (ECCV). Springer, 2020, pp. 473–492

  129. [137]

    Guiding local feature matching with surface curvature,

    S. Wang, J. Kannala, M. Pollefeys, and D. Barath, “Guiding local feature matching with surface curvature,” inProc. Int. Conf. Comput. Vis. (ICCV), 2023, pp. 17 981–17 991

  130. [138]

    LiftFeat: 3D geometry-aware local feature matching,

    Y . Liu, W. Lai, Z. Zhao, Y . Xiong, J. Zhu, J. Cheng, and Y . Xu, “LiftFeat: 3D geometry-aware local feature matching,” inProc. IEEE Int. Conf. Robot. Autom. (ICRA), 2025, pp. 11 714–11 720

  131. [139]

    Semantic-aware representation learning for homography estimation,

    Y . Liu, Q. Huang, S. Hui, J. Fu, S. Zhou, K. Wu, P. Li, and J. Wang, “Semantic-aware representation learning for homography estimation,” inProc. ACM Int. Conf. Multimedia (ACM MM), 2024, pp. 2506–2514

  132. [140]

    Leveraging semantic cues from foundation vision models for enhanced local feature correspondence,

    F. Cadar, G. Potje, R. Martins, C. Demonceaux, and E. R. Nascimento, “Leveraging semantic cues from foundation vision models for enhanced local feature correspondence,” inProc. Asian Conf. Comput. Vis. (ACCV), 2025, pp. 54–70

  133. [141]

    DINO-VO: A feature-based visual odometry leveraging a visual foundation model,

    M. B. Azhari and D. H. Shim, “DINO-VO: A feature-based visual odometry leveraging a visual foundation model,”IEEE Robot. Autom. Lett., vol. 10, no. 9, pp. 9152–9159, 2025

  134. [142]

    SigMa: Semantic similarity-guided semi- dense feature matching,

    X. Fang, Z. Li, and J. Ma, “SigMa: Semantic similarity-guided semi- dense feature matching,”IEEE Trans. Image Process., vol. 35, pp. 872– 887, 2026

  135. [143]

    SGAT: learning feature matching with singularity-enhanced graph attention network,

    Y . Zhang, K. Sun, C. Tang, Y . Liu, and X. Li, “SGAT: learning feature matching with singularity-enhanced graph attention network,” inProc. AAAI Conf. Artif. Intell. (AAAI), 2026, pp. 12 952–12 960

  136. [144]

    DistillMatch: Leveraging knowledge distillation from vision foundation model for multimodal image matching,

    M. Yang, F. Fan, Z. Li, S. Deng, Y . Ma, and J. Ma, “DistillMatch: Leveraging knowledge distillation from vision foundation model for multimodal image matching,”arXiv preprint arXiv:2509.16017, 2025

  137. [145]

    TextFM: Robust semi-dense feature matching with language guidance,

    Z. Zheng, J. Feng, N. Savaliya, Z.-H. Yeh, B. Lang, and M. C. Chuah, “TextFM: Robust semi-dense feature matching with language guidance,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026, pp. 16 635–16 644

  138. [146]

    RIM: A retrieval-in-matching framework for cross-domain global visual localization of uavs,

    X. Li, S. Duan, S. Wang, Z. Mao, B. Hu, and G. Zhang, “RIM: A retrieval-in-matching framework for cross-domain global visual localization of uavs,”arXiv preprint arXiv:2607.20116, 2026

  139. [147]

    MV-RoMa: From pairwise matching into multi-view track reconstruction,

    J. Lee, S. Kang, and S. Yoo, “MV-RoMa: From pairwise matching into multi-view track reconstruction,”arXiv preprint arXiv:2603.27542, 2026

  140. [148]

    DMESA: Densely matching everything by segmenting anything,

    Y . Zhang and X. Zhao, “DMESA: Densely matching everything by segmenting anything,”arXiv preprint arXiv:2408.00279, 2024

  141. [149]

    Segment anything model is a good teacher for local feature learning,

    J. Wu, R. Xu, Z. Wood-Doughty, C. Wang, S. Xu, and E. Y . Lam, “Segment anything model is a good teacher for local feature learning,” IEEE Trans. Image Process., vol. 34, pp. 2097–2111, 2025

  142. [150]

    SAMatcher: Co-visibility modeling with segment anything for robust feature match- ing,

    X. Pan, Q. Ma, M. Dong, H. Chen, W. Ji, and X. Zheng, “SAMatcher: Co-visibility modeling with segment anything for robust feature match- ing,”arXiv preprint arXiv:2606.03406, 2026

  143. [151]

    Emergent correspondence from image diffusion,

    L. Tang, M. Jia, Q. Wang, C. P. Phoo, and B. Hariharan, “Emergent correspondence from image diffusion,” inProc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2023, pp. 1363–1389

  144. [152]

    MATCHA: Towards matching anything,

    F. Xue, S. Elflein, L. Leal-Taix ´e, and Q. Zhou, “MATCHA: Towards matching anything,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025, pp. 27 081–27 091

  145. [153]

    DUSt3R: Geometric 3D vision made easy,

    S. Wang, V . Leroy, Y . Cabon, B. Chidlovskii, and J. Revaud, “DUSt3R: Geometric 3D vision made easy,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024, pp. 20 697–20 709. 20

  146. [154]

    Depth anything v2,

    L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything v2,” inProc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2024, pp. 21 875–21 911

  147. [155]

    Sgpfeat: Semantic and geometric priors for multi-modal image matching,

    Y . Deng, B. Wang, K. Zhang, H. Zhang, and J. Ma, “Sgpfeat: Semantic and geometric priors for multi-modal image matching,” inProc. AAAI Conf. Artif. Intell. (AAAI), 2026, pp. 3569–3577

  148. [156]

    PMatch: Paired masked image modeling for dense geometric matching,

    S. Zhu and X. Liu, “PMatch: Paired masked image modeling for dense geometric matching,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023, pp. 21 909–21 918

  149. [157]

    GIM: Learning generalizable image matcher from internet videos,

    X. Shen, Z. Cai, W. Yin, M. M ¨uller, Z. Li, K. Wang, X. Chen, and C. Wang, “GIM: Learning generalizable image matcher from internet videos,”arXiv preprint arXiv:2402.11095, 2024

  150. [158]

    MINIMA: Modality invariant image matching,

    J. Ren, X. Jiang, Z. Li, D. Liang, X. Zhou, and X. Bai, “MINIMA: Modality invariant image matching,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025, pp. 23 059–23 068

  151. [159]

    MatchAnything: Universal cross-modality image matching with large- scale pre-training,

    X. He, H. Yu, S. Peng, D. Tan, Z. Shen, H. Bao, and X. Zhou, “MatchAnything: Universal cross-modality image matching with large- scale pre-training,”arXiv preprint arXiv:2501.07556, 2025

  152. [160]

    Learning dense feature matching via lifting single 2D image to 3D space,

    Y . Liang, Y . Hu, W. Shao, and Y . Fu, “Learning dense feature matching via lifting single 2D image to 3D space,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), October 2025, pp. 6621–6631

  153. [161]

    P. J. Huber and E. M. Ronchetti,Robust Statistics, ser. Wiley Series in Probability and Statistics. Hoboken, NJ: John Wiley & Sons, Inc., 2009

  154. [162]

    Diakonikolas and D

    I. Diakonikolas and D. M. Kane,Algorithmic high-dimensional robust statistics. Cambridge, United Kingdom: Cambridge University Press, 2023

  155. [163]

    Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography,

    M. A. Fischler and R. C. Bolles, “Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography,”Commun. ACM, vol. 24, no. 6, pp. 381–395, 1981

  156. [164]

    RANSAC for robotic applications: A survey,

    J. M. Mart ´ınez-Otzeta, I. Rodr ´ıguez-Moreno, I. Mendialdua, and B. Sierra, “RANSAC for robotic applications: A survey,”Sensors, vol. 23, no. 1, pp. 327.1–327.26, 2023

  157. [165]

    Robust parameterization and computation of the trifocal tensor,

    P. H. Torr and A. Zisserman, “Robust parameterization and computation of the trifocal tensor,”Image Vis. Comput., vol. 15, no. 8, pp. 591–605, 1997

  158. [166]

    MLESAC: A new robust estimator with application to estimating image geometry,

    P. H. S. Torr and A. Zisserman, “MLESAC: A new robust estimator with application to estimating image geometry,”Comput. Vis. Image Underst., vol. 78, no. 1, pp. 138–156, 2000

  159. [167]

    Bayesian model estimation and selection for epipolar geometry and generic manifold fitting,

    P. H. S. Torr, “Bayesian model estimation and selection for epipolar geometry and generic manifold fitting,”Int. J. Comput. Vis., vol. 50, no. 1, pp. 35–61, 2002

  160. [168]

    Locally optimized RANSAC,

    O. Chum, J. Matas, and J. Kittler, “Locally optimized RANSAC,” in Joint Pattern Recognition Symposium, 2003, pp. 236–243

  161. [169]

    ANSAC: Adaptive non-minimal sample and consensus,

    P. S. Victor Fragoso, Christopher Sweeney and M. Turk, “ANSAC: Adaptive non-minimal sample and consensus,” inProc. Br. Mach. Vis. Conf. (BMVC), September 2017, pp. 43.1–43.11

  162. [170]

    BANSAC: A dynamic Bayesian network for adaptive sample consensus,

    V . Piedade and P. Miraldo, “BANSAC: A dynamic Bayesian network for adaptive sample consensus,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023, pp. 3715–3724

  163. [171]

    A polynomial-time bound for matching and registration with outliers,

    C. Olsson, O. Enqvist, and F. Kahl, “A polynomial-time bound for matching and registration with outliers,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2008

  164. [172]

    An exact penalty method for locally convergent maximum consensus,

    H. Le, T.-J. Chin, and D. Suter, “An exact penalty method for locally convergent maximum consensus,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2017, pp. 1888–1896

  165. [173]

    Deterministic consensus max- imization with biconvex programming,

    Z. Cai, T.-J. Chin, H. Le, and D. Suter, “Deterministic consensus max- imization with biconvex programming,” inProc. Eur. Conf. Comput. Vis. (ECCV), 2018, pp. 685–700

  166. [174]

    Graph-cut RANSAC,

    D. Barath and J. Matas, “Graph-cut RANSAC,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018, pp. 6733–6741

  167. [175]

    Graph-Cut RANSAC: Local optimiza- tion on spatially coherent structures,

    Barath, Daniel and Matas, Jiri, “Graph-Cut RANSAC: Local optimiza- tion on spatially coherent structures,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 9, pp. 4961–4974, 2022

  168. [176]

    Robust estimation via robust gradient estimation,

    A. Prasad, A. S. Suggala, S. Balakrishnan, and P. Ravikumar, “Robust estimation via robust gradient estimation,”J. R. Stat. Soc. Ser. B Stat. Methodol., vol. 82, no. 3, pp. 601–627, 2020

  169. [177]

    DSAC - differentiable RANSAC for cam- era localization,

    E. Brachmann, A. Krull, S. Nowozin, J. Shotton, F. Michel, S. Gumhold, and C. Rother, “DSAC - differentiable RANSAC for cam- era localization,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2017, pp. 6684–6692

  170. [178]

    Learning to find good models in RANSAC,

    D. Barath, L. Cavalli, and M. Pollefeys, “Learning to find good models in RANSAC,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022, pp. 15 744–15 753

  171. [179]

    Progressive correspondence pruning by consensus learning,

    C. Zhao, Y . Ge, F. Zhu, R. Zhao, H. Li, and M. Salzmann, “Progressive correspondence pruning by consensus learning,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021, pp. 6464–6473

  172. [180]

    Neural-guided RANSAC: Learning where to sample model hypotheses,

    E. Brachmann and C. Rother, “Neural-guided RANSAC: Learning where to sample model hypotheses,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2019, pp. 4322–4331

  173. [181]

    Generalized differentiable RANSAC,

    T. Wei, Y . Patel, A. Shekhovtsov, J. Matas, and D. Barath, “Generalized differentiable RANSAC,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023, pp. 17 649–17 660

  174. [182]

    Unsupervised learning of consensus maximization for 3D vision problems,

    T. Probst, D. P. Paudel, A. Chhatkuli, and L. V . Gool, “Unsupervised learning of consensus maximization for 3D vision problems,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019, pp. 929–938

  175. [183]

    YFCC100M: The new data in multimedia research,

    B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li, “YFCC100M: The new data in multimedia research,”Commun. ACM, vol. 59, no. 2, pp. 64–73, 2016

  176. [184]

    Learning two-view correspondences and geometry using order-aware network,

    J. Zhang, D. Sun, Z. Luo, A. Yao, L. Zhou, T. Shen, Y . Chen, L. Quan, and H. Liao, “Learning two-view correspondences and geometry using order-aware network,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2019, pp. 5845–5854

  177. [185]

    Structure-from-motion revisited,

    J. L. Schonberger and J.-M. Frahm, “Structure-from-motion revisited,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2016, pp. 4104–4113

  178. [186]

    MegaDepth: Learning single-view depth pre- diction from internet photos,

    Z. Li and N. Snavely, “MegaDepth: Learning single-view depth pre- diction from internet photos,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018, pp. 2041–2050

  179. [187]

    ScanNet: Richly-annotated 3D reconstructions of indoor scenes,

    A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “ScanNet: Richly-annotated 3D reconstructions of indoor scenes,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2017, pp. 5828–5839

  180. [188]

    SUN3D: A database of big spaces reconstructed using SfM and object labels,

    J. Xiao, A. Owens, and A. Torralba, “SUN3D: A database of big spaces reconstructed using SfM and object labels,” inProc. Int. Conf. Comput. Vis. (ICCV), 2013, pp. 1625–1632

  181. [189]

    Hartley and A

    R. Hartley and A. Zisserman,Multiple view geometry in computer vision. Cambridge University Press, 2003

  182. [190]

    HPatches: A benchmark and evaluation of handcrafted and learned local descrip- tors,

    V . Balntas, K. Lenc, A. Vedaldi, and K. Mikolajczyk, “HPatches: A benchmark and evaluation of handcrafted and learned local descrip- tors,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2017, pp. 5173–5182

  183. [191]

    Bench- marking 6DoF outdoor visual localization in changing conditions,

    T. Sattler, W. Maddern, C. Toft, A. Torii, L. Hammarstrand, E. Sten- borg, D. Safari, M. Okutomi, M. Pollefeys, J. Sivicet al., “Bench- marking 6DoF outdoor visual localization in changing conditions,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018, pp. 8601–8610

  184. [192]

    InLoc: Indoor visual localization with dense matching and view synthesis,

    H. Taira, M. Okutomi, T. Sattler, M. Cimpoi, M. Pollefeys, J. Sivic, T. Pajdla, and A. Torii, “InLoc: Indoor visual localization with dense matching and view synthesis,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018, pp. 7199–7209

  185. [193]

    From coarse to fine: Robust hierarchical localization at large scale,

    P.-E. Sarlin, C. Cadena, R. Siegwart, and M. Dymczyk, “From coarse to fine: Robust hierarchical localization at large scale,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019, pp. 12 716–12 725

  186. [194]

    NetVLAD: CNN architecture for weakly supervised place recognition,

    R. Arandjelovi ´c, P. Gronat, A. Torii, T. Pajdla, and J. Sivic, “NetVLAD: CNN architecture for weakly supervised place recognition,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 40, no. 6, pp. 1437–1451, 2018

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.