REVIEW 3 major objections 6 minor 53 references
Beyond Cartesian Representations for Local Descriptors
T0 review · 3 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Sampling image patches in log-polar coordinates rather than on a Cartesian grid lets learned local descriptors match keypoints even when the detected scales differ by a factor of 2-3x, with usable performance at 3-4x.
desk verdict Log-polar sampling of raw pixels is a genuinely new and well-supported idea; the headline scale range is measured on a self-built, COLMAP-dependent test set, so treat the exact 2-4x numbers as provisional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the log-polar sampling pattern, implemented as a polar transformer sampler. Instead of pooling features computed on a Cartesian grid, as earlier log-polar descriptors did, the sampler warps the raw pixel intensities so that the target patch is indexed by log radius and angle. A rotation of the scene becomes a shift along the angular axis, and a scale change becomes a shift along the radial axis, making mismatched-scale patches look similar both to the eye and to the network. This is what lets the same seven-layer convolutional network with a hardest-in-batch triplet loss learn scale-invariant descriptors without any change to the architecture.
What would settle it
Take image pairs with known homographies and deliberately extract keypoints whose scales differ by 2x and 3x; if the log-polar descriptors fail to match them at the same rate as matched-scale keypoints, the central scale-invariance claim fails. Concretely, compute rank-1 accuracy on such controlled mismatched-scale pairs and compare the degradation pattern with the paper's Fig. 4.
Extended reading notes
Core claim
The central claim is that the representation itself, not the network architecture or training loss, is what confers scale invariance. The authors extract a 32x32 log-polar-warped patch around each SIFT keypoint using a polar transformer sampler, while keeping the same convolutional descriptor architecture and hardest-in-batch triplet loss used for Cartesian patches. On a training set built from real photo-tourism images with ground-truth depth maps, they deliberately match keypoints whose detected scales are not in correspondence. Their log-polar models achieve the best false-positive-rate-at-95%-recall on all test sequences, tolerate scale changes up to 2-3x with negligible drop and 3-4x with usable performance, and improve as the support-region size grows up to a radius 8 times larger than the Cartesian optimum, whereas Cartesian models degrade sharply. The same models transfer without fine-tuning to public benchmarks and rank near the top on a pose-estimation challenge.
Load-bearing premise
The ground-truth matches used for training come from depth maps and camera poses produced by a reconstruction pipeline, so if those reconstructions are noisy or biased toward certain scales, the measured scale-invariance gains may be partly an artifact of the training data.
Editorial extensions
If this is right
- Scale errors from keypoint detectors stop being fatal: matching can proceed even when the two detected scales differ by 2-3x, without re-detecting or re-scaling keypoints.
- Support regions can be 8 times larger in radius, or 64 times larger in area, than the Cartesian optimum, so descriptors can use more context without being derailed by occlusions or background motion.
- Exposing a network to mismatched-scale Cartesian patches is not enough; the invariance comes from the log-polar warp itself, implying that future invariant descriptors can be designed at the sampling stage.
- The models transfer directly to other benchmarks and to pose estimation, so the gains are not confined to the authors' training data.
- The same representation should simplify end-to-end pipelines that currently learn scale detection separately, because scale is encoded as a shift rather than as a parameter to estimate.
Reading between the lines
- A natural extension, not pursued in the paper, is to make rotation invariance explicit: the log-polar angular axis suggests that full rotation invariance could be learned or even read off directly, building on the orientation jitter already used in training.
- The same warp could be used as a scale estimator: because scale changes become shifts, a small correlation layer could recover the scale ratio between matched patches, turning a known failure mode into a measurable quantity.
- The boundary problems observed beyond a support radius of 96 point to a concrete improvement: padding or masking the warped patch rather than relying on mirror-padded images could extend the gains to even larger support regions.
- More generally, the result invites testing other coordinate warps, such as affine, cylindrical, or spherical sampling, as the input layer for learned descriptors rather than only log-polar.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes replacing the standard Cartesian sampling of local patch support regions with direct log-polar sampling via a Polar Transformer Network (PTN), followed by a HardNet-style CNN that produces 128-dimensional descriptors. The authors build a new training set from YFCC photo-tourism images processed with COLMAP, deliberately collecting SIFT keypoint pairs whose scales are not matched, and train with a hardest-in-batch triplet loss. They report FPR95 on their own test set, rank-1 retrieval on HPatches and AMOS, and pose-estimation mAP on the CVPR 2019 PhotoTourism challenge. The central claims are that log-polar descriptors tolerate scale mismatches up to 2-4x with negligible degradation, that they can exploit much larger support regions (up to lambda=96) without the performance collapse seen for Cartesian patches, and that this yields state-of-the-art results on three datasets.
Significance. If the claims hold, the paper makes a useful and potentially influential contribution: it decouples descriptor matching from the accuracy of detector scale estimation and shows that a simple change in sampling geometry, combined with standard deep learning, can enlarge effective support regions. Strengths include the public release of code, models, and training data; the controlled lambda ablation that keeps the architecture and training pipeline identical between Cartesian and log-polar variants; and validation on three external benchmarks (HPatches, AMOS, PhotoTourism), which goes beyond a single in-house evaluation. The principal weakness is that the quantitative 2-4x scale-invariance claim is established almost entirely on an internally generated test set whose ground-truth correspondences and scale-ratio histogram derive from the same COLMAP depth/pose estimates used to create the training data; an independent, scale-controlled evaluation is needed to make the central claim fully credible.
major comments (3)
- [Section 4.1.2 and Fig. 4] The central claim that log-polar models show 'a negligible drop in performance under scale changes up to 2-3x, and remain useful even at 3-4x' is evaluated only on the authors' COLMAP-based test set. Ground-truth correspondences in Section 4.1.1 are generated by projecting SIFT keypoints through COLMAP depth maps and estimated poses with a 1.5-pixel threshold and a cyclic consistency check, and the scale-ratio histogram in Fig. 4 is computed from these same depth/pose estimates. If COLMAP depth errors are correlated with appearance or with scale ratio, the positive pairs in the largest-scale bins could be systematically easier for log-polar patches for reasons unrelated to genuine scale invariance. Because this directly supports the abstract's 'much wider range of scales' claim, the authors should add an independent scale-controlled evaluation, for example using synthetic images with known homographies and zoom factors, or stratifying HPatches by the scale ratio induced by the ground-truth homography, and report matching accuracy per scale-ratio bin.
- [Abstract; Section 4.1.2 and Table 2] The abstract and Section 4.1.2 state that log-polar support regions can be made much larger 'without suffering from occlusions,' but no experiment directly measures occlusion robustness. Table 2 shows that log-polar FPR95 improves with lambda while Cartesian degrades, yet this is measured on the general COLMAP test set, where a larger lambda also exposes more scale variation and more background content; the improvement could be driven by scale equivariance rather than by occlusion tolerance. An explicit occlusion experiment (for example, masking or occluding controlled fractions of the support region, or reporting performance separately for keypoints with known occlusion masks) is required to substantiate this load-bearing claim.
- [Tables 1-4] No error bars, repeated runs, or statistical significance tests are reported for any FPR95 or rank-1 numbers. For example, in Table 3 the viewpoint split shows Ours-LogPol lambda=96 at 0.847 versus lambda=64 at 0.849, and several baseline differences are on the order of 0.001-0.01, so wording such as 'performance increases with lambda, until it saturates' and some comparative statements could depend on variance. Reporting per-sequence standard deviations across training runs or a significance test over sequences would make the empirical conclusions more robust, especially where the reported gaps are small.
minor comments (6)
- [Section 3.1, Eq. (1)] The expression e^{log(r_i)} in Eq. (1) should be simplified to r_i; as written it introduces a redundant exponential/log pair and obscures the intended radial coordinate.
- [Section 3.1, Eq. (2)] The second line of Eq. (2) writes y_t = y_i + x_s sin(theta_i) sigma_i/W + y_s sin(theta_i) sigma_i/H; the second sine should likely be a cosine to represent a standard rotation in the Cartesian sampler. Please check and correct the formula.
- [Section 4.1.1] The text refers to COLMAP outputs as 'ground truth camera poses' and later to 'ground truth correspondences'; since these are estimated quantities, the wording should be softened to 'estimated poses and depth' to avoid implying independent ground truth.
- [Abstract and Section 4.4, Table 5] The abstract's 'state-of-the-art results on three different datasets' overstates Table 5, where the method ranks second on both PhotoTourism tracks, although it is first by average rank. Consider describing the result as 'top-performing on average' or 'second on both tracks' for accuracy.
- [Fig. 4 caption] Several scale/orientation bins are sparsely populated, as the caption acknowledges. Adding sample counts per bin, or suppressing bins with very few matches, would help readers judge which parts of the 2-4x scale-invariance curve are reliable.
- [Section 4.1.2, Table 1] The text says 'the small gap between HardNet and Ours-Cartesian,' but the average FPR95 values are 0.98 and 0.72, respectively; 'small relative to the other baselines' would be more accurate.
Circularity Check
No circularity: the scale-invariance claim is an empirical result tested on held-out and external benchmarks, not a reduction to the training objective.
full rationale
The paper's central claim is that log-polar sampling of patches, combined with a standard descriptor-learning network, yields descriptors that tolerate larger scale mismatches than Cartesian sampling. This is an empirical finding, not a mathematical derivation that reduces to its inputs. The authors train on a new COLMAP-derived dataset with deliberately mismatched scales, but they evaluate on held-out test sequences from the same data collection procedure and, crucially, on three external benchmarks (HPatches, AMOS patches, and the PhotoTourism challenge) without fine-tuning. The comparison between log-polar and Cartesian models is controlled: both are trained under identical settings, and the Cartesian model does not achieve the same scale invariance, so the improvement is not merely an artifact of training on mismatched scales. The hyperparameter lambda is selected by performance, but the claim that larger support regions help log-polar models is corroborated by monotonic trends on external datasets and by the contrasting degradation of Cartesian models, so no fitted parameter is being renamed as a prediction. The few self-citations, such as using the HardNet architecture and triplet loss, are not load-bearing: those methods are established, externally validated baselines, and the paper's contribution does not depend on an unverified self-citation chain. The use of COLMAP-derived depth and poses for both training and test correspondences is a legitimate concern about benchmark quality and generalization, but it is not a circularity in the sense of a derivation that assumes what it claims to prove. No equation in the paper reduces to another by construction, and no claimed result is definitionally equivalent to its inputs.
Assumptions & free parameters
free parameters (5)
- lambda (support region scale multiplier) =
96
- orientation jitter standard deviation =
25 degrees
- correspondence projection threshold =
1.5 pixels
- orientation consistency threshold =
25 degrees
- minimum keypoint spacing =
7 pixels
assumptions (3)
- standard math Rotations in Cartesian space correspond to shifts along the polar axis, and scale changes correspond to shifts along the log-radius axis in log-polar coordinates.
- domain assumption COLMAP's depth maps and camera poses are accurate enough to establish reliable correspondences under a 1.5-pixel projection threshold and cyclic consistency check.
- domain assumption SIFT keypoints with inaccurate scale estimates occur frequently enough in real images to make scale-invariant descriptors useful.
Cite this review
Pith. "Pith review of Beyond Cartesian Representations for Local Descriptors." pith.science (2026). https://pith.science/paper/ITKZHAYT
@misc{pith2026190805547,
author = {Pith},
title = {Pith review of: Beyond Cartesian Representations for Local Descriptors},
year = {2026},
howpublished = {\url{https://pith.science/paper/ITKZHAYT}},
note = {Machine review of arXiv:1908.05547}
}
read the original abstract
The dominant approach for learning local patch descriptors relies on small image regions whose scale must be properly estimated a priori by a keypoint detector. In other words, if two patches are not in correspondence, their descriptors will not match. A strategy often used to alleviate this problem is to "pool" the pixel-wise features over log-polar regions, rather than regularly spaced ones. By contrast, we propose to extract the "support region" directly with a log-polar sampling scheme. We show that this provides us with a better representation by simultaneously oversampling the immediate neighbourhood of the point and undersampling regions far away from it. We demonstrate that this representation is particularly amenable to learning descriptors with deep networks. Our models can match descriptors across a much wider range of scales than was possible before, and also leverage much larger support regions without suffering from occlusions. We report state-of-the-art results on three different datasets.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
https://image-matching-workshop
Phototourism Challenge, CVPR 2019 Image Matching Workshop. https://image-matching-workshop. github.io. Accessed August 1, 2019. 5, 8
work page 2019
-
[2]
Alexandre Alahi, Raphael Ortiz, and Pierre Vandergheynst. FREAK: Fast Retina Keypoint. In CVPR, 2012. 3
work page 2012
-
[3]
Hpatches: A Benchmark and Evaluation of Handcrafted and Learned Local Descriptors
Vassileios Balntas, Karel Lenc, Andrea Vedaldi, and Krys- tian Mikolajczyk. Hpatches: A Benchmark and Evaluation of Handcrafted and Learned Local Descriptors. In CVPR,
-
[4]
Learning Local Feature Descriptors with Triplets and Shallow Convolutional Neural Networks
Vassileios Balntas, Edgar Riba, Daniel Ponsa, and Krys- tian Mikolajczyk. Learning Local Feature Descriptors with Triplets and Shallow Convolutional Neural Networks. In BMVC, 2016. 2, 3, 4, 5
work page 2016
-
[5]
SURF: Speeded Up Robust Features
Herbert Bay, Tinne Tuytelaars, and Luc Van Gool. SURF: Speeded Up Robust Features. In ECCV, 2006. 1, 2
work page 2006
-
[6]
Shape Matching and Object Recognition Using Shape Contexts
Serge Belongie, Jitendra Malik, and Jan Puzicha. Shape Matching and Object Recognition Using Shape Contexts. PAMI, 24(24):509–522, April 2002. 3
work page 2002
-
[7]
Discrimi- native Learning of Local Image Descriptors
Matthew Brown, Gang Hua, and Simon Winder. Discrimi- native Learning of Local Image Descriptors. PAMI, 2011. 2, 5, 7
work page 2011
-
[8]
From handcrafted to deep local features
Gabriela Csurka and Martin Humenberger. From hand- crafted to deep local invariant features. In arXiv preprint arXiv:1807.10254, 2018. 2
work page Pith review arXiv 2018
Show all 53 references
-
[9]
Superpoint: Self-Supervised Interest Point Detec- tion and Description
Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabi- novich. Superpoint: Self-Supervised Interest Point Detec- tion and Description. CVPR Workshop on Deep Learning for Visual SLAM, 2018. 2
2018
-
[10]
D2-Net: A Trainable CNN for Joint Detection and Description of Lo- cal Features
Mihai Dusmanu, Ignacio Rocco, Tomas Pajdla, Marc Polle- feys, Josef Sivic, Akihiko Torii, and Torsten Sattler. D2-Net: A Trainable CNN for Joint Detection and Description of Lo- cal Features. In CVPR, 2019. 1, 2
2019
-
[11]
Polar Transformer Networks
Carlos Esteves, Christine Allen-Blanchette, Xiaowei Zhou, and Kostas Daniilidis. Polar Transformer Networks. In ICLR, 2018. 3
2018
-
[12]
Xufeng Han, Thomas Leung, Yangqing Jia, Rahul Suk- thankar, and Alexander C. Berg. MatchNet: Unifying Fea- ture and Metric Learning for Patch-Based Matching. In CVPR, 2015. 1, 2, 3
2015
-
[13]
Harris and Mike .J
Christopher G. Harris and Mike .J. Stephens. A Combined Corner and Edge Detector. In F ourth Alvey Vision Confer- ence, 1988. 2
1988
-
[14]
On SIFTs and Their Scales
Tal Hassner, Viki Mayzels, and Lihi Zelnik-Manor. On SIFTs and Their Scales. In CVPR, 2012. 2
2012
-
[15]
Local descriptors opti- mized for average precision
Kun He, Yan Lu, and Stan Sclaroff. Local descriptors opti- mized for average precision. In CVPR, 2018. 1, 2, 4
2018
-
[16]
Schoenberger, Enrique Dunn, and Jan-Michael Frahm
Jared Heinly, Johannes L. Schoenberger, Enrique Dunn, and Jan-Michael Frahm. Reconstructing the World in Six Days. In CVPR, 2015. 5
2015
-
[17]
Spatial Transformer Networks
Max Jaderberg, Karen Simonyan, Andrew Zisserman, and Koray Kavukcuoglu. Spatial Transformer Networks. In NIPS, pages 2017–2025, 2015. 3
2017
-
[18]
PCA-SIFT: A More Distinc- tive Representation for Local Image Descriptors
Yan Ke and Rahul Sukthankar. PCA-SIFT: A More Distinc- tive Representation for Local Image Descriptors. In CVPR, pages 111–119, 2000. 2
2000
-
[19]
Learning Deep Descriptors with Scale- Aware Triplet Networks
Michel Keller, Zetao Chen, Fabiola Maffra, Patrik Schmuck, and Margarita Chli. Learning Deep Descriptors with Scale- Aware Triplet Networks. In CVPR, 2018. 1, 2, 4
2018
-
[20]
Dense Scale Invariant Descriptors for Images and Surfaces
Iasonas Kokkinos, Michael Bronstein, and Alan Yuille. Dense Scale Invariant Descriptors for Images and Surfaces. Technical report, INRIA, 2012. 2
2012
-
[21]
BRISK: Binary Robust Invariant Scalable Keypoints
Stefan Leutenegger, Margarita Chli, and Roland Siegwart. BRISK: Binary Robust Invariant Scalable Keypoints. In ICCV, 2011. 3
2011
-
[22]
SIFT Flow: Dense Correspondence Across Scenes and Its Applications
Ce Liu, Jenny Yuen, and Antonio Torralba. SIFT Flow: Dense Correspondence Across Scenes and Its Applications. In ECCV, 2008. 2
2008
-
[23]
Distinctive Image Features from Scale- Invariant Keypoints
David Lowe. Distinctive Image Features from Scale- Invariant Keypoints. IJCV, 20(2):91–110, Nov 2004. 1, 2, 3, 5
2004
-
[24]
ContextDesc: Local Descriptor Augmentation with Cross-Modality Context
Zixin Luo, Tianwei Shen, Lei Zhou, Jiahui Zhang, Yao Yao, Shiwei Li, Tian Fang, and Long Quan. ContextDesc: Local Descriptor Augmentation with Cross-Modality Context. In CVPR, 2019. 1, 2
2019
-
[25]
Geodesc: Learning Local Descriptors by Integrating Geometry Constraints
Zixin Luo, Tianwei Shen, Lei Zhou, Siyu Zhu, Runze Zhang, Yao Yao, Tian Fang, and Long Quan. Geodesc: Learning Local Descriptors by Integrating Geometry Constraints. In ECCV, 2018. 1, 2, 5
2018
-
[26]
A Performance Evaluation of Local Descriptors
Krystian Mikolajczyk and Cordelia Schmid. A Performance Evaluation of Local Descriptors. PAMI, 27(10):1615–1630,
-
[27]
A Comparison of Affine Region Detectors
Krystian Mikolajczyk, Tinne Tuytelaars, Cordelia Schmid, Andrew Zisserman, Jiri Matas, Frederik Schaffalitzky, Timor Kadir, and Luc Van Gool. A Comparison of Affine Region Detectors. IJCV, 65(1/2):43–72, 2005. 5
2005
-
[28]
Working Hard to Know Your Neighbor’s Margins: Local Descriptor Learning Loss
Anastasiia Mishchuk, Dmytro Mishkin, Filip Radenovic, and Jiri Matas. Working Hard to Know Your Neighbor’s Margins: Local Descriptor Learning Loss. In NIPS, 2017. 1, 2, 3, 4, 5
2017
-
[29]
Lf-Net: Learning Local Features from Images
Yuki Ono, Eduard Trulls, Pascal Fua, and Kwang Moo Yi. Lf-Net: Learning Local Features from Images. In NIPS,
-
[30]
Leveraging outdoor webcams for local descriptor learning
Milan Pultar, Dmytro Mishkin, and Jiri Matas. Leveraging outdoor webcams for local descriptor learning. In Computer Vision Winter Workshop, 2019. 5, 8
2019
-
[31]
R2D2: Repeatable and Reliable Detector and De- scriptor
Jerome Revaud, Philippe Weinzaepfel, C ´esar De Souza, Noe Pion, Gabriela Csurka, Yohann Cabon, and Martin Humen- berger. R2D2: Repeatable and Reliable Detector and De- scriptor. In arXiv Preprint, 2019. 1, 2
2019
-
[32]
ORB: An Efficient Alternative to SIFT or SURF
Ethan Rublee, Vincent Rabaud, Kurt Konolige, and Gary Bradski. ORB: An Efficient Alternative to SIFT or SURF. In ICCV, 2011. 2
2011
-
[33]
Benchmarking 6DOF Outdoor Visual Local- ization in Changing Conditions
Torsten Sattler, Will Maddern, Carl Toft, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, Fredrik Kahl, and Tomas Pajdla. Benchmarking 6DOF Outdoor Visual Local- ization in Changing Conditions. In CVPR, 2018. 1
2018
-
[34]
Sch ¨onberger and Jan-Michael Frahm
Johannes L. Sch ¨onberger and Jan-Michael Frahm. Structure- From-Motion Revisited. In CVPR, 2016. 5
2016
-
[35]
Sch ¨onberger, Hans Hardmeier, Torsten Sat- tler, and Marc Pollefeys
Johannes L. Sch ¨onberger, Hans Hardmeier, Torsten Sat- tler, and Marc Pollefeys. Comparative Evaluation of Hand- Crafted and Learned Local Features. In CVPR, 2017. 8
2017
-
[36]
Qi Shan, Changchang Wu, Brian Curless, Yasutaka Fu- rukawa, Carlos Hernandez, and Steven M. Seitz. Accurate Geo-registration by Ground-to-Aerial Image Matching. In 3DV, 2014. 2
2014
-
[37]
Matching Local Self- Similarities Across Images and Videos
Eli Shechtman and Michal Irani. Matching Local Self- Similarities Across Images and Videos. CVPR, 2007. 3
2007
-
[38]
Dis- criminative Learning of Deep Convolutional Feature Point Descriptors
Edgar Simo-Serra, Eduard Trulls, Luis Ferraz, Iasonas Kokkinos, Pascal Fua, and Francesc Moreno-noguer. Dis- criminative Learning of Deep Convolutional Feature Point Descriptors. In ICCV, 2015. 1, 2, 3, 4
2015
-
[39]
Learning Local Feature Descriptors Using Convex Optimi- sation
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Learning Local Feature Descriptors Using Convex Optimi- sation. PAMI, 2014. 1, 2
2014
-
[40]
LDAHash: Improved Matching with Smaller Descriptors
Christoph Strecha, Alex Bronstein, Michael Bronstein, and Pascal Fua. LDAHash: Improved Matching with Smaller Descriptors. PAMI, 34(1), January 2012. 2
2012
-
[41]
L2-Net: Deep Learn- ing of Discriminative Patch Descriptor in Euclidean Space
Yurun Tian, Bin Fan, and Fuchao Wu. L2-Net: Deep Learn- ing of Discriminative Patch Descriptor in Euclidean Space. In CVPR, 2017. 1, 2, 3, 5
2017
-
[42]
A Fast Local Descriptor for Dense Matching
Engin Tola, Vincent Lepetit, and Pascal Fua. A Fast Local Descriptor for Dense Matching. In CVPR, 2008. 1, 3
2008
-
[43]
Dense Segmentation-Aware De- scriptors
Eduard Trulls, Iasonas Kokkinos, Alberto Sanfeliu, and Francesc Moreno-Noguer. Dense Segmentation-Aware De- scriptors. Dense Image Correspondences for Computer Vi- sion, 2015. 2
2015
-
[44]
Im- proved Texture Networks: Maximizing Quality and Diver- sity in Feed-Forward Stylization and Texture Synthesis
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. Im- proved Texture Networks: Maximizing Quality and Diver- sity in Feed-Forward Stylization and Texture Synthesis. In CVPR, 2017. 4
2017
-
[45]
Kernelized Subspace Pooling for Deep Local Descriptors
Xing Wei, Yue Zhang, Yihong Gong, and Nanning Zheng. Kernelized Subspace Pooling for Deep Local Descriptors. In CVPR, 2018. 1, 2
2018
-
[46]
Learning Local Image Descriptors
Simon Winder and Matthew Brown. Learning Local Image Descriptors. In CVPR, June 2007. 1, 3
2007
-
[47]
MORB: A Multi-Scale Binary Descriptor
Alessio Xompero, Oswald Lanz, and Andrea Cavallaro. MORB: A Multi-Scale Binary Descriptor. In ICIP, 2018. 2
2018
-
[48]
LIFT: Learned Invariant Feature Transform
Kwang Moo Yi, Eduard Trulls, Vincent Lepetit, and Pascal Fua. LIFT: Learned Invariant Feature Transform. In ECCV,
-
[49]
Learning to Find Good Correspondences
Kwang Moo Yi, Eduard Trulls, Yuki Ono, Vincent Lepetit, Mathieu Salzmann, and Pascal Fua. Learning to Find Good Correspondences. In CVPR, 2018. 1
2018
-
[50]
Learning to Com- pare Image Patches via Convolutional Neural Networks
Sergey Zagoruyko and Nikos Komodakis. Learning to Com- pare Image Patches via Convolutional Neural Networks. In CVPR, 2015. 1, 2, 3
2015
-
[51]
Eigendecomposition-Free Training of Deep Networks with Zero Eigenvalue-Based Losses
Dang Zheng, Kwang Moo Yi, Yinlin Hu, Fei Wang, Pas- cal Fua, and Mathieu Salzmann. Eigendecomposition-Free Training of Deep Networks with Zero Eigenvalue-Based Losses. In ECCV, 2018. 1
2018
-
[52]
Progressive Large Scale-Invariant Image Matching In Scale Space
Lei Zhou, Siyu Zhu, Tianwei Shen, Jinglu Wang, Tian Fang, and Long Quan. Progressive Large Scale-Invariant Image Matching In Scale Space. In ICCV, 2017. 2
2017
-
[53]
Regarding the Dataset In order to train scale-invariant models with real data relevant to wide-baseline stereo, it was necessary to collect training data
Beyond Cartesian Representations for Local Descriptors: Supplementary Material 6.1. Regarding the Dataset In order to train scale-invariant models with real data relevant to wide-baseline stereo, it was necessary to collect training data. For this we rely on public collections...
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.