Pith. sign in

REVIEW 3 major objections 5 minor 60 references

GLAMpoints: Greedily Learned Accurate Match points

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A keypoint detector trained on a reward for correct matches, not repeatability, outperforms classical and learned detectors on retinal image registration.

desk verdict The detector training idea is genuinely useful and the slitlamp results are convincing, but the FIRE headline is weakened by the 15% downscaling and the paper overclaims generalization. read the letter →

arxiv 1908.06812 v3 pith:Q74LAI5M submitted 2019-08-19 cs.CV

classification cs.CV
keywords keypointdetectionfeaturematchingretinalimageregistrationsemi-supervisedlearningreinforcement-learning-stylerewardU-Netslitlampfundusimageshomographyestimation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that a keypoint detector should be trained for the final matching result, not for an intermediate proxy such as repeatability, and that this is feasible even though matching and homography estimation are non-differentiable. The authors introduce GLAMpoints, a U-Net trained in a semi-supervised way with a reinforcement-learning-style reward: a detected point is rewarded only if it produces a correct match under the ground-truth homography of a synthetically warped image pair. On two retinal registration benchmarks, this scheme yields 94.78% acceptable registrations on FIRE and 68.45% on pre-processed slitlamp frames, surpassing SIFT, KAZE, SuperPoint, LIFT and LF-NET, with the same gain holding on raw, unpreprocessed data. The paper's central claim is that optimizing directly for matching accuracy produces a detector that is both more accurate and more robust than detectors optimized for repeatability.

What carries the argument

The mechanism is a reward-based training loss for a pixel-wise keypoint probability map. For a synthetically warped pair $(I, I')$, the network outputs score maps; non-maximum suppression extracts keypoints; a fixed root-SIFT descriptor describes them; bidirectional brute-force matching and the ground-truth homography $H$ decide which keypoints are true positives using $\|H*x - x'\| \leq \varepsilon$ with $\varepsilon = 3$ px. The reward matrix $R$ is 1 at true positives and 0 elsewhere, and the loss $L(\theta,I)=\sum (f_\theta(I)-R)^2 \cdot M / \sum M$ back-propagates only through all true positives and a mined subset of false positives (mask $M$), countering the extreme class imbalance. The network is a four-level U-Net with sigmoid output; only the score-map prediction is differentiable, so the reward acts like a delayed reinforcement signal. This design is what lets the detector be optimized for the final matching objective rather than a proxy.

What would settle it

Re-run the FIRE and slitlamp evaluations on test pairs whose true homographies fall outside the training envelope, for example rotations beyond 25 degrees, scaling below 0.7, or images from cameras and pathologies absent from the 10 training patients, and compare acceptable-registration rates. If the GLAMpoints advantage over SIFT and LIFT disappears or reverses on such out-of-envelope pairs, the reported superiority is a property of the matched synthetic-to-real distribution rather than of the matching-based training objective itself.

Watch

Extended reading notes

Core claim

The central discovery is that a detector trained to maximize correct matches, rather than repeatability or corner-like structure, extracts keypoints that are dense, uniformly spread, and stable in low-texture retinal images. Trained on synthetic pairs generated from slitlamp frames by random homographies and appearance augmentations, GLAMpoints receives a positive reward only for keypoints that survive bidirectional nearest-neighbor matching and fall within 3 pixels of the ground-truth warped position; all other pixels receive zero reward. With sample mining to balance true and false positives, the U-Net learns to suppress clustered or ambiguous points. Paired with a fixed root-SIFT descriptor, GLAMpoints improves registration success over every baseline detector with the same descriptor, raises acceptable registrations on FIRE to 94.78% (33.6 points over SIFT), achieves 63.59% acceptable registrations on raw slitlamp images and 68.45% after preprocessing, and registers on average 9.98 consecutive video frames before failure versus 1.04 for SIFT. The paper also reports that the same model, trained only on slitlamp images, registers 75.38% of natural-image pairs acceptably, second to rotation-invariant SIFT.

Load-bearing premise

The load-bearing premise is that synthetic image pairs created from 10 patients' slitlamp frames, with rotations up to 25 degrees and scaling from 0.7 to 1.3, resemble real test pairs; the paper itself notes (supplementary section A.2) that the training transform ranges were chosen to match the test set, so pairs outside that envelope could shrink the reported gains.

Editorial extensions

If this is right

  • Any fixed descriptor benefits: GLAMpoints with ORB or BRISK beats the original ORB or BRISK detector, and with root-SIFT it beats SIFT, showing the gain comes from point selection, not the descriptor.
  • Preprocessing becomes optional: GLAMpoints loses only a few percentage points between pre-processed and raw slitlamp images, while SIFT and SuperPoint drop 20 to 30 points.
  • Repeatability is the wrong optimization target: LF-NET, trained for repeatability, has high repeatability but low matching scores, while GLAMpoints has lower repeatability and the best matching and registration.
  • Same-model generalization: a model trained only on slitlamp fundus images registers 75.38% of natural-image pairs acceptably, second only to rotation-invariant SIFT among tested detectors.
  • Faster pipeline: GLAMpoints detection takes about 27 ms per image with no preprocessing, versus about 45 ms plus a 16 ms preprocessing step for SIFT.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves implicit that the same reward could be used to train the descriptor jointly rather than fixing root-SIFT; this would likely reduce the gap on rotation-heavy natural-image sets, where the authors note their rotation-dependent descriptor is at a disadvantage.
  • Because training requires only synthetic homographies and no manual labels, the method transfers to other low-texture imaging domains, such as endoscopy or histopathology, if the augmentation ranges are matched to those domains.
  • The uniform spread of GLAMpoints can be interpreted as learned non-maximum suppression: clustered candidates compete for the same reward, so the network unlearns them; this could be verified by measuring coverage fraction over a sweep of test-time NMS window sizes.
  • A testable variant would give partial reward to repeatable but unmatched points, to see whether a smoother reward changes the repeatability-matching tradeoff.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes GLAMpoints, a U-Net-based keypoint detector trained with a reward derived from matching success rather than from repeatability. Training pairs are generated from slitlamp fundus images by random homographies and appearance augmentations; the network predicts a per-pixel score map, keypoints are extracted by non-maximum suppression, described by root-SIFT, matched bidirectionally, and true positives under the known homography provide a sparse reward signal. The detector is evaluated by registration success on a 206-pair slitlamp test set, on the FIRE benchmark downscaled to 15% of native resolution, and on natural-image datasets. The authors report that GLAMpoints yields the highest acceptable-registration rates on slitlamp (68.45% pre-processed, 63.59% raw) and FIRE (94.78%), and that the gains over SIFT persist when the same descriptor is used.

Significance. The core idea of directly optimizing detection for downstream matching is valuable and contrasts with the repeatability-based training of LF-NET and SuperPoint. Strengths include the released training code and weights, the random-grid and descriptor ablations showing that uniform coverage alone or descriptor choice does not explain the results, and the explicit comparison of repeatability versus matching metrics. If the FIRE results are reproduced at native resolution, the method would be a strong practical detector for low-texture medical images. However, the current public-benchmark evidence is weakened by the 15% downscaling protocol and by the acknowledged tuning of training augmentation ranges to the test distribution; the absence of uncertainty estimates also makes the 'significantly outperforms' claim in the abstract unsupported as stated.

major comments (3)
  1. [Section 4.1, Table 2] The FIRE images are downscaled to 15% of native resolution 'to match the resolution of the training set,' and no native-resolution results are reported. Since GLAMpoints is a fixed network trained on 256x256 crops from 300-700px slitlamp images, this protocol shifts the FIRE benchmark into the model's preferred scale range while removing high-frequency vascular detail that scale-invariant baselines such as SIFT, KAZE, and LIFT could exploit; the 33.6-point margin over SIFT in Table 2 may therefore reflect resolution alignment rather than detector quality. Please evaluate on native-resolution FIRE (2912x2912) or, if that is not feasible, justify the downscaling as the intended deployment setting and explicitly restrict the claim to low-resolution retinal images.
  2. [Supplementary A.2, Table 5] The text states that the geometric transformation ranges used to synthesize training pairs were chosen so that 'the resulting synthetic training set resembles the test set,' with the test-set rotation range explicitly used to limit training rotations. This makes the reported gains a measure of performance under a domain-matched augmentation schedule rather than of a general detector advantage; it also creates a risk of overfitting to the test distribution for the slitlamp and FIRE evaluations. Please report the sensitivity of the main results to the augmentation ranges, or fix the schedule a priori and evaluate on independent test sets.
  3. [Section 4.2, Tables 1 and 2] The acceptable/inaccurate classification thresholds (MEE < 10, MAE < 30) are described as 'found empirically by post-viewing the results,' and all success rates are reported as point estimates with no confidence intervals or significance tests. Given that the abstract claims the method 'significantly outperforms' baselines, the authors should provide bootstrap confidence intervals or pairwise tests for the key comparisons, and show that the ranking is stable under reasonable variations of the thresholds.
minor comments (5)
  1. [Table 3] Table 3 reports 'CNN:16.28± 96.86' and 'Total 27.48 ± 98.74'; these standard deviations are implausibly large and likely a typographical error, and the table omits the 'ms' unit in the numeric cells.
  2. [Section 4.1 vs. Supplementary A.2] Section 4.1 states that the slitlamp test pairs have rotations 'up to 15 degrees,' while Supplementary A.2 refers to rotation 'up to 30 degrees' and Table 5 caps training rotation at 25 degrees; please reconcile these numbers.
  3. [Figure 4] Figure 4 uses dual axes for four metrics but the caption does not state which curve or axis corresponds to which metric, making the plot difficult to interpret.
  4. [Sections 4.2-4.3] The manuscript uses 'RanSaC' and 'RANSAC' interchangeably; please standardize the spelling.
  5. [Section 4.6 and Supplementary Figure 9] Section 4.6 reports only the aggregate natural-image success rates; pointing to the per-dataset breakdown in Supplementary Figure 9 would help the reader assess where the method fails (e.g., Viewpoint rotations).

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the detector is trained on synthetic pairs with a matching reward and evaluated on held-out real images and public benchmarks; the self-citations are not load-bearing.

full rationale

The paper's central chain is a supervised/semi-supervised training loop: Eq. (3) minimizes (f_theta(I)-R)^2 over a reward R defined by true-positive SIFT matches against a ground-truth homography on synthetic pairs (Sec. 3.1-3.2). The test evaluations use held-out slitlamp pairs from different patients and the external FIRE, Oxford, EF, Webcam, and ViewPoint datasets with the same root-SIFT descriptor (Sec. 4.1-4.6). Because the reward is computed on synthetically warped training images and never on test pairs, the reported matching and registration numbers are predictions, not re-statements of the training objective. The 'random grid (SIFT)' baseline (Table 1b) and the same-descriptor comparisons (GLAMpoints vs SIFT on FIRE) further show that the claimed gain is not an artifact of descriptor choice or uniform coverage alone. Self-citations exist (Truong et al. [47] for prior detector evaluation; De Zanet et al. [19] for pre-processing and blending), but none carries the central claim: the baselines are run in-paper and public benchmarks provide external evidence. Two evaluation choices deserve attention as correctness risks, not as circularity: FIRE images were downscaled to 15% of native resolution 'to match the resolution of the training set' (Sec. 4.1), and the acceptable-registration thresholds were 'found empirically by post-viewing the results' (Sec. 4.2). Neither is an equation-level reduction of the reported result to a fitted parameter; both are protocol validity concerns. The training augmentation ranges were also chosen to resemble the test set (Supp. A.2), a domain-alignment choice that could inflate results but is not a circular derivation. Overall, no load-bearing step reduces by construction to its inputs.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central claim depends on the choice of NMS window (10 px), the true-positive epsilon (3 px), the empirically chosen registration thresholds, and the constrained homography augmentation ranges. These are not fitted to the test results in the usual sense, but they are hand-chosen and, in the case of the augmentation ranges, tuned to match the test distribution. No new entities are introduced.

free parameters (4)
  • NMS window size w = 10 px
    Non-maximum suppression window used to extract keypoints from score maps; training and testing with NMS10 (Section 4.3).
  • Match inlier threshold epsilon = 3 px
    Epsilon-neighborhood for true positive match under ground truth homography (Section 3.2).
  • Acceptable registration thresholds = MEE < 10 px, MAE < 30 px
    Thresholds for classifying registrations, found empirically by post-viewing results (Section 4.2).
  • Synthetic homography parameter ranges = rotation max 25 deg, scaling 0.7 to 1.3, translation 100 px, shearing +/- 0.2
    Ranges sampled to generate training pairs, chosen to resemble the test set (Supplementary Table 5).
assumptions (4)
  • domain assumption Planar scene assumption for slitlamp and fundus images
    Used to justify generating ground truth homographies for training and test pairs (Section 4.1, citing [13, 24]).
  • domain assumption SIFT descriptor is an appropriate fixed descriptor for reward and evaluation
    The reward and all comparisons use root-SIFT; the detector is therefore specialized to this descriptor (Sections 3.2 and 4.3).
  • domain assumption Manual annotations of at least 5 points yield accurate ground truth homographies
    Test ground truth homographies are estimated from manual point correspondences (Section 4.1).
  • ad hoc to paper Synthetic homography distribution matches test distribution
    Training pairs are generated with constrained geometric ranges chosen to match the test set, as stated in Supplementary A.2.

how reviews work

0 comments
Cite this review

Pith. "Pith review of GLAMpoints: Greedily Learned Accurate Match points." pith.science (2026). https://pith.science/paper/Q74LAI5M

@misc{pith2026190806812,
  author       = {Pith},
  title        = {Pith review of: GLAMpoints: Greedily Learned Accurate Match points},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q74LAI5M}},
  note         = {Machine review of arXiv:1908.06812}
}
read the original abstract

We introduce a novel CNN-based feature point detector - GLAMpoints - learned in a semi-supervised manner. Our detector extracts repeatable, stable interest points with a dense coverage, specifically designed to maximize the correct matching in a specific domain, which is in contrast to conventional techniques that optimize indirect metrics. In this paper, we apply our method on challenging retinal slitlamp images, for which classical detectors yield unsatisfactory results due to low image quality and insufficient amount of low-level features. We show that GLAMpoints significantly outperforms classical detectors as well as state-of-the-art CNN-based methods in matching and registration quality for retinal images. Our method can also be extended to other domains, such as natural images. Training code and model weights are available at https://github.com/PruneTruong/GLAMpoints_pytorch.

Figures

Figures reproduced from arXiv: 1908.06812 by the authors.

Figure 1
Figure 1. Keypoints detected by SIFT and GLAMpoints [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. a) Training steps for an image pair Ii and I 0 i at epoch i created from a particular base image B. Ii and I 0 i are created by warping B according to homographies gi and g 0 i respectively. ai and a 0 i refer to the additional appearance augmentations applied to each image. b) Loss computation corresponding to situation a. c) Schematic representation of Unet-4. Let T denote the set of true positive key points. If a… view at source ↗
Figure 3
Figure 3. Examples of images from the slit lamp dataset showing challenging conditions for registration. From left to right: low vascularization and over-exposure leading to weak contrasts and lack of corners, motion blur, focus blur, acquisition artifacts and reflections. Examples are shown in [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Summary of detector/descriptor performance metrics evaluated over 206 pairs of the [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Interest points detected by a) SIFT and b) GLAMpoints and corresponding matches for a pair of images from the [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Mosaics obtained from registration of consecutive [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Examples of image pairs from the Oxford dataset. [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: Matches on the F IRE dataset. Detected points are in white, green lines are true positive matches while red ones are false positive. 14 [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: Summary of detector/descriptor performance metrics evaluated over 195 pairs of natural images. [PITH_FULL_IMAGE:figures/full_fig_p015_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

60 extracted references · 58 canonical work pages

  1. [1]

    OpenCV: cv::BFMatcher Class Reference. 3

  2. [2]

    OpenCV: cv::xfeatures2d::SIFT Class Reference. 2, 6

  3. [3]

    TensorFlow: Large- Scale Machine Learning on Heterogeneous Systems, 2015

    Martn Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Man, Rajat Monga, Sherry Moore, Derek Murray, Chris Ol...

  4. [4]

    FREAK: Fast Retina Keypoint

    Alexandre Alahi, Raphal Ortiz, and Pierre Vandergheynst. FREAK: Fast Retina Keypoint. In Conference on Computer Vision and Pattern Recognition, 2012. 3

  5. [5]

    Pablo Fern ´andez Alcantarilla, Adrien Bartoli, and Andrew J. Davison. KAZE Features. In European Conference on Com- puter Vision, 2012. 2, 6

  6. [6]

    In the Saddle: Chasing Fast and Repeatable Features

    Javier Aldana-Iuit, Dmytro Mishkin, Ondrej Chum, and Jiri Matas. In the Saddle: Chasing Fast and Repeatable Features. In International Conference on Pattern Recognition , pages 675–680, 2016. 5

  7. [7]

    Learning to Match Aerial Images with Deep Attentive Architecture

    Hani Altwaijry, Eduard Trulls, Serge Belongie, James Hays, and Pascal Fua. Learning to Match Aerial Images with Deep Attentive Architecture. In Conference on Computer Vision and Pattern Recognition, 2016. 3

  8. [8]

    Learning to Detect and Match Keypoints with Deep Architectures

    Hani Altwaijry, Andreas Veit, and Serge Belongie. Learning to Detect and Match Keypoints with Deep Architectures. In British Machine Vision Conference, 2016. 3

Show all 60 references
  1. [9]

    Three things everyone should know to improve object retrieval

    Relja Arandjelovic and Andrew Zisserman. Three things everyone should know to improve object retrieval. In Con- ference on Computer Vision and Pattern Recognition, pages 2911–2918, 2012. 3, 11

  2. [10]

    PN-Net: Conjoined Triple Deep Network for Learning Local Image Descriptors

    Vassileios Balntas, Edward Johns, Lilian Tang, and Krystian Mikolajczyk. PN-Net: Conjoined Triple Deep Network for Learning Local Image Descriptors. CoRR, abs/1601.05030,

  3. [11]

    HPatches: A Benchmark and Evaluation of Handcrafted and Learned Local Descriptors

    Vassileios Balntas, Karel Lenc, Andrea Vedaldi, and Krys- tian Mikolajczyk. HPatches: A Benchmark and Evaluation of Handcrafted and Learned Local Descriptors. In Confer- ence on Computer Vision and Pattern Recognition , pages 3852–3861, 2017. 3

  4. [12]

    Surf: Speeded up robust features

    Herbert Bay, Tinne Tuytelaars, and Luc Van Gool. Surf: Speeded up robust features. In European Conference on Computer Vision, pages 404–417, 2006. 2

  5. [13]

    Cattin, Herbert Bay, Luc Van Gool, and G ´abor Sz´ekely

    Philippe C. Cattin, Herbert Bay, Luc Van Gool, and G ´abor Sz´ekely. Retina Mosaicing Using Local Features. In Confer- ence on Medical Image Computing and Computer Assisted Intervention, pages 185–192, 2006. 2, 5

  6. [14]

    J. Chen, J. Tian, N. Lee, J. Zheng, R. T. Smith, and A. F. Laine. A Partial Intensity Invariant Feature Descriptor for Multimodal Retinal Image Registration. IEEE Transactions on Biomedical Engineering, 57(7):1707–1718, 2010. 5

  7. [15]

    Smith, and Andrew F

    Jian Chen, Jie Tian, Noah Lee, Jian Zheng, Theodore R. Smith, and Andrew F. Laine. A Partial Intensity Invari- ant Feature Descriptor for Multimodal Retinal Image Reg- istration. IEEE Transactions on Biomedical Engineering , 57(7):1707–1718, 2010. 2, 5

  8. [16]

    Cideciyan

    Artur V . Cideciyan. Registration of Ocular Fundus Images: an Algorithm Using Cross-correlation of Triple Invariant Im- age Descriptors. IEEE Engineering in Medicine and Biology Magazine, 14(1):52–58, 1995. 2

  9. [17]

    Dahl, Henrik Aanæs, and Kim S

    Anders L. Dahl, Henrik Aanæs, and Kim S. Pedersen. Find- ing the Best Feature Detector-Descriptor Combination. InIn- ternational Conference on 3D Imaging, Modeling, Process- ing, Visualization and Transmission, pages 318–325, 2011. 5

  10. [18]

    Real-Time Simultaneous Localisation and Mapping with a Single Camera

    Andrew Davison. Real-Time Simultaneous Localisation and Mapping with a Single Camera. In International Conference on Computer Vision, 2003. 2

  11. [19]

    Retinal Slit Lamp Video Mosaicking

    Sandro De Zanet, Tobias Rudolph, Rogerio Richa, Christoph Tappeiner, and Raphael Sznitman. Retinal Slit Lamp Video Mosaicking. International Journal of Computer Assisted Ra- diology and Surgery, 11(6):1035–1041, 2016. 5, 8

  12. [20]

    SuperPoint: Self-Supervised Interest Point Detec- tion and Description

    Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabi- novich. SuperPoint: Self-Supervised Interest Point Detec- tion and Description. In Conference on Computer Vision and Pattern Recognition Workshops, pages 224–236, 2018. 2, 3, 6

  13. [21]

    KCNN: Extremely-Efficient Hardware Keypoint Detection With a Compact Convolutional Neural Network

    Paolo Di Febbo, Carlo Dal Mutto, Kinh Tieu, and Stefano Mattoccia. KCNN: Extremely-Efficient Hardware Keypoint Detection With a Compact Convolutional Neural Network. In Conference on Computer Vision and Pattern Recognition Workshops, 2018. 3

  14. [22]

    De- scriptor Matching with Convolutional Neural Networks: a Comparison to SIFT

    Philipp Fischer, Alexey Dosovitskiy, and Thomas Brox. De- scriptor Matching with Convolutional Neural Networks: a Comparison to SIFT . Technical Report 1405.5769, arXiv, May 2014. 2

  15. [23]

    Fischler and Robert C

    Martin A. Fischler and Robert C. Bolles. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography.Communications of the ACM, 24(6):381–395, June 1981. 2

  16. [24]

    Textureless Macula Swelling Detection With Multiple Retinal Fundus Images

    Luca Giancardo, Fabrice Meriaudeau, Thomas Karnowski, Tobin Kenneth W., Jr, Enrico Grisan, Paolo Favaro, Al- fredo Ruggeri, and Edward Chaum. Textureless Macula Swelling Detection With Multiple Retinal Fundus Images. IEEE Transactions on Biomedical Engineering , 58(3):795– 799...

  17. [25]

    Retinal Image Registration Based on the Feature of Bifurcation Point

    Yiliu Hang, Xiaofeng Zhang, Yeqin Shao, Huiqun Wu, and Wei Sun. Retinal Image Registration Based on the Feature of Bifurcation Point. In International Congress on Image and Signal Processing, BioMedical Engineering and Infor- matics, 2017. 2

  18. [26]

    A Combined Corner and Edge Detector

    Chris Harris and Mike Stephens. A Combined Corner and Edge Detector. In Fourth Alvey Vision Conference, 1988. 2

  19. [27]

    FIRE : Fundus Image Registration dataset

    Carlos Hernandez-Matas, Xenophon Zabulis, Areti Tri- antafyllou, Panagiota Anyfanti, Stella Douma, and Antonis Argyros. FIRE : Fundus Image Registration dataset. Journal for Modeling in Ophthalmology, 4:16–28, 2017. 2, 5 9

  20. [28]

    Consis- tent Temporal Variations in Many Outdoor Scenes

    Nathan Jacobs, Nathaniel Roman, and Robert Pless. Consis- tent Temporal Variations in Many Outdoor Scenes. In Con- ference on Computer Vision and Pattern Recognition, 2007. 5, 12

  21. [29]

    Phase Correlation-based Iris Image Registration Model

    Li Ma Jun-Zhou Huang, Tie-Niu Tan and Yun-Hong Wang. Phase Correlation-based Iris Image Registration Model. Journal of Computer Science and Technology , 20(3):419– 425, 2005. 2

  22. [30]

    Diederik. P. Kingma and Jimmy Ba. Adam: A Method for Stochastic Optimisation. In International Conference on Learning Representations, 2015. 6

  23. [31]

    Improving Accuracy and Efficiency of Mutual Informa- tion for Multi-modal Retinal Image Registration using Adap- tive Probability Density Estimation

    Phil Legg, Paul Rosin, David Marshall, and James Mor- gan. Improving Accuracy and Efficiency of Mutual Informa- tion for Multi-modal Retinal Image Registration using Adap- tive Probability Density Estimation. Computerized Medical Imaging and Graphics, 37(7-8):597–606, 2013. 2

  24. [32]

    BRISK: Binary Robust Invariant Scalable Keypoints

    Stefan Leutenegger, Margarita Chli, and Roland Siegwart. BRISK: Binary Robust Invariant Scalable Keypoints. In In- ternational Conference on Computer Vision, 2011. 3

  25. [33]

    P. Li, Q. Chen, W. Fan, and S. Yuan. Registration of OCT Fundus Images with Color Fundus Images Based on Invari- ant Features. In Cloud Computing and Security, pages 471– 482, 2017. 2

  26. [34]

    David G. Lowe. Distinctive Image Features from Scale- Invariant Keypoints. International Journal of Computer Vi- sion, 20(2):91–110, Nov 2004. 2, 5, 11

  27. [35]

    A Performance Evaluation of Local Descriptors

    Krystian Mikolajczyk and Cordelia Schmid. A Performance Evaluation of Local Descriptors. IEEE Transactions on Pat- tern Analysis and Machine Intelligence , 27(10):1615–1630,

  28. [36]

    A Comparison of Affine Re- gion Detectors

    Krystian Mikolajczyk, Tinne Tuytelaars, Cordelia Schmid, Andrew Zisserman, Jiri Matas, Frederik Schaffalitzky, Timor Kadir, and Luc Van Gool. A Comparison of Affine Re- gion Detectors. International Journal of Computer Vision , 65(1/2):43–72, 2005. 5, 12

  29. [37]

    LF-Net: Learning Local Features from Images

    Yuki Ono, Eduard Trulls, Pascal Fua, and Kwang Moo Yi. LF-Net: Learning Local Features from Images. In Ad- vances in Neural Information Processing Systems , pages 6237–6247, 2018. 2, 3, 6

  30. [38]

    Josien P. W. Pluim, J. B. Antoine Maintz, and Max A. Viergever. Mutual Information Based Registration of Medi- cal Images: A Survey. IEEE Transactions on Medical Imag- ing, 22(8):986–1004, 2003. 2

  31. [39]

    Feature-Based Retinal Image Registration Using D-Saddle Feature.Journal of Healthcare Engineering, 2017:1–15, 10 2017

    Roziana Ramli, Mohd Yamani Idna Idris, Khairunnisa Hasikin, Noor Khairiah A Karim, Ainuddin Wahid Abdul Wahab, Ismail Ahmedy, Fatimah Ahmedy, Nahrizul Adib Kadri, and Hamzah Arof. Feature-Based Retinal Image Registration Using D-Saddle Feature.Journal of Healthcare Engineering...

  32. [40]

    U- Net: Convolutional Networks for Biomedical Image Seg- mentation

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- Net: Convolutional Networks for Biomedical Image Seg- mentation. In Conference on Medical Image Computing and Computer Assisted Intervention, pages 234–241, 2015. 4

  33. [41]

    ORB: An Efficient Alternative to SIFT or SURF

    Ethan Rublee, Vincent Rabaud, Kurt Konolige, and Gary Bradski. ORB: An Efficient Alternative to SIFT or SURF. In International Conference on Computer Vision, 2011. 3

  34. [42]

    Blumenthal, Parag A

    C ´esar A S ´anchez-Galeana, Christopher Bowd, Eytan Z. Blumenthal, Parag A. Gokhale, Linda M. Zangwill, and Robert N. Weinreb. Using Optical Imaging Summary Data to Detect Glaucoma. Opthamology, pages 1812–1818, 2001. 1

  35. [43]

    Dis- criminative Learning of Deep Convolutional Feature Point Descriptors

    Edgar Simo-Serra, Eduard Trulls, Luis Ferraz, Iasonas Kokkinos, Pascal Fua, and Franscesc Moreno-Noguer. Dis- criminative Learning of Deep Convolutional Feature Point Descriptors. In International Conference on Computer Vi- sion, 2015. 4

  36. [44]

    McLauchlan, Richard I

    Bill Triggs, Philip F. McLauchlan, Richard I. Hartley, and Andrew W. Fitzgibbon. Bundle Adjustment – A Modern Synthesis. In Vision Algorithms: Theory and Practice, pages 298–372, 2000. 2

  37. [45]

    GLAMpoints : Github project page in PyTorch

    Prune Truong. GLAMpoints : Github project page in PyTorch. https://github.com/PruneTruong/ GLAMpoints_pytorch, 2019. 2

  38. [46]

    GLAMpoints : GitLab project page in Tensor- Flow

    Prune Truong, Stefanos Apostolopoulos, Agata Mosin- ska, Samuel Stucky, Carlos Ciller, and Sandro De Zanet. GLAMpoints : GitLab project page in Tensor- Flow. https://gitlab.com/retinai_sandro/ glampoints, 2019. 2

  39. [47]

    Comparison of Feature Detectors for Retinal Image Alignment

    Prune Truong, Sandro De Zanet, and Stefanos Apostolopou- los. Comparison of Feature Detectors for Retinal Image Alignment. In ARVO, 2019. 3, 6

  40. [48]

    TILDE: A Temporally Invariant Learned DEtec- tor

    Yannick Verdie, Kwang Moo Yi, Pascal Fua, and Vincent Lepetit. TILDE: A Temporally Invariant Learned DEtec- tor. Conference on Computer Vision and Pattern Recogni- tion, pages 5279–5288, 2015. 5, 12

  41. [49]

    Robust Point Matching Method for Multimodal Reti- nal Image Registration

    Gang Wang, Zhicheng Wang, Yufei Chen, and Weidong Zhao. Robust Point Matching Method for Multimodal Reti- nal Image Registration. Biomedical Signal Processing and Control, 19:68–76, 2015. 2, 5

  42. [50]

    Learning Local Image Descriptors

    Simon Winder and Matthew Brown. Learning Local Image Descriptors. In Conference on Computer Vision and Pattern Recognition, June 2007. 5

  43. [51]

    Simon Winder, Gang Hua, and Matthew. Brown. Picking the Best DAISY. In Conference on Computer Vision and Pattern Recognition, pages 178–185, 2009. 5

  44. [52]

    LIFT: Learned Invariant Feature Transform

    Kwang Moo Yi, Eduard Trulls, Vincent Lepetit, and Pascal Fua. LIFT: Learned Invariant Feature Transform. In Euro- pean Conference on Computer Vision, pages 467–483, 2016. 2, 6, 7

  45. [53]

    Learning to Assign Orientations to Feature Points

    Kwang Moo Yi, Yannick Verdie, Pascal Fua, and Vincent Lepetit. Learning to Assign Orientations to Feature Points. In Conference on Computer Vision and Pattern Recognition,

  46. [54]

    Rzeszotarski, Lawrence J

    Liang Zhou, Mark S. Rzeszotarski, Lawrence J. Singerman, and Jeanne M. Chokreff. The Detection and Quantification of Retinopathy Using Digital Angiograms. IEEE Transac- tions on Medical Imaging, 13(4):619–626, 1994. 1

  47. [55]

    Edge Foci Interest Points

    Larry Zitnick and Krishnan Ramnath. Edge Foci Interest Points. In International Conference on Computer Vision ,

  48. [57]

    The dataset contains various imaging changes in- cluding viewpoint, rotation, blur, illumination, scale, JPEG compression changes

    Oxford dataset [36]: 8 sequences with 45 pairs in to- tal. The dataset contains various imaging changes in- cluding viewpoint, rotation, blur, illumination, scale, JPEG compression changes. We evaluated on six of these sequences, excluding the ones showing rotation (boat and b...

  49. [58]

    It exhibits large viewpoint changes and in-plane rotations up to 45 degrees

    ViewPoint dataset [53]: 5 sequences with 25 pairs in total. It exhibits large viewpoint changes and in-plane rotations up to 45 degrees

  50. [59]

    The dataset exhibits drastic lighting changes as well as daytime changes and viewpoint changes

    EF dataset [55]: 3 sequences with 17 pairs in total. The dataset exhibits drastic lighting changes as well as daytime changes and viewpoint changes

  51. [60]

    It shows seasonal changes as well as day time changes of scenes taken from far away

    Webcam dataset [48, 28]: 6 sequences with 124 pairs in total. It shows seasonal changes as well as day time changes of scenes taken from far away. For all of the aforementioned datasets, the images pairs are related by homography transforms. Indeed, the scenes are either plana...

  52. [2011]

    We then give additional qualitative and quantitative evalua- tion results on fundus images in Section B

    5, 12 10 Supplementary material In this supplementary material, we first provide addi- tional details on the training methodology in Section A. We then give additional qualitative and quantitative evalua- tion results on fundus images in Section B. Finally, in Sec- tion C, we s...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.