REVIEW 4 major objections 6 minor 37 references
COph100: A comprehensive fundus image registration dataset from infants constituting the "RIDIRP" database
T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read COph100, the first infant-focused retinal registration dataset, shows current algorithms still misalign key vessels in retinopathy-of-prematurity follow-up exams.
desk verdict A useful pediatric retinal registration benchmark whose manual ground truth needs an error bar before the numbers can be taken at face value. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the dataset itself, COph100, which functions as a controlled benchmark by pairing each query image with a later-examination reference image and supplying 10 manual control-point correspondences as ground truth. The evaluation machinery follows the established fundus registration protocol, classifying registrations as failed, inaccurate, or acceptable using median and maximum point-error thresholds, and then summarizing performance with RMSE and area under the curve. The vessel segmentation masks serve a second role, as an alternate input modality to test whether registration methods improve when vascular structure is the only guide. This combination of manual points, masks, and a fixed protocol is what lets the authors turn a collection of messy clinical images into a measurable registration challenge.
What would settle it
Have two independent graders annotate the 10 corresponding points on a random sample of 50 image pairs and measure the mean distance between graders; if that distance is comparable to or larger than the roughly 4.8-pixel acceptable root-mean-square error, the benchmark's accuracy rankings become unreliable.
Extended reading notes
Core claim
On the paper's own terms, the contribution is a new resource: COph100, a subset of 325 images drawn from a public ROP dataset of 6,004 images, organized into 491 pairs from 100 eyes with 2 to 9 examination sessions each. For every pair, the authors provide 10 manually placed control points at vessel intersections and a vessel segmentation mask produced by a model trained on a public fundus vessel dataset. They benchmark twelve registration methods, including general-purpose feature matchers and retina-specific ones, both with and without segmentation input. The headline numbers are that SuperGlue achieves the best mean area under the curve (mAUC) of 0.809 and the best acceptable registrations have a root-mean-square error near 4.8 pixels, while the traditional method GDB-ICP improves to a 2.24 percent inaccurate rate when given vessel masks; several retina-trained methods fail or degrade on this data. The paper concludes that current registration algorithms do not yet align infant ROP images reliably, and that COph100 provides a challenging benchmark for closing that gap.
Load-bearing premise
The manual 10-point ground-truth correspondences are accurate enough to serve as the evaluation target, yet the paper reports no inter-observer variability, no repeat labeling, and no independent verification of those points.
Editorial extensions
If this is right
- If COph100 becomes a standard benchmark, registration algorithms will have to handle the blur, obstruction, illumination shifts, and limited overlap that are typical of infant examinations.
- The longitudinal structure of 2 to 9 sessions per eye makes it possible to evaluate registration as a tool for tracking lesion and vessel changes over time.
- Because vessel masks are provided, researchers can test whether feeding segmentation maps to feature matchers improves accuracy, which the paper reports for most methods.
- The gap between the best current result and perfect vessel alignment quantifies the headroom for future registration work on pediatric images.
Reading between the lines
- Beyond the paper's claims, the same selection pipeline could be applied to the higher-resolution sections of the source dataset, letting researchers test whether current algorithms improve with more image detail.
- The absence of repeated or multi-observer ground-truth labeling suggests that adding an inter-observer variability study would sharpen the reliability of the benchmark.
- Because the source dataset also contains lesion segmentations, a natural extension is to use COph100 pairs for joint registration and disease-progression measurement rather than registration alone.
- The finding that segmentation masks hurt some retina-trained models but help natural-image-trained models hints that domain shift, not segmentation per se, drives much of the performance gap.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces COph100, a retinal image registration dataset built from 100 infant eyes (491 image pairs) selected from the public Timkovic ROP dataset. For each image pair the authors provide 10 manually marked control-point correspondences, automatic vessel segmentation masks, and vessel overlay images. They evaluate 12 existing traditional and deep-learning registration methods on the dataset, reporting failure/inaccurate rates, RMSE, and mAUC, and show that current methods perform substantially worse than on the adult FIRE dataset. The paper claims this is the first retinal registration dataset specifically focused on disease progression in infants.
Significance. If the annotations are reliable, COph100 is a valuable new public benchmark: it is larger than most existing retinal registration datasets (491 pairs, 100 eyes), targets a pediatric population, and contains realistic image-quality challenges such as blur, obstruction, and illumination change. The public release of images, ground-truth points, vessel masks, and evaluation code supports reproducible benchmarking. The evaluation is largely non-circular because it uses pre-trained models without fine-tuning to the target dataset, avoiding parameter fitting on the benchmark itself.
major comments (4)
- [Registration Groundtruth] The reliability of the 10 manual control-point pairs per image pair is not assessed. No inter-observer variability, repeated labeling, or independent verification is reported, even though the author contributions list several people involved in labelling and ground-truth generation. Because every quantitative metric in Table 4 (mAUC, RMSE, failed/inaccurate rates) is computed against these points, unknown label noise could be comparable to the sub-pixel to few-pixel differences among top methods (for example, SuperGlue RMSE 5.079 vs SuperPoint RMSE 5.075). Please add a reproducibility study, such as having two or more annotators re-mark a random subset of pairs and reporting the mean and standard deviation of point distances, and ideally analyze how metric rankings change under realistic label noise.
- [Registration evaluation, Table 4] The RMSE metric is computed only on the subset of registrations classified as 'acceptable' (MEE < 20 and MAE < 50), so different methods are compared on different subsets of image pairs. For example, LK-SuperRetina has 40.53% failed registrations without segmentation and an RMSE of 4.841, lower than SuperGlue's 5.079, but this likely reflects an easier subset rather than superior accuracy. This biases cross-method comparisons. Please report RMSE on a common subset of pairs (for example, pairs that all methods register acceptably) or use an error measure that accounts for all pairs, and clarify how failed registrations are incorporated into the mAUC.
- [Registration evaluation, Table 4 and surrounding text] The discussion of segmentation-based results is internally inconsistent. The text states that 'methods trained in natural image environments have demonstrated improved accuracy with the addition of segmentation, except for SuperGlue', but Table 4 shows that SuperJunction's failure rate jumps from 0% to 100% with segmentation, and that SuperRetina, Swin U-SuperRetina, and LK-SuperRetina also degrade. Please correct the narrative to match the table and provide a concrete explanation for why segmentation information can be harmful for these methods, particularly the complete failure of SuperJunction.
- [Registration Groundtruth and Figure 2] The annotation protocol is underspecified. It is unclear whether the 10 control points are the same set tracked across all examinations of an eye, as Figure 2 suggests for the 9-examination case, or whether they are selected independently for each image pair. The paper also does not state how points in obstructed or blurred regions, which affect 213/325 and 100/325 images respectively, are handled. Without this description, users cannot assess the consistency and validity of the ground-truth correspondences.
minor comments (6)
- [Introduction] The phrase 'Minimal apprarance variablity' contains typos and should read 'Minimal appearance variability'.
- [Throughout] Please correct 'opthalmologists' to 'ophthalmologists' and 'Root Square Error' to 'Root Mean Square Error (RMSE)'.
- [Table 1] The publish time for COph100 is listed as 2024, but the paper is dated 2025; please update for consistency.
- [Title and Abstract] The acronym RIDIRP is used but never defined; please spell out the name and explain its relationship to the Timkovic et al. dataset.
- [Data Records] The GitHub link is truncated in the printed text; please ensure the full URL is provided.
- [Introduction] The claim that COph100 includes a 'diverse patient population' is not supported by any demographic data about the 100 eyes; please either provide such data or temper the claim.
Circularity Check
No circular derivation; dataset construction and evaluation are independent of the paper's conclusions.
full rationale
The paper is a dataset contribution and does not derive a result from fitted parameters or self-referential definitions. It selects 491 image pairs from a public ROP dataset, manually labels control points, trains a vessel segmentation model on the external FIVES dataset, and evaluates registration methods using their official pre-trained weights. The reported metrics (mAUC, RMSE, failure/inaccuracy rates) are empirical outcomes of applying existing methods to new data, not quantities forced by construction. The paper's self-citations (review [6], segmentation backbone [21], and the dataset record [24]) are not load-bearing: [21] is trained on FIVES, [6] is a literature review, and [24] is the dataset release itself. The lack of inter-observer variability analysis for the manual ground truth is an annotation-quality limitation, not circularity, because the benchmark conclusions are not defined in terms of the annotations; they are measured outcomes. The central novelty claim about being the first infant-focused retinal registration dataset is a factual claim about prior datasets, not a derivation. Overall, no specific circular step can be quoted, so the paper is essentially self-contained against external benchmarks.
Assumptions & free parameters
free parameters (4)
- MEE threshold =
20 pixels
- MAE threshold =
50 pixels
- Minimum feature matches =
4
- Ground truth points per pair =
10
assumptions (5)
- domain assumption The RIDIRP/Timkovic dataset is a valid and correctly characterized source of infant fundus images.
- domain assumption The manually labeled point correspondences are correct and consistent.
- domain assumption The vessel segmentation model trained on the adult FIVES dataset generalizes to infant RetCam images.
- domain assumption The pre-trained models and official implementations used in evaluation are appropriately configured and representative.
- ad hoc to paper Excluding eyes without consistent lesion visibility yields a dataset that still represents the target population.
Cite this review
Pith. "Pith review of COph100: A comprehensive fundus image registration dataset from infants constituting the "RIDIRP" database." pith.science (2026). https://pith.science/paper/XJVWRW6K
@misc{pith2026250102800,
author = {Pith},
title = {Pith review of: COph100: A comprehensive fundus image registration dataset from infants constituting the "RIDIRP" database},
year = {2026},
howpublished = {\url{https://pith.science/paper/XJVWRW6K}},
note = {Machine review of arXiv:2501.02800}
}
read the original abstract
Retinal image registration is vital for diagnostic therapeutic applications within the field of ophthalmology. Existing public datasets, focusing on adult retinal pathologies with high-quality images, have limited number of image pairs and neglect clinical challenges. To address this gap, we introduce COph100, a novel and challenging dataset known as the Comprehensive Ophthalmology Retinal Image Registration dataset for infants with a wide range of image quality issues constituting the public "RIDIRP" database. COph100 consists of 100 eyes, each with 2 to 9 examination sessions, amounting to a total of 491 image pairs carefully selected from the publicly available dataset. We manually labeled the corresponding ground truth image points and provided automatic vessel segmentation masks for each image. We have assessed COph100 in terms of image quality and registration outcomes using state-of-the-art algorithms. This resource enables a robust comparison of retinal registration methodologies and aids in the analysis of disease progression in infants, thereby deepening our understanding of pediatric ophthalmic conditions.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Timkoviˇc, J. et al. Retinal image dataset of infants and retinopathy of prematurity. Sci. Data 11, 814, https://doi.org/10. 1038/s41597-024-03409-7 (2024)
work page 2024
-
[2]
Avants, B. B., Epstein, C. L., Grossman, M. & Gee, J. C. Symmetric diffeomorphic image registration with cross- correlation: evaluating automated labeling of elderly and neurodegenerative brain. Med. image analysis 12, 26–41, https://doi.org/10.1016/j.media.2007.06.004 (2008)
-
[3]
Alam, F., Rahman, S. U., Ullah, S. & Gulati, K. Medical image registration in image guided surgery: Issues, challenges and research opportunities. Biocybern. Biomed. Eng. 38, 71–89, https://doi.org/10.1016/j.bbe.2017.10.001 (2018)
-
[4]
Javaid, F. Z., Brenton, J., Guo, L. & Cordeiro, M. F. Visual and ocular manifestations of alzheimer’s disease and their use as biomarkers for diagnosis and progression. Front. neurology7, 55, https://doi.org/10.3389/fneur.2016.00055 (2016)
-
[5]
James, A. P. & Dasarathy, B. V . Medical image fusion: A survey of the state of the art. Inf. fusion 19, 4–19, https: //doi.org/10.1016/j.inffus.2013.12.002 (2014)
-
[6]
Nie, Q., Zhang, X., Hu, Y ., Gong, M. & Liu, J. Medical image registration and its application in retinal images: a review. Vis. Comput. for Ind. Biomed. Art 7, 21, https://doi.org/10.1186/s42492-024-00173-8 (2024)
- [7]
-
[8]
Adal, K. M., van Etten, P. G., Martinez, J. P., van Vliet, L. J. & Vermeer, K. A. Accuracy assessment of intra-and intervisit fundus image registration for diabetic retinopathy screening. Investig. ophthalmology & visual science 56, 1805–1812, https://doi.org/10.1167/iovs.14-15949 (2015)
Show all 37 references
-
[9]
Decenciere, E. et al. Teleophta: Machine learning and image processing methods for teleophthalmology.Irbm 34, 196–203, https://doi.org/10.1016/j.irbm.2013.01.010 (2013)
2013 doi
-
[10]
G., Rouco, J., Barreira, N
Ortega, M., Penedo, M. G., Rouco, J., Barreira, N. & Carreira, M. J. Retinal verification using a feature points-based biometric pattern. EURASIP J. on Adv. Signal Process. 2009, 1–13, https://doi.org/10.1155/2009/235746 (2009)
2009 doi
-
[11]
Lee, J. A. et al. Registration of color and oct fundus images using low-dimensional step pattern analysis. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part II 18, 214...
2015 doi
-
[12]
Almasi, R. et al. Registration of fluorescein angiography and optical coherence tomography images of curved retina via scanning laser ophthalmoscopy photographs. Biomed. Opt. Express 11, 3455–3476, https://doi.org/10.1364/BOE.395784 (2020)
2020 doi
-
[13]
E., Ramchandran, R
Ding, L., Kuriyan, A. E., Ramchandran, R. S., Wykoff, C. C. & Sharma, G. Weakly-supervised vessel detection in ultra-widefield fundus photography via iterative multi-modal registration and learning.IEEE Transactions on Med. Imaging 40, 2748–2758, https://doi.org/10.1109/TMI.20...
2020
-
[14]
J., Cancelas, D., Novo, J
Martínez-Río, J., Carmona, E. J., Cancelas, D., Novo, J. & Ortega, M. Robust multimodal registration of fluorescein angiography and optical coherence tomography angiography images using evolutionary algorithms.Comput. Biol. Medicine 134, 104529, https://doi.org/10.1016/j.compb...
2021
-
[15]
Memo: dataset and methods for robust multimodal retinal image registration with large or small vessel density differences
Wang, C.-Y .et al. Memo: dataset and methods for robust multimodal retinal image registration with large or small vessel density differences. Biomed. Opt. Express 15, 3457–3479, https://doi.org/10.1364/BOE.516481 (2024). 10/12
2024 doi
-
[16]
Ding, L., Kang, T., Kuriyan, A. et al. Flori21: Fluorescein angiography longitudinal retinal image registration dataset. IEEE Dataport https://doi.org/10.21227/ydp8-zf19 (2021)
2021 doi
- [17]
-
[18]
Hernandez-Matas, C. et al. Fire: fundus image registration dataset. Model. Artif. Intell. Ophthalmol. 1, 16–28, https: //doi.org/10.35119/maio.v1i4.42 (2017)
2017 doi
-
[19]
Superjunction: Learning-based junction detection for retinal image registration
Wang, Y .et al. Superjunction: Learning-based junction detection for retinal image registration. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, 292–300 (2024)
2024
-
[20]
& Kotecha, K
Agrawal, R., Kulkarni, S., Walambe, R. & Kotecha, K. Assistive framework for automatic detection of all the zones in retinopathy of prematurity using deep learning.J. Digit. Imaging 34, 932–947, https://doi.org/10.1007/s10278-021-00477-8 (2021)
2021 doi
-
[21]
Qiu, Z. et al. Rethinking dual-stream super-resolution semantic learning in medical image segmentation.IEEE Transactions on Pattern Analysis Mach. Intell. 46, 451–464, https://doi.org/10.1109/TPAMI.2023.3322735 (2024)
2024
-
[22]
& Rockett, P
Owler, J. & Rockett, P. Influence of background preprocessing on the performance of deep learning retinal vessel detection. J. Med. Imaging 8, 064001–064001, https://doi.org/10.1117/1.JMI.8.6.064001 (2021)
2021 doi
-
[23]
Jin, K. et al. Fives: A fundus image dataset for artificial intelligence based vessel segmentation. Sci. data 9, 475, https://doi.org/10.1038/s41597-022-01564-3 (2022)
2022 doi
-
[24]
Coph100: A comprehensive fundus image registration dataset from infants constituting the "ridirp" database
Hu, Y .et al. Coph100: A comprehensive fundus image registration dataset from infants constituting the "ridirp" database. figshare https://doi.org/10.6084/m9.figshare.27061084 (2024)
2024 doi
-
[25]
Fu, H. et al. Evaluation of retinal image quality assessment networks in different color-spaces. InMedical Image Computing and Computer Assisted Intervention–MICCAI 2019: 22nd International Conference, Shenzhen, China, October 13–17, 2019, Proceedings, Part I 22, 48–56, https:...
2019 doi
-
[26]
Liu, L. et al. Deepfundus: a flow-cytometry-like image quality classifier for boosting the whole life cycle of medical artificial intelligence. Cell Reports Medicine 4, https://doi.org/10.1016/j.xcrm.2022.100912 (2023)
2023
-
[27]
Domain-invariant interpretable fundus image quality assessment
Shen, Y .et al. Domain-invariant interpretable fundus image quality assessment. Med. image analysis 61, 101654, https://doi.org/10.1016/j.media.2020.101654 (2020)
2020
-
[28]
Truong, P. et al. Glampoints: Greedily learned accurate match points. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 10732–10741, https://doi.org/10.1109/ICCV .2019.01083 (2019)
2019
-
[29]
Lowe, D. G. Distinctive image features from scale-invariant keypoints. Int. journal computer vision 60, 91–110, https://doi.org/10.1023/B:VISI.0000029664.99615.94 (2004)
2004
-
[30]
V ., Sofka, M
Yang, G., Stewart, C. V ., Sofka, M. & Tsai, C.-L. Registration of challenging image pairs: Initialization, estimation, and decision. IEEE transactions on pattern analysis machine intelligence 29, 1973–1989, https://doi.org/10.1109/TPAMI.2007. 1116 (2007)
2007 doi
-
[31]
& Argyros, A
Hernandez-Matas, C., Zabulis, X. & Argyros, A. A. Rempe: Registration of retinal images through eye modelling and pose estimation. IEEE journal biomedical health informatics 24, 3362–3373, https://doi.org/10.1109/JBHI.2020.2984483 (2020)
2020
-
[32]
& Rabinovich, A
DeTone, D., Malisiewicz, T. & Rabinovich, A. Superpoint: Self-supervised interest point detection and description. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, 224–236, https://doi.org/10. 1109/CVPRW.2018.00060 (2018)
2018
- [33]
-
[34]
& Rabinovich, A
Sarlin, P.-E., DeTone, D., Malisiewicz, T. & Rabinovich, A. Superglue: Learning feature matching with graph neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 4938–4947, https: //doi.org/10.1109/CVPR42600.2020.00499 (2020)
2020
-
[35]
& Zhou, X
Sun, J., Shen, Z., Wang, Y ., Bao, H. & Zhou, X. Loftr: Detector-free local feature matching with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 8922–8931, https://doi.org/10.1109/ CVPR46437.2021.00881 (2021)
2021
-
[36]
& Ding, D
Liu, J., Li, X., Wei, Q., Xu, J. & Ding, D. Semi-supervised keypoint detector and descriptor for retinal image matching. In European Conference on Computer Vision, 593–609, https://doi.org/10.1007/978-3-031-19803-8_35 (Springer, 2022). 11/12
-
[37]
Climbing Program
Nasser, S. A., Gupte, N. & Sethi, A. Reverse knowledge distillation: Training a large model using a small one for retinal image matching on limited data. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 7778–7787, https://doi.org/10.1109/W A...
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.