Pith. sign in

REVIEW 3 major objections 6 minor 64 references

A Clinician-Friendly Platform for Ophthalmic Image Analysis Without Technical Barriers

T0 review · 3 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read GlobeReady's central claim is that training-free retrieval from a labeled image library, augmented with local features, can diagnose eye diseases across countries and modalities without retraining.

desk verdict GlobeReady has a solid retrieval-based core and substantial new datasets, but its headline confidence-threshold and OOD numbers are inflated by choosing θ* on the test set. read the letter →

arxiv 2504.15928 v2 pith:XSO6NIOT submitted 2025-04-22 cs.CV cs.AI

classification cs.CVcs.AI
keywords ophthalmicfoundationmodelretrieval-augmenteddiagnosisdomainshiftfundusphotographyopticalcoherencetomographyconfidencequantificationout-of-distributiondetectiontraining-freedeployment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

GlobeReady is an attempt to remove the main obstacle to using AI in eye clinics: the need to retrain or fine-tune a model for each new hospital, device, or patient population. The paper claims that a single pretrained feature extractor, combined with retrieval from a labeled library of fundus and OCT images, can diagnose 11 disease categories from color fundus photographs (93.9-98.5% Top-1 accuracy) and 15 categories from OCT scans (87.2-92.7% Top-1 accuracy) with no task-specific retraining. In external centers, adding locally extracted features to the library raised accuracy substantially, and a Bayesian confidence mechanism with threshold-based flagging further improved accuracy while detecting diseases never seen in the reference set. If the claim holds, deploying ophthalmic AI would become a matter of uploading images and reference data, not a machine-learning project.

What carries the argument

The load-bearing mechanism is a feature-matching diagnostic engine. Its backbone is a ViT-L/16 vision transformer pretrained with DINOv2 self-supervision on synthetic ophthalmic images, then aligned to clinical text with a CLIP-style contrastive objective on over 475,000 image-text pairs. Diagnosis happens by retrieval: the query embedding is compared with a labeled reference library, and the labels of the top matches supply the prediction. Domain adaptation is achieved without weight updates by appending locally extracted features to the library. Confidence comes from 100 stochastic forward passes with Monte Carlo dropout: the consistency of predictions across passes yields a confidence score, and Youden's index selects a threshold for flagging low-confidence cases for clinician review.

What would settle it

Re-run the JLHW11 and JSIEC-OCT15 flagging experiments with the confidence cutoff chosen on an independent validation split instead of on the test set; if Top-1 accuracy after flagging falls materially below the reported 99.4% (CFP) and 96.2% (OCT), the deployment-time benefit of the confidence mechanism is not supported.

Watch

Extended reading notes

Core claim

The paper's central claim is that ophthalmic disease diagnosis can be cast as a nearest-neighbor retrieval problem over a fixed feature library rather than as a trained classifier. GlobeReady builds the library by extracting features with a vision transformer pretrained first on 38 million synthetic fundus and OCT images via DINOv2-style self-supervision, then on 475,845 real image-text pairs via CLIP-style contrastive learning. At inference, a query image is matched against the library and the labels of the top matches determine the diagnosis. The authors report Top-1 accuracies of 93.9-98.5% on 11 CFP categories and 87.2-92.7% on 15 OCT categories, and show that augmenting the library with local features lifts cross-center accuracy in Vietnam from 51.0-72.7% to 86.3-96.9% and in the UK from 33.6-55.9% to 90.2-98.9%. A Bayesian variant using Monte Carlo dropout produces confidence scores, and thresholding on those scores raises Top-1 accuracy to 99.4% on CFPs and 96.2% on OCT while flagging out-of-distribution diseases with 86.3% and 90.6% detection rates respectively.

Load-bearing premise

The post-flagging accuracy and OOD detection numbers assume the confidence cutoff is chosen using the test set answers themselves (Youden's index on the test distribution), so those numbers may not hold for a cutoff fixed before deployment.

Editorial extensions

If this is right

  • A single pretrained feature extractor can serve both fundus photography and OCT diagnosis across 11 and 15 disease categories without task-specific retraining.
  • Hospitals in new regions can adapt the system by adding their own images to the reference library; the largest accuracy gains appear where baseline performance was lowest.
  • Confidence scoring with a threshold allows low-confidence diagnoses to be deferred to clinicians, raising Top-1 accuracy to 99.4% on CFPs and 96.2% on OCT after flagging.
  • The same feature library supports case retrieval, letting clinicians find similar images in unlabeled local databases without retraining.
  • Clinician usability ratings (SUS 86.4, helpfulness 4.6/5) indicate the platform is operable without programming expertise.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If replicated, the retrieval-based design shifts the deployment bottleneck from model training to data curation and library maintenance; a natural next test is how accuracy scales with library size, label noise, and the mix of local versus global reference images.
  • Because local feature augmentation works without retraining, the same mechanism could be tested on severity grading or on other imaging modalities such as ultrasound or pathology, not just disease classification.
  • The reported post-flagging accuracies depend on the threshold being chosen from the test distribution, so a practical extension would be per-site threshold selection on a small labeled sample, which the paper does not evaluate.
  • Retrieval-based diagnosis exposes the evidence behind each prediction, so a follow-up study could measure whether showing clinicians the retrieved similar cases changes their trust in or agreement with the system.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces GlobeReady, a retrieval-based ophthalmic image analysis platform that performs disease diagnosis without retraining or fine-tuning. The system is pretrained with DINOv2 on 38 million synthetic ophthalmic images and with CLIP on 475,845 image-text pairs, then used to extract features for k-nearest-neighbor diagnosis against labeled reference libraries. The authors report Top-1 accuracies of 93.9% on an 11-class CFP dataset and 87.2% on a 15-class OCT dataset, improved cross-center accuracy after local feature augmentation (e.g., 73.4–91.0% in Singapore, 86.3–96.9% in Vietnam, 90.2–98.9% in the UK), and enhanced post-threshold accuracy (99.4% CFP, 96.2% OCT) using a Bayesian confidence mechanism with Monte Carlo dropout. The paper also describes OOD detection for 49 CFP and 13 OCT disease categories, feature-based case retrieval evaluated by three ophthalmologists, and a usability survey of seven clinicians. The central methodological claim is that training-free retrieval, augmented with locally extracted features, is sufficient for deployment across centers and populations.

Significance. If the results hold, GlobeReady would be a practically valuable contribution: it combines large-scale pretraining, a code-free retrieval interface, and local feature augmentation to address domain shift without model retraining. The base retrieval accuracies are internally plausible for a k-NN classifier, and the prospective retrieval evaluation by three ophthalmologists is a genuine strength. However, the headline 'confidence-quantifiable' and OOD results rest on a threshold-selection procedure that uses test-set labels, making those specific numbers optimistic estimates of deployment performance. The cross-center adaptation results are also dependent on labeled local reference libraries, which should be clearly positioned as a supervised adaptation cost. Overall, the paper's core idea is sound and likely publishable, but the confidence-threshold claims need a methodological fix before acceptance.

major comments (3)
  1. [Sec. 4.12, Eq. (3); Sec. 2.3] The confidence threshold θ* is defined in Eq. (3) as the maximizer of Youden's index computed on the test dataset, and Sec. 2.3 then reports post-flagging accuracy of 99.4% (CFP) and 96.2% (OCT) on that same test set. Because the threshold is fitted to the test labels, these numbers are in-sample optimized values, not predictions of deployment performance. The improvement over the unthresholded baseline (93.9% and 87.2%) is therefore partly an artifact of test-set adaptation, and the reported P values comparing thresholded and unthresholded results are invalid for inferring a real gain. Please re-evaluate with a threshold chosen on a separate validation set or fixed a priori, and report the corresponding performance and uncertainty.
  2. [Sec. 2.4] The OOD detection rates (86.3% for CFPOOD49, 90.6% for OCT-OOD13, and the low-quality image detection rates) are reported using 'confidence-threshold filtering,' but the manuscript does not specify how the threshold was selected for these OOD evaluations. If θ* is again tuned on the OOD test distribution, the same circularity as in Sec. 4.12 applies. Please state explicitly the threshold-selection protocol for all OOD results, including whether the threshold was held out from the OOD test set, and provide detection rates along with confidence intervals.
  3. [Sec. 2.3, Supplementary Fig. 3] The comparison with RETFound and VisionFM reports that GlobeReady showed 'comparable or superior Top-1' accuracy and superior thresholded accuracy, but the threshold for the fine-tuned baselines is not described. For a fair comparison, the confidence-thresholding protocol must be identical across methods: either a fixed threshold for all methods or a separately tuned threshold per method, with the tuning set disjoint from the test set. Otherwise, the superiority claim for thresholded accuracy may reflect test-set tuning rather than a genuine advantage.
minor comments (6)
  1. [Sec. 2.1 heading] The heading reads 'Ocular disease diagnosis by GlobaFree' and should be 'GlobeReady'; this typo appears to be a simple misspelling.
  2. [Sec. 2.3, JSIEC-OCT15 result] The phrase 'P> 0.001' following the OCT post-flagging recall likely should be 'P < 0.001' to match the direction of the reported improvement; please verify and correct.
  3. [Introduction, paragraph 1] The sentence 'We then curated 475,845 real image-text pairs from these global datasets.' is duplicated verbatim; one occurrence should be removed.
  4. [Throughout] The abbreviation for color fundus photographs is inconsistent: the abstract and Section 2.1 use 'CPFs' while most of the text uses 'CFPs' (e.g., 'CFPOOD49'). Please standardize to a single abbreviation.
  5. [Sec. 4.12, Eq. (2)] The definition of Sensitivity and Specificity in Eq. (2) is unclear: it should state explicitly that these are computed with respect to the binary event 'prediction is correct' as the positive class, not with respect to a disease label. Please clarify the notation.
  6. [Sec. 4.11] The number of stochastic forward passes (100) and the retrieval neighborhood size k are presented as fixed choices, but no sensitivity analysis is provided; a brief statement on how these were chosen would help reproducibility.

Circularity Check

1 steps flagged · score 6.0 of 10

Post-flagging accuracy and OOD detection rates rely on a confidence threshold θ* selected on the test-set labels (Eq. 3), so the headline 99.4%/96.2% and 86.3%/90.6% figures are partly fitted rather than predicted.

  1. fitted input called prediction [Section 4.12, Eqs. (1)-(3); results reported in Sections 2.3-2.4]
    "In this study, to balance sensitivity and specificity in retinal disease diagnosis, we adopt Youden's index J [41] to determine the confidence threshold θ∗ based on the distribution of confidence scores in the test dataset: J (θ) = Sensitivity (θ) + Specificity (θ) − 1 ... θ∗ = arg max θ (J (θ))."

    Sensitivity(θ) and Specificity(θ) in Eq. (2) are computed from the test-set labels, so Eq. (3) fits θ* to maximize Youden's index on exactly the evaluation set. The paper then reports post-flagging Top-1 accuracy of 99.4% (JLHW11) and 96.2% (JSIEC-OCT15), and OOD/low-quality detection rates (86.3%, 90.6%, 94.6%, 92.8%) on that same test distribution. These numbers measure the fitted threshold's own training objective, not an independent deployment prediction: a clinician at a new site would not know the label-optimal threshold, and the reported P values for thresholded comparisons are not valid because the same labels determine both θ* and the outcome.

full rationale

The core retrieval-based diagnosis is self-contained: the JLHW11/JSIEC-OCT15 reference subsets are disjoint from the test subsets, local retrieval augmentation is evaluated on held-out local test images, and the cross-center results are external validation rather than an artifact of training. The main circularity is confined to the confidence-quantification and OOD claims, where θ* is selected by Youden's index on the test distribution (Section 4.12) and post-flagging accuracy/OOD rates are reported on the same set. That step reduces the headline 'confidence-quantifiable' and OOD figures to an in-sample fit. Self-citations (FundusGAN, image-text pair curation, UIOS/FMUE25) are not load-bearing for the central diagnostic claim; they support pretraining data or related prior work without forbidding alternatives. Overall score reflects partial circularity in one prominent sub-claim while the base diagnostic and retrieval evidence retain independent content.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The central system relies on a pretrained feature extractor and a labeled reference library; no new theoretical constants are introduced. The only explicit fitted parameter in the experimental pipeline is the confidence threshold θ*, selected by Youden's index on test data. The k value and number of dropout passes are design choices rather than fitted constants. The axioms are domain assumptions about feature transferability and label reliability, which are standard in empirical medical imaging papers but are not validated by ablations here.

free parameters (3)
  • Confidence threshold θ* = Per-dataset argmax of Youden's index over test set (JLHW11, JSIEC-OCT15)
    Chosen using test-set labels via Eq. 3; directly determines post-flagging accuracy and OOD detection rates.
  • Number of MC dropout forward passes = 100
    Hand-chosen; the confidence score is the consistency across 100 stochastic passes (Sec. 4.11).
  • Retrieval neighborhood size k = Top-1 used for primary diagnosis; Top-3/5/10 reported for retrieval
    The diagnostic method uses nearest-neighbor labels; k is varied across metrics and not optimized, but its choice affects the reported numbers.
assumptions (3)
  • domain assumption Pretrained features from DINOv2+CLIP on synthetic and real ophthalmic images are sufficiently discriminative and domain-invariant for k-NN diagnosis across centers and modalities.
    The entire system rests on this; no ablation is provided to isolate its contribution (Sec. 4.7-4.10).
  • domain assumption Reference library labels (from public datasets and local hospitals) are correct and complete.
    Diagnosis is label propagation from retrieved cases; label noise would propagate (Sec. 4.1-4.3).
  • domain assumption Random partitioning of JLHW11 and JSIEC-OCT15 produces patient-disjoint reference and test sets.
    The paper states 'randomly partitioned' without confirming no patient or device overlap (Sec. 4.2).

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Clinician-Friendly Platform for Ophthalmic Image Analysis Without Technical Barriers." pith.science (2026). https://pith.science/paper/XSO6NIOT

@misc{pith2026250415928,
  author       = {Pith},
  title        = {Pith review of: A Clinician-Friendly Platform for Ophthalmic Image Analysis Without Technical Barriers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XSO6NIOT}},
  note         = {Machine review of arXiv:2504.15928}
}
read the original abstract

Artificial intelligence (AI) shows remarkable potential in medical imaging diagnostics, yet most current models require retraining when applied across different clinical settings, limiting their scalability. We introduce GlobeReady, a clinician-friendly AI platform that enables fundus disease diagnosis that operates without retraining, fine-tuning, or the needs for technical expertise. GlobeReady demonstrates high accuracy across imaging modalities: 93.9-98.5% for 11 fundus diseases using color fundus photographs (CPFs) and 87.2-92.7% for 15 fundus diseases using optic coherence tomography (OCT) scans. By leveraging training-free local feature augmentation, GlobeReady platform effectively mitigates domain shifts across centers and populations, achieving accuracies of 88.9-97.4% across five centers on average in China, 86.3-96.9% in Vietnam, and 73.4-91.0% in Singapore, and 90.2-98.9% in the UK. Incorporating a bulit-in confidence-quantifiable diagnostic mechanism further enhances the platform's accuracy to 94.9-99.4% with CFPs and 88.2-96.2% with OCT, while enabling identification of out-of-distribution cases with 86.3% accuracy across 49 common and rare fundus diseases using CFPs, and 90.6% accuracy across 13 diseases using OCT. Clinicians from countries rated GlobeReady highly for usability and clinical relevance (average score 4.6/5). These findings demonstrate GlobeReady's robustness, generalizability and potential to support global ophthalmic care without technical barriers.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

64 extracted references · 52 canonical work pages

  1. [1]

    J. D. Steinmetz, R. R. Bourne, P. S. Briant, S. R. Flaxman, H. R. Taylor, J. B. Jonas, A. A. Abdoli, W. A. Abrha, A. Abualhasan, E. G. Abu-Gharbiehet al., “Causes of blindness and vision impairment in 2020 and trends over 30 years, and prevalence of avoidable blindness in relation to vision 2020: the right to sight: an analysis for the global burden of di...

  2. [2]

    The lancet global health commission on global eye health: vision beyond 2020,

    M. J. Burton, J. Ramke, A. P. Marques, R. R. Bourne, N. Congdon, I. Jones, B. A. A. Tong, S. Arunga, D. Bachani, C. Bascaran et al., “The lancet global health commission on global eye health: vision beyond 2020,”The Lancet Global Health, vol. 9, no. 4, pp. e489–e551, 2021

  3. [3]

    Application of a deep-learning marker for morbidity and mortality prediction derived from retinal photographs: a cohort development and validation study,

    S. Nusinovici, T. H. Rim, H. Li, M. Yu, M. Deshmukh, T. C. Quek, G. Lee, C. C. Y. Chong, Q. Peng, C. C. Xueet al., “Application of a deep-learning marker for morbidity and mortality prediction derived from retinal photographs: a cohort development and validation study,”The lancet Healthy longevity, vol. 5, no. 10, 2024

  4. [4]

    A generalist vision–language foundation model for diverse biomedical tasks,

    K. Zhang, R. Zhou, E. Adhikarla, Z. Yan, Y. Liu, J. Yu, Z. Liu, X. Chen, B. D. Davison, H. Renet al., “A generalist vision–language foundation model for diverse biomedical tasks,”Nature Medicine, pp. 1–13, 2024

  5. [5]

    A deep learning system for detecting diabetic retinopathy across the disease spectrum,

    L. Dai, L. Wu, H. Li, C. Cai, Q. Wu, H. Kong, R. Liu, X. Wang, X. Hou, Y. Liuet al., “A deep learning system for detecting diabetic retinopathy across the disease spectrum,”Nature communications, vol. 12, no. 1, p. 3242, 2021

  6. [6]

    Automatic staging for retinopathy of prematurity with deep feature fusion and ordinal classification strategy,

    Y. Peng, W. Zhu, Z. Chen, M. Wang, L. Geng, K. Yu, Y. Zhou, T. Wang, D. Xiang, F. Chenet al., “Automatic staging for retinopathy of prematurity with deep feature fusion and ordinal classification strategy,”IEEE transactions on medical imaging, vol. 40, no. 7, pp. 1750–1762, 2021

  7. [7]

    A deep network deepopacitynet for detection of cataracts from color fundus photographs,

    A. Elsawy, T. D. Keenan, Q. Chen, A. T. Thavikulwat, S. Bhandari, T. C. Quek, J. H. L. Goh, Y.-C. Tham, C.-Y. Cheng, E. Y. Chewet al., “A deep network deepopacitynet for detection of cataracts from color fundus photographs,” Communications Medicine, vol. 3, no. 1, p. 184, 2023

  8. [8]

    A foundation model for generalizable disease detection from retinal images,

    Y. Zhou, M. A. Chia, S. K. Wagner, M. S. Ayhan, D. J. Williamson, R. R. Struyven, T. Liu, M. Xu, M. G. Lozano, P. Woodward-Courtet al., “A foundation model for generalizable disease detection from retinal images,”Nature, vol. 622, no. 7981, pp. 156–163, 2023

Show all 64 references
  1. [9]

    Visionfm: a multi-modal multi-task vision foundation model for generalist ophthalmic artificial intelligence,

    J. Qiu, J. Wu, H. Wei, P. Shi, M. Zhang, Y. Sun, L. Li, H. Liu, H. Liu, S. Houet al., “Visionfm: a multi-modal multi-task vision foundation model for generalist ophthalmic artificial intelligence,”arXiv preprint arXiv:2310.04992, 2023

  2. [10]

    A guide to deep learning in healthcare,

    A. Esteva, A. Robicquet, B. Ramsundar, V. Kuleshov, M. DePristo, K. Chou, C. Cui, G. Corrado, S. Thrun, and J. Dean, “A guide to deep learning in healthcare,”Nature medicine, vol. 25, no. 1, pp. 24–29, 2019

  3. [11]

    On the opportunities and risks of foundation models,

    R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskillet al., “On the opportunities and risks of foundation models,”arXiv preprint arXiv:2108.07258, 2021. GlobeReady 15

  4. [12]

    On the opportunities and risks of foundation models for natural language processing in radiology,

    W. F. Wiggins and A. S. Tejani, “On the opportunities and risks of foundation models for natural language processing in radiology,”Radiology: Artificial Intelligence, vol. 4, no. 4, p. e220119, 2022

  5. [13]

    Dinov2: Learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Noubyet al., “Dinov2: Learning robust visual features without supervision,”arXiv preprint arXiv:2304.07193, 2023

  6. [14]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clarket al., “Learning transferable visual models from natural language supervision,” inInternational conference on machine learning. PmLR, 2021, pp. 8748–8763

  7. [15]

    Retrieval-augmented generation for knowledge-intensive nlp tasks,

    P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Kuttler, M. Lewis, W.-t. Yih, T. Rocktaschel et al., “Retrieval-augmented generation for knowledge-intensive nlp tasks,”Advances in neural information processing systems, vol. 33, pp. 9459–9474, 2020

  8. [16]

    Query rewriting in retrieval-augmented large language models,

    X. Ma, Y. Gong, P. He, H. Zhao, and N. Duan, “Query rewriting in retrieval-augmented large language models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 5303–5315

  9. [17]

    P. W. Jordan, B. Thomas, I. L. McClelland, and B. Weerdmeester,Usability evaluation in industry. CRC press, 1996

  10. [18]

    User acceptance of information technology: Toward a unified view,

    V. Venkatesh, M. G. Morris, G. B. Davis, and F. D. Davis, “User acceptance of information technology: Toward a unified view,” MIS quarterly, pp. 425–478, 2003

  11. [19]

    Visualization of supervised and self-supervised neural networks via attribution guided factorization,

    S. Gur, A. Ali, and L. Wolf, “Visualization of supervised and self-supervised neural networks via attribution guided factorization,” inProceedings of the AAAI conference on artificial intelligence, vol. 35, no. 13, 2021, pp. 11545–11554

  12. [20]

    Code-free deep learning glaucoma detection on color fundus images,

    D. Milad, F. Antaki, D. Mikhail, A. Farah, J. El-Khoury, S. Touma, G. M. Durr, T. Nayman, C. Playout, P. A. Keane et al., “Code-free deep learning glaucoma detection on color fundus images,”Ophthalmology Science, vol. 5, no. 4, p. 100721, 2025

  13. [21]

    Development and international validation of custom-engineered and code-free deep-learning models for detection of plus disease in retinopathy of prematurity: a retrospective study,

    S. K. Wagner, B. Liefers, M. Radia, G. Zhang, R. Struyven, L. Faes, J. Than, S. Balal, C. Hennings, C. Kilduffet al., “Development and international validation of custom-engineered and code-free deep-learning models for detection of plus disease in retinopathy of prematurity: ...

  14. [22]

    Bisonget al., Building machine learning and deep learning models on Google cloud platform

    E. Bisonget al., Building machine learning and deep learning models on Google cloud platform. Springer, 2019

  15. [23]

    Barnes,Microsoft Azure essentials Azure machine learning

    J. Barnes,Microsoft Azure essentials Azure machine learning. Microsoft Press, 2015

  16. [24]

    Towards a general-purpose foundation model for computational pathology,

    R. J. Chen, T. Ding, M. Y. Lu, D. F. Williamson, G. Jaume, A. H. Song, B. Chen, A. Zhang, D. Shao, M. Shaban et al., “Towards a general-purpose foundation model for computational pathology,”Nature Medicine, vol. 30, no. 3, pp. 850–862, 2024

  17. [25]

    A visual–language foundation model for pathology image analysis using medical twitter,

    Z. Huang, F. Bianchi, M. Yuksekgonul, T. J. Montine, and J. Zou, “A visual–language foundation model for pathology image analysis using medical twitter,”Nature medicine, vol. 29, no. 9, pp. 2307–2316, 2023

  18. [26]

    Uncertainty-inspired open set learning for retinal anomaly identification,

    M. Wang, T. Lin, L. Wang, A. Lin, K. Zou, X. Xu, Y. Zhou, Y. Peng, Q. Meng, Y. Qianet al., “Uncertainty-inspired open set learning for retinal anomaly identification,”Nature Communications, vol. 14, no. 1, p. 6757, 2023

  19. [27]

    Enhancing ai reliability: A foundation model with uncertainty estimation for optical coherence tomography-based retinal disease diagnosis,

    Y. Peng, A. Lin, M. Wang, T. Lin, L. Liu, J. Wu, K. Zou, T. Shi, L. Feng, Z. Lianget al., “Enhancing ai reliability: A foundation model with uncertainty estimation for optical coherence tomography-based retinal disease diagnosis,” Cell Reports Medicine, vol. 6, no. 1, 2025

  20. [28]

    Deep triplet hashing network for case-based medical image retrieval,

    J. Fang, H. Fu, and J. Liu, “Deep triplet hashing network for case-based medical image retrieval,”Medical image analysis, vol. 69, p. 101981, 2021

  21. [29]

    Automated assessment of diabetic retinopathy severity using content-based image retrieval in multimodal fundus photographs,

    G. Quellec, M. Lamard, G. Cazuguel, L. Bekri, W. Daccache, C. Roux, and B. Cochener, “Automated assessment of diabetic retinopathy severity using content-based image retrieval in multimodal fundus photographs,”Investigative ophthalmology and visual science, vol. 52, no. 11, pp...

  22. [30]

    Zero-shot text-to-image generation,

    A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in International conference on machine learning. Pmlr, 2021, pp. 8821–8831

  23. [31]

    Hierarchical text-conditional image generation with clip latents,

    A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen, “Hierarchical text-conditional image generation with clip latents,” arXiv preprint arXiv:2204.06125, vol. 1, no. 2, p. 3, 2022

  24. [32]

    Improving image generation with better captions,

    J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guoet al., “Improving image generation with better captions,”Computer Science., vol. 2, no. 3, p. 8, 2023

  25. [33]

    Fundusgan: A hierarchical feature-aware generative framework for high-fidelity fundus image generation,

    Q. Hou, M. Wang, P. Cao, Z. Ke, X. Liu, H. Fu, and O. R. Zaiane, “Fundusgan: A hierarchical feature-aware generative framework for high-fidelity fundus image generation,”arXiv preprint arXiv:2503.17831, 2025

  26. [34]

    Common and rare fundus diseases identification using vision-language foundation model with knowledge of over 400 diseases,

    M. Wang, T. Lin, A. Lin, K. Yu, Y. Peng, L. Wang, C. Chen, K. Zou, H. Liang, M. Chenet al., “Common and rare fundus diseases identification using vision-language foundation model with knowledge of over 400 diseases,”arXiv preprint arXiv:2406.09317, 2024

  27. [35]

    Cohort profile: the singapore epidemiology of eye diseases study (seed),

    S. Majithia, Y.-C. Tham, M.-L. Chee, S. Nusinovici, C. L. Teo, M.-L. Chee, S. Thakur, Z. D. Soh, N. Kumari, E. Lamoureuxet al., “Cohort profile: the singapore epidemiology of eye diseases study (seed),”International journal of epidemiology, vol. 50, no. 1, pp. 41–52, 2021

  28. [36]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gellyet al., “An image is worth 16x16 words: Transformers for image recognition at scale,”arXiv preprint arXiv:2010.11929, 2020

  29. [37]

    Lora: Low-rank adaptation of large language models

    E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chenet al., “Lora: Low-rank adaptation of large language models.”ICLR, vol. 1, no. 2, p. 3, 2022

  30. [38]

    Publicly available clinical bert embeddings,

    E. Alsentzer, J. R. Murphy, W. Boag, W.-H. Weng, D. Jin, T. Naumann, and M. McDermott, “Publicly available clinical bert embeddings,”arXiv preprint arXiv:1904.03323, 2019

  31. [39]

    What uncertainties do we need in bayesian deep learning for computer vision?

    A. Kendall and Y. Gal, “What uncertainties do we need in bayesian deep learning for computer vision?”Advances in neural information processing systems, vol. 30, 2017. 16 M. Wang et al

  32. [40]

    Dropout as a bayesian approximation: Representing model uncertainty in deep learning,

    Y. Gal and Z. Ghahramani, “Dropout as a bayesian approximation: Representing model uncertainty in deep learning,” in international conference on machine learning. PMLR, 2016, pp. 1050–1059

  33. [41]

    Youden index and associated cut-points for three ordinal diagnostic groups,

    J. Luo and C. Xiong, “Youden index and associated cut-points for three ordinal diagnostic groups,”Communications in Statistics-Simulation and Computation, vol. 42, no. 6, pp. 1213–1234, 2013

  34. [42]

    Determining what individual sus scores mean: Adding an adjective rating scale,

    A. Bangor, P. Kortum, and J. Miller, “Determining what individual sus scores mean: Adding an adjective rating scale,” Journal of usability studies, vol. 4, no. 3, pp. 114–123, 2009

  35. [43]

    Cnns for automatic glaucoma assessment using fundus images: an extensive validation,

    A. Diaz-Pinto, S. Morales, V. Naranjo, T. Kohler, J. M. Mossi, and A. Navea, “Cnns for automatic glaucoma assessment using fundus images: an extensive validation,”Biomedical engineering online, vol. 18, pp. 1–19, 2019

  36. [44]

    Deep learning-based glaucoma detection with cropped optic cup and disc and blood vessel segmentation,

    M. T. Islam, S. T. Mashfu, A. Faisal, S. C. Siam, I. T. Naheen, and R. Khan, “Deep learning-based glaucoma detection with cropped optic cup and disc and blood vessel segmentation,”Ieee Access, vol. 10, pp. 2828–2841, 2021

  37. [45]

    Deepdrid: Diabetic retinopathy—grading and image quality estimation challenge,

    R. Liu, X. Wang, Q. Wu, L. Dai, X. Fang, T. Yan, J. Son, S. Tang, J. Li, Z. Gaoet al., “Deepdrid: Diabetic retinopathy—grading and image quality estimation challenge,”Patterns, vol. 3, no. 6, 2022

  38. [46]

    Advancing bag-of-visual-words representations for lesion classification in retinal images,

    R. Pires, H. F. Jelinek, J. Wainer, E. Valle, and A. Rocha, “Advancing bag-of-visual-words representations for lesion classification in retinal images,”PloS one, vol. 9, no. 6, p. e96814, 2014

  39. [47]

    Teleophta: Machine learning and image processing methods for teleophthalmology,

    E. Decenciere, G. Cazuguel, X. Zhang, G. Thibault, J.-C. Klein, F. Meyer, B. Marcotegui, G. Quellec, M. Lamard, R. Dannoet al., “Teleophta: Machine learning and image processing methods for teleophthalmology,”Irbm, vol. 34, no. 2, pp. 196–203, 2013

  40. [48]

    Airogs: artificial intelligence for robust glaucoma screening challenge,

    C. De Vente, K. A. Vermeer, N. Jaccard, H. Wang, H. Sun, F. Khader, D. Truhn, T. Aimyshev, Y. Zhanibekuly, T.-D. Leet al., “Airogs: artificial intelligence for robust glaucoma screening challenge,”IEEE transactions on medical imaging, 2023

  41. [49]

    Deepopht: medical report generation for retinal images via deep models and visual explanation,

    J.-H. Huang, C.-H. H. Yang, F. Liu, M. Tian, Y.-C. Liu, T.-W. Wu, I. Lin, K. Wang, H. Morikawa, H. Changet al., “Deepopht: medical report generation for retinal images via deep models and visual explanation,” inProceedings of the IEEE/CVF winter conference on applications of c...

  42. [50]

    Fives: A fundus image dataset for artificial intelligence based vessel segmentation,

    K. Jin, X. Huang, J. Zhou, Y. Li, Y. Yan, Y. Sun, Q. Zhang, Y. Wang, and J. Ye, “Fives: A fundus image dataset for artificial intelligence based vessel segmentation,”Scientific Data, vol. 9, no. 1, p. 475, 2022

  43. [51]

    G1020: A benchmark retinal fundus image dataset for computer-aided glaucoma detection,

    M. N. Bajwa, G. A. P. Singh, W. Neumeier, M. I. Malik, A. Dengel, and S. Ahmed, “G1020: A benchmark retinal fundus image dataset for computer-aided glaucoma detection,” in2020 International Joint Conference on Neural Networks (IJCNN). IEEE, 2020, pp. 1–7

  44. [52]

    Image processing based automatic diagnosis of glaucoma using wavelet features of segmented optic disc from fundus image,

    A. Singh, M. K. Dutta, M. ParthaSarathi, V. Uher, and R. Burget, “Image processing based automatic diagnosis of glaucoma using wavelet features of segmented optic disc from fundus image,”Computer methods and programs in biomedicine, vol. 124, pp. 108–120, 2016

  45. [53]

    An adaptive threshold based image processing technique for improved glaucoma detection and classification,

    A. Issac, M. P. Sarathi, and M. K. Dutta, “An adaptive threshold based image processing technique for improved glaucoma detection and classification,”Computer methods and programs in biomedicine, vol. 122, no. 2, pp. 229–244, 2015

  46. [54]

    Idrid: Diabetic retinopathy–segmentation and grading challenge,

    P. Porwal, S. Pachade, M. Kokare, G. Deshmukh, J. Son, W. Bae, L. Liu, J. Wang, X. Liu, L. Gaoet al., “Idrid: Diabetic retinopathy–segmentation and grading challenge,”Medical image analysis, vol. 59, p. 101561, 2020

  47. [55]

    Applying artificial intelligence to disease staging: Deep learning for improved staging of diabetic retinopathy,

    H. Takahashi, H. Tampo, Y. Arai, Y. Inoue, and H. Kawashima, “Applying artificial intelligence to disease staging: Deep learning for improved staging of diabetic retinopathy,”PloS one, vol. 12, no. 6, p. e0179790, 2017

  48. [56]

    Refuge challenge: A unified framework for evaluating automated methods for glaucoma assessment from fundus photographs,

    J. I. Orlando, H. Fu, J. B. Breda, K. Van Keer, D. R. Bathula, A. Diaz-Pinto, R. Fang, P.-A. Heng, J. Kim, J. Lee et al., “Refuge challenge: A unified framework for evaluating automated methods for glaucoma assessment from fundus photographs,” Medical image analysis, vol. 59, ...

  49. [57]

    Origa-light: An online retinal fundus image database for glaucoma analysis and research,

    Z. Zhang, F. S. Yin, J. Liu, W. K. Wong, N. M. Tan, B. H. Lee, J. Cheng, and T. Y. Wong, “Origa-light: An online retinal fundus image database for glaucoma analysis and research,” in2010 Annual international conference of the IEEE engineering in medicine and biology. IEEE, 201...

  50. [58]

    Dataset from fundus images for the study of diabetic retinopathy,

    V. E. C. Benitez, I. C. Matto, J. C. M. Roman, J. L. V. Noguera, M. Garcia-Torres, J. Ayala, D. P. Pinto-Roa, P. E. Gardel-Sotomayor, J. Facon, and S. A. Grillo, “Dataset from fundus images for the study of diabetic retinopathy,” Data in brief, vol. 36, p. 107068, 2021

  51. [59]

    Improving medical images classi- fication with label noise using dual-uncertainty estimation,

    L. Ju, X. Wang, L. Wang, D. Mahapatra, X. Zhao, Q. Zhou, T. Liu, and Z. Ge, “Improving medical images classi- fication with label noise using dual-uncertainty estimation,”IEEE transactions on medical imaging, vol. 41, no. 6, pp. 1533–1546, 2022

  52. [60]

    Auto- matic detection of 39 fundus diseases and conditions in retinal photographs using deep neural networks,

    L.-P. Cen, J. Ji, J.-W. Lin, S.-T. Ju, H.-J. Lin, T.-P. Li, Y. Wang, J.-F. Yang, Y.-F. Liu, S. Tanet al., “Auto- matic detection of 39 fundus diseases and conditions in retinal photographs using deep neural networks,”Nature communications, vol. 12, no. 1, p. 4828, 2021

  53. [61]

    Retinal fundus multi-disease image dataset (rfmid): A dataset for multi-disease detection research,

    S. Pachade, P. Porwal, D. Thulkar, M. Kokare, G. Deshmukh, V. Sahasrabuddhe, L. Giancardo, G. Quellec, and F. Meriaudeau, “Retinal fundus multi-disease image dataset (rfmid): A dataset for multi-disease detection research,” Data, vol. 6, no. 2, p. 14, 2021

  54. [62]

    Brset:abrazilian multilabel ophthalmological dataset of retina fundus photos,

    L.F.Nakayama,D.Restrepo,J.Matos,L.Z.Ribeiro,F.K.Malerbi,L.A.Celi,andC.S.Regatieri,“Brset:abrazilian multilabel ophthalmological dataset of retina fundus photos,”PLOS Digital Health, vol. 3, no. 7, p. e0000454, 2024

  55. [63]

    Accuracy assessment of intra- and intervisit fundus image registration for diabetic retinopathy screening,

    K. M. Adal, P. G. van Etten, J. P. Martinez, L. J. van Vliet, and K. A. Vermeer, “Accuracy assessment of intra- and intervisit fundus image registration for diabetic retinopathy screening,”Investigative ophthalmology and visual science, vol. 56, no. 3, pp. 1805–1812, 2015

  56. [64]

    Identifying medical diagnoses and treatable diseases by image-based deep learning,

    D. S. Kermany, M. Goldbaum, W. Cai, C. C. Valentim, H. Liang, S. L. Baxter, A. McKeown, G. Yang, X. Wu, F. Yanet al., “Identifying medical diagnoses and treatable diseases by image-based deep learning,”cell, vol. 172, no. 5, pp. 1122–1131, 2018

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.