Pith. sign in

REVIEW 5 major objections 5 minor 1 cited by

Predicting Diabetic Macular Edema Treatment Responses Using OCT: Dataset and Methods of APTOS Competition

T0 review · 5 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This paper introduces a public OCT dataset from 2,000 diabetic macular edema patients and shows that deep learning models can predict anti-VEGF treatment response, with the best model reaching 80.06% AUC for the decision to continue…

desk verdict The OCT4DME dataset is a genuinely useful public resource, but the headline 80.06% AUC is misattributed to the top team and mislabeled as AUC, so the paper needs a corrected results table before the number can be trusted. read the letter →

arxiv 2505.05768 v1 pith:F72XOXTN submitted 2025-05-09 eess.IV cs.AIcs.CV

classification eess.IVcs.AIcs.CV
keywords opticalcoherencetomographydiabeticmacularedemaanti-VEGFtherapytreatmentresponsepredictiondeeplearningmedicalimagingcompetitionpatientstratificationpublicdataset
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Diabetic macular edema responds to anti-VEGF injections unevenly, and clinicians lack a reliable way to predict who will benefit. This paper addresses that gap by releasing a public dataset of tens of thousands of OCT images from 2,000 DME patients, labelled for four prediction tasks: retinal fluid biomarkers, central subfield thickness, visual acuity, and the decision to continue injection. The authors organized a two-stage competition around these tasks and report that the best model reached an AUC of 80.06% for the Continue Injection decision, with higher AUCs for detecting individual fluid signs. The paper claims this is the first pre-treatment stratification study for DME treatment response. If the results generalize, a single baseline OCT scan could help personalize anti-VEGF therapy before the next injection is given.

What carries the argument

The load-bearing object is the OCT4DME dataset: paired pre- and post-treatment OCT scans from 2,000 DME patients, each eye labelled with four fluid biomarkers (intraretinal fluid, subretinal fluid, pigment epithelial detachment, hyperreflective foci), central subfield thickness, visual acuity, and the clinical Continue Injection decision. The competition pairs the dataset with a two-round evaluation in which participants predict the biomarkers as classification, CST and visual acuity as regression, and Continue Injection as the integrative clinical endpoint. The winning solutions treat biomarker classification as multiple-instance learning, in which each OCT B-scan is an instance and the eye is the bag, using a vision transformer or convolutional backbone, and then feed the predicted probabilities, thicknesses, and patient metadata into ensemble regressors for the Continue Injection decision.

What would settle it

Retrain the top model architecture on the OCT4DME training set only, hold out a fresh clinical cohort whose Continue Injection labels are kept secret until submission, and allow each team exactly one evaluation; if the AUC falls well below 80.06%, the reported performance reflects leaderboard probing rather than generalizable prediction.

Watch

Extended reading notes

Core claim

The central claim is that pre-treatment OCT images contain enough signal to predict whether a diabetic macular edema patient will need continued anti-VEGF injection, and that standard deep learning models can extract that signal when given enough labelled data. The evidence is the OCT4DME dataset, with pre- and post-treatment OCT scans from 2,000 patients, and the results of the competition held on it. On held-out test sets, the best model scored 80.06% AUC for Continue Injection, 97.96% AUC for pigment epithelial detachment, and over 93% AUC for subretinal and intraretinal fluid detection. The authors state this is the first exploration of pre-treatment stratification for DME treatment response and position the dataset as a public benchmark for that task.

Load-bearing premise

The load-bearing premise is that the leaderboard scores measure true predictive skill; because teams could submit many times per day and tune to the private test set, the reported 80.06% AUC may not transfer to new patients.

Editorial extensions

If this is right

  • A clinician could use a pre-treatment OCT scan to stratify DME patients before the second anti-VEGF injection, replacing trial-and-error treatment with an evidence-based continue-or-stop decision.
  • The public OCT4DME dataset gives the research community a common benchmark for DME treatment-response prediction, addressing the previous scarcity of large public OCT datasets in this area.
  • A single multi-task model can output fluid biomarkers, central subfield thickness, and visual acuity alongside the Continue Injection decision, providing intermediate clinical signals that could be checked against standard measurements.
  • The success of generic deep learning architectures in the competition suggests that data scale and labelling quality, rather than novel model design, were the main drivers of predictive performance.
  • If the 80.06% AUC for Continue Injection reproduces in a prospective setting, automated triage of DME patients for follow-up injection becomes feasible with a non-invasive imaging test.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An ablation that removes the OCT images and uses only patient metadata would quantify how much of the Continue Injection signal is actually carried by the scan; the paper's own searches for simple metadata rules found none, suggesting the image contribution is substantial.
  • The reported scores are specific to an Asian cohort collected on one OCT device; validation on other populations and devices is the natural next step and is flagged by the paper itself.
  • Because the competition allowed many daily submissions, the absolute AUC values may overstate real-world performance; a fixed split with a single locked submission per team would give a more trustworthy estimate.
  • The dataset's two annotation styles, per-eye labels in the first stage and per-image labels in the second, create a ready-made test bed for weakly supervised and label-efficient learning methods.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The manuscript reports on the 2021 APTOS Big Data Competition, which used a newly released public OCT dataset, OCT4DME, to predict diabetic macular edema (DME) treatment outcomes after anti-VEGF therapy. It describes the dataset collection and annotation (approximately 2,000 patients from Thailand, India, and China), the four competition sub-tasks (IRF/SRF/PED/HRF presence, CST regression, VA prediction, and continuation-of-injection), the evaluation metrics, and the top-three winning solutions. The headline claim is that the best model reached an AUC of 80.06% for the Continue Injection decision, and the authors state this is the first study to explore pre-treatment stratification for DME treatment response. The paper is largely descriptive: it is a competition report with dataset statistics, rules, and method summaries rather than a methods paper presenting a single validated model.

Significance. The dataset contribution is genuinely valuable: a large, publicly accessible, multi-task OCT benchmark for DME treatment response with labeled data, an official leaderboard, and detailed descriptions of the winning methods. The annotation quality-control procedure (two retina fellows with specialist audit) adds credibility. If the reported numbers are corrected, the paper would be a useful resource for benchmarking treatment-response prediction and for comparing deep-learning approaches on a clinical task that is underrepresented in public benchmarks. However, the central numeric claim is currently internally inconsistent, the metric labels are conflated, and the 'first' claim is overstated in light of the manuscript's own citations. No code or confidence intervals are provided, so the headline result cannot be independently verified from the manuscript.

major comments (5)
  1. [Abstract and Table 5] The headline 'top-performing team achieved an AUC of 80.06%' is not supported by the manuscript's own results. In Table 5, the value 0.8006 is DarkStyle's CIstage2 score, while DarkStyle is ranked third overall (0.7354); the first-place team BlueSky has CIstage2 = 0.7828. In addition, Section 6.3 describes the 80.06% as 'decision consistency' for Continued Injection, not as AUC. The abstract must be corrected to specify the exact team, task, and metric, or the number should be removed.
  2. [Section 6.3, text around Table 5] Two attributions in the results prose are reversed. The statement that BlueSky showed 'strength in CST during the first round, where they achieved a high score of 0.7026' is incorrect: 0.7026 is BlueSky's CIstage1 score, not a CST score (their CSTstage1 score is 0.5906). Similarly, the statement that DarkStyle 'predicted the second stage CST with a score of 0.8006' is incorrect: 0.8006 is DarkStyle's CIstage2 score, while their CSTstage2 score is 0.643. This misreading of Table 5 makes the section's ranking of team strengths unreliable and must be corrected.
  3. [Section 6.3 final paragraph and Section 4.4.1] The text reports 'CST prediction before treatment and after treatment' with AUCs of 68.71% and 69.30%, and 'VA prediction after treatment' with an AUC of 55.56%. However, Section 4.4.1 defines CST and VA scoring as tolerance-window accuracy (e.g., ±7.5% for CST), not AUC; AUC is specified only for CI, IRF, SRF, and HRF. The 68.71 and 69.30 values are BlueSky's preCSTstage2 and CSTstage2 scores in Table 5, which are tolerance-window scores. Reporting regression scores as AUC is a category error and should be fixed in both Section 6.3 and the abstract.
  4. [Abstract, Introduction, Section 8, and reference [46]] The claim 'this study is the first to explore pre-treatment stratification for predicting DME treatment responses' is contradicted by the manuscript's own citation of Feng et al. [46], a 2020 study predicting anti-VEGF injection effectiveness from OCT images using deep learning. Please qualify the claim to something like 'first large-scale public competition and dataset' for this task, or otherwise reconcile it with the cited prior work.
  5. [Section 4.2 and Section 7.4] The competition allowed up to 10 submissions per day in the preliminary round and 3 per day in the final round, so the final leaderboard scores may reflect repeated probing of the private test set. The manuscript provides no confidence intervals, no code, and no external validation cohort, yet Section 7.4 interprets the results as supporting viability for clinical decision-making. Please add uncertainty estimates and explicitly discuss this generalization limitation, or temper the clinical-application claim.
minor comments (5)
  1. [Section 7.2] The text refers to 'ResNet35' as the backbone used by LightRain, but Section 5.1.2 correctly identifies it as ResNet34; this is likely a typo.
  2. [Section 4.4.1] The task list includes 'preHED' and 'HED', but the annotation table (Table 2) and all other sections use PED; the HED spelling should be corrected.
  3. [Table 5] Table 5 reports a PEDstage2 column but no PEDstage1 column, even though PED is one of the first-stage classification tasks in Section 4.3.1; please add the missing column or explain why it is omitted.
  4. [Section 3.2 and Section 3.3] The abstract says 'tens of thousands of OCT images from 2,000 patients,' but the data description gives eye counts (2,366 training eyes in stage 1, 361 test eyes, etc.) and per-eye scan counts; please state the total number of images and clarify whether 'patients' means distinct individuals or eyes.
  5. [Data Availability] The data availability statement refers to an 'APTOS2021 Dataset' without a working URL; please provide a direct link and the exact access steps.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the paper is a competition report whose claimed 80.06% AUC is an external benchmark result, not a quantity constructed from its own inputs.

full rationale

The paper's central claims are empirical: it releases a dataset (OCT4DME) and reports results from a benchmark competition in which 41 finalist teams submitted models evaluated on held-out labels. There is no derivation chain in which a fitted input is renamed as a prediction. The CI score of 0.8006 in Table 5 is computed by comparing participant predictions with ophthalmologist labels, and Section 4.4 specifies AUC as the scoring metric for Continue Injection; this is an external evaluation, not a tautology. Self-citations to prior APTOS competitions ([34], [38]) are contextual and are not load-bearing for the dataset or the benchmark results. Two non-circular weaknesses should be flagged as correctness risks rather than circularity. First, the abstract attributes 'an AUC of 80.06%' to 'the top-performing team,' while Table 5 shows that 0.8006 is DarkStyle's CIstage2 score and DarkStyle is ranked third overall; Section 6.3 also calls this value 'decision consistency' rather than AUC, so the headline numeric claim is internally inconsistent and needs correction or clarification. Second, the abstract's 'first to explore pre-treatment stratification' claim conflicts with the paper's own citation of Feng et al. [46], a 2020 study predicting anti-VEGF effectiveness from OCT images; this is a novelty overclaim, not a circular argument. Repeated-submission leaderboard probing (Section 4.2) may inflate the reported test scores, but that is a validity limitation of the benchmark, not a reduction of the prediction to its inputs. Overall, the benchmark evaluation is self-contained against held-out data, so there is no significant circularity.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central benchmark rests on clinical assumptions about OCT biomarkers, on the reliability of the Continue Injection label, on patient-disjoint splits that are not explicitly documented, and on arbitrary scoring thresholds. None of these assumptions is independently validated in the paper.

free parameters (2)
  • CST tolerance beta = 0.075
    Used in the competition scoring rule (Section 4.4.1, Eq. 2b) to award full credit for CST predictions within ±7.5%; the threshold is arbitrary and directly determines leaderboard scores.
  • VA scoring window = ±0.05 logMAR or ±7.5% relative
    Section 4.4.1: VA predictions score 1 if within ±0.05 for VA < 1 or within ±7.5% for VA > 1; chosen by organizers, no justification given.
assumptions (4)
  • domain assumption OCT-based biomarkers (IRF, SRF, PED, HRF) and CST are valid and sufficient indicators of anti-VEGF treatment response.
    Defines the dataset labels and tasks (Section 3.4, Table 2); no comparison with functional outcome or external standard is provided.
  • domain assumption The Continue Injection label, generated from clinical records, is a reliable ground truth for treatment response.
    The paper shows in Section 7.3 that simple metadata rules do not predict CI, suggesting the label encodes unmeasured clinician judgment; no inter-rater reliability is reported.
  • domain assumption Training and test splits are patient-independent with no data leakage.
    Section 3.2 reports eye counts per stage but does not state whether any patient's eyes appear in both training and test sets; competition scores depend on this.
  • ad hoc to paper The arbitrary scoring thresholds define a meaningful measure of model quality.
    Section 4.4.1 sets ±7.5% CST tolerance and ±0.05/±7.5% VA windows without external anchoring; rankings could change under different thresholds.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Predicting Diabetic Macular Edema Treatment Responses Using OCT: Dataset and Methods of APTOS Competition." pith.science (2026). https://pith.science/paper/F72XOXTN

@misc{pith2026250505768,
  author       = {Pith},
  title        = {Pith review of: Predicting Diabetic Macular Edema Treatment Responses Using OCT: Dataset and Methods of APTOS Competition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F72XOXTN}},
  note         = {Machine review of arXiv:2505.05768}
}
read the original abstract

Diabetic macular edema (DME) significantly contributes to visual impairment in diabetic patients. Treatment responses to intravitreal therapies vary, highlighting the need for patient stratification to predict therapeutic benefits and enable personalized strategies. To our knowledge, this study is the first to explore pre-treatment stratification for predicting DME treatment responses. To advance this research, we organized the 2nd Asia-Pacific Tele-Ophthalmology Society (APTOS) Big Data Competition in 2021. The competition focused on improving predictive accuracy for anti-VEGF therapy responses using ophthalmic OCT images. We provided a dataset containing tens of thousands of OCT images from 2,000 patients with labels across four sub-tasks. This paper details the competition's structure, dataset, leading methods, and evaluation metrics. The competition attracted strong scientific community participation, with 170 teams initially registering and 41 reaching the final round. The top-performing team achieved an AUC of 80.06%, highlighting the potential of AI in personalized DME treatment and clinical decision-making.

Figures

Figures reproduced from arXiv: 2505.05768 by the authors.

Figure 1
Figure 1. An example from the dataset OCT4DME. coexisting ocular conditions that could affect visual acuity (e.g., advanced glau￾coma, diabetic retinopathy, or significant media opacity). OCT images with a resolution of 768 × 768 pixels were acquired using a 25-line raster scan protocol on the Heidelberg Spectralis OCT2 plus system (Heidelberg Engineering, Germany). In instances where raster scans were un￾available, 6-line ra… view at source ↗
Figure 2
Figure 2. The overall organization workflow of the 2021 Asia Pacific Tele-Ophthalmology [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Three sub-tasks of the competition, which are diagnosing the symptoms of the eye, [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Solution for Task 1 by the champion team BlueSky. This illustrates the architecture [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]
Figure 5
Figure 5. Figure 5: Solution for Task 1 by the runner-up team LightRain. They used a weakly supervised [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 6
Figure 6. Figure 6: The flow of data augmentation, proposed by the third-place team DarkStyle. [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]
Figure 7
Figure 7. Figure 7: Solution for Task 2 by the champion team BlueSky. They used a two-step training [PITH_FULL_IMAGE:figures/full_fig_p017_7.png]
Figure 8
Figure 8. Figure 8: The ensemble strategy for predicting visual acuity and treatment continuation [PITH_FULL_IMAGE:figures/full_fig_p019_8.png]
Figure 9
Figure 9. Figure 9: Solution for predicting continued injection by the third-place team DarkStyle. They [PITH_FULL_IMAGE:figures/full_fig_p020_9.png]
Figure 10
Figure 10. Figure 10: The distribution of data for IRF (Intraretinal Fluid), SRF (Subretinal Fluid), PED [PITH_FULL_IMAGE:figures/full_fig_p020_10.png]
Figure 11
Figure 11. Figure 11: Gama transform of the OCT images. Binary Classification Tasks related to IRF, SRF, PED, and HRF. Despite the simplicity, their approach ranked first in the preliminary round and second in the final. In the realm of medical image analysis, the design of task-specific m…
Figure 12
Figure 12. Figure 12: The correlation matrix for all metadata columns in the dataset. [PITH_FULL_IMAGE:figures/full_fig_p025_12.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. APTOS-2024 challenge report: Generation of synthetic 3D OCT images from fundus photographs

    cs.CV 2025-06 reject novelty 5.0 of 10

    The APTOS-2024 challenge benchmark shows that current generative models can produce 3D OCT volumes from fundus photos, but the primary metric does not beat a simple random-crop baseline.

Reference graph

Works this paper leans on

71 extracted references · 61 canonical work pages · cited by 1 Pith paper

  1. [46]

    D. Feng, X. Chen, Z. Zhou, H. Liu, Y. Wang, L. Bai, S. Zhang, X. Mou, A preliminary study of predicting effectiveness of anti-vegf injection using oct images based on deep learning, in: 2020 42nd annual international conference of the IEEE engineering in medicine & biology society (EMBC), IEEE, 2020, pp. 5428–5431

  2. [1]

    Hashemi, F

    H. Hashemi, F. Rezvan, R. Pakzad, A. Ansaripour, S. Heydarian, A. Yekta, H. Ostadimoghaddam, M. Pakbin, M. Khabazkhoob, Global and regional prevalence of diabetic retinopathy; a comprehensive systematic review and meta-analysis, Semin Ophthalmol 37 (3) (2022) 291–306. doi:10.1080/ 08820538.2021.1962920

  3. [2]

    Daruich, A

    A. Daruich, A. Matet, A. Moulin, L. Kowalczuk, M. Nicolas, A. Sellam, P. R. Rothschild, S. Omri, E. G´ eliz´ e, L. Jonet, K. Delaunay, Y. De Kozak, M. Berdugo, M. Zhao, P. Crisanti, F. Behar-Cohen, Mechanisms of macular edema: Beyond the surface, Prog Retin Eye Res 63 (2018) 20–68. doi: 10.1016/j.preteyeres.2017.10.006

  4. [3]

    S. D. Solomon, E. Chew, E. J. Duh, L. Sobrin, J. K. Sun, B. L. VanderBeek, C. C. Wykoff, T. W. Gardner, Diabetic retinopathy: A position statement by the american diabetes association, Diabetes Care 40 (3) (2017) 412–418. doi:10.2337/dc16-2641

  5. [4]

    Zhang, J

    J. Zhang, J. Zhang, C. Zhang, J. Zhang, L. Gu, D. Luo, Q. Qiu, Diabetic macular edema: Current understanding, molecular mechanisms and thera- peutic implications, Cells 11 (21) (2022). doi:10.3390/cells11213362

  6. [5]

    N. M. Bressler, W. T. Beaulieu, A. R. Glassman, K. J. Blinder, S. B. Bressler, L. M. Jampol, M. Melia, r. Wells, J. A., Persistent macular thick- ening following intravitreous aflibercept, bevacizumab, or ranibizumab for central-involved diabetic macular edema with vision impairment: A sec- ondary analysis of a randomized clinical trial, JAMA Ophthalmol 1...

  7. [6]

    Olson, P

    J. Olson, P. Sharp, K. Goatman, G. Prescott, G. Scotland, A. Fleming, S. Philip, C. Santiago, S. Borooah, D. Broadbent, V. Chong, P. Dodson, S. Harding, G. Leese, C. Styles, K. Swa, H. Wharton, Improving the eco- nomic value of photographic screening for optical coherence tomography- 29 detectable macular oedema: a prospective, multicentre, uk study, Heal...

  8. [7]

    N. K. Waheed, J. S. Duker, Oct in the management of diabetic macular edema, Current Ophthalmology Reports 1 (3) (2013) 128–133. doi:10. 1007/s40135-013-0019-z . URL https://doi.org/10.1007/s40135-013-0019-z

Show all 71 references
  1. [8]

    Koozekanani, C

    D. Koozekanani, C. Roberts, S. E. Katz, E. E. Herderick, Intersession re- peatability of macular thickness measurements with the humphrey 2000 oct, Invest Ophthalmol Vis Sci 41 (6) (2000) 1486–91

  2. [9]

    D. A. Salz, A. J. Witkin, Imaging in diabetic retinopathy, Middle East Afr J Ophthalmol 22 (2) (2015) 145–50. doi:10.4103/0974-9233.151887

  3. [10]

    D. Shi, W. Zhang, S. He, Y. Chen, F. Song, S. Liu, R. Wang, Y. Zheng, M. He, Translation of color fundus photography into fluorescein angiogra- phy using deep learning for enhanced diabetic retinopathy screening, Oph- thalmology science 3 (4) (2023) 100401

  4. [11]

    X. Chen, W. Zhang, P. Xu, Z. Zhao, Y. Zheng, D. Shi, M. He, Ffa-gpt: an automated pipeline for fundus fluorescein angiography interpretation and question-answer, NPJ digital medicine 7 (1) (2024) 111

  5. [12]

    R. Chen, W. Zhang, F. Song, H. Yu, D. Cao, Y. Zheng, M. He, D. Shi, Translating color fundus photography to indocyanine green angiography using deep-learning for age-related macular degeneration screening, NPJ digital medicine 7 (1) (2024) 34

  6. [13]

    J. Sun, D. Wei, L. Wang, Y. Zheng, Lesion guided explainable few weak- shot medical report generation, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, 2022

  7. [14]

    F. Tang, X. Wang, A. R. Ran, C. K. M. Chan, M. Ho, W. Yip, A. L. Young, J. Lok, S. Szeto, J. Chan, F. Yip, R. Wong, Z. Tang, D. Yang, D. S. Ng, 30 L. J. Chen, M. Brel´ en, V. Chu, K. Li, T. H. T. Lai, G. S. Tan, D. S. W. Ting, H. Huang, H. Chen, J. H. Ma, S. Tang, T. Leng, S. ...

  8. [15]

    Midena, L

    E. Midena, L. Toto, L. Frizziero, G. Covello, T. Torresin, G. Midena, L. Danieli, E. Pilotto, M. Figus, C. Mariotti, M. Lupidi, Validation of an automated artificial intelligence algorithm for the quantification of ma- jor oct parameters in diabetic macular edema, J Clin Med 1...

  9. [16]

    Tripathi, P

    A. Tripathi, P. Kumar, V. Mayya, A. Tulsani, Generating oct b-scan dme images using optimized generative adversarial networks (gans), Heliyon 9 (8) (2023) e18773. doi:10.1016/j.heliyon.2023.e18773

  10. [17]

    F. Y. Tang, D. S. Ng, A. Lam, F. Luk, R. Wong, C. Chan, S. Mohamed, A. Fong, J. Lok, T. Tso, et al., Determinants of quantitative optical coher- ence tomography angiography metrics in patients with diabetes, Scientific reports 7 (1) (2017) 2575

  11. [18]

    P. P. Srinivasan, L. A. Kim, P. S. Mettu, S. W. Cousins, G. M. Comer, J. A. Izatt, S. Farsiu, Fully automated detection of diabetic macular edema and dry age-related macular degeneration from optical coherence tomography images, Biomedical optics express 5 (10) (2014) 3568–3577

  12. [19]

    Niemeijer, B

    M. Niemeijer, B. Van Ginneken, M. J. Cree, A. Mizutani, G. Quellec, C. I. S´ anchez, B. Zhang, R. Hornero, M. Lamard, C. Muramatsu, et al., Retinopathy online challenge: automatic detection of microaneurysms in digital color fundus photographs, IEEE transactions on medical ima...

  13. [20]

    Dugas, Jared, Jorge, W

    E. Dugas, Jared, Jorge, W. Cukierski, Diabetic retinopa- thy detection, https://kaggle.com/competitions/ diabetic-retinopathy-detection, kaggle (2015)

  14. [21]

    S. G. Kobat, N. Baygin, E. Yusufoglu, M. Baygin, P. D. Barua, S. Dogan, O. Yaman, U. Celiker, H. Yildirim, R.-S. Tan, et al., Automated diabetic retinopathy detection using horizontal and vertical patch division-based pre-trained densenet with digital fundus images, Diagnostic...

  15. [22]

    Porwal, S

    P. Porwal, S. Pachade, M. Kokare, G. Deshmukh, J. Son, W. Bae, L. Liu, J. Wang, X. Liu, L. Gao, et al., Idrid: Diabetic retinopathy–segmentation and grading challenge, Medical image analysis 59 (2020) 101561

  16. [23]

    R. Liu, X. Wang, Q. Wu, L. Dai, X. Fang, T. Yan, J. Son, S. Tang, J. Li, Z. Gao, et al., Deepdrid: Diabetic retinopathy—grading and image quality estimation challenge, Patterns 3 (6) (2022)

  17. [24]

    B. Qian, H. Chen, X. Wang, Z. Guan, T. Li, Y. Jin, Y. Wu, Y. Wen, H. Che, G. Kwon, et al., Drac 2022: A public benchmark for diabetic retinopathy analysis on ultra-wide optical coherence tomography angiography images, Patterns 5 (3) (2024)

  18. [25]

    J. I. Orlando, H. Fu, J. B. Breda, K. Van Keer, D. R. Bathula, A. Diaz- Pinto, R. Fang, P.-A. Heng, J. Kim, J. Lee, et al., Refuge challenge: A unified framework for evaluating automated methods for glaucoma assess- ment from fundus photographs, Medical image analysis 59 (2020) 101570

  19. [26]

    H. Fang, F. Li, J. Wu, H. Fu, X. Sun, J. Son, S. Yu, M. Zhang, C. Yuan, C. Bian, et al., Refuge2 challenge: A treasure trove for multi- dimension analysis and evaluation in glaucoma screening, arXiv preprint arXiv:2202.08994 (2022)

  20. [27]

    De Vente, K

    C. De Vente, K. A. Vermeer, N. Jaccard, H. Wang, H. Sun, F. Khader, D. Truhn, T. Aimyshev, Y. Zhanibekuly, T.-D. Le, et al., Airogs: Artificial 32 intelligence for robust glaucoma screening challenge, IEEE transactions on medical imaging 43 (1) (2023) 542–557

  21. [28]

    H. G. Lemij, C. de Vente, C. I. S´ anchez, K. A. Vermeer, Characteristics of a large, labeled data set for the training of artificial intelligence for glaucoma screening with fundus photographs, Ophthalmology Science 3 (3) (2023) 100300

  22. [29]

    Peking university international competition on ocular disease intelligent recognition (odir-2019), URL:https://odir2019.grand-challenge.org/ (2019)

  23. [30]

    Pachade, P

    S. Pachade, P. Porwal, M. Kokare, G. Deshmukh, V. Sahasrabuddhe, Z. Luo, F. Han, Z. Sun, L. Qihan, S.-i. Kamata, et al., Rfmid: Retinal image analysis for multi-disease detection challenge, Medical Image Anal- ysis 99 (2025) 103365

  24. [31]

    H. Fang, F. Li, J. Wu, H. Fu, X. Sun, J. I. Orlando, H. Bogunovi´ c, X. Zhang, Y. Xu, Open fundus photograph dataset with pathologic myopia recogni- tion and anatomical structure annotation, Scientific Data 11 (1) (2024) 99

  25. [32]

    H. Fang, F. Li, H. Fu, X. Sun, X. Cao, F. Lin, J. Son, S. Kim, G. Quellec, S. Matta, et al., Adam challenge: Detecting age-related macular degener- ation from fundus images, IEEE transactions on medical imaging 41 (10) (2022) 2828–2847

  26. [33]

    H. Fu, F. Li, X. Sun, X. Cao, J. Liao, J. I. Orlando, X. Tao, Y. Li, S. Zhang, M. Tan, et al., Age challenge: angle closure glaucoma evaluation in anterior segment optical coherence tomography, Medical Image Analysis 66 (2020) 101798

  27. [34]

    Zhang, P

    W. Zhang, P. Chotcomwongse, X. Chen, F. H. Chung, F. Song, X. Zhang, M. He, D. Shi, P. Ruamviboonsuk, Angiographic report generation for the 33 3rd aptos’s competition: Dataset and baseline methods, medRxiv (2023) 2023–11

  28. [35]

    Bogunovi´ c, F

    H. Bogunovi´ c, F. Venhuizen, S. Klimscha, S. Apostolopoulos, A. Bab- Hadiashar, U. Bagci, M. F. Beg, L. Bekalo, Q. Chen, C. Ciller, et al., Retouch: The retinal oct fluid detection and segmentation benchmark and challenge, IEEE transactions on medical imaging 38 (8) (2019) 1858–1874

  29. [36]

    Rasti, H

    R. Rasti, H. Rabbani, Retinal oct classification challenge (rocc), URL:https://rocc.grand-challenge.org/ (2017)

  30. [37]

    J. Wu, H. Fang, F. Li, H. Fu, F. Lin, J. Li, Y. Huang, Q. Yu, S. Song, X. Xu, et al., Gamma challenge: glaucoma grading from multi-modality images, Medical Image Analysis 90 (2023) 102938

  31. [38]

    org/big-data-competition/ (2024)

    2024 aptos big data competition, URL: https://2024.asiateleophth. org/big-data-competition/ (2024)

  32. [39]

    M. H. Akpinar, A. Sengur, O. Faust, L. Tong, F. Molinari, U. R. Acharya, Artificial intelligence in retinal screening using oct images: A review of the last decade (2013–2023), Computer Methods and Programs in Biomedicine 254 (2024) 108253. doi:https://doi.org/10.1016/j.cmpb. ...

  33. [40]

    Sonobe, H

    T. Sonobe, H. Tabuchi, H. Ohsugi, H. Masumoto, N. Ishitobi, S. Morita, H. Enno, D. Nagasato, Comparison between support vector machine and deep learning, machine-learning technologies for detecting epiretinal mem- brane using 3d-oct, International Ophthalmology 39 (8) (2019) 1...

  34. [41]

    Mihalache, R

    A. Mihalache, R. S. Huang, M. M. Popovic, N. S. Patil, B. U. Pandya, R. Shor, A. Pereira, J. M. Kwok, P. Yan, D. T. Wong, P. J. Kertes, R. H. Muni, Accuracy of an artificial intelligence chatbot’s interpretation of clin- ical ophthalmic images, JAMA Ophthalmology 142 (4) (2024...

  35. [42]

    M. A. Hussain, A. Bhuiyan, C. D. Luu, R. Theodore Smith, R. H. Guymer, H. Ishikawa, J. S. Schuman, K. Ramamohanarao, Classification of healthy and diseased retina using sd-oct imaging and random forest algorithm, PLOS ONE 13 (6) (2018) 1–17. doi:10.1371/journal.pone.0198281

  36. [43]

    Chen, M.-L

    H.-Y. Chen, M.-L. Huang, P.-T. Hung, Logistic regression analysis for glaucoma diagnosis using stratus optical coherence tomography, Optometry and Vision Science 83 (7) (2006) 00023. doi:10.1097/ 00006324-200607000-00023

  37. [44]

    Seeb¨ ock, S

    P. Seeb¨ ock, S. M. Waldstein, S. Klimscha, H. Bogunovic, T. Schlegl, B. S. Gerendas, R. Donner, U. Schmidt-Erfurth, G. Langs, Unsupervised iden- tification of disease marker candidates in retinal oct imaging data, IEEE transactions on medical imaging 38 (4) (2018) 1037–1047

  38. [45]

    Awais, H

    M. Awais, H. M¨ uller, T. B. Tang, F. Meriaudeau, Classification of sd- oct images using a deep learning approach, in: 2017 IEEE International Conference on Signal and Image Processing Applications (ICSIPA), IEEE, 2017, pp. 489–492

  39. [47]

    J. Wang, G. Deng, W. Li, Y. Chen, F. Gao, H. Liu, Y. He, G. Shi, Deep learning for quality assessment of retinal oct images, Biomedical optics express 10 (12) (2019) 6057–6072

  40. [48]

    R. M. Kamble, G. C. Chan, O. Perdomo, M. Kokare, F. A. Gonzalez, H. M¨ uller, F. M´ eriaudeau, Automated diabetic macular edema (dme) analysis using fine tuning with inception-resnet-v2 on oct images, in: 2018 IEEE-EMBS Conference on Biomedical Engineering and Sciences (IECBES...

  41. [49]

    Y. Shen, J. Li, W. Zhu, K. Yu, M. Wang, Y. Peng, Y. Zhou, L. Guan, X. Chen, Graph attention u-net for retinal layer surface detection and choroid neovascularization segmentation in oct images, IEEE Transactions on Medical Imaging 42 (11) (2023) 3140–3154

  42. [50]

    Ndipenoch, A

    N. Ndipenoch, A. Miron, Y. Li, Performance evaluation of retinal oct fluid segmentation, detection, and generalization over variations of data sources, IEEE Access 12 (2024) 31719–31735

  43. [51]

    D. Shi, W. Zhang, X. Chen, Y. Liu, J. Yang, S. Huang, Y. C. Tham, Y. Zheng, M. He, Eyefound: a multimodal generalist foundation model for ophthalmic imaging, arXiv preprint arXiv:2405.11338 (2024)

  44. [52]

    Zhang, S

    W. Zhang, S. Huang, J. Yang, R. Chen, Z. Ge, Y. Zheng, D. Shi, M. He, Fundus2video: Cross-modal angiography video generation from static fun- dus photography with clinical knowledge guidance, in: International Con- ference on Medical Image Computing and Computer-Assisted Inter...

  45. [53]

    D. Shi, W. Zhang, J. Yang, S. Huang, X. Chen, M. Yusufu, K. Jin, S. Lin, S. Liu, Q. Zhang, et al., Eyeclip: A visual-language foundation model for multi-modal ophthalmic image analysis, arXiv preprint arXiv:2409.06644 (2024)

  46. [54]

    P. Xu, X. Chen, Z. Zhao, D. Shi, Unveiling the clinical incapabilities: a benchmarking study of gpt-4v (ision) for ophthalmic multimodal image analysis, British Journal of Ophthalmology 108 (10) (2024) 1384–1389

  47. [55]

    Cappellani, K

    F. Cappellani, K. R. Card, C. L. Shields, J. S. Pulido, J. A. Haller, Reliabil- ity and accuracy of artificial intelligence chatgpt in providing information on ophthalmic diseases and management to patients, Eye 38 (7) (2024) 1368–1373

  48. [56]

    Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, B. Guo, Swin trans- former: Hierarchical vision transformer using shifted windows, in: Proceed- 36 ings of the IEEE/CVF international conference on computer vision, 2021, pp. 10012–10022

  49. [57]

    Carbonneau, V

    M.-A. Carbonneau, V. Cheplygina, E. Granger, G. Gagnon, Multiple in- stance learning: A survey of problem characteristics and applications, Pat- tern Recognition 77 (2018) 329–353

  50. [58]

    K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recog- nition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778

  51. [59]

    G. K. Wallace, The jpeg still picture compression standard, IEEE transac- tions on consumer electronics 38 (1) (1992) xviii–xxxiv

  52. [60]

    Sultani, C

    W. Sultani, C. Chen, M. Shah, Real-world anomaly detection in surveil- lance videos, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 6479–6488

  53. [61]

    Shanmugam, D

    D. Shanmugam, D. Blalock, G. Balakrishnan, J. Guttag, Better aggregation in test-time augmentation, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 1214–1223

  54. [62]

    M. Tan, Q. Le, Efficientnet: Rethinking model scaling for convolutional neural networks, in: International conference on machine learning, PMLR, 2019, pp. 6105–6114

  55. [63]

    Sandler, A

    M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, L.-C. Chen, Mobilenetv2: Inverted residuals and linear bottlenecks, in: Proceedings of the IEEE con- ference on computer vision and pattern recognition, 2018, pp. 4510–4520

  56. [64]

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large- scale hierarchical image database, in: 2009 IEEE conference on computer vision and pattern recognition, Ieee, 2009, pp. 248–255

  57. [65]

    Huang, Z

    G. Huang, Z. Liu, L. Van Der Maaten, K. Q. Weinberger, Densely con- nected convolutional networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 4700–4708. 37

  58. [66]

    J. H. Friedman, Greedy function approximation: a gradient boosting ma- chine, Annals of statistics (2001) 1189–1232

  59. [67]

    M. R. Segal, Machine learning benchmarks and random forest regression (2004)

  60. [68]

    Panozzo, E

    G. Panozzo, E. Gusson, B. Parolini, A. Mercanti, Role of oct in the diagno- sis and follow up of diabetic macular edema, in: Seminars in ophthalmology, Vol. 18, Taylor & Francis, 2003, pp. 74–81

  61. [69]

    X. Wang, F. Tang, H. Chen, C. Y. Cheung, P.-A. Heng, Deep semi- supervised multiple instance learning with self-correction for dme classi- fication from oct images, Medical Image Analysis 83 (2023) 102673

  62. [70]

    Padmasini, R

    N. Padmasini, R. Umamaheswari, M. Y. Sikkandar, M. D. Sindal, Com- parison of deep CNNs in the identification of DME structural changes in retinal OCT scans, Elsevier, 2023, pp. 35–51

  63. [71]

    Z. Qi, S. Khorram, F. Li, Visualizing deep networks by optimizing with integrated gradients, in: CVPR Workshops, Vol. 2, 2019, pp. 1–4. 38 Appendix A. APTOS 2021 T eam Introduction BlueSky: 1. Ruijie Yao (ruijie.yao@duke.edu) is a first-year Ph.D. stu- dent in the Mechanical E...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.