{"id":"9f0c24f3-4b3d-4efb-824a-ca76e47077ad","arxiv_id":"2505.05768","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new public benchmark dataset, OCT4DME, plus competition results for predicting DME treatment response from OCT images, with top scores around 80% AUC on the continuation-decision subtask.","lead":"This paper describes a public dataset of OCT eye scans from 2,000 patients with diabetic macular edema and a 2021 competition that used it to predict responses to anti-VEGF injections. The best reported result, an 80.06% AUC for one subtask, is hard to verify because the paper's tables and text disagree and no code or model weights are released.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 'AUC of 80.06%' does not match the paper's own Table 5 attribution, so the headline performance claim is internally inconsistent and unverifiable without correction.","rationale":"The reader's verdict is CONDITIONAL, and I agree that the paper as written cannot be verified internally. However, the reader's weakest assumption focused on leaderboard probing (repeated submissions up to 10/day preliminary, 3/day final) as the main reason the 80.06% AUC might not reflect real-world performance. That is a valid concern about generalization, but it is not the most load-bearing issue for the paper's central claim. The more immediate, checkable problem is that the abstract's headline number is not traceable to any specific result in the paper's own results section. The abstract says 'the top-performing team achieved an AUC of 80.06%,' but Table 5 shows 0.8006 for DarkStyle's CIstage2, and DarkStyle is ranked third overall. Section 6.3 describes the CI result as 'decision consistency' and says 'AI aligning with ophthalmologists in 80.06% of cases,' which is an accuracy/agreement framing, not necessarily an AUC. Section 4.4.1 says AUC is used for CI scoring, but the final round score formula (average of 14 indices) and the reported numbers do not make clear whether the 0.8006 is AUC, balanced accuracy, or something else. This is a smaller, more definite inconsistency than the leaderboard-probing concern, and it undermines the central numeric claim as stated. The 'first' claim is also weakened by the paper's own citation [46], a 2020 study predicting anti-VEGF effectiveness from OCT. The dataset itself (OCT4DME) is a genuine contribution, and the paper provides useful details on the competition, so I do not think REJECT is warranted. The correct verdict remains CONDITIONAL, with the condition being that the authors correct or clarify the abstract's 80.06% claim and specify the exact metric for each reported score.","tokens_in":19229,"tokens_out":2144,"duration_ms":17510,"concrete_test":"Independently cross-tabulate every number in the abstract, Section 6.3, and Tables 4-5. Specifically, verify whether any team's score in Table 5 is an AUC as defined in Section 4.4.1, and whether the abstract's 'AUC of 80.06%' matches the CIstage2 column of DarkStyle (0.8006). If the CIstage2 scores are not AUC, the abstract and Section 6.3 must specify the metric (e.g., decision consistency) or the table must be recomputed with AUC. Also check the exact definition of the CI evaluation against the competition's official scoring formula to determine whether AUC or accuracy/consistency was used.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The abstract's headline result, 'the top-performing team achieved an AUC of 80.06%,' does not match any entry in Table 5 or the surrounding text. Table 5 shows DarkStyle's second-stage CI (Continue Injection) score as 0.8006, but the abstract describes this as an AUC achieved by 'the top-performing team' for treatment response. In fact, DarkStyle placed third overall, and the CIstage2 column shows 0.8006 for DarkStyle while BlueSky (rank 1) has 0.7828 and LightRain (rank 2) has 0.7800. The paper also refers to the CI metric as 'decision consistency' in Section 6.3 and 'AUC' in Section 4.4 and the abstract, without clarifying which metric was actually used to produce the 0.8006/80.06% number. Because the paper's own Table 5 contradicts the abstract's attribution, the 80.06% AUC figure is not internally consistent and cannot be verified from the manuscript. Additionally, the abstract calls this 'the first to explore pre-treatment stratification for predicting DME treatment responses,' yet the paper cites Feng et al. [46], a 2020 study predicting anti-VEGF injection effectiveness from OCT images, which undercuts the 'first' claim as stated. The dataset contribution (OCT4DME) is real and potentially valuable, but the central numeric performance claim needs correction or clarification.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports on the 2021 APTOS Big Data Competition, which used a newly released public OCT dataset, OCT4DME, to predict diabetic macular edema (DME) treatment outcomes after anti-VEGF therapy. It describes the dataset collection and annotation (approximately 2,000 patients from Thailand, India, and China), the four competition sub-tasks (IRF/SRF/PED/HRF presence, CST regression, VA prediction, and continuation-of-injection), the evaluation metrics, and the top-three winning solutions. The headline claim is that the best model reached an AUC of 80.06% for the Continue Injection decision, and the authors state this is the first study to explore pre-treatment stratification for DME treatment response. The paper is largely descriptive: it is a competition report with dataset statistics, rules, and method summaries rather than a methods paper presenting a single validated model.","tokens_in":19447,"tokens_out":7695,"duration_ms":72043,"significance":"The dataset contribution is genuinely valuable: a large, publicly accessible, multi-task OCT benchmark for DME treatment response with labeled data, an official leaderboard, and detailed descriptions of the winning methods. The annotation quality-control procedure (two retina fellows with specialist audit) adds credibility. If the reported numbers are corrected, the paper would be a useful resource for benchmarking treatment-response prediction and for comparing deep-learning approaches on a clinical task that is underrepresented in public benchmarks. However, the central numeric claim is currently internally inconsistent, the metric labels are conflated, and the 'first' claim is overstated in light of the manuscript's own citations. No code or confidence intervals are provided, so the headline result cannot be independently verified from the manuscript.","major_comments":[{"comment":"The headline 'top-performing team achieved an AUC of 80.06%' is not supported by the manuscript's own results. In Table 5, the value 0.8006 is DarkStyle's CIstage2 score, while DarkStyle is ranked third overall (0.7354); the first-place team BlueSky has CIstage2 = 0.7828. In addition, Section 6.3 describes the 80.06% as 'decision consistency' for Continued Injection, not as AUC. The abstract must be corrected to specify the exact team, task, and metric, or the number should be removed.","section":"Abstract and Table 5"},{"comment":"Two attributions in the results prose are reversed. The statement that BlueSky showed 'strength in CST during the first round, where they achieved a high score of 0.7026' is incorrect: 0.7026 is BlueSky's CIstage1 score, not a CST score (their CSTstage1 score is 0.5906). Similarly, the statement that DarkStyle 'predicted the second stage CST with a score of 0.8006' is incorrect: 0.8006 is DarkStyle's CIstage2 score, while their CSTstage2 score is 0.643. This misreading of Table 5 makes the section's ranking of team strengths unreliable and must be corrected.","section":"Section 6.3, text around Table 5"},{"comment":"The text reports 'CST prediction before treatment and after treatment' with AUCs of 68.71% and 69.30%, and 'VA prediction after treatment' with an AUC of 55.56%. However, Section 4.4.1 defines CST and VA scoring as tolerance-window accuracy (e.g., ±7.5% for CST), not AUC; AUC is specified only for CI, IRF, SRF, and HRF. The 68.71 and 69.30 values are BlueSky's preCSTstage2 and CSTstage2 scores in Table 5, which are tolerance-window scores. Reporting regression scores as AUC is a category error and should be fixed in both Section 6.3 and the abstract.","section":"Section 6.3 final paragraph and Section 4.4.1"},{"comment":"The claim 'this study is the first to explore pre-treatment stratification for predicting DME treatment responses' is contradicted by the manuscript's own citation of Feng et al. [46], a 2020 study predicting anti-VEGF injection effectiveness from OCT images using deep learning. Please qualify the claim to something like 'first large-scale public competition and dataset' for this task, or otherwise reconcile it with the cited prior work.","section":"Abstract, Introduction, Section 8, and reference [46]"},{"comment":"The competition allowed up to 10 submissions per day in the preliminary round and 3 per day in the final round, so the final leaderboard scores may reflect repeated probing of the private test set. The manuscript provides no confidence intervals, no code, and no external validation cohort, yet Section 7.4 interprets the results as supporting viability for clinical decision-making. Please add uncertainty estimates and explicitly discuss this generalization limitation, or temper the clinical-application claim.","section":"Section 4.2 and Section 7.4"}],"minor_comments":[{"comment":"The text refers to 'ResNet35' as the backbone used by LightRain, but Section 5.1.2 correctly identifies it as ResNet34; this is likely a typo.","section":"Section 7.2"},{"comment":"The task list includes 'preHED' and 'HED', but the annotation table (Table 2) and all other sections use PED; the HED spelling should be corrected.","section":"Section 4.4.1"},{"comment":"Table 5 reports a PEDstage2 column but no PEDstage1 column, even though PED is one of the first-stage classification tasks in Section 4.3.1; please add the missing column or explain why it is omitted.","section":"Table 5"},{"comment":"The abstract says 'tens of thousands of OCT images from 2,000 patients,' but the data description gives eye counts (2,366 training eyes in stage 1, 361 test eyes, etc.) and per-eye scan counts; please state the total number of images and clarify whether 'patients' means distinct individuals or eyes.","section":"Section 3.2 and Section 3.3"},{"comment":"The data availability statement refers to an 'APTOS2021 Dataset' without a working URL; please provide a direct link and the exact access steps.","section":"Data Availability"}],"recommendation":"major_revision","confidential_remarks":"The dataset and competition infrastructure are real contributions, and the flaws identified are correctable in revision rather than fatal. The biggest concern is that the abstract's headline number will be quoted independently of the manuscript's own tables; I would ask the editor to ensure the revised version makes the metric and team attribution explicit in the abstract. There is also a potential conflict-of-interest disclosure issue: winning-team members are co-authors on the paper, and the competing-interest statement is generic; the editor may wish to have the authors add a specific disclosure about competition participation and authorship."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — the thing worth knowing: the real contribution here is the OCT4DME dataset, and that part is solid. The advertised result, however, is not. The abstract says the top-performing team achieved an AUC of 80.06%, but Table 5 shows 0.8006 as DarkStyle's second-stage Continue Injection score, and DarkStyle placed third overall. Section 6.3 also calls the 80.06% a \"decision consistency\" figure, not an AUC. So the headline number is misattributed and mislabeled. That needs fixing before anyone quotes it.\n\nWhat is actually new: a 2,000-patient OCT dataset with labels across four subtasks — fluid signs, CST, VA, and continue-injection decision — for pre-treatment DME response prediction. Prior public DME OCT data is tiny; the gap is real. The competition setup is described in careful detail, including the timeline, the evaluation protocol, and the top three teams' methods. The methods themselves are standard — Swin Transformer with MIL, ResNet, EfficientNet, ensembling — which is fine for a benchmark report, and the paper does not oversell them as novel. Credit is due for the dataset and the multi-task benchmark itself, both of which are reproducible contributions.\n\nSoft spots, in proportion. First, the performance reporting is internally inconsistent: besides the 80.06% problem, Section 6.3 reports CST and VA predictions as AUC even though Section 4.4 defines those as tolerance-window scores. The metric labels need to be corrected table by table. Second, the abstract's \"first to explore pre-treatment stratification\" claim is undercut by the paper's own reference [46], Feng et al. 2020, which predicts anti-VEGF effectiveness from OCT. That claim should be toned down. Third, the leaderboard design allowed up to 10 submissions per day in the preliminary round and 3 per day in the final round, so test-set scores may reflect leaderboard probing rather than pure generalization. The top-5 code defense is a partial mitigation, but there are no confidence intervals and no evaluation code, so the reported numbers are not independently verifiable from the manuscript. The dataset's value does not depend on those numbers, though.\n\nWho this is for: anyone building or benchmarking models for DME treatment response. The paper deserves a serious referee because the dataset is valuable enough that the reporting problems should be fixed, not the paper rejected. I would bring it to reading group mainly to discuss what a competition report must disclose about test-set probing and metric definitions.\n\nRecommendation: send it to peer review, but only after the organizers correct the abstract/Table 5 mismatch and label every metric exactly. If that correction does not happen, treat the headline as unverified and cite only the dataset.","headline":"The OCT4DME dataset is a genuinely useful public resource, but the headline 80.06% AUC is misattributed to the top team and mislabeled as AUC, so the paper needs a corrected results table before the number can be trusted.","tokens_in":20105,"tokens_out":2472,"would_cite":true,"duration_ms":24693,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces a public OCT dataset from 2,000 diabetic macular edema patients and shows that deep learning models can predict anti-VEGF treatment response, with the best model reaching 80.06% AUC for the decision to continue…","keywords":["optical coherence tomography","diabetic macular edema","anti-VEGF therapy","treatment response prediction","deep learning","medical imaging competition","patient stratification","public dataset"],"falsifier":"Retrain the top model architecture on the OCT4DME training set only, hold out a fresh clinical cohort whose Continue Injection labels are kept secret until submission, and allow each team exactly one evaluation; if the AUC falls well below 80.06%, the reported performance reflects leaderboard probing rather than generalizable prediction.","tokens_in":18985,"feed_emoji":"👁️","tokens_out":8023,"duration_ms":79909,"temperature":0.7,"pith_summary":"Diabetic macular edema responds to anti-VEGF injections unevenly, and clinicians lack a reliable way to predict who will benefit. This paper addresses that gap by releasing a public dataset of tens of thousands of OCT images from 2,000 DME patients, labelled for four prediction tasks: retinal fluid biomarkers, central subfield thickness, visual acuity, and the decision to continue injection. The authors organized a two-stage competition around these tasks and report that the best model reached an AUC of 80.06% for the Continue Injection decision, with higher AUCs for detecting individual fluid signs. The paper claims this is the first pre-treatment stratification study for DME treatment response. If the results generalize, a single baseline OCT scan could help personalize anti-VEGF therapy before the next injection is given.","feed_headline":"Eye scans predict DME treatment response at 80% AUC","feed_subtitle":"A 2,000-patient public OCT dataset lets models pick who should continue therapy after the first injection.","key_machinery":"The load-bearing object is the OCT4DME dataset: paired pre- and post-treatment OCT scans from 2,000 DME patients, each eye labelled with four fluid biomarkers (intraretinal fluid, subretinal fluid, pigment epithelial detachment, hyperreflective foci), central subfield thickness, visual acuity, and the clinical Continue Injection decision. The competition pairs the dataset with a two-round evaluation in which participants predict the biomarkers as classification, CST and visual acuity as regression, and Continue Injection as the integrative clinical endpoint. The winning solutions treat biomarker classification as multiple-instance learning, in which each OCT B-scan is an instance and the eye is the bag, using a vision transformer or convolutional backbone, and then feed the predicted probabilities, thicknesses, and patient metadata into ensemble regressors for the Continue Injection decision.","core_discovery":"The central claim is that pre-treatment OCT images contain enough signal to predict whether a diabetic macular edema patient will need continued anti-VEGF injection, and that standard deep learning models can extract that signal when given enough labelled data. The evidence is the OCT4DME dataset, with pre- and post-treatment OCT scans from 2,000 patients, and the results of the competition held on it. On held-out test sets, the best model scored 80.06% AUC for Continue Injection, 97.96% AUC for pigment epithelial detachment, and over 93% AUC for subretinal and intraretinal fluid detection. The authors state this is the first exploration of pre-treatment stratification for DME treatment response and position the dataset as a public benchmark for that task.","pith_inferences":["An ablation that removes the OCT images and uses only patient metadata would quantify how much of the Continue Injection signal is actually carried by the scan; the paper's own searches for simple metadata rules found none, suggesting the image contribution is substantial.","The reported scores are specific to an Asian cohort collected on one OCT device; validation on other populations and devices is the natural next step and is flagged by the paper itself.","Because the competition allowed many daily submissions, the absolute AUC values may overstate real-world performance; a fixed split with a single locked submission per team would give a more trustworthy estimate.","The dataset's two annotation styles, per-eye labels in the first stage and per-image labels in the second, create a ready-made test bed for weakly supervised and label-efficient learning methods."],"forward_implications":["A clinician could use a pre-treatment OCT scan to stratify DME patients before the second anti-VEGF injection, replacing trial-and-error treatment with an evidence-based continue-or-stop decision.","The public OCT4DME dataset gives the research community a common benchmark for DME treatment-response prediction, addressing the previous scarcity of large public OCT datasets in this area.","A single multi-task model can output fluid biomarkers, central subfield thickness, and visual acuity alongside the Continue Injection decision, providing intermediate clinical signals that could be checked against standard measurements.","The success of generic deep learning architectures in the competition suggests that data scale and labelling quality, rather than novel model design, were the main drivers of predictive performance.","If the 80.06% AUC for Continue Injection reproduces in a prospective setting, automated triage of DME patients for follow-up injection becomes feasible with a non-invasive imaging test."],"supporting_citations":[{"why":"Documents the wide variation in DME treatment response, with 31.6% to 65.6% of patients showing persistent edema after four injections, which motivates the need for pre-treatment stratification.","marker":"[5]"},{"why":"The small public DME OCT dataset with only 15 subjects that the paper cites as insufficient for robust deep learning, motivating the release of OCT4DME.","marker":"[18]"},{"why":"Prior multitask deep learning system for DME classification on OCT, the state of the art that this paper extends toward treatment-response prediction.","marker":"[14]"},{"why":"The prior OCT fluid detection challenge whose task design informs the biomarker subtasks of this competition.","marker":"[35]"},{"why":"Swin Transformer, the vision transformer backbone used by the champion team for fluid biomarker classification.","marker":"[56]"},{"why":"Multiple instance learning survey that supplies the bag-of-B-scans framework for combining per-image predictions into per-eye decisions.","marker":"[57]"},{"why":"ResNet, the convolutional backbone used by the champion and runner-up teams for CST regression and classification tasks.","marker":"[58]"},{"why":"EfficientNet, the backbone selected by the third-place team after comparing several architectures on the competition tasks.","marker":"[62]"}],"fun_headline_variants":["OCT images predict DME anti-VEGF response with 80% AUC","Pre-treatment OCT scans forecast DME therapy outcome","Dataset of 2,000 patients enables DME response prediction","AI model uses OCT to decide if DME patients need repeat injections","First study: OCT predicts DME treatment response before therapy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the leaderboard scores measure true predictive skill; because teams could submit many times per day and tune to the private test set, the reported 80.06% AUC may not transfer to new patients.","fun_headline_variants_meta":{"raw":{"variants":["OCT images predict DME anti-VEGF response with 80% AUC","Pre-treatment OCT scans forecast DME therapy outcome","Dataset of 2,000 patients enables DME response prediction","AI model uses OCT to decide if DME patients need repeat injections","First study: OCT predicts DME treatment response before therapy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000995,"raw_usage":{"total_tokens":4180,"prompt_tokens":878,"completion_tokens":3302,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":3216}},"tokens_in":494,"tokens_out":3302,"duration_ms":25470,"temperature":1.0,"reasoning_tokens":3216,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:56:42.222776+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the top model architecture on the OCT4DME training set only, hold out a fresh clinical cohort whose Continue Injection labels are kept secret until submission, and allow each team exactly one evaluation; if the AUC falls well below 80.06%, the reported performance reflects leaderboard probing rather than generalizable prediction.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the wide variation in DME treatment response, with 31.6% to 65.6% of patients showing persistent edema after four injections, which motivates the need for pre-treatment stratification."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The small public DME OCT dataset with only 15 subjects that the paper cites as insufficient for robust deep learning, motivating the release of OCT4DME."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior multitask deep learning system for DME classification on OCT, the state of the art that this paper extends toward treatment-response prediction."},{"cited_title":"Bogunovi´ c, F","cited_arxiv_id":null,"evidence_quote":"The prior OCT fluid detection challenge whose task design informs the biomarker subtasks of this competition."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Swin Transformer, the vision transformer backbone used by the champion team for fluid biomarker classification."},{"cited_title":"Carbonneau, V","cited_arxiv_id":null,"evidence_quote":"Multiple instance learning survey that supplies the bag-of-B-scans framework for combining per-image predictions into per-eye decisions."}],"review_version":1}