{"id":"7454ec03-0052-4801-9fda-ca7643061fc1","arxiv_id":"2507.01278","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"GPT-4 shows moderate skill at binary diabetic retinopathy referral from textual fundus descriptions but near-zero skill at glaucoma referral, and metadata does not affect its predictions.","lead":"GPT-4 was tested on text descriptions of retinal photos for diabetes and glaucoma screening decisions; it did moderately on simple referral tasks but failed on fine grading and glaucoma. Adding patient metadata, real or synthetic, did not change its answers, so the model relies on the image text, not the clinical context.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Description provenance is undisclosed; if descriptions were derived from labels, all reported performance and metadata-invariance results are uninterpretable. The unexplained 300-to-265 image drop compounds this.","rationale":"The reader's weakest assumption correctly identified the provenance and label-independence of the textual descriptions as the most load-bearing issue, and I agree. The paper's abstract and results present performance numbers for GPT-4 on tasks derived from structured text; if those texts were generated from, or contaminated by, the outcome labels, then the ICDR and referral metrics are inflated and the metadata-null result becomes trivial. The Methods give no information about the description pipeline, so the published numbers cannot be independently audited. The sample-size discrepancy (300 selected, 265 analyzed) is a separate but reinforcing problem: even if descriptions are clean, unexplained exclusions could bias the reported accuracy and kappa. I do not interpret the missing provenance as evidence of misconduct or of a false result; it is an unverified precondition. The reader's CONDITIONAL verdict already appropriately requires clarification before the results can be fully trusted, so no verdict change is needed. The concrete test above would settle the concern: an independent blinded re-description and release of the original description corpus would reveal whether label leakage is present and whether the excluded cases alter the conclusions.","tokens_in":8349,"tokens_out":3696,"duration_ms":50570,"concrete_test":"Contact the authors for the full set of textual descriptions and an itemized account of the 300-to-265 case drop. Then run the image-only condition on 100 randomly selected cases using descriptions authored by an ophthalmologist who never saw the ground-truth labels; compare ICDR kappa/macro-F1 and glaucoma F1 to the reported values. A material shift (e.g., |Δkappa| > 0.05) would indicate that the reported numbers are an artifact of the unknown description pipeline rather than of the model's inherent ability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that GPT-4, given structured textual descriptions of fundus images, attains the reported ICDR kappa/F1, DR referral F1, and glaucoma F1, and that metadata does not change predictions. The load-bearing precondition is that the descriptions are accurate, complete, and independent of the ground-truth labels. Section 2.4 states only that 'all image content was conveyed through structured textual descriptions'; it never identifies who wrote them, from what source, whether the writers saw the labels, or whether any validation of description accuracy was performed. Without this, the performance numbers may be circular: if a description says 'laser scars present' (which the Appendix prompt instructs the model to treat as ICDR level 4), the model's correct classification simply recovers a label-derived phrase, not reasoning from images. If descriptions omit optic-disc detail, the poor glaucoma result is likewise an artifact of description content, not a generic LLM limitation. The metadata-invariance claim also depends on this: if descriptions carry the label signal, metadata should have no effect because the model already has the answer. Additionally, Section 3's class counts sum to 265, not the stated 300 images (198+27+22+15+3=265; binary counts 220+45=265, 206+59=265). The 35 excluded images are never accounted for, so all denominators and kappa estimates rest on an unexplained subset. These two issues determine whether the headline numbers mean anything.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a retrospective evaluation of GPT-4 as a text-only simulator of ophthalmic decision-making. The authors selected 300 fundus photographs from mBRSET, converted each image into a structured textual description, and prompted GPT-4 under three conditions: image description alone, image description plus real patient metadata, and image description plus synthetic metadata. The model was asked to output an ICDR severity score, a diabetic retinopathy referral decision, and a cup-to-disc-ratio based glaucoma referral decision. The main findings are moderate agreement for ICDR grading (kappa 0.25, macro F1 0.33), improved but still limited performance for binary DR referral (F1 0.54, kappa 0.44), essentially no skill for glaucoma referral (F1 below 0.04), and no statistically significant or practically large effect of adding metadata. The authors conclude that GPT-4 is not clinically reliable for these tasks but may be useful for educational, documentation, or annotation workflows.","tokens_in":8610,"tokens_out":2845,"duration_ms":36793,"significance":"If the central empirical claims hold, the paper provides a useful early benchmark: it shows that a state-of-the-art LLM, fed only with textual descriptions of retinal images, achieves only moderate performance on simple DR referral and essentially no skill on glaucoma referral, and that adding patient metadata does not change predictions. The study design has genuine strengths: the three-condition comparison is clear, the use of paired McNemar tests and change-rate analysis is appropriate for the research question, and the Discussion is appropriately cautious about clinical applicability. The synthetic metadata condition is a sensible perturbation test. However, the value of the results depends entirely on two unverified preconditions: that the textual image descriptions are a faithful and label-independent representation of the photographs, and that the analyzed 265-image subset is representative of the stated 300-image sample.","major_comments":[{"comment":"The central precondition of the study is that the structured textual descriptions accurately and completely convey the fundus image content without incorporating the ground-truth labels. Section 2.4 states only that 'all image content was conveyed through structured textual descriptions' and does not identify who authored the descriptions, whether they were derived from clinical reports, automated captioning, or the labels themselves, or whether the authors were blinded to the outcome. This matters because the Appendix prompt instructs the model to treat panretinal laser scars as ICDR level 4; if the descriptions were written with knowledge of the labels, then a correct ICDR 4 classification would simply recover a label-derived phrase rather than demonstrate reasoning. The same issue affects the glaucoma result: if the descriptions omit or include optic-disc detail in a label-dependent way, the poor F1 could be an artifact of the description content. The authors must provide the full description-generation protocol and, ideally, sample descriptions with the corresponding images and labels, or at least a clear statement that the descriptions were produced by graders blinded to the reference standard and validated for accuracy and completeness.","section":"§2.4"},{"comment":"The manuscript states that 300 images were selected, but all reported class counts in Section 3 sum to 265, not 300: ICDR counts 198+27+22+15+3 = 265, DR referral counts 220+45 = 265, and glaucoma referral counts 206+59 = 265. The 35 excluded images are never mentioned or explained. If the exclusions are related to image quality, missing metadata, or description-generation failures, the reported accuracy, kappa, and McNemar results may be biased. The authors must report the number of excluded images, the reasons for exclusion, and ideally a comparison of included versus excluded cases on available variables. Without this, all denominators and agreement estimates rest on an unexplained subset, and the 67.5% ICDR accuracy and other headline figures cannot be taken at face value.","section":"§3"},{"comment":"The evaluation appears to use a single GPT-4 response per prompt, accessed through the ChatGPT platform with 'temperature and sampling parameters left unchanged.' This gives no information about run-to-run variability, and the paper reports no confidence intervals around the point estimates. The central null finding that metadata does not change predictions (McNemar p > 0.05, change rates under 7%) is only meaningful if the underlying predictions are stable. With a stochastic model, a single run cannot distinguish a true null effect from sampling noise, especially for the small change rates reported. The authors should either run each prompt multiple times and report the distribution of metrics and prediction-change rates, or justify why single-run outputs are sufficient for the claims made.","section":"§2.4/§2.6"}],"minor_comments":[{"comment":"There is a typographical error in 'thestatsmodels library'; it should read 'the statsmodels library.'","section":"§2.7"},{"comment":"The output format templates contain mismatched braces and parentheses, e.g., 'Yes/No { with a brief explanation' appears twice; these should be corrected to avoid ambiguity in the prompt specification.","section":"Appendix"},{"comment":"The synthetic metadata generation rule assigns 20% of ages as '≥90' and the rest as random integers between 20 and 89; the paper should state whether this distribution was chosen to match mBRSET or for another reason, and how the remaining variables were sampled independently.","section":"§2.5"},{"comment":"For the McNemar tests that yield p-values of 1.00, the paper should report the number of discordant pairs; a p-value of 1.00 can arise from very few changes, and without the discordant cell counts the reader cannot assess the power of the comparison.","section":"§3"},{"comment":"Reference [15] is titled 'Evaluating Large Language Models for Simulated Ophthalmic Decision-Making,' which closely resembles the current manuscript's title; please confirm this is a distinct prior publication and clarify the relationship.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The decisive issue for me is the provenance of the textual image descriptions in §2.4. This is not a stylistic omission: every reported accuracy, kappa, and metadata-invariance result is conditional on the descriptions being faithful, complete, and label-independent. If the authors cannot provide evidence on this point—for example, a description-generation protocol with blinding and a validation step—the paper's conclusions would be unsupported. I would also ask the editor to require an accounting of the 300-to-265 discrepancy before any revision is considered. These concerns are potentially addressable, so I do not recommend rejection at this stage, but the revision must be substantive rather than cosmetic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Fairly straightforward read. The paper gives GPT-4 structured textual descriptions of fundus photos and asks for ICDR grade, DR referral, and glaucoma referral, with and without real or synthetic metadata. The headline finding—that adding metadata changes almost nothing—is a useful negative result, and comparing real versus synthetic metadata is the one genuinely novel piece. The reporting is also refreshingly honest: they publish the low kappa and F1 values instead of spinning them.\n\nWhere it goes soft: the Methods never say who wrote the image descriptions, from what source, or whether the writers saw the labels. That is not a minor omission. If the descriptions were derived from the ground truth, the reported kappa and the metadata invariance are both trivially explained—the model already has the answer in the text. The Appendix shows the prompt structure but not the description content itself. Second, the paper says 300 images were selected, but all class counts sum to 265. The missing 35 are never explained, so every denominator sits on an unexplained subset. Third, a single stochastic run per prompt with default temperature is treated as the answer; no seeds, no replication, no confidence intervals. These three issues are fixable, but without the fixes the central claims are conditional.\n\nOn the positive side, the statistical tests used (McNemar, change rates, kappa on prediction pairs) are appropriate. The discussion is appropriately cautious, and they cite ref [15], which has an almost identical title, rather than hiding it. The novelty concern is real but limited: the broader finding that LLMs underperform on text-only ophthalmic tasks is already out there; the metadata comparison is the delta.\n\nWho is this for? People building LLM-based documentation or annotation workflows, and anyone designing similar LLM evaluations. It deserves a serious referee because the topic is timely and the flaws are correctable. I'd send it to review but with a clear request for a major revision: describe the description-generation pipeline, reconcile the sample size, and report more than one run.\n\nNet: the direction of the results is plausible, but the missing provenance and missing images prevent me from trusting the numbers as they stand.","headline":"Honest negative result, but the paper must disclose how the textual descriptions were made before its headline numbers mean anything.","tokens_in":9150,"tokens_out":2058,"would_cite":false,"duration_ms":23898,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports that GPT-4, reading structured text descriptions of retinal fundus photographs, reaches only moderate agreement on diabetic retinopathy severity and essentially no skill at glaucoma referral, and that adding patient…","keywords":["GPT-4","large language models","diabetic retinopathy screening","glaucoma screening","retinal fundus photographs","structured text prompts","clinical metadata","ICDR severity grading"],"falsifier":"Inspect the description-generation step: if the structured text fed to GPT-4 was authored from the reference labels or by unblinded raters, the central claim collapses. A clean test is to have blinded graders or an automated captioning system write fresh descriptions for the same 300 images and rerun the identical prompts; if ICDR $\\kappa$ or glaucoma $F_1$ rises sharply, the original numbers were an artifact of the descriptions, and if they stay near 0.25 and 0.03, the paper's conclusion holds.","tokens_in":8109,"feed_emoji":"👁️","tokens_out":7891,"duration_ms":76256,"temperature":0.7,"pith_summary":"The authors are trying to establish that a text-only large language model, GPT-4, cannot replace clinical screening for diabetic retinopathy or glaucoma when it is fed structured descriptions of retinal photographs rather than the images themselves. They report moderate performance only on the coarsest distinction: ICDR severity grading reached 67.5% accuracy but $\\kappa=0.25$ and macro $F_1=0.33$, driven by correct detection of normal cases, while mild and severe DR classes scored $F_1=0.00$. Reframing the task as binary referral for diabetic retinopathy improved results (accuracy 82.3%, $F_1=0.54$, $\\kappa=0.44$), but glaucoma referral based on cup-to-disc ratio remained near zero across every condition ($F_1<0.04$, $\\kappa<0.03$). Adding real or synthetic patient metadata changed predictions in fewer than 7% of ICDR cases and under 3% of DR referral cases, with no significant McNemar differences, so the authors conclude the model leans on image descriptions and pretrained medical priors, not patient context. If the results hold, text-only LLMs are not clinically usable for these screening tasks, though they may still serve educational, documentation, or annotation workflows.","feed_headline":"GPT-4 misses glaucoma referrals from text-only eye exam summaries","feed_subtitle":"In 300 retinal images, added patient metadata changed nothing: DR grading stayed weak and glaucoma referral near zero.","key_machinery":"The mechanism is a structured prompt that converts each fundus photograph into a text-only clinical vignette and asks the model to perform four linked tasks: list diabetic retinopathy signs, assign an ICDR score from 0 to 4 with panretinal photocoagulation scars forcing level 4, recommend DR referral when the score is at least 2 or macular edema is present, and estimate the cup-to-disc ratio, referring for glaucoma when the ratio exceeds 0.6. The same prompt is run under three conditions, image description alone, with real patient metadata, and with synthetic metadata, and the three outputs are compared through accuracy, $F_1$, Cohen's $\\kappa$, McNemar's test, and pairwise change rates. The design isolates the contribution of metadata by keeping the description fixed and varying only the demographic and clinical context.","core_discovery":"The central claim is that GPT-4 can simulate only the coarse parts of ophthalmic screening from text, and that its predictions are insensitive to clinical metadata. In 300 annotated fundus images from mBRSET, the model assigned ICDR severity levels with limited agreement with the reference standard ($\\kappa=0.250$, accuracy 67.5%, macro $F_1=0.330$, weighted $F_1=0.672$), and its accuracy was concentrated in the normal class (class 0 $F_1=0.82$), with $F_1=0.00$ for mild and severe non-proliferative DR. The binary DR referral task fared better (accuracy 82.3%, $F_1=0.54$, $\\kappa=0.436$), while glaucoma referral from estimated cup-to-disc ratio was essentially absent ($F_1<0.04$, $\\kappa<0.03$). Neither real nor synthetic metadata produced a statistically significant shift in any task ($p>0.05$ by McNemar's test), and the authors read this stability as evidence that the model is not integrating patient-specific context into its decisions.","pith_inferences":["Editorial extension: if the structured image descriptions were written from the ground-truth labels or by raters who had seen them, the reported $\\kappa$ and $F_1$ values would partly measure label leakage through the descriptions rather than GPT-4's clinical reasoning; the paper does not disclose who authored the descriptions or whether they were blinded.","Editorial extension: a direct next experiment is to feed the same prompt template to a genuinely multimodal model that reads the pixels, which would separate the loss caused by the text bottleneck from the intrinsic difficulty of the grading tasks.","Editorial extension: because the model's decisions are nearly deterministic across prompt conditions, it could serve as a cheap first-pass annotator for easy normal cases, with human review concentrated on the small fraction of cases where the predicted label shifts.","Editorial extension: the glaucoma failure may be a task-format artifact as much as a model limitation; prompting for quantified optic-disc features such as neuroretinal rim width instead of a single cup-to-disc ratio could give a language model a fairer chance."],"forward_implications":["A text-only GPT-4 pipeline should not be used for standalone screening or triage in diabetic retinopathy or glaucoma, because its near-zero performance on referable DR classes and on glaucoma referral would miss exactly the patients who need follow-up.","Adding demographic or clinical metadata to prompts is not a shortcut: because predictions move in under 7% of cases, efforts to improve LLM screening should focus on how the image content is represented rather than on enriching the patient context.","Binary referral decisions are a more realistic target for text-only LLMs than fine-grained severity grading, since the model's DR referral performance was substantially stronger than its ICDR multiclass performance.","The model's near-total blindness to metadata has a fairness implication: its decisions do not adjust for age, sex, or comorbidities, so any use in annotation or documentation would inherit this context-free behavior."],"supporting_citations":[{"why":"Supplies the mBRSET retinal photographs, demographic data, and reference labels from which the 300-image subset was drawn.","marker":"[19]"},{"why":"Defines the ICDR severity scale that the model is asked to assign and that the dataset's ground truth uses.","marker":"[20]"},{"why":"Describes GPT-4, the single language model under test, accessed through the ChatGPT platform.","marker":"[16]"},{"why":"Establishes the image-based deep learning benchmark for diabetic retinopathy detection that motivates the referral task.","marker":"[11]"},{"why":"Provides the multiethnic deep learning DR grading system whose difficulty on subtle signs the authors cite when interpreting their misclassifications.","marker":"[12]"},{"why":"Supplies the cup-to-disc ratio context that underlies the >0.6 glaucoma referral threshold used in the prompt.","marker":"[6]"},{"why":"Motivates adding glaucoma screening to diabetic retinopathy telemedicine programs, the clinical rationale for the glaucoma referral task.","marker":"[5]"},{"why":"Documents interobserver variability in optic disc assessment, which the authors use to explain the poor glaucoma performance.","marker":"[13]"}],"fun_headline_variants":["GPT-4 flunks glaucoma referral from text eye exams","Metadata doesn't boost GPT-4's eye screening accuracy","GPT-4 misses glaucoma cases in simulated retina screening","Text-only fundus descriptions stump GPT-4 for glaucoma","LLM eye screening: weak DR grading, absent glaucoma detection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation stands or falls on the structured textual descriptions of the 300 images being accurate, complete, and independent of the reference labels — if those descriptions came from the ground-truth diagnoses or from raters who saw them, the reported scores would reflect label leakage rather than GPT-4's decision-making.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4 flunks glaucoma referral from text eye exams","Metadata doesn't boost GPT-4's eye screening accuracy","GPT-4 misses glaucoma cases in simulated retina screening","Text-only fundus descriptions stump GPT-4 for glaucoma","LLM eye screening: weak DR grading, absent glaucoma detection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000518,"raw_usage":{"total_tokens":2595,"prompt_tokens":1116,"completion_tokens":1479,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":732,"completion_tokens_details":{"reasoning_tokens":1398}},"tokens_in":732,"tokens_out":1479,"duration_ms":108248,"temperature":1.0,"reasoning_tokens":1398,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:55:41.569204+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the description-generation step: if the structured text fed to GPT-4 was authored from the reference labels or by unblinded raters, the central claim collapses. A clean test is to have blinded graders or an automated captioning system write fresh descriptions for the same 300 images and rerun the identical prompts; if ICDR $\\kappa$ or glaucoma $F_1$ rises sharply, the original numbers were an artifact of the descriptions, and if they stay near 0.25 and 0.03, the paper's conclusion holds.","supporting_citations":[{"cited_title":"A portable retina fundus photos dataset for clinical, demographic, and diabetic retinopathy prediction","cited_arxiv_id":null,"evidence_quote":"Supplies the mBRSET retinal photographs, demographic data, and reference labels from which the 300-image subset was drawn."},{"cited_title":"Proposed international clinical diabetic retinopathy and diabetic macular edema disease severity scales","cited_arxiv_id":null,"evidence_quote":"Defines the ICDR severity scale that the model is asked to assign and that the dataset's ground truth uses."},{"cited_title":"Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs.JAMA","cited_arxiv_id":null,"evidence_quote":"Establishes the image-based deep learning benchmark for diabetic retinopathy detection that motivates the referral task."},{"cited_title":"Development and validation of a deep learning system for diabetic retinopathy and related eye diseases using retinal images from multiethnic populations with diabetes","cited_arxiv_id":null,"evidence_quote":"Provides the multiethnic deep learning DR grading system whose difficulty on subtle signs the authors cite when interpreting their misclassifications."},{"cited_title":"The effect of optic disc diameter on ver- tical cup to disc ratio percentiles in a population based cohort: the Blue Mountains Eye Study","cited_arxiv_id":null,"evidence_quote":"Supplies the cup-to-disc ratio context that underlies the >0.6 glaucoma referral threshold used in the prompt."},{"cited_title":"Evaluating the outcome of screening for glau- coma using colour fundus photography-based referral criteria in a teleophthalmology screening programme for diabetic retinopathy","cited_arxiv_id":null,"evidence_quote":"Motivates adding glaucoma screening to diabetic retinopathy telemedicine programs, the clinical rationale for the glaucoma referral task."},{"cited_title":"Interobserver agree- ment for the clinical assessment of optic discs","cited_arxiv_id":null,"evidence_quote":"Documents interobserver variability in optic disc assessment, which the authors use to explain the poor glaucoma performance."}],"review_version":1}