{"id":"0ff362a2-1c99-40b8-b4e6-bec7364cd664","arxiv_id":"1908.06399","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An AI diabetic retinopathy detector transferred well to handheld fundus images for proliferative disease but showed a significant drop in accuracy for referable disease, partly due to grading scheme mismatch.","lead":"This study tested an AI system for diabetic retinopathy against images from a handheld portable fundus camera in a real-world Mexican screening cohort, and compared its performance with a curated desktop camera dataset. The AI matched desktop-level performance for proliferative disease but showed a large, statistically significant drop for referable disease, though grading scheme differences complicate the comparison.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The RDR performance drop is attributed to the handheld camera, but the grading-scheme mismatch between MAILOR (Scottish) and IDRiD (ICDR) confounds the comparison; the paper itself calls this the biggest contributing factor.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the observed RDR performance difference is confounded by the grading-scheme mismatch and other dataset differences, so the abstract's causal attribution of the drop to the handheld device is not established. This is the correct central issue because it directly undermines the paper's headline conclusion, not a peripheral methodological detail. The paper's own Discussion factor 2 makes the concern explicit, and the PDR result—where grading schemes coincide and no significant difference is found—actually supports the interpretation that grading mismatch, rather than device transferability, drives the RDR gap. A concrete re-analysis with harmonized ICDR labels would settle the question. I do not recommend rejecting the paper: the empirical measurements on a large real-world handheld cohort are valuable, and the limitations are acknowledged, but the central claim needs to be reframed as a difference in measured performance across two non-comparable label definitions, not as a device-induced degradation. The reader's CONDITIONAL verdict is appropriate, and my read does not move it.","tokens_in":12635,"tokens_out":4740,"duration_ms":51095,"concrete_test":"Regrade a random sample (or all) of the MAILOR images that are Scottish non-referable but contain haemorrhages or exudates meeting ICDR referable criteria, specifically four or more haemorrhages in one hemifield or exudates within one disc diameter of the fovea, and relabel them as ICDR-referable. Recompute the RDR AUROC on the harmonized labels. If the AUROC rises from 89.4% toward the 98.5% benchmark, or the difference becomes non-significant, the device-attribution claim collapses; if the AUROC remains near 89.4%, the handheld-camera effect is supported. A simpler computational version is to apply the paper's own discordance rules to the existing MAILOR grades and re-run the ROC analysis; the key quantity is the AUROC after label harmonization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim is that RDR AUROC falls from 98.5% (IDRiD) to 89.4% (MAILOR) 'when using the handheld device.' This causal attribution is not supported because the RDR labels in the two datasets are defined by different grading schemes. The paper states that Pegasus outputs ICDR grades, while the MAILOR reference standard is the Scottish scheme (R2 or above, or DM1). It then explicitly describes the discordance: images with haemorrhages not reaching four in a hemifield, or exudates farther than one disc diameter from the fovea, are ICDR-referable but not Scottish-referable. Since Pegasus is optimized for ICDR, such cases are scored as false positives against the Scottish clinical reference standard, mechanically lowering the measured AUROC and specificity on MAILOR independent of any device effect. The Discussion itself acknowledges this: 'the biggest contributing factor is the mismatch in grading systems,' and the PDR result, where the grading schemes coincide and no significant difference is seen, is used as supporting evidence. The abstract conclusion, however, still attributes the RDR decrease to the handheld device. The load-bearing assumption, therefore, is that the grading mismatch did not produce the gap; the paper provides no quantitative test of that assumption. Additional internal inconsistencies (Discussion CIs for sensitivity/specificity do not match Results; Table 1 category counts do not sum to the stated N) further reduce confidence in the reported comparisons.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a retrospective evaluation of a proprietary deep learning system (Pegasus, Visulytix) for diabetic retinopathy (DR) detection on images from a handheld portable fundus camera (Volk Pictor Plus) used in a Mexican screening cohort (MAILOR, ~5,752 patients, 22,180 images). The authors compare Pegasus's performance on MAILOR against its performance on the public IDRiD benchmark (516 desktop-camera images) for referable DR (RDR) and proliferative DR (PDR). They find a statistically significant drop in AUROC for RDR between IDRiD (98.5%) and MAILOR (89.4%), while PDR AUROC is not significantly different (92.2% vs 94.3%). The abstract concludes that Pegasus generalizes well to the handheld device for PDR but shows a substantial decrease for RDR when using the handheld device. The Discussion lists four possible explanatory factors, of which the authors identify the mismatch between the Scottish grading scheme used as the MAILOR reference standard and the ICDR scheme output by Pegasus as the 'biggest contributing factor.'","tokens_in":13030,"tokens_out":4859,"duration_ms":44507,"significance":"If the reported performance figures are reliable, this is a practically relevant evaluation: real-world clinical deployment of AI DR screening on low-cost handheld cameras is an important use case for telemedicine, and independent evidence on such transferability is scarce. The study uses a large, naturalistic cohort, a clearly described protocol, and standard AUROC analysis with bootstrap CIs; the external IDRiD benchmark provides a common reference. The measured performance numbers are empirical and the paper does not fit parameters to the test data. However, the study is sponsored and analyzed by the system's developer, and the central causal interpretation—that the RDR drop is due to the handheld device—is not supported by the study design, as the comparison is confounded by multiple simultaneous differences between the two datasets. The manuscript itself acknowledges the largest confounder (grading-scheme mismatch), which makes the abstract's device-specific conclusion overreach.","major_comments":[{"comment":"The abstract conclusion attributes the RDR performance decrease to 'using the handheld device,' but the Discussion (factor 2) explicitly states that 'the biggest contributing factor is the mismatch in grading systems' between the Scottish scheme used for the MAILOR reference standard and the ICDR scale output by Pegasus. Because the two datasets also differ in camera, population, mydriasis, fields of view, and image quality, the measured AUROC difference cannot be attributed to the device. The abstract should be revised to state that the RDR drop is an observed difference between two unadjusted cohorts, with the grading-system mismatch as a leading potential cause, rather than a device effect.","section":"Abstract and Discussion"},{"comment":"The disease severity distribution in Table 1 is internally inconsistent. For MAILOR, 5,017 + 595 + 60 = 5,672, not 5,752, and the percentages sum to 98.6%. For IDRiD, 168 + 323 + 62 = 553, not 516, and the listed percentages (32.6%, 52.6%, 12.0%) do not match the counts (323/516 = 62.6%, not 52.6%). The authors should clarify whether the categories are mutually exclusive (and if RDR includes PDR, state that explicitly) and correct the table so that counts and percentages are consistent and sum to the stated totals.","section":"Table 1"},{"comment":"The sensitivity and specificity confidence intervals reported for Pegasus in this paragraph do not match those reported in the Results (Table 2). For RDR on MAILOR, the text quotes '81.6% (95% CI: 83.9-90.2)' for sensitivity and '81.7% (95% CI: 85.7-89.7)' for specificity, whereas Table 2 lists 81.6% (79.0-84.2) and 81.7% (80.9-82.6). The subsequent sentence 'In terms of RDR prediction, Rajalakshmi et al. report ... compared to the 86.6% ... sensitivity and 87.7% ... specificity obtained by Pegasus' uses PDR values. These appear to be copy-paste errors that must be corrected for the manuscript to be considered reliable.","section":"Discussion, comparison to Rajalakshmi et al."},{"comment":"The paper identifies the grading-scheme mismatch as the 'biggest contributing factor' to the RDR drop, yet provides no quantitative estimate of its effect. The authors should either (a) re-grade the MAILOR images according to the ICDR scheme (or have a subset re-graded) and recompute the comparison, or (b) at minimum quantify the fraction of Pegasus false positives that fall into the discordant region (e.g., cases with haemorrhages in fewer than four hemifields, or exudates farther than one disc diameter from the fovea). Without such a quantitative assessment, the claim that the device is responsible for the RDR drop is unsubstantiated.","section":"Discussion, factor 2"},{"comment":"The comparison of AUROC between the MAILOR and IDRiD cohorts is unadjusted for any confounders (e.g., image quality, field type, patient demographics, and grading scheme). Given the authors themselves list four differing factors, the permutation test only establishes that the two independent samples have different AUROC values; it does not isolate the device contribution. The authors should temper the causal language and either perform a stratified or matched analysis (e.g., restricting both datasets to macula-centered, good-quality images) or explicitly state that the comparison is descriptive and hypothesis-generating.","section":"Statistical Analysis / Results"}],"minor_comments":[{"comment":"Reference numbering is disturbed in the Discussion: the 'One of the systems evaluated could not handle disc-centred images... 10' citation appears to refer to Tufail et al. (ref. 12), and the ICDR scale is also ref. 10. Please renumber and re-check all in-text citations for accuracy.","section":"References"},{"comment":"The claim that this is 'the first to evaluate ... a fully portable, handheld device' is questionable because the Discussion cites Rajalakshmi et al. using a smartphone-based fundus camera, which is also a portable handheld imaging approach. Please qualify the claim (e.g., a dedicated handheld fundus camera) or adjust the literature statement.","section":"Conclusion"},{"comment":"There is a typo in the final paragraph: 'clinical practise' should be 'clinical practice'.","section":"Discussion"},{"comment":"The abbreviation 'CRS' in the Table 1 header is not defined in the table footnote; define it (clinical reference standard) or spell out fully.","section":"Table 1"},{"comment":"The sentence 'There was approximately a 12-13% disparity in the RDR sensitivity/specificity performances between the handheld and benchmark desktop devices' is slightly imprecise; the sensitivity drop is 11.8 percentage points and the specificity drop is 12.5 percentage points. Using exact values would be more rigorous.","section":"Results"}],"recommendation":"major_revision","confidential_remarks":"The paper is written by the developers of Pegasus and funded by Visulytix. There is an explicit conflict-of-interest statement, which is good, but the abstract's conclusion is noticeably more favorable to the product than the Discussion's own assessment supports. I recommend that the editor verify that the revised abstract and conclusions reflect the confounding of the RDR comparison, and consider whether the study's descriptive design is sufficient for the journal's standard. The internal data inconsistencies (Table 1) and the mismatched confidence intervals in the Discussion should be treated as requiring correction before any reconsideration."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: worth reading, but the central claim overreaches. The paper evaluates a proprietary AI (Pegasus) on 5,752 patients imaged with a handheld fundus camera in Mexico, compares to the IDRiD desktop benchmark. That is genuinely new – first evaluation I know of on a fully portable handheld in a real-world cohort, and the numbers are large. PDR detection transfers well (AUROC 94.3% vs 92.2%, n.s.), which is a useful result because the grading definitions coincide for PDR.\n\nThe soft spot is RDR. The abstract says the drop from 98.5% to 89.4% is \"when using the handheld device.\" But the datasets differ in camera, field of view, image quality, population, and crucially the reference standard: MAILOR uses Scottish grading, Pegasus outputs ICDR. The authors themselves say the biggest contributing factor is the grading mismatch, and they show the RDR false positives are often cases that are ICDR-referable but not Scottish-referable. That means the measured AUROC drop is at least partly an artifact of label mismatch, not a pure device effect. The paper should have either re-graded the MAILOR images to ICDR, restricted the RDR analysis to cases where the schemes agree, or at minimum rephrased the conclusion to say \"performance was lower when evaluated against a Scottish-graded clinical standard\" rather than attributing the drop to the device.\n\nThere are also internal numeric inconsistencies: the CIs quoted in the Discussion for sensitivity/specificity don't match Table 2, and the severity counts in Table 1 don't sum to the stated N for either cohort. Minor but sloppy and undermines confidence in the reported comparisons. No code or data for MAILOR, though the benchmark is public.\n\nThe Discussion is honest about limitations, which counts in its favor. The PDR result stands regardless of the RDR framing. This deserves peer review – it addresses a real deployment question and has a large clinically relevant dataset. But it needs major revision: a re-analysis or re-framing of the RDR claim, correction of the numeric errors, and ideally a re-graded subset to quantify the grading-scheme effect. I would not desk-reject it.","headline":"Useful real-world evaluation of AI on handheld fundus images, but the headline RDR drop is confounded by the grading-scheme mismatch the authors themselves identify, so the causal attribution to the device does not hold.","tokens_in":13437,"tokens_out":2028,"would_cite":true,"duration_ms":20965,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An AI diabetic-retinopathy system kept proliferative-disease accuracy on handheld-camera images but lost significant ground on referable disease, from 98.5% to 89.4% AUROC.","keywords":["diabetic retinopathy","deep learning","handheld fundus camera","referable diabetic retinopathy","proliferative diabetic retinopathy","real-world validation","grading scale mismatch","AUROC"],"falsifier":"Regrade the handheld-camera images using the same ICDR scale the AI outputs and re-run the frozen system; if the referable-DR AUC returns to near the 98.5% desktop benchmark, the grading-system mismatch is the driver. If the gap persists after regrading, image quality or device differences are responsible.","tokens_in":12454,"feed_emoji":"👁️","tokens_out":7100,"duration_ms":63186,"temperature":0.7,"pith_summary":"The paper tests whether an AI diabetic-retinopathy detector trained and validated on high-quality desktop fundus images keeps its accuracy when fed images from a lightweight handheld fundus camera used in real-world screening. It reports that the system's ability to detect proliferative diabetic retinopathy transferred well, with AUROC statistically unchanged (94.3% versus 92.2% on a curated benchmark). But for referable diabetic retinopathy, the AUROC dropped significantly from 98.5% to 89.4%, with sensitivity and specificity both near 82%. The result matters because handheld cameras are the practical way to reach remote populations, and AI screening is only useful if it works on the images those devices produce. The paper's own explanation points mainly to a mismatch between the grading scheme the software uses and the scheme used to label the real-world images, not necessarily to a failure of the AI itself.","feed_headline":"AI diabetic-eye screening slips on handheld camera","feed_subtitle":"Referable-retinopathy AUC fell from 0.985 to 0.894; proliferative disease detection held steady.","key_machinery":"The argument is carried by a paired, out-of-the-box evaluation design. The same frozen AI system—a deep-learning tool (Pegasus) that outputs a diabetic-retinopathy grade on the International Clinical Diabetic Retinopathy (ICDR) scale—is run without any adaptation on two image sets: a curated public benchmark of 516 desktop-camera photographs, and a real-world cohort of 5,752 patients imaged with a handheld portable non-mydriatic camera, yielding 22,180 images. The real-world reference standard is the Scottish DR grading scheme, and the paper compares AUROC at the referable (RDR) and proliferative (PDR) thresholds, using bootstrap confidence intervals and permutation tests for significance. The grading-scheme mismatch is a deliberate part of the machinery: it is the paper's main candidate explanation for the RDR gap, because PDR definitions coincide across schemes while RDR definitions do not.","core_discovery":"The central claim is that transferability from curated desktop-camera images to real-world handheld-camera images is disease-severity dependent. For proliferative diabetic retinopathy, the system's AUROC on the handheld cohort was 94.3% (95% CI 91.0-96.9), statistically indistinguishable from the 92.2% (95% CI 89.4-94.8) on the desktop benchmark (p=0.172). For referable diabetic retinopathy, the AUROC fell to 89.4% (95% CI 88.0-90.7) from 98.5% (95% CI 97.8-99.2) on the benchmark (p<0.001), with sensitivity and specificity of about 82% at the equal-error operating point. The paper concludes that the system transfers well for PDR but that RDR performance drops substantially, and attributes the RDR gap mainly to the mismatch between the software's ICDR output and the Scottish reference standard, with image quality and field type as contributing factors.","pith_inferences":["Because PDR lesions are coarse while RDR hinges on small haemorrhages and microaneurysms, the paper's pattern suggests that the transferability gap is concentrated in precisely the features most sensitive to resolution, focus, and grading definitions; the paper does not disentangle these, but its false-positive figures point that way.","A direct test of the paper's main explanation would be to regrade the handheld images with the ICDR scale and rerun the frozen AI; narrowing of the RDR gap would implicate the grading mismatch, while a persistent gap would implicate image quality or device differences.","If the grading-scheme mismatch is the dominant factor, the practical fix is a mapping layer that translates the AI output into the local referral rules, which could recover much of the lost RDR performance without retraining the network.","The comparison also implies that high benchmark AUCs in the literature are weak evidence of field readiness for referable diabetic retinopathy; local device-specific validation is the deciding test."],"forward_implications":["Proliferative diabetic retinopathy screening with a handheld camera and an unmodified desktop-trained AI is plausible in this type of real-world cohort, since PDR accuracy did not degrade.","Referable diabetic retinopathy screening should not be assumed to transfer; a validation against the actual device and grading protocol is needed before deployment.","Aligning the AI's grading scale with the local screening program's referral definitions could remove a major source of apparent false positives, per the paper's own analysis.","Curated public benchmark results for referable DR should be treated as optimistic upper bounds, not expected field performance.","Using macula-centred fields rather than disc-centred fields improved referable-DR AUC by 2.3%, so image field selection affects screening performance."],"supporting_citations":[{"why":"Provides the curated desktop-camera benchmark dataset that defines the comparison condition for the handheld evaluation.","marker":"7"},{"why":"Documents the quality gap between handheld and stand-alone non-mydriatic cameras, which the paper uses to frame its image-quality explanation.","marker":"8"},{"why":"Defines the Scottish DR grading protocol used as the clinical reference standard for the real-world cohort.","marker":"9"},{"why":"Defines the ICDR severity scale that the AI system outputs and the paper's RDR and PDR threshold definitions.","marker":"10"},{"why":"Reports a smartphone-based portable-camera AI study whose sensitivity and specificity numbers the paper compares against its own.","marker":"11"},{"why":"Reports clinical-setting performance of multiple desktop-camera AI systems, giving the benchmark for real-world sensitivity and specificity comparisons.","marker":"12"},{"why":"Large multi-ethnic deep-learning validation on conventional images, used to contrast curated benchmark performance with the handheld results.","marker":"13"},{"why":"Development and validation of a deep-learning algorithm on conventional-camera datasets with clinician-removed poor-quality images, used as a comparison point for the desktop benchmark.","marker":"14"}],"fun_headline_variants":["Handheld camera stumbles AI eye screening for referable retinopathy","AI DR detection holds for proliferative, slips for referable on handheld","AI eye test: handheld cuts referable DR accuracy, not proliferative","Handheld AI: proliferative retinopathy detection holds, referable drops","AI retinopathy screening: handheld gap only for referable disease"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the observed difference in referable-DR performance between the two datasets is attributable to the handheld camera and real-world conditions, rather than to systematic differences between the datasets—above all, the mismatch between the Scottish grading scheme used for the real-world reference standard and the ICDR scheme the AI outputs.","fun_headline_variants_meta":{"raw":{"variants":["Handheld camera stumbles AI eye screening for referable retinopathy","AI DR detection holds for proliferative, slips for referable on handheld","AI eye test: handheld cuts referable DR accuracy, not proliferative","Handheld AI: proliferative retinopathy detection holds, referable drops","AI retinopathy screening: handheld gap only for referable disease"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000832,"raw_usage":{"total_tokens":3726,"prompt_tokens":1135,"completion_tokens":2591,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":751,"completion_tokens_details":{"reasoning_tokens":2498}},"tokens_in":751,"tokens_out":2591,"duration_ms":16985,"temperature":1.0,"reasoning_tokens":2498,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:45:50.210076+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Regrade the handheld-camera images using the same ICDR scale the AI outputs and re-run the frozen system; if the referable-DR AUC returns to near the 98.5% desktop benchmark, the grading-system mismatch is the driver. If the gap persists after regrading, image quality or device differences are responsible.","supporting_citations":[{"cited_title":"Quality and learning curve of handheld versus stand-alone non-mydriatic cameras","cited_arxiv_id":null,"evidence_quote":"Documents the quality gap between handheld and stand-alone non-mydriatic cameras, which the paper uses to frame its image-quality explanation."},{"cited_title":"Grading diabetic retinopathy (DR) using the Scottish grading protocol","cited_arxiv_id":null,"evidence_quote":"Defines the Scottish DR grading protocol used as the clinical reference standard for the real-world cohort."},{"cited_title":"Automated diabetic retinopathy detection in smartphone-based fundus photography using artificial intelligence","cited_arxiv_id":null,"evidence_quote":"Reports a smartphone-based portable-camera AI study whose sensitivity and specificity numbers the paper compares against its own."},{"cited_title":"Automated diabetic retinopathy image assessment software: diagnostic accuracy and cost-effectiveness compared with human graders","cited_arxiv_id":null,"evidence_quote":"Reports clinical-setting performance of multiple desktop-camera AI systems, giving the benchmark for real-world sensitivity and specificity comparisons."},{"cited_title":"Development and validation of a deep learning system for diabetic retinopathy and related eye diseases using retinal images from multiethnic populations with diabetes","cited_arxiv_id":null,"evidence_quote":"Large multi-ethnic deep-learning validation on conventional images, used to contrast curated benchmark performance with the handheld results."},{"cited_title":"Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs","cited_arxiv_id":null,"evidence_quote":"Development and validation of a deep-learning algorithm on conventional-camera datasets with clinician-removed poor-quality images, used as a comparison point for the desktop benchmark."}],"review_version":1}