{"id":"4c8d1374-369d-4a19-a742-e363aad5cb4c","arxiv_id":"1908.01428","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A 3D CNN infers aggregate visual field metrics (VFI and MD) directly from raw OCT volumes, outperforming classical machine learning on standard OCT features (PC 0.88 vs 0.74 for VFI).","lead":"This paper trains an artificial intelligence model to estimate visual field test scores directly from raw 3D eye scans, reaching a correlation of 0.88 with the measured visual field index on optic nerve head scans. If the result holds in real clinics, glaucoma monitoring could become faster and cheaper, with fewer time-consuming visual field tests.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline p<0.01 superiority claim is not statistically supported: scan-level repeated measures from only 579 patients are treated as independent, with no clustering-corrected test described.","rationale":"I read the paper in good faith. The core result—a 3D CNN regressing aggregate VFI/MD from raw OCT—is plausible, and the patient-wise fold split is a real strength that avoids the most obvious identity-leakage artifact. The reader's weakest assumption, temporal alignment between OCT and VFT, is a legitimate reporting gap, but it would most likely add noise and bias the correlation downward rather than inflate it. The more load-bearing weakness is statistical: the headline superiority claim (p<0.01) is made on scan-level data with repeated measures from the same patients, and no valid clustering-corrected test is described. With 4155 scans from 579 patients, within-patient correlation in structural and functional status is substantial, so any test that treats scans as independent is anti-conservative. The reported standard deviations over folds are not confidence intervals for the CNN-versus-RFR difference. This directly affects the abstract's comparative claim and the paper's novelty relative to classical baselines. The concrete patient-level re-analysis would settle whether the comparison survives. Because the point estimates may still be informative, I do not recommend rejection; the verdict remains conditional on the missing statistical analysis and, ideally, external validation.","tokens_in":8295,"tokens_out":10323,"duration_ms":119290,"concrete_test":"Recompute the CNN-versus-RFR comparison with the patient as the unit of analysis. For each of the five folds, average the model predictions and true VFI/MD over all visits of each patient in the test fold, then compute fold-level patient-averaged PC and RMSE for both models. Test the difference with a paired test over the five folds or a cluster bootstrap that resamples patients rather than scans, and report a confidence interval for the PC difference. If the cluster-corrected p-value is ≥0.05, or the patient-level PC gap is much smaller than 0.14, the abstract's p<0.01 superiority claim is unsupported and the paper should be conditioned on stronger statistical evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 (Tables 2 and 3) reports a CNN VFI Pearson correlation of 0.88±0.029 versus 0.74±0.090 for the best classical ML method and states the difference is 'significantly higher' (p<0.01). The data consist of 4155 OCT/VFT pairs from only 579 patients (Section 2.2), roughly 7 visits per patient, and all scans of a patient are placed in the same fold. The paper never states how p<0.01 was computed. If it comes from a paired test over the five fold-level PCs, n=5 is small and the result depends on the covariance of fold-level differences, which is not reported; with the reported fold standard deviations the implied t-statistic would be borderline at best. If it comes from pooling all test scans, the effective sample size is inflated by within-patient correlation in both OCT structure and VFI/MD trajectory, making the test anti-conservative. Either way, the abstract's central comparative claim—that the CNN is significantly better than classical ML—is not established. A patient-level or cluster-robust analysis is required before the 0.88-versus-0.74 gap can be interpreted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a 3D convolutional neural network that regresses the glaucoma severity metrics Visual Field Index (VFI) and Mean Deviation (MD) directly from raw OCT volumes of the optic nerve head or macula, and compares it with ten classical machine-learning regressors trained on 22 segmentation-based OCT features. Using 4155 OCT/VFT pairs from 579 patients with patient-wise 5-fold cross-validation, the authors report a test Pearson correlation of 0.88 ± 0.029 for VFI on ONH scans versus 0.74 ± 0.090 for the best classical baseline (a random forest), and claim the difference is significant (p < 0.01). They additionally report MD results, RMSE values, and class activation maps. The central claim is that aggregate visual field measurements can be inferred from a single raw OCT volume more accurately than from hand-crafted OCT features.","tokens_in":8470,"tokens_out":7995,"duration_ms":79937,"significance":"If the reported results hold, the work is a useful empirical contribution: it demonstrates that regression of global VFT indices from raw OCT is feasible, avoids explicit layer segmentation and registration, and the use of patient-wise folds addresses a common leakage problem. The comparison against classical baselines on the same data is informative, and the discussion of the VFT retest upper bound (r ≈ 0.88) provides sensible context. The paper's strengths include a clinically relevant question, a moderately sized real-world dataset, and transparent reporting of fold-level means and standard deviations. However, the statistical evidence for the headline 'significantly higher' claim is incomplete, and the cross-validation protocol is described ambiguously; these issues must be resolved before the comparative claim can be accepted at face value.","major_comments":[{"comment":"The claim that the CNN's PC of 0.88 is 'significantly higher' (p < 0.01) than the random forest's 0.74 is not supported by any statistical test described in the manuscript. The data contain 4155 scan-VFT pairs from only 579 patients (about 7 visits per patient), and although the authors state that folds are patient-wise, the paper does not report how the p-value was computed. A paired test over five fold-level correlations would have n = 5 and requires the covariance of the fold-level differences, which is not reported; pooling all test scans would inflate the effective sample size because of within-patient correlation in both OCT structure and VFI/MD trajectories. Please report a cluster-robust test or a patient-level bootstrap, or remove the significance claim. The same issue applies to the statement that the macula PC of 0.86 is 'still significantly higher' and to the claim in Section 3 that there is no significant difference between VFI and MD for the same region.","section":"Section 3, Tables 2 and 3, Abstract"},{"comment":"The description of the data split is internally inconsistent. Section 2.2 states that the data set was divided into an 80% training, 10% validation, and 10% test split, while Sections 2.3 and 2.4 report all results as 5-fold cross-validation with means and standard deviations over five test folds. Under a standard 5-fold protocol the test portion would be 20% of the data per fold (with validation carved from the training portion), whereas the stated 80/10/10 split implies a single test set. Please specify exactly how the five folds were constructed, whether the five test folds are disjoint at the patient level, and how validation data were created within each fold. If the test folds are not disjoint, the reported fold-level standard deviations and any fold-based significance tests are not valid.","section":"Section 2.2 and Sections 2.3/2.4"},{"comment":"The temporal alignment between each OCT volume and its 'corresponding' visual field test is not specified. If the OCT and VFT were not acquired at the same visit, glaucoma progression between the two acquisitions would add uncorrelated noise to the structure-function relationship and would attenuate the achievable correlation; if they were acquired at the same visit, this should be stated explicitly. Please report the time window between OCT and VFT acquisition (e.g., same visit, or median and range in days) and discuss how the interval affects the interpretation of the reported Pearson correlations.","section":"Section 2.2"}],"minor_comments":[{"comment":"The RMSE formula is incorrect as printed: it should be RMSE = sqrt((1/n) * sum_i (x_i - y_i)^2). The current expression omits the division by n inside the square root and includes an extra leading factor of 1/n. Please correct the equation and confirm that the RMSE values in Table 3 were computed with the correct formula.","section":"Section 2.1, Eq. (1)"},{"comment":"There are several typographical errors: Table 1's heading 'VIF' should be 'VFI', Table 2's 'Person Correlation' should be 'Pearson Correlation', and the text contains 'aquired', 'Similarily', 'metricies', and 'An method'. These should be corrected.","section":"Tables 1 and 2"},{"comment":"The scatter plot reports PC = 0.93 on the entire data set, including training and validation samples. The text explains this, but the figure and caption should label it clearly as a whole-data correlation, not a test-set performance measure, to avoid misleading readers.","section":"Section 3, Figure 2"},{"comment":"The class activation maps (CAMs) are said to be computed following Zhou et al., but the standard CAM formulation is defined for classification networks with a score before softmax. Since the proposed network is a regressor with a tanh output, please describe how the CAMs were computed for the regression setting and whether the highlighted regions were validated quantitatively.","section":"Section 3.1"},{"comment":"The classical baselines underwent hyperparameter tuning via 100 random samples per training fold, while the CNN appears to use a fixed architecture and training schedule. Please state whether the CNN hyperparameters were also tuned on the validation set or whether the comparison is between a tuned classical pipeline and a fixed CNN architecture; this context affects how the performance gap should be interpreted.","section":"Sections 2.3 and 2.4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses a clinically relevant question, and the patient-wise fold design is a genuine strength. The main barrier to acceptance is that the headline superiority claim (p < 0.01) is not verifiable from the reported methods, and the cross-validation protocol is ambiguous. I would ask the authors for a clustering-corrected significance test (or a softened claim), a precise description of the fold construction, and the OCT-VFT time interval. No concerns about circularity or data integrity are apparent from the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's actual contribution is modest but real: a 3D CNN regresses aggregate VFI and MD directly from raw ONH or macula OCT volumes, and in patient-wise 5-fold cross-validation it beats classical ML trained on the usual 22 segmentation features (PC 0.88 vs 0.74 for VFI on ONH). The comparison is internally consistent, the classical baselines are reasonably extensive, and the patient-wise fold split is the right call. The discussion of the retest variability upper bound (Abramoff's 0.88) is honest and puts the result in proper context. The CAM analysis is a nice addition, though not load-bearing.\n\nThe soft spots are real. The headline 'significantly higher (p < 0.01)' is not supported by any described statistical test. With 4155 scans from only 579 patients, roughly seven visits per patient, within-patient correlation in both OCT and VFI/MD is unavoidable. The stress-test note is right: if the p-value comes from a paired test on five fold-level PCs, n=5 is tiny and the covariance is unreported; if it comes from pooling all scans, the effective sample size is inflated. Either way, the comparative claim is not established as stated. This is the paper's main flaw.\n\nTwo softer issues: the temporal alignment between OCT and VFT is never specified, and if the gap is weeks or months it adds noise to both the regression and any clinical interpretation. And the RMSE of 12 on VFI (0-100 scale) is clinically large; calling the estimates 'accurate' oversells them. The paper acknowledges the retest correlation ceiling but not the error tolerance question.\n\nThe novelty relative to prior work (Bogunovic, Guo, Sugimoto) is real but incremental: they predict pointwise thresholds from segmented maps, this predicts aggregate indices from raw volumes. That is a reasonable next step. The exclusion of map-based baselines from the comparison is a limitation but not a fatal one.\n\nFor peer review: yes, this deserves serious refereeing, because the question matters and the design is mostly sound. But it needs a major revision: a cluster-robust or bootstrap significance test, a statement of OCT-VFT time alignment, and a more measured conclusion. If those are fixed, the core finding—that a CNN on raw OCT captures structure-function correlation near the retest ceiling—is worth publishing.","headline":"Plausible internal result with a real statistical hole: the p<0.01 superiority claim is not supported as described, and the VFI RMSE of 12 undercuts the 'accurate' language.","tokens_in":9098,"tokens_out":2074,"would_cite":false,"duration_ms":23846,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 3D CNN estimates visual field indices from one raw OCT volume with 0.88 correlation.","keywords":["glaucoma","optical coherence tomography","visual field test","deep learning","3D convolutional neural network","Visual Field Index","Mean Deviation","structure-function correlation"],"falsifier":"Take a new dataset in which each OCT volume is acquired within one day of its visual field test and retrain and evaluate the same CNN; if the test-fold Pearson correlation falls to the 0.74 random-forest level or below, the reported 0.88 is an artifact of temporal mismatch rather than genuine structure-function inference.","tokens_in":8056,"feed_emoji":"👁️","tokens_out":5683,"duration_ms":58006,"temperature":0.7,"pith_summary":"The paper asks whether two aggregate glaucoma measures from visual field tests—Visual Field Index (VFI) and Mean Deviation (MD)—can be estimated from a single raw OCT volume instead of a full perimetry exam. It proposes a 3D convolutional neural network that processes down-sampled optic nerve head or macula scans without layer segmentation or multi-field registration. On optic nerve head scans the network reaches a Pearson correlation of 0.88 with measured VFI, compared with 0.74 for the best classical regressor trained on standard segmentation-derived features. If the result holds, a quick OCT scan could substitute for a slow, variable visual field test in routine glaucoma monitoring, reducing examination time and cost.","feed_headline":"One OCT scan predicts visual field loss at 0.88 correlation","feed_subtitle":"Glaucoma's VFI and MD inferred directly from raw scan volumes, near the test-retest limit of perimetry.","key_machinery":"The central object is a five-layer 3D convolutional neural network with batch normalization, spatial dropout, global average pooling, and a tanh regression head, applied to raw OCT volumes down-sampled to 64x64x128 voxels. The global average pooling layer both produces the regression output and supports class activation maps. The comparison machinery is a set of classical regressors, led by a random forest, trained on the 22 optic nerve head features extracted by the OCT scanner, such as peripapillary retinal nerve fiber layer thickness at clock-hours and quadrants, cup-to-disc ratios, and cup volume.","core_discovery":"The paper reports that aggregate functional status in glaucoma can be read directly from structure: a 3D CNN regresses VFI and MD from a single raw OCT volume. On optic nerve head scans it achieves a Pearson correlation of 0.88 for both VFI and MD, and on macula scans 0.86 and 0.85 respectively, all above the best classical baseline. The paper also notes that this 0.88 matches the retest variability of visual field tests, suggesting the network is operating near the noise ceiling of its training labels. Class activation maps show the network concentrating on the ganglion cell and inner plexiform layers, with more focal attention in eyes with greater vision loss.","pith_inferences":["Inference: if the OCT and visual field measurements in this dataset were typically acquired weeks or months apart, the reported 0.88 correlation would overstate the true structure-function accuracy, and a same-day paired dataset should be collected to measure the gap.","Inference: matching the 0.88 retest ceiling suggests further gains will come less from bigger networks than from more reliable functional endpoints or from longitudinal models that pool multiple perimetry sessions.","Inference: the class activation map pattern can be turned into a testable prediction: masking the highlighted GCIPL regions should degrade VFI inference much more than masking other retinal layers.","Inference: combining optic nerve head and macula volumes, which the paper says it did not evaluate, is the most direct route to seeing whether 0.88 is a ceiling or just a single-scan limit."],"forward_implications":["When the model is trained on optic nerve head scans, inferred VFI and MD both reach a Pearson correlation of 0.88, which the paper notes is close to the retest variability of visual field tests.","Macula scans support nearly the same accuracy (VFI 0.86, MD 0.85), so the approach does not depend on a particular scan location.","Because the CNN operates on raw down-sampled volumes, it avoids layer segmentation and nine-field registration that earlier threshold-estimation methods require.","The 3D CNN is clearly more accurate than every classical regressor trained on the 22 standard segmentation-based OCT features, whose best VFI correlation is 0.74 (random forest).","Class activation maps show the network attending to the ganglion cell and inner plexiform layers, with more focal attention in eyes with greater vision loss."],"supporting_citations":[{"why":"Supplies the base convolutional architecture from which the proposed 3D CNN is derived.","marker":"[31]"},{"why":"Provides the classical machine-learning implementations used as comparison baselines.","marker":"[25]"},{"why":"Defines the 22 segmentation-based OCT features used to train the classical baseline regressors.","marker":"[24]"},{"why":"Supplies the visual field test retest variability of 0.88 that the paper uses as an accuracy ceiling.","marker":"[16]"},{"why":"Earlier OCT-based visual field threshold prediction that requires nine-field registration and segmentation, which the paper contrasts with its raw-volume approach.","marker":"[17]"},{"why":"Earlier SVM-based threshold estimation from RNFL and GCIPL thickness maps, another structure-function baseline.","marker":"[15]"},{"why":"Provides the global average pooling and class activation map method used to visualize which OCT regions drive VFI inference.","marker":"[34]"},{"why":"Defines the Visual Field Index and its range, grounding the target variable.","marker":"[21]"},{"why":"Sets the reliability criteria used to discard unreliable visual field tests from the dataset.","marker":"[23]"}],"fun_headline_variants":["Deep learning predicts visual field loss from a single OCT scan","OCT scan alone yields 0.88 correlation with visual field index","AI maps glaucoma vision loss directly from optic nerve head scans","One OCT volume infers VFI and MD: correlation hits 0.88","3D CNN reads glaucoma function from raw OCT, near test-retest"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes each OCT volume and its paired visual field measurement describe the same disease state at essentially the same time, but it does not report the time interval between the two tests.","fun_headline_variants_meta":{"raw":{"variants":["Deep learning predicts visual field loss from a single OCT scan","OCT scan alone yields 0.88 correlation with visual field index","AI maps glaucoma vision loss directly from optic nerve head scans","One OCT volume infers VFI and MD: correlation hits 0.88","3D CNN reads glaucoma function from raw OCT, near test-retest"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000218,"raw_usage":{"total_tokens":1415,"prompt_tokens":893,"completion_tokens":522,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":430}},"tokens_in":509,"tokens_out":522,"duration_ms":5901,"temperature":1.0,"reasoning_tokens":430,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:13:44.164752+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a new dataset in which each OCT volume is acquired within one day of its visual field test and retrain and evaluate the same CNN; if the test-fold Pearson correlation falls to the 0.74 random-forest level or below, the reported 0.88 is an artifact of temporal mismatch rather than genuine structure-function inference.","supporting_citations":[{"cited_title":"Maetschke, B","cited_arxiv_id":null,"evidence_quote":"Supplies the base convolutional architecture from which the proposed 3D CNN is derived."},{"cited_title":"Pedregosa, G","cited_arxiv_id":null,"evidence_quote":"Provides the classical machine-learning implementations used as comparison baselines."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the 22 segmentation-based OCT features used to train the classical baseline regressors."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the visual field test retest variability of 0.88 that the paper uses as an accuracy ceiling."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier OCT-based visual field threshold prediction that requires nine-field registration and segmentation, which the paper contrasts with its raw-volume approach."},{"cited_title":"Bogunovi´ c, Y","cited_arxiv_id":null,"evidence_quote":"Earlier SVM-based threshold estimation from RNFL and GCIPL thickness maps, another structure-function baseline."},{"cited_title":"Bengtsson, A","cited_arxiv_id":null,"evidence_quote":"Defines the Visual Field Index and its range, grounding the target variable."},{"cited_title":"Yaqub, Visual ﬁelds interpretation in glaucoma: a focus on static auto- mated perimetry, Community eye health 25 (79-80) (2012) 1","cited_arxiv_id":null,"evidence_quote":"Sets the reliability criteria used to discard unreliable visual field tests from the dataset."}],"review_version":1}