{"id":"80b2561e-702f-4f99-a1df-384e9e3024a2","arxiv_id":"2506.17262","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Adding IOP-induced strain to morphology improved AI classification of one superior visual field defect pattern (arcuate) by 0.04 AUC, and saliency maps localized the predictive regions to the inferior optic nerve rim.","lead":"A study of 237 glaucoma patients used 3D eye scans taken at two eye pressures to train an AI that classifies three patterns of vision loss and to highlight which optic nerve head regions the AI relies on. The AI gained a small predictive boost from adding tissue strain for one defect pattern, and its attention concentrated on the lower rim of the optic nerve rather than on the deeper lamina cribrosa.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Saliency maps average gradients over strain, thickness, and spatial features, so labeling high-gradient rim regions as 'strain-sensitive' is unsupported; a strain-only gradient control would settle it.","rationale":"The central quantitative ablation (AUC 0.87 vs 0.83, p<0.05) is a legitimate but fragile result; the reader's concerns about small positives, no validation set, and final-iteration evaluation are valid. I do not find a separate fatal flaw there beyond what the reader noted. The most load-bearing issue is the semantic gap between mixed-feature saliency and 'strain-sensitive regions.' This is not just an interpretability nicety: the paper's second objective and its clinical conclusion (rim strain dominates over LC) depend on attributing saliency to strain. The authors' own limitation statement concedes the model cannot disentangle strain from thickness. The test I propose directly checks whether the published saliency pattern is reproducible from strain gradients alone. Because the manuscript is honest about this limitation and the AUC ablation is reported, the result should remain CONDITIONAL rather than being rejected outright; if the control fails, the localization and conclusion should be revised, while the predictive ablation could stand. Agreement with reader: their weakest_assumption is the same concern.","tokens_in":10504,"tokens_out":3827,"duration_ms":46956,"concrete_test":"Compute saliency maps using gradients w.r.t. the strain feature only (set gradients for spatial coordinates and tissue thickness to zero) and compare with the combined maps in Figure 2d-f. If the strain-only maps do not reproduce the inferior/inferotemporal rim arc, the 'strain-sensitive region' claim fails. As a complementary control, run the identical saliency pipeline on a morphology-only PointNet (no strain inputs); if its maps are substantially similar to the published ones, the localization is morphology-driven, not strain-driven.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's second central claim—that the inferior/inferotemporal neuroretinal rim is the key 'strain-sensitive' region—is not supported by the saliency evidence as presented. In Methods (Explainable AI), saliency maps are 'gradient magnitudes across spatial, structural, and strain features averaged' per point. High gradient magnitude in such a mixed-feature map cannot be attributed to strain: it can equally arise from tissue-thickness variation, boundary shape, or spatial-coordinate gradients. The authors acknowledge this in Discussion, limitations: PointNet 'has limited ability to disentangle the regional contributions of individual features, such as strain versus tissue thickness, within the saliency map.' Yet the Results and Abstract convert these mixed saliency maps into 'strain-sensitive regions' and conclude 'strain at the rim could play a dominant role.' There is also a units/scale issue: simply averaging gradient magnitudes over spatial coordinates (mm), thickness (mm), and dimensionless strain is not scale-invariant; whichever feature has the largest raw gradients will dominate the map. If rim thickness (which is tightly correlated with the same VF labels) drives the saliency, the localization conclusion collapses even though the AUC ablation for arcuate prediction may remain true. The cross-sectional design further means 'progressive expansion' is across severity groups, not measured progression.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a geometric deep learning pipeline (PointNet) that takes 3D optic nerve head (ONH) point clouds, with morphological features and IOP-induced effective strain from digital volume correlation, to classify three patterns of superior visual field loss in 237 glaucoma subjects. The authors report AUCs of 0.77–0.88, and a five-split ablation for the superior arcuate task shows a small but statistically significant improvement when strain features are included (0.87±0.02 vs. 0.83±0.02, p<0.05). Saliency maps are used to identify high-gradient regions, which the paper labels as \"strain-sensitive regions,\" and the authors conclude that the inferior and inferotemporal neuroretinal rim is the most critical region, with the arc of importance expanding across nasal step, arcuate, and hemifield defect groups.","tokens_in":10662,"tokens_out":4235,"duration_ms":48673,"significance":"If the claims are validated, the work would provide a clinically plausible link between acute IOP-induced ONH deformation and region-specific functional loss, and it would point to the neuroretinal rim rather than the lamina cribrosa as a biomechanically vulnerable site. The study has genuine strengths: a well-characterized clinical cohort, direct measurement of IOP-induced strain via DVC, a morphology-only ablation that is not circular, and a spatially resolved analysis of model predictions. However, the main strain-benefit evidence is limited to one of three tasks and rests on a fragile statistical protocol, and the central localization claim is undermined by the mixed-feature saliency computation, which the authors themselves acknowledge cannot separate strain from morphology. The cross-sectional design also does not support the \"progressive expansion\" language. These issues currently limit the strength of the conclusions.","major_comments":[{"comment":"The saliency maps are computed as gradient magnitudes averaged across spatial, structural, and strain features, and the manuscript then defines \"regions of high gradient\" as \"strain-sensitive regions.\" Because the gradient is a mixed-feature quantity, a high value in the inferior rim could reflect tissue-thickness variation, boundary shape, or coordinate-scale effects rather than strain; the authors explicitly acknowledge this in the Discussion (\"PointNet has limited ability to disentangle the regional contributions of individual features, such as strain versus tissue thickness, within the saliency map\"). Consequently, the second central claim—that the inferior and inferotemporal neuroretinal rim is the key strain-sensitive region—is not supported by the presented evidence. Please provide strain-only attribution maps or per-feature saliency, or restrict the conclusions to \"regions important for prediction\" rather than \"strain-sensitive regions.\"","section":"Methods, Explainable AI"},{"comment":"The claim that \"ONH strain improved VF loss prediction beyond morphology alone\" is supported only for the superior arcuate task (AUC 0.87±0.02 vs. 0.83±0.02, p<0.05), not for the nasal step or hemifield tasks, yet the abstract and conclusion present it as a general result. The protocol also reports no validation set, hyperparameters inherited from prior work, and final-iteration weights, and the p-value is based on only 62 positive cases across five splits. Please report strain ablations for all three tasks, specify the exact statistical test and dispersion (e.g., paired test, confidence intervals), and qualify the claim to the arcuate pattern unless additional evidence is provided.","section":"Results, Incorporation of Effective Strains"},{"comment":"The study is cross-sectional, but the Results state that the arc \"increased in length as the defect progressed from a nasal step to an arcuate pattern and, ultimately, to full hemifield loss,\" and the Discussion uses \"as glaucoma progresses.\" Because the three defect groups are severity categories, not longitudinal observations, \"progressive expansion\" and progression-related causal language are not supported by the data. Please replace these with severity-associated language, such as \"the arc length was larger in more advanced defect categories.\"","section":"Results, Explainable AI Reveals an Arching Pattern"}],"minor_comments":[{"comment":"The abstract reports peak AUCs of 0.77–0.88, while the Results first report 0.88 for superior partial arcuate defects and later report 0.87±0.02 for the same task; please reconcile these numbers and make clear whether 0.88 is a single-split value or a mean.","section":"Abstract and Results"},{"comment":"References 17 and 38 appear to cite the same paper (Chuangsuwanich et al., \"Biomechanics-Function in Glaucoma\"); please deduplicate or clarify if they are intended as distinct works.","section":"References"},{"comment":"The caption states \"Red regions in the nerve fiber defect maps denote areas of nerve loss,\" but the panels appear to show visual field pattern deviation maps; please clarify the anatomical or functional nature of the displayed maps and ensure the color scheme is legible in print.","section":"Figure 2 caption"},{"comment":"The text uses \"ophthalmo-dynamometry\" in the abstract and \"ophthalmodynamometer\" in the Methods; please standardize the terminology. Also, \"primarily gaze\" should be \"primary gaze.\"","section":"Methods, Classification of Visual Field Defects"}],"recommendation":"major_revision","confidential_remarks":"The paper's central claims are interesting and potentially impactful, but the localization claim needs a per-feature attribution analysis or a substantial softening, and the strain-benefit claim needs a more rigorous evaluation protocol. The authors' own limitation paragraph already concedes the key attribution problem, so the revision path is clear but requires new experiments rather than text changes alone."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper's real contribution is the strain ablation for the arcuate visual field defect (AUC 0.87 vs 0.83 across five splits, p<0.05). The saliency localization to the inferior/inferotemporal rim is presented as the other headline result, but the saliency maps average gradients over spatial, structural, and strain features, and the authors admit in the Discussion that PointNet cannot separate strain from tissue thickness. Calling the high-gradient rim regions 'strain-sensitive' is therefore not supported by the evidence as presented. That is a soft spot, not a fatal one—the classification ablation stands on its own.\n\nCredit where due: the cohort is non-trivial for this niche (237 subjects, OCT under acute IOP elevation via ophthalmodynamometry), the writing is clear, and the limitations section is unusually candid. They also correctly note the cross-sectional design, so 'progressive expansion' with severity is a between-group comparison, not measured progression.\n\nSoft spots in proportion. First, the saliency issue is the main one. A strain-only gradient control would settle whether the maps actually reflect strain. Without it, the 'strain-sensitive regions' headline should be downgraded to 'regions that contribute to the classifier.' There is also a scale problem: averaging gradient magnitudes over spatial coordinates (mm), thickness (mm), and dimensionless strain is not scale-invariant, so whichever feature has the largest raw gradients will dominate the map. Second, the ML protocol is fragile: no validation set, hyperparameters inherited from prior work, and final-iteration weights evaluated on the test set. The arcuate ablation uses 62 positive cases; the other two tasks (26 and 25 positives) have no reported ablation and no error bars—just 'peak AUCs' of 0.77 and 0.87. That needs fixing before the broader claims about three defect types can be trusted.\n\nIs there a load-bearing flaw? Not in the arcuate ablation itself, but the localization claim is overstated in the abstract and conclusion. That should be toned down or supported with a better saliency method.\n\nWho should read this: glaucoma researchers interested in biomechanics-function relationships, and anyone working on interpretability for point-cloud medical data. The paper deserves a serious referee; the data is hard to acquire, the authors are honest, and the ablation is a genuine empirical increment, even if the protocol needs tightening.\n\nRecommendation: send it to peer review, with the expectation that the authors add error bars for all tasks, run the ablation for all three defect types, and either add a strain-only saliency control or soften the 'strain-sensitive' language.","headline":"Solid increment with one overstated headline: the arcuate strain ablation is real, but the saliency maps don't isolate strain.","tokens_in":11354,"tokens_out":3589,"would_cite":false,"duration_ms":36432,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"IOP-induced strain in the optic nerve head sharpens AI prediction of glaucoma vision loss.","keywords":["glaucoma","optic nerve head biomechanics","effective strain","digital volume correlation","geometric deep learning","PointNet","explainable AI","visual field loss"],"falsifier":"Train the same PointNet classifier using only morphological features, generate saliency maps, and compare them with the strain-inclusive maps: if the inferior and inferotemporal rim remains the dominant high-saliency region without any strain input, then the strain-sensitive localization conclusion does not follow from the saliency analysis.","tokens_in":10183,"feed_emoji":"👁","tokens_out":3492,"duration_ms":41406,"temperature":0.7,"pith_summary":"The paper claims that adding intraocular-pressure-induced biomechanical strain of the optic nerve head (ONH) to structural morphology improves AI prediction of specific glaucoma visual field loss patterns, with the clearest gain for superior partial arcuate defects (AUC 0.87 versus 0.83 without strain, p<0.05). It further claims that explainable AI consistently localizes the class-driving evidence to the inferior and inferotemporal neuroretinal rim, where the high-importance region expands as defects progress from nasal step to arcuate to full hemifield loss. If true, this means local tissue deformation, not just tissue shape, carries clinically relevant information about where glaucoma injures vision, and the neuroretinal rim, rather than the lamina cribrosa, emerges as the biomechanically vulnerable site. A sympathetic reader would care because it points toward biomechanics-informed risk assessment and region-specific monitoring of glaucomatous damage.","feed_headline":"Strain data sharpens glaucoma vision-loss AI predictions","feed_subtitle":"Adding IOP-induced tissue strain lifts arcuate-defect AUC from 0.83 to 0.87 (p<0.05), with the rim doing the work.","key_machinery":"The central machinery is a pipeline that converts OCT segmentations of the ONH into 3D point clouds, computes IOP-induced effective strain at each point using digital volume correlation, and feeds these point clouds into separate PointNet classifiers for each visual field defect type. Saliency is quantified by averaging gradient magnitudes across spatial, structural, and strain features, and the resulting 3D saliency values are sum-projected onto an en-face BMO-centered grid to define strain-sensitive regions.","core_discovery":"The paper establishes that ONH tissue strain enhances prediction of glaucomatous visual field loss patterns beyond morphology alone, and that the neuroretinal rim, rather than the lamina cribrosa, is the most critical region contributing to these model predictions. This is supported by a sensitivity analysis showing significantly higher AUC for superior arcuate defects with strain features included, and by saliency maps that show an arching pattern in the inferior and inferotemporal rim that lengthens with increasing disease severity.","pith_inferences":["A direct test of the localization claim would be to train the same model without strain features and compare saliency maps: if the inferior rim remains dominant, the maps reflect morphology rather than strain sensitivity.","Because the paper uses only effective strain, it cannot distinguish tension from compression or shear; future models including the full strain tensor or stress fields may relocate or refine the key regions.","The cross-sectional design leaves open whether strain maps predict concurrent rather than future damage; longitudinal biomechanical testing would be needed to establish a causal or progressive link.","The patient-specific saliency maps could in principle be combined with the Garway-Heath map to generate individualized structure-function predictions, but this is an extension the paper does not itself pursue."],"forward_implications":["Adding effective strain to morphology yields a statistically significant AUC gain for superior arcuate defects, showing biomechanical features have predictive value beyond structure.","The inferior and inferotemporal neuroretinal rim is repeatedly identified as the most salient region across all three classification tasks, which suggests a consistent anatomical focus for early axonal injury.","The high-saliency arc lengthens as defects progress from nasal step to arcuate to full hemifield loss, implying that more severe damage recruits broader regions of the rim.","The comparatively low saliency of the lamina cribrosa indicates that either its biomechanical role is less directly tied to these visual field patterns or that current OCT resolution does not capture the relevant LC features."],"supporting_citations":[{"why":"Supplies the established digital volume correlation protocol for computing IOP-induced effective strain in the ONH.","marker":"[3]"},{"why":"Earlier work by the authors showing that IOP-induced neural strains improve visual field predictions; this paper extends that approach to specific defect patterns.","marker":"[17]"},{"why":"Provides the visual field defect classification scheme used to group patients into the four categories.","marker":"[18]"},{"why":"Establishes the geometric deep learning and point-cloud saliency framework for identifying critical structural features of the optic nerve head.","marker":"[25]"},{"why":"The Garway-Heath map is used as the external structure-function reference against which the data-driven saliency patterns are compared.","marker":"[33]"}],"fun_headline_variants":["Rim strain, not lamina, predicts glaucoma vision loss","AI pinpoints rim strain behind glaucoma functional loss","Optic nerve rim strain boosts glaucoma loss predictions","Inferior rim strain key to glaucoma defect forecasting","Strain-sensitive rim regions flag glaucoma progression"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that the identified regions are truly strain-sensitive rests on treating gradient magnitudes averaged across spatial, structural, and strain features as a measure that isolates biomechanical strain, even though the PointNet architecture cannot fully disentangle strain from morphology such as tissue thickness.","fun_headline_variants_meta":{"raw":{"variants":["Rim strain, not lamina, predicts glaucoma vision loss","AI pinpoints rim strain behind glaucoma functional loss","Optic nerve rim strain boosts glaucoma loss predictions","Inferior rim strain key to glaucoma defect forecasting","Strain-sensitive rim regions flag glaucoma progression"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000163,"raw_usage":{"total_tokens":1284,"prompt_tokens":1025,"completion_tokens":259,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":641,"completion_tokens_details":{"reasoning_tokens":185}},"tokens_in":641,"tokens_out":259,"duration_ms":3977,"temperature":1.0,"reasoning_tokens":185,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:24:20.022185+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same PointNet classifier using only morphological features, generate saliency maps, and compare them with the strain-inclusive maps: if the inferior and inferotemporal rim remains the dominant high-saliency region without any strain input, then the strain-sensitive localization conclusion does not follow from the saliency analysis.","supporting_citations":[{"cited_title":"Differing Associations between Optic Nerve Head Strains and Visual Field Loss in Normal-and High-Tension Glaucoma Subjects","cited_arxiv_id":null,"evidence_quote":"Supplies the established digital volume correlation protocol for computing IOP-induced effective strain in the ONH."},{"cited_title":"Biomechanics-Function in Glaucoma: Improved Visual Field Predictions from IOP-Induced Neural Strains","cited_arxiv_id":null,"evidence_quote":"Earlier work by the authors showing that IOP-induced neural strains improve visual field predictions; this paper extends that approach to specific defect patterns."},{"cited_title":"Pattern of visual field loss in primary angle- closure glaucoma across different severity levels","cited_arxiv_id":null,"evidence_quote":"Provides the visual field defect classification scheme used to group patients into the four categories."},{"cited_title":"Geometric deep learning to identify the critical 3D structural features of the optic nerve head for glaucoma diagnosis","cited_arxiv_id":null,"evidence_quote":"Establishes the geometric deep learning and point-cloud saliency framework for identifying critical structural features of the optic nerve head."},{"cited_title":"Mapping the visual field to the optic disc in normal tension glaucoma eyes","cited_arxiv_id":null,"evidence_quote":"The Garway-Heath map is used as the external structure-function reference against which the data-driven saliency patterns are compared."}],"review_version":1}