{"id":"2a04d852-8a70-493d-a0ec-99451c5c9491","arxiv_id":"2505.09993","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A fine-tuned ResNet-101 classifies ovarian tissue brightfield images into five FIGO-like stages with 97.62% test accuracy, but patient-level independence and reproducibility are unverified.","lead":"This paper trains a ResNet-101 neural network to sort microscope images of ovarian tissue into five cancer stages, reporting 97.62% accuracy on a 126-image test set. The result hints that deep learning could support pathologists, but the small, imbalanced dataset and missing code make the claim hard to verify.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline accuracy cannot be read as evidence of staging until patient-level independence and the absence of batch shortcuts are established; no patient or core metadata is provided.","rationale":"The reader's verdict is CONDITIONAL with high correctness risk, and their weakest-assumption statement centers on whether FIGO stage is learnable from a thin core. This pass identifies a sharper, testable version of the same concern: the 97.62% figure is uninterpretable until patient-level independence is demonstrated. This is load-bearing because every downstream claim, including morphological feature learning and clinical utility, depends on the test accuracy reflecting stage-associated tissue content rather than image-level artifacts. The manuscript provides no patient-level metadata and no external validation, and the limitations section itself concedes that multi-institutional validation is needed. The proposed check is feasible because the authors retain the raw data; they can produce the manifest and rerun a patient-disjoint evaluation. If the accuracy survives at a similar level, the result would justify the conditional acceptance; if it drops substantially, the main claim is unsupported. This does not change the reader's verdict, but it identifies the concrete experiment that would settle the central question.","tokens_in":8719,"tokens_out":7740,"duration_ms":79400,"concrete_test":"Obtain the full data manifest (patient ID, TMA block/core ID, scan batch, image source) from the authors. Recompute accuracy under a strict patient-disjoint split: put all images from each patient either in training or in test, and select hyperparameters only on a patient-disjoint validation fold. If the patient-disjoint accuracy is materially below 97.62%, or if no manifest can be produced, the headline result is not evidence of automated staging. Additionally, for each test patient with multiple cores/images, compute per-patient majority-vote accuracy; large within-patient disagreement would show the model is not capturing patient-level stage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central result is the 97.62% test accuracy (Sections 2.2 and 3.1). For that number to support the paper's claim that the model has learned stage-associated morphological features, the 126-image test set must be independent of the 2418 training images at the patient/core level, and the model must not be exploiting acquisition or staining covariates. Neither is established. Section 3.1 reports only image counts and describes the primary collection as split 80/20 into training and validation (Appendix A); no patient, block, core, or scan-session identifiers are reported. Since each TMA slide contains 24 cores and the scanner captures many images per slide in an automated run, images from the same core, adjacent sections, or the same batch can easily straddle a nominal image-level train/test split. A CNN can then reach near-perfect accuracy by memorizing tile-specific or batch-specific texture rather than stage-determining morphology. The test set is separately sourced, which helps, but without a demonstrated patient-disjoint split the decisive confound remains. The authors' own limitations paragraph concedes that multi-institutional validation is still needed, yet the conclusion is stated categorically. This is the load-bearing weakness: if patient independence fails, the staging claim collapses; if it holds, the biological learnability question must still be tested externally.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a deep learning framework for automated ovarian cancer staging (classes 0, I, II, III, IV) from transmission optical microscopy bright-field images of thin tissue cores. Using a ResNet-101 initialized with ImageNet weights, the authors fine-tune on 2418 images (80/20 stratified into training/validation) and evaluate on an independent test set of 126 images, reporting an overall accuracy of 97.62% (123/126 correct). The training pipeline includes data augmentation, weighted random sampling and class-weighted cross-entropy loss, with learning rate, dropout rate, and weight decay tuned via a genetic algorithm. The paper reports per-class precision, recall, F1-scores, a confusion matrix, and describes misclassifications as primarily occurring between adjacent stages.","tokens_in":8890,"tokens_out":5138,"duration_ms":51786,"significance":"If the central result holds, the paper would contribute a practical automated staging tool for ovarian cancer histopathology, a domain where most prior work has focused on grading, subtyping, or survival prediction rather than direct FIGO stage classification from bright-field images. The use of a held-out test set, per-class metrics, explicit class-imbalance handling, and transparent descriptions of the augmentation and hyperparameter search are strengths. However, the significance is tempered by two load-bearing concerns: the independence of the test set from the training set at the patient and tissue-core level is not established, and the biological plausibility of deriving surgical FIGO stage from the morphology of a single 1.5 mm core is not justified. These issues must be resolved before the reported accuracy can be interpreted as evidence of genuine stage-associated morphological learning.","major_comments":[{"comment":"The independence of the test set is not demonstrated at the patient or tissue-core level. The manuscript reports only image counts and states that the training/validation and test images come from 'two distinct collections,' but no patient, block, core, or slide identifiers are provided. Because each TMA slide contains 24 cores and scanning is automated, images from the same core, adjacent sections, or the same scanning session could straddle the train/test boundary. Without a patient-disjoint split, the reported 97.62% accuracy (123/126) may reflect memorization of tile-specific or batch-specific texture rather than stage-associated morphology. The authors should provide patient-level metadata or re-analyze with a patient-exclusive split to support the claim of an 'independent' test set.","section":"Section 3.1 and Appendix A"},{"comment":"The text states that misclassifications 'primarily occurred between adjacent stages,' but it also reports that one actual Class 3 sample was misclassified as Class 1, which is a non-adjacent error (a jump of two stages). This internal inconsistency affects the interpretation of the confusion matrix and should be corrected, either by revising the statement or by re-checking the confusion matrix.","section":"Section 2.3"},{"comment":"The premise that FIGO surgical stage is learnable from the morphology of a single 1.5 mm tissue core imaged in brightfield is biologically questionable. FIGO staging is defined by the extent of tumor spread found at surgery (e.g., pelvic versus upper abdominal involvement), which is not necessarily present in a thin core. If the model is not actually capturing stage-defining morphology, the high accuracy may reflect spurious correlations with acquisition or staining covariates. The manuscript should either provide evidence that stage-discriminating morphology is present in these cores (e.g., comparison with pathologist assessment or supporting references specific to ovarian core samples) or temper the claim to 'image-based classification of biopsy tissue' rather than 'staging.' This is a correctness-risk concern that directly bears on the external validity of the central claim.","section":"Sections 3.1 and 3.3"}],"minor_comments":[{"comment":"The title mentions 'grading and staging,' but the model predicts only stage (five classes, with class 0 as control); no separate grading output is reported. Consider revising the title to 'staging' or clarifying that grading is subsumed in the staging assessment.","section":"Title"},{"comment":"The description of the independent test set is unclear: 'obtained from BX61' presumably refers to the same Olympus BX61 microscope used for the primary collection, but it should specify whether the test images come from different TMA slides, different patients, or different scanning sessions, and how many patients are represented in each split.","section":"Section 3.1"},{"comment":"The manuscript reports only image counts and does not provide the number of patients, cores, or images per core. This information is necessary for the reader to assess the effective sample size and the degree of data dependence in the splits.","section":"Section 3.1 and Appendix A"},{"comment":"The optimal dropout rate (0.303) lies at the boundary of the GA search range [0.3, 0.7]. This suggests that the optimum may lie outside the defined range, and the search space should be extended to ensure that the hyperparameter choice is not an artifact of the chosen bounds.","section":"Section 3.5"},{"comment":"The manuscript does not provide code, trained model weights, or a detailed data availability statement beyond 'available from the corresponding author upon reasonable request.' Adding a repository or explicit data-sharing plan would improve reproducibility, which is particularly important given the need to verify the patient-level independence of the splits.","section":"Section 4 and Data Availability"}],"recommendation":"major_revision","confidential_remarks":"The central issue is the lack of any patient-level or core-level identifiers; this is a common but serious omission in histopathology AI work. If the authors can provide or re-analyze the data with a patient-disjoint split, the paper could be publishable. The biological plausibility concern is also worth raising in the revision, but it may be addressable by softening the claim to 'classification of biopsy images' rather than 'staging' in a clinical sense. The manuscript is otherwise straightforward and would be of interest to the computational pathology community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a routine fine-tuning of ResNet-101 on ovarian tissue brightfield images, and the main number — 97.62% on 126 test images — is not supported as evidence of automated staging until two things are shown. First, that the independent test set is disjoint from the training set at the patient or core level. The paper reports image counts only. Since each TMA slide holds 24 cores and scanning is automated, adjacent tiles from the same core or batch can easily straddle a nominal split. A CNN can hit near-perfect accuracy by memorizing tile- or batch-specific texture rather than stage-defining morphology. Second, that a 1.5 mm core actually contains the morphological information that defines FIGO stage. FIGO staging is defined by the extent of tumor spread found at surgery; a thin core may not show that. If either point fails, the 97.62% is not staging.\n\nWhat the paper does well: it uses a sensible transfer-learning setup, with augmentation, class weighting, and a genetic algorithm for hyperparameters. It reports a confusion matrix and per-class metrics, and the errors cluster between adjacent stages, which is plausible. The authors also acknowledge the need for multi-institutional validation in their limitations paragraph. That counts.\n\nThe soft spots are proportionate to the claim. The central claim is categorical: \"deep learning can achieve high accuracy in discriminating between different stages.\" But the evidence does not establish discriminative features of stage as opposed to correlated artifacts or patient identity. The test set is small and imbalanced — 61 of 126 are Stage I — and there are no error bars. No code or data is released, so the result is not reproducible. The citation pattern is fine; no red flags there.\n\nThis paper is for anyone working in computational pathology who wants to see a standard pipeline applied to a new task. It deserves peer review because the claim is testable and the experiment is reported with enough detail to be checked — but only if reviewers insist on patient-level splitting and external validation. The authors should be asked to show that the test set is patient-disjoint and that the model generalizes across institutions or scanners.\n\nMy recommendation: send it for review with a strong request for those revisions. It's not a desk reject, but it's not close to being accepted as is.","headline":"A standard transfer-learning result whose headline accuracy is not yet evidence of staging; the missing patient-level split and the questionable link between a single tissue core and FIGO stage are the load-bearing weaknesses.","tokens_in":9492,"tokens_out":2258,"would_cite":false,"duration_ms":22749,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A ResNet-101 fine-tuned on brightfield microscopy images of thin ovarian tissue cores assigns FIGO stage (0–IV) with 97.62% accuracy on an independent test set of 126 images.","keywords":["deep learning","brightfield microscopy","histopathology","ovarian cancer staging","FIGO staging","ResNet-101","transfer learning","digital pathology"],"falsifier":"Take the trained model to a cohort scanned with a different microscope, stain batch, or institution, with pathologist-confirmed FIGO labels, and compare accuracy: a collapse toward chance would show the 97.62% result came from acquisition-specific cues rather than stage morphology. A second, cheaper check is to inspect saliency maps: if the model attends to glass edges, background, or stain color patches rather than tissue structures, the stage-learning claim is falsified.","tokens_in":8437,"feed_emoji":"🔬","tokens_out":7623,"duration_ms":72043,"temperature":0.7,"pith_summary":"This paper claims that a convolutional neural network can read the stage of ovarian cancer directly from ordinary brightfield microscopy images of thin tissue cores, without molecular or clinical inputs. On an independent test set of 126 images spanning control tissue and FIGO Stages I–IV, the fine-tuned ResNet-101 classified 123 images correctly, an accuracy of 97.62%, with all three errors falling between adjacent stages. The claim matters because ovarian cancer treatment and prognosis depend heavily on stage, and manual staging is time-consuming and subject to inter-observer variability. If the result holds, automated staging could offer a fast, consistent first-pass read in digital pathology workflows, flagging cases for pathologist review rather than replacing the pathologist.","feed_headline":"AI stages ovarian cancer from tissue images at 97.62%","feed_subtitle":"A fine-tuned CNN labels five FIGO stages from thin brightfield cores, missing only 3 of 126 test images.","key_machinery":"The load-bearing object is a ResNet-101 convolutional neural network pre-trained on ImageNet and fine-tuned for five-way stage classification. Its classification head replaces the original last layer with a dropout layer (rate $0.303$) followed by a linear layer to five outputs and a softmax. The other mechanisms are data augmentation during training (flips, small rotations, color jitter, affine transforms), class-imbalance handling through a weighted random sampler and weighted cross-entropy loss, early stopping monitored on validation accuracy, and a genetic algorithm that selected the learning rate ($2.00\\times10^{-4}$), dropout rate, and weight decay ($3.72\\times10^{-5}$). Together these let a small histopathology dataset adapt a generic visual-feature extractor to stage-associated tissue morphology.","core_discovery":"The central discovery is that the FIGO surgical stage label, mapped to five classes (0 for control, 1–4 for Stages I–IV), is recoverable from the morphology of a single 1.5 mm-diameter, 5 µm-thick tissue core imaged in transmission brightfield. The trained model, optimized with data augmentation, weighted sampling, weighted cross-entropy loss, and hyperparameters found by a genetic algorithm, achieved perfect precision and recall on the extreme classes (control and Stage IV) and F1-scores of 0.98, 0.91, and 0.97 on the middle stages. The authors interpret the adjacent-stage error pattern as consistent with the genuinely subtle morphological boundaries between Stages I, II, and III, and conclude that deep learning can act as an assistive staging tool.","pith_inferences":["Because the independent test set was acquired with the same BX61 microscope and imaging protocol as the training data, the 97.62% figure is best read as a within-protocol ceiling; a cross-scanner or multi-stain cohort would probably lower it, and that gap is exactly what the paper's own call for multi-institutional validation targets.","If the model is truly reading stage morphology, its saliency maps should concentrate on epithelial structures, stroma, and immune infiltrate rather than on background glass, tissue edges, or stain-color patches; this is a cheap check that could be run before any clinical pilot.","A reader checking reproducibility should note that the architecture citation 'ResNet-101 [22]' has no matching entry in the reference list, which ends at [19]; the paper's bibliography therefore does not by itself supply the standard provenance for the backbone.","The near-perfect separation of the extreme classes suggests a continuous ordinal morphology score may underlie the five discrete labels; regressing on tumor burden or stage as an ordered variable could give finer-grained utility than classification alone."],"forward_implications":["Control and Stage IV are perfectly separable in this test set (F1 = 1.00), so the practical ambiguity lies between the adjacent middle stages, where even morphologically similar cases are hard for human readers.","An automated stage prediction could be used as a triage and consistency-check tool in digital pathology, prioritizing cases with low-confidence predictions for closer pathologist review.","The accuracy was reached with only 2,418 training images, suggesting ImageNet-pretrained backbones plus augmentation and class weighting can handle small-domain histopathology staging tasks.","The model as tested sees single tissue cores; the paper argues that combining patch-level predictions across whole-slide images with clinical data may further improve separation of adjacent stages."],"supporting_citations":[{"why":"Defines FIGO staging by extent of spread, the target the model is built to predict.","marker":"[5]"},{"why":"Systematic review of AI in ovarian cancer histopathology that identifies the gap this study addresses.","marker":"[11]"},{"why":"Prior deep-learning work on ovarian pathology that the paper extends from subtype classification to staging.","marker":"[12]"},{"why":"Supplies the brightfield imaging and automated scanning protocol used to produce the dataset.","marker":"[16]"},{"why":"Documents generalization limits of deep learning in digital pathology and motivates the paper's call for multi-institutional validation.","marker":"[17]"},{"why":"Supports the paper's proposed next step of analysing whole-slide images rather than individual cores.","marker":"[18]"}],"fun_headline_variants":["AI stages ovarian cancer at 97.6% from a single tissue core","Deep learning stages ovarian cancer 97.6% on brightfield slides","Automated cancer staging: CNN achieves 97.6% on ovarian brightfield","Ovarian cancer stage: AI 97.6% from a single brightfield core","Deep learning reads biopsy brightfield, stages ovarian cancer 97.6%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a single 1.5 mm tissue core imaged in brightfield contains enough morphological information to determine the FIGO surgical stage, which is defined by the anatomic extent of tumor spread found at surgery.","fun_headline_variants_meta":{"raw":{"variants":["AI stages ovarian cancer at 97.6% from a single tissue core","Deep learning stages ovarian cancer 97.6% on brightfield slides","Automated cancer staging: CNN achieves 97.6% on ovarian brightfield","Ovarian cancer stage: AI 97.6% from a single brightfield core","Deep learning reads biopsy brightfield, stages ovarian cancer 97.6%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001198,"raw_usage":{"total_tokens":4934,"prompt_tokens":932,"completion_tokens":4002,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":3898}},"tokens_in":548,"tokens_out":4002,"duration_ms":27966,"temperature":1.0,"reasoning_tokens":3898,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:19:08.709157+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the trained model to a cohort scanned with a different microscope, stain batch, or institution, with pathologist-confirmed FIGO labels, and compare accuracy: a collapse toward chance would show the 97.62% result came from acquisition-specific cues rather than stage morphology. A second, cheaper check is to inspect saliency maps: if the model attends to glass edges, background, or stain color patches rather than tissue structures, the stage-learning claim is falsified.","supporting_citations":[{"cited_title":"Javadi, D","cited_arxiv_id":null,"evidence_quote":"Defines FIGO staging by extent of spread, the target the model is built to predict."},{"cited_title":"Breen, K","cited_arxiv_id":null,"evidence_quote":"Systematic review of AI in ovarian cancer histopathology that identifies the gap this study addresses."},{"cited_title":"Breen, K","cited_arxiv_id":null,"evidence_quote":"Prior deep-learning work on ovarian pathology that the paper extends from subtype classification to staging."},{"cited_title":"Sengupta, M","cited_arxiv_id":null,"evidence_quote":"Supplies the brightfield imaging and automated scanning protocol used to produce the dataset."},{"cited_title":"Jarkman, M","cited_arxiv_id":null,"evidence_quote":"Documents generalization limits of deep learning in digital pathology and motivates the paper's call for multi-institutional validation."},{"cited_title":"Greeley, L","cited_arxiv_id":null,"evidence_quote":"Supports the paper's proposed next step of analysing whole-slide images rather than individual cores."}],"review_version":1}