{"id":"36c89e7a-7ef8-4403-a715-7273ab1aa4a3","arxiv_id":"2608.03145","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"H&E-based AI risk heatmaps guided spatial proteomics to find cell-cycle-enriched high-risk vs immune-enriched low-risk tumor niches in TNBC, with a 13-protein composite that adds moderate recurrence discrimination to image scores.","lead":"A team used an AI model to map recurrence risk on H&E-stained breast cancer slides, then cut out and mass-analyzed the high- and low-risk regions to see what molecules differ there. The molecular contrast, cell-cycle programs in high-risk regions and immune activity in low-risk regions, suggests outcome-trained AI can serve as a spatial guide for biomarker discovery.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed 0.06 C-index improvement from adding the spatial protein composite to the H&E risk score is estimated on 12 recurrence events with no confidence interval and no external protein-level validation, so the incremental-value claim may be sampling noise.","rationale":"The reader identified the two-patient spatial discovery and the patch-label inheritance as the weakest assumptions. I partially agree: the label-inheritance issue is real but partially mitigated by the within-patient high- versus low-risk spatial comparison, which controls for patient-level confounders. The two-patient discovery is acknowledged, but the process-level concordance across many regions provides limited support. However, the most load-bearing missing piece is the absence of any honest held-out estimate of the combined model's incremental value. The Fig. 5e-f OOB bootstrap is an internal resampling procedure on a cohort with only 12 events, not an external validation. The METABRIC transcript analysis is reassuring but does not validate the protein composite or its complementarity with morphology. The paper itself concedes limited power from the modest event count, yet no confidence intervals are reported for the key C-index improvement. This does not invalidate the entire paper—the bulk proteomic associations and spatial proteomic contrasts are coherent and the METABRIC direction is consistent—but it means the strongest clinical-translational claim rests on a statistically fragile estimate. The reader's CONDITIONAL verdict is therefore appropriate; requiring confidence intervals and external validation is the right condition. I do not propose changing the verdict, hence UNCHANGED.","tokens_in":16969,"tokens_out":5741,"duration_ms":73980,"concrete_test":"Recompute the out-of-bag evaluation in Fig. 5e using 2,000 stratified bootstrap resamples and report the 95% percentile/bootstrap confidence interval for the paired difference in C-index between RS+composite and RS alone. If the confidence interval includes zero, the stated improvement from 0.679 to 0.739 is not distinguishable from sampling noise and the incremental-value claim should be downgraded. If the confidence interval excludes zero, the concern is mitigated, but independent protein-level validation in a separate TNBC cohort with more recurrence events would still be needed to confirm generalizability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central translational claim is that the 13-protein spatial composite adds prognostic information beyond H&E morphology. The evidence for this is the out-of-bag bootstrap C-index increase from 0.679 to 0.739 in the expanded cohort (Fig. 5e-f). This is not a held-out evaluation: the composite was derived from spatial contrasts in two patients, selected for detectability in the bulk test cohort, and then evaluated on that same test cohort plus 47 additional patients (total n=96, 12 recurrence events). With only 12 events, Harrell's C-index has wide sampling variability, and the 0.06 increment may easily be within noise. No confidence intervals are reported for the C-index or for the difference, despite the authors explicitly noting that 'the modest number of recurrence events limits the precision of the performance estimates.' The METABRIC transcript-based validation supports the biological direction of the markers, but it does not validate the protein-level composite, the DIA-MS assay, or the incremental value over the H&E score. If the 0.739 value is optimistic, the strongest claim—that the spatially derived protein signature adds prognostic information beyond morphology—is unsupported, even though the spatial biological findings may still hold.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an outcome-informed spatial pathology framework for triple-negative breast cancer (TNBC). A patch-level H&E classifier trained with weak patient-level labels produces recurrence-risk heatmaps; histogram-plus-top-k aggregation yields a patient-level risk score with AUC and C-index of 0.77 in a 49-patient test cohort. Bulk proteomics of the test cohort associates high-risk status with cell-cycle/genome-maintenance programs and low-risk status with immune programs. Within-slide analysis shows high- and low-risk patches coexisting in the same tissue compartments, with distinct nuclear/architectural features. Using AI heatmaps as guides, the authors physically isolate and profile 46 tumor regions from two recurrence patients by mass spectrometry, reporting concordant mitotic vs immune/antigen-presentation programs. A 13-protein composite derived from these spatial contrasts is evaluated in an expanded 96-patient cohort, where adding it to the H&E score improves an out-of-bag C-index from 0.679 to 0.739; a transcript-based version of the composite stratifies recurrence in the independent METABRIC TNBC cohort.","tokens_in":17311,"tokens_out":3380,"duration_ms":42438,"significance":"This is a conceptually novel and technically ambitious integration of outcome-trained deep learning with laser-capture-mass-spectrometry spatial proteomics. If the central claims hold, the framework would convert black-box H&E risk predictions into physically addressable molecular hypotheses, a valuable step for computational pathology. The paper also contains a substantial set of internal controls: patch-level distribution analyses, tissue-compartment deconvolution, bulk proteomic enrichment, and an external transcript-level validation. These strengths make the core idea compelling even though the quantitative incremental-value claim is not yet established at the evidence level presented.","major_comments":[{"comment":"The central claim that the spatially derived 13-protein composite adds prognostic information beyond H&E morphology rests on a C-index increase from 0.679 to 0.739 in an expanded cohort with only 12 recurrence events. No confidence intervals are reported for either C-index or the difference, despite the authors acknowledging the modest event count. Moreover, the expanded cohort includes the original 49-patient test cohort, which contains the two discovery patients (P1, P2) and the bulk samples used to characterize risk groups; the composite's constituent markers and their directions were selected using those same samples. The OOB bootstrap procedure does not make this a held-out evaluation. The authors should report CIs, a bootstrap test of the increment, and ideally a sensitivity analysis with the discovery patients excluded, or temper the claim accordingly.","section":"Section 6, Fig. 5e–f"},{"comment":"The spatial discovery is based on two deliberately selected patients: P1 (greatest within-slide variance) and P2 (highest patient-level risk). This selection maximizes the chance of finding contrasting molecular programs, but it does not establish that those programs reflect a generalizable high- vs low-risk axis. The paper itself reports limited protein-level overlap between the two patients (21 overlapping DEPs, of which 19 concordant), with substantial inter-patient heterogeneity. The text in the abstract and Results states that spatial profiling 'revealed a concordant molecular contrast,' but this is a process-level enrichment in two extreme cases, not a population-level discovery. The authors should explicitly frame this as hypothesis-generating and provide more patients or external spatial validation before making generalizable claims.","section":"Section 5, Fig. 4b"},{"comment":"The patch-level classifier is trained with each patch inheriting the patient's recurrence label. This weak supervision scheme means patch scores may reflect patient-level confounders (e.g., staining batch, treatment differences, tumor size) rather than localized recurrence drivers. The claim that high- and low-risk patches coexist within the same tissue compartment depends on the assumption that patch-level scores are biologically local. Although tissue-compartment analysis partially addresses composition, the paper does not report any analysis of within-patient patch-score variability against between-patient variability, nor any correction for known confounders that differ between development and test cohorts (Table 1: tumor size, T stage, neoadjuvant/adjuvant chemotherapy, radiotherapy are significantly different). This omission leaves open the possibility that the observed 'intratumor","section":"Methods, Model development; Results, Section 2"},{"comment":"The METABRIC validation is transcript-based, whereas the discovered signature is protein-based and measured by DIA-MS. The transcript-level stratification of RFS supports the biological direction of the markers, but it does not validate the protein composite itself, the DIA-MS assay, the precursor selection, or the incremental prognostic value over H&E. Given that the C-index improvement is already estimated on a small, partly overlapping cohort, the external validation should be at the protein level (or clearly labeled as only supporting biological plausibility) for the incremental-value claim to be credible.","section":"Section 6, Fig. 5c–d"}],"minor_comments":[{"comment":"'Out-of-bag' (OOB) is used without definition in the abstract and main text; define the procedure when first used, since it is not a standard term for all readers.","section":"Abstract / Section 6"},{"comment":"The C-index comparison is presented with point estimates only; adding bootstrap confidence intervals or a distribution plot would make the precision of the estimates visible.","section":"Results, Section 6, Fig. 5e"},{"comment":"The standardization of protein abundances 'using the corresponding low-risk regions as the reference' is described informally; a precise formula in the main text or Supp. Methods would improve reproducibility.","section":"Results, Section 5, Fig. 4e"},{"comment":"Minor grammar issues: 'Patients with recurrence  had higher mean fractions' contains a duplicate space; also the sentence beginning 'Patches from patients with recurrence were relatively more represented...' could be rephrased for clarity.","section":"Results, Section 2"},{"comment":"References 24 and 30 are dated 2026; if these are preprints or in-press articles, the citation style should be made consistent with the journal's guidelines.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper's most prominent quantitative claim—the C-index improvement from 0.679 to 0.739—is based on 12 events and an evaluation cohort that overlaps with the discovery cohort. This is a load-bearing issue for the abstract and discussion. If the authors can provide confidence intervals or an external protein-level validation, the claim may survive; otherwise, the incremental-value conclusion should be explicitly softened to a hypothesis-generating observation. The spatial biology and technical novelty are strong enough to justify a revised version rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing you should know: this paper has a genuinely new workflow. Most AI-guided spatial omics picks regions by cell type or morphology; here they use an outcome-trained patch-level risk heatmap as the coordinate map for mass-spec spatial proteomics. That is a real step beyond Mund et al., and the idea alone makes the paper worth reading.\n\nThe pipeline is coherent, and the bulk proteomics shows the expected biology: high-risk tumors look proliferative, low-risk tumors look immune-enriched. The intratumoral finding—high- and low-risk patches coexisting in the same tumor compartment with distinct nuclear shape and architecture—is supported by the dip test and the histology examples. The METABRIC transcript validation of the 13-marker axis is a useful external check, even though it is not protein-level.\n\nThe soft spots are serious. The headline C-index improvement (0.679 to 0.739) comes from 12 recurrence events in an expanded cohort, with no confidence intervals. That increment is within sampling noise at this event count, and the authors admit the precision is limited. Second, the 13-protein composite was selected using data from the test cohort, then evaluated in an expanded cohort that includes that same test cohort. The METABRIC result supports the biological direction of the markers, but it does not validate the protein-level assay or the incremental value over H&E. Third, the spatial discovery is two deliberately selected patients; that is fine for hypothesis generation, but the paper sometimes implies a more generalizable contrast. Finally, the patch-level classifier gives each patch the patient's recurrence label, so patch scores could reflect patient-level confounders or tissue composition rather than truly localized risk. The compartment analysis helps, but the caveat stands. No data or code are released, and the SLACS platform is proprietary, so reproducibility is limited.\n\nTo their credit, the authors clearly acknowledge many of these limitations. The core biological observation—mitotic versus immune programs in AI-defined risk niches—is plausible and interesting. But the strongest quantitative claim, that the spatial protein composite adds real prognostic information beyond morphology, is not supported by the evidence as presented.\n\nI would send this to a serious referee. The idea deserves to be in the literature, but the paper overclaims the complementarity. A revision that presents the spatial findings as exploratory, reports confidence intervals, and either obtains external protein-level validation or narrows the claim would be publishable. In a reading group, it would invite a good discussion about outcome-guided spatial sampling and why small event counts sink incremental-value claims.","headline":"A genuinely novel workflow—AI risk heatmaps as coordinates for spatial proteomics—with suggestive biology, but the headline complementary-prognosis claim rests on very few events and in-sample selection; worth a careful referee, not acceptance as is.","tokens_in":17884,"tokens_out":3318,"would_cite":true,"duration_ms":38697,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"H&E-based AI recurrence heatmaps localize molecular niches, and a 13-protein composite adds prognostic signal in TNBC.","keywords":["triple-negative breast cancer","recurrence prediction","whole-slide imaging","deep learning","spatial proteomics","tumor microenvironment","mass spectrometry","biomarker discovery"],"falsifier":"Profile AI-defined high- and low-risk tumor regions across, say, 20 or more recurrence patients spanning the full risk-score range; if the concordant direction (mitotic enrichment in high-risk regions, immune and antigen enrichment in low-risk regions) does not reappear in the majority, or if adding the 13-protein composite to the H&E risk score does not beat the H&E-only C-index of 0.679 in an independent cohort, the central claim would be falsified.","tokens_in":16885,"feed_emoji":"🔬","tokens_out":8394,"duration_ms":89521,"temperature":0.7,"pith_summary":"The paper sets out to show that the recurrence-risk heatmaps produced by an H&E-based deep-learning model label real, spatially localized biology rather than arbitrary statistical patterns. By using those heatmaps as coordinates to physically isolate high- and low-risk tumor patches and profiling them with mass-spectrometry proteomics, the authors found a concordant molecular contrast in both profiled recurrence patients: mitotic programs dominate high-risk patches, while immune and antigen-presentation programs dominate low-risk patches. They then built a 13-protein composite from that contrast and showed that adding it to the image risk score improved recurrence discrimination, increasing the out-of-bag C-index from 0.679 to 0.739. This matters because it turns an opaque prediction into a spatially explicit molecular map that can guide sampling and biomarker discovery in TNBC.","feed_headline":"H&E AI plus spatial proteomics lifts recurrence C-index to 0.739","feed_subtitle":"Protein signatures from AI-highlighted regions add late-recurrence information H&E misses.","key_machinery":"The load-bearing object is the patch-level recurrence-associated risk score (RRS), assigned by a weakly supervised patch classifier to roughly 150-micrometer H&E patches and then remapped to slide coordinates to form recurrence-risk heatmaps. Distribution-based aggregation, specifically a histogram of the ten highest-scoring patches followed by Lasso regression, converts patch scores into a patient-level risk score. The second mechanism is AI-guided spatial isolation: laser-based microdissection physically cuts out RRS-defined tumor patches for mass-spectrometry proteomics, yielding the spatial protein contrasts that seed the 13-protein tumor composite.","core_discovery":"The central discovery is that outcome-trained patch-level risk scores on H&E slides separate the tumor into spatially distinct molecular states. In a 156-patient development/test design, aggregating the top-scoring patches with a histogram reached an AUC of 0.77 and a C-index of 0.77, and heatmaps showed high- and low-risk patches coexisting within the same tumor compartment. Guided by those heatmaps, the authors isolated 46 tumor regions from two recurrent patients; high-risk regions were concordantly enriched in mitotic and cell-cycle programs, low-risk regions in immune and antigen-presentation programs. The concordant proteins were condensed into a 13-protein tumor composite whose bulk-t","pith_inferences":["Extension: if the mitotic-versus-immune contrast replicates in larger cohorts, outcome-guided spatial proteomics could be applied to other tumor types whose recurrence drivers are poorly understood.","Extension: because each patch inherits its patient's recurrence label, part of the risk signal may reflect patient-level confounders; a cross-validated design with patient-level splitting would test how much of the heatmap is genuinely local.","Extension: the temporal complementarity, with H&E strong early and the protein composite stronger later, suggests a combined score could be evaluated as a dynamic surveillance tool."],"forward_implications":["The ten highest-scoring patches aggregated as a 20-bin histogram distinguish recurrent from non-recurrent TNBC patients with an AUC of 0.77 and a C-index of 0.77 in an independent test cohort.","High- and low-risk patches coexist within the same tumor compartment, so the risk heatmap captures intratumoral heterogeneity beyond tissue-compartment identity.","Spatially isolated high-risk tumor regions are enriched in mitotic programs and low-risk regions in immune and antigen-presentation programs, with the contrast concordant across the two profiled patients.","Adding the 13-protein tumor composite to the H&E risk score raises the out-of-bag C-index from 0.679 to 0.739 and improves time-dependent discrimination at 3 and 5 years.","A transcript-based version of the composite stratifies recurrence-free survival in an independent TNBC cohort."],"supporting_citations":[{"why":"Prior work combining image-based region selection with ultrasensitive mass-spec proteomics; the paper distinguishes its outcome-informed sampling from this phenotype-based approach.","marker":"[25]"},{"why":"Supplies the DIA mass-spec data-processing method used to quantify bulk and spatial proteomic samples.","marker":"[28]"},{"why":"Provides the deconvolution method used to validate immune-score and tumor-purity contrasts between AI risk groups.","marker":"[29]"},{"why":"Provides the tissue-compartment classifier used to assign each patch to tumor, stroma, necrosis, or immune compartments.","marker":"[30]"},{"why":"Supplies the statistical test used to show risk-score distributions are non-unimodal within tumor and stromal compartments.","marker":"[33]"},{"why":"Describes the laser-based cell-sorting platform used to physically isolate AI-defined tissue regions for proteomics.","marker":"[34]"},{"why":"Supplies the independent transcriptomic TNBC cohort used to test the survival stratification of the transcript-based composite.","marker":"[35]"}],"fun_headline_variants":["AI-guided spatial proteomics maps TNBC recurrence niches","H&E risk heatmaps reveal protein states behind recurrence","Spatial proteomics + H&E AI: C-index hits 0.739","Outcome AI directs proteomics to find recurrence drivers","Risk-patch AI reveals immune vs mitotic tumor niches"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The spatial molecular contrast is derived from only two deliberately selected recurrence patients, so the claim that AI-defined high- versus low-risk regions carry a generalizable mitotic-versus-immune biology assumes those two cases stand in for TNBC as a whole.","fun_headline_variants_meta":{"raw":{"variants":["AI-guided spatial proteomics maps TNBC recurrence niches","H&E risk heatmaps reveal protein states behind recurrence","Spatial proteomics + H&E AI: C-index hits 0.739","Outcome AI directs proteomics to find recurrence drivers","Risk-patch AI reveals immune vs mitotic tumor niches"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000273,"raw_usage":{"total_tokens":1520,"prompt_tokens":836,"completion_tokens":684,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":580,"completion_tokens_details":{"reasoning_tokens":600}},"tokens_in":580,"tokens_out":684,"duration_ms":7425,"temperature":1.0,"reasoning_tokens":600,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T00:42:39.786255+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Profile AI-defined high- and low-risk tumor regions across, say, 20 or more recurrence patients spanning the full risk-score range; if the concordant direction (mitotic enrichment in high-risk regions, immune and antigen enrichment in low-risk regions) does not reappear in the majority, or if adding the 13-protein composite to the H&E risk score does not beat the H&E-only C-index of 0.679 in an independent cohort, the central claim would be falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior work combining image-based region selection with ultrasensitive mass-spec proteomics; the paper distinguishes its outcome-informed sampling from this phenotype-based approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the DIA mass-spec data-processing method used to quantify bulk and spatial proteomic samples."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the deconvolution method used to validate immune-score and tumor-purity contrasts between AI risk groups."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the tissue-compartment classifier used to assign each patch to tumor, stroma, necrosis, or immune compartments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the statistical test used to show risk-score distributions are non-unimodal within tumor and stromal compartments."},{"cited_title":"& Mann, M","cited_arxiv_id":null,"evidence_quote":"Describes the laser-based cell-sorting platform used to physically isolate AI-defined tissue regions for proteomics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the independent transcriptomic TNBC cohort used to test the survival stratification of the transcript-based composite."}],"review_version":1}