{"id":"70dd10d6-4884-4812-8870-25271ff312d4","arxiv_id":"2412.10392","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A narrative review maps AI approaches for predicting breast cancer biomarkers, including genomic, transcriptomic, proteomic, and metabolomic profiles, from H&E histopathology slides.","lead":"This paper reviews artificial intelligence methods that predict breast cancer molecular markers from routine H&E stained tissue images. It organizes dozens of published studies by biomarker type and highlights data and validation gaps that block clinical use.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table transcription errors, exemplified by the [60] gene-count discrepancy, threaten the review's reliability as a map of methods and performance.","rationale":"The reader identified the same load-bearing concern—accuracy of transcribed performance numbers—and the [60] contradiction is concrete evidence. I would add that the methodological absence of a search protocol is a second, related weakness, but the transcription issue is the more immediate threat because it is already demonstrable. The check proposed would settle the concern by distinguishing a one-off wording error from a systematic reporting problem. If the spot-check shows few mismatches, the review could be provisionally useful; if it shows many, the review's gap analysis and clinical conclusions are not supported. Either way, the current evidence does not allow the paper to be accepted as a reliable synthesis, so the verdict should stay UNVERDICTED.","tokens_in":26528,"tokens_out":5429,"duration_ms":47421,"concrete_test":"Re-extract the results of study [60] from Schmauch et al. (Nature Communications, 2020), recording the number of genes significantly predicted in the pan-cancer model and in the breast cancer cohort separately. If the original reports 10,000 pan-cancer and 2,902 for BRCA, Section 5.2.1 is wrong. Then extend the check to a random sample of 20 rows from Tables 5–8, comparing each quoted metric (AUC, accuracy, R, gene counts, dataset sizes) against the cited primary source. If more than two of the twenty rows contain a mismatch in the reported value or its definition, the review's quantitative synthesis is not reliable enough to support its conclusions, and the verdict should remain UNVERDICTED pending a corrected and systematically searched version.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The review's central claim—that it provides a reliable map of AI biomarker-prediction methods and a trustworthy gap analysis—depends entirely on the fidelity of Tables 3–8 to the primary literature. That fidelity is already broken in one visible place: Section 5.2.1 credits study [60] with predicting 'approximately 10,000 genes with adjusted p-values below 0.05 specifically for breast cancer,' while Table 6 reports only 2,902 significantly well-predicted genes for the same study. The original paper likely reports ~10,000 genes across 28 cancer types and 2,902 for the breast cohort; the review's text conflates the two. This is not a cosmetic slip: if other rows misattribute numbers the same way, the paper's stated conclusion that omic biomarker prediction is 'limited by dataset availability' may rest on a distorted view of what has actually been demonstrated. A second red flag is Table 3 row [26], whose Method column lists 'CycleGAN to normalize staining variations, Resnet34 network pretrained on the ImageNet dataset,' while the text's own description of [26] is a tissue-fingerprint pretraining approach with no mention of CycleGAN. The absence of a documented search strategy means the reader cannot tell whether the included studies are representative, but the more immediate threat to the central claim is the unverified transcription of quantitative results into the tables.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is a narrative review of artificial-intelligence and computational methods for predicting breast cancer molecular biomarkers from routine H&E-stained histopathology images. It categorizes work into non-omic biomarkers (ER, PR, HER2, Ki67, PD-L1), HER2 scoring, and omic biomarkers (genomics, transcriptomics, proteomics, metabolomics), and summarizes each study in tables that list datasets, patient counts, objectives, performance metrics, and methods. The review argues that AI can predict multiple molecular biomarkers from H&E images, that receptor and protein biomarkers are relatively well explored, and that omic biomarker prediction remains limited, largely by dataset availability. It closes with discussion of datasets, annotation, architectures, interpretability, and future directions.","tokens_in":26747,"tokens_out":4422,"duration_ms":43268,"significance":"If the table transcriptions are reliable, the review would be a useful entry-level map of this fast-moving field: it organizes a large set of primary studies by biomarker type, includes datasets and performance numbers, and its discussion of weakly supervised learning, domain-specific pretraining, and explainability is sensible. The stated focus on medical relevance and the explicit comparison with prior reviews in Table 2 give it a plausible niche. However, the review's value as a reliable map depends on the fidelity of Tables 3 through 8 to the primary literature, and the manuscript currently contains at least two visible transcription problems. Because the paper does not report a systematic search or inclusion protocol, readers also cannot audit whether the included studies are representative enough to support the \"comprehensive review\" claim. The subject is important and the paper is readable, but the evidence base for its main claims needs strengthening before it can be accepted as a dependable reference.","major_comments":[{"comment":"The review claims comprehensiveness but reports no literature-search strategy: no databases, search strings, date range, inclusion/exclusion criteria, or screening process are described. Section 3 and Table 2 position the work as covering \"Dataset Details\" and \"Computational-Clinical Link\" in a way that implies a systematic comparison, yet the selection of the 90-odd cited studies is not reproducible. Without a documented protocol, the central claim of a comprehensive review cannot be audited, and the reader cannot distinguish an intended representative sample from an arbitrary one.","section":"Sections 2–3"},{"comment":"There is an internal contradiction in the reported gene count for the HER2NA study. Section 5.2.1 states that the model \"effectively predicted approximately 10,000 genes with adjusted p-values below 0.05 specifically for breast cancer,\" whereas Table 6 reports for the same study that \"2902 genes are significantly well predicted.\" These two numbers cannot both describe the same breast-cancer-specific result; one is likely the pan-cancer count and one the breast-cancer count, but the review does not say so. Because Table 6 is the main evidence for the transcriptomics section, this discrepancy must be resolved and the same check applied to all table rows.","section":"Section 5.2.1 and Table 6, study [60]"},{"comment":"The Method column of Table 3 for reference [26] says \"CycleGAN to normalize staining variations, Resnet34 network pretrained on the ImageNet dataset for classification,\" but the text immediately above the table describes [26] as a tissue-fingerprint pretraining approach that learns to pair left/right halves of pathologic images, with no mention of CycleGAN. If CycleGAN is used in the primary paper, the text should say so; if it is not, the table row is a transcription error. Either way, the discrepancy between the table and the narrative in the same section indicates that the tables have not been systematically verified against the cited sources.","section":"Section 4 and Table 3, row [26]"},{"comment":"The main gap analysis—that protein biomarkers are well explored while omic biomarkers are limited by dataset availability—is built directly on the census of studies in Tables 3 through 8. Given the confirmed transcription issues in Table 3 and Table 6, the current manuscript does not yet establish that the remaining rows are accurate enough to support this conclusion. The authors should either perform a systematic verification of all performance numbers and dataset descriptions against the primary papers or qualify the claimed comprehensiveness and the quantitative parts of the gap analysis.","section":"Section 6 and Section 7"}],"minor_comments":[{"comment":"The table gives both \"TCGA: 939, ABCTB: 2535\" in the Dataset column and \"TCGA: 1014, ABCTB: 2535\" in the Size column for the same study; the review should clarify which figure is the number of patients and which is the number of slides or WSIs.","section":"Table 3, row [29]"},{"comment":"The sentence \"super-resolution techniques in spatial transcriptomic profiling [69]\" appears to cite the wrong reference: [69] is the VGG16-based study by Monjo et al., while the super-resolution spatial-transcriptomics prediction method is reference [70].","section":"Section 6"},{"comment":"There are numerous formatting and typographical errors, including missing spaces in the running text (e.g., \"Theworkpresentedin[25]aimstodiscriminate\" and \"identifyspecificbiomarkersthatarerelatedtothediseaseoccurrence\"), which reduce readability and should be corrected.","section":"Throughout"},{"comment":"The checkmark/cross symbols in Table 2 are not defined in a legend; the reader must infer that they denote presence or absence of a feature, and a legend would make the comparison unambiguous.","section":"Table 2"},{"comment":"The abbreviation HER2NA is introduced for study [60] but the method is described in Table 6 only as a \"50-layer ResNet pretrained on the ImageNet\"; the table should mention the HER2NA name or explain the relationship.","section":"Section 5.2.1"},{"comment":"The discussion of weak supervision notes that [38] reported significant differences between patch-level and slide-level labels, but it does not give the quantitative performance gap; adding the ER/PR/Ki67/HER2 numbers would make the point more concrete.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a narrative review and, as such, does not necessarily need a formal systematic-review protocol; however, once the authors claim comprehensiveness and build a gap analysis on a detailed table census, the absence of search documentation and the visible table errors become load-bearing. If the authors can correct the two named transcription problems and audit the remaining rows, the paper would be a serviceable review for its intended audience. I would not reject on the basis of disagreement with the field's consensus; the issues are internal-consistency and documentation problems that are within the scope of a major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely useful review for anyone entering the breast cancer H&E biomarker prediction space, but you cannot trust the performance tables without checking the primary papers. The organizational scheme is good and the gap analysis is sensible; the errors are concentrated in the transcription of numbers into tables.\n\nWhat it does well: It is the first review I know of that is specifically breast-cancer-focused, H&E-only, and includes dataset details and performance numbers in the tables. That is a real service. The categorization into non-omic (ER/PR/HER2/Ki67/PD-L1) versus omic (genomic, transcriptomic, proteomic, metabolomic) biomarkers is clear, and the discussion of challenges—data availability, annotation, model explainability, domain-specific pretraining—is balanced and informed. The gap analysis (protein biomarkers well explored, omic biomarkers limited by dataset availability) is plausible and matches what people in the field would say.\n\nThe soft spots: There are at least two concrete transcription errors. Study [60] is credited with predicting 'approximately 10,000 genes' in the text, but Table 6 says 2,902; the original paper probably reports both numbers for different scopes (pan-cancer vs breast), and the review conflates them. Similarly, Table 3 attributes a CycleGAN normalization step to [26], while the text correctly describes that paper as a tissue-fingerprint approach. These are not cosmetic: the review's value is precisely the curated tables, and if numbers are misattributed, the reliability of the map is compromised. There is also no reported search strategy or inclusion criteria, which makes it a narrative review and leaves representativeness unclear.\n\nNone of this sinks the central argument. The conclusion that AI can predict receptor and protein biomarkers well, and that omic prediction is still constrained by data, would survive even if every row of the tables were corrected. But for a review that sells itself on dataset details and performance, accuracy is the product.\n\nWho this is for: graduate students and clinicians wanting an overview, or researchers looking for a starting bibliography. It will not change practice or settle a technical debate.\n\nRecommendation: it deserves a serious referee, but the outcome should be major revision. The authors need to fix the internal contradictions, add a search strategy or at least state explicit inclusion criteria, and report a verification pass over the tables against primary sources. I would accept it for review, with that expectation.","headline":"A useful but imperfect review: solid organization and gap analysis, with table transcription errors that need fixing before it can be trusted as a reference map.","tokens_in":27261,"tokens_out":2807,"would_cite":true,"duration_ms":26245,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review argues that artificial intelligence can extract a large part of the breast cancer molecular profile—hormone receptor status, HER2, Ki-67, PD-L1, gene mutations, gene expression levels, molecular subtypes, and protein…","keywords":["breast cancer","molecular profiling","histopathology","H&E images","deep learning","biomarkers","omics","precision medicine"],"falsifier":"Open the original Nature Communications paper for study [60] and count the genes significantly associated with RNA-seq expression in breast cancer: if the true number is 2,902 rather than roughly 10,000, the review's text overstates the result by a factor of three, and a spot-check of the other table entries, say the AUCs claimed for ER and HER2 predictions, would be the next decisive test of whether the review's map of the field is trustworthy.","tokens_in":26324,"feed_emoji":"🧬","tokens_out":6331,"duration_ms":53030,"temperature":0.7,"pith_summary":"This review argues that artificial intelligence can extract a large part of the breast cancer molecular profile—hormone receptor status, HER2, Ki-67, PD-L1, gene mutations, gene expression levels, molecular subtypes, and protein abundance—directly from routine H&E-stained tissue sections, without the special stains, sequencing, or other assays those measurements normally require. It organizes the field into non-omic biomarkers, the proteins detected in the clinic by IHC, and omic biomarkers such as genomic, transcriptomic, proteomic, and metabolomic markers, and for each category it catalogs the methods, datasets, and reported performance. The assembled evidence points to receptor and protein biomarkers as well explored, while omic biomarkers remain limited mainly by the scarcity of datasets that pair histopathology images with molecular profiles. If the map is accurate, it gives researchers and clinicians a concrete picture of which H&E-based predictions are close to usable and where the remaining bottlenecks are.","feed_headline":"AI predicts breast cancer biomarkers from H&E slides","feed_subtitle":"Receptor status is mature; gene-expression and mutation prediction remain dataset-limited.","key_machinery":"The load-bearing mechanism is the digitized H&E whole-slide image cut into tiles, processed by convolutional networks or vision transformers, usually trained with weakly supervised multiple-instance learning (MIL), in which only the slide-level molecular label is known and the model learns which tissue regions carry the signal. The review's organizing device is a biomarker taxonomy that separates non-omic protein biomarkers (ER, PR, HER2, Ki-67, PD-L1), detectable today by IHC/FISH, from omic biomarkers (genomic mutations and copy-number changes, transcriptomic expression and subtypes, proteomic abundance, metabolomic features), which normally require sequencing or mass spectrometry. Within that taxonomy, the evidence for the central claim is carried by tabulated studies that predict molecular states from H&E alone, including pan-cancer single-model predictors and transcriptome-wide expression regression, together with domain-specific pretraining and explainability techniques that locate the morphological features driving each prediction.","core_discovery":"The paper's central claim is that modern deep learning models can predict multiple breast cancer biomarkers from H&E whole-slide images with clinically meaningful accuracy, and that the same approach can be pushed beyond single markers to omic-scale inference, including expression of thousands of genes, mutation and copy-number status of key genes, PAM50 molecular subtypes, and protein levels. The authors report that ER, PR, HER2, and Ki-67 predictions are comparatively mature, with many studies reporting AUCs above 0.75, while PD-L1 and Ki-67 are noticeably under-explored, HER2 scoring has been validated on a single benchmark cohort, and omic biomarker prediction is dominated by TCGA-based studies whose generalizability across populations has not been established. They conclude that the bottleneck is not algorithmic but data-related: models succeed when paired histopathology-molecular datasets exist, and progress will accelerate as those datasets grow, diversify, and become accessible.","pith_inferences":["If this line of work matures, a consequence the authors do not spell out is that archived H&E slides, collected for decades without molecular data, could be re-analyzed retrospectively to screen for gene-expression signatures or mutations, effectively creating molecular cohorts from existing tissue banks.","The catalog implies a predictability gradient: biomarkers with a visible morphological footprint, such as CDH1-mutant lobular cancer, are predicted more accurately, which suggests that gene selection for future image-based assays should prioritize mutations with strong phenotypic effects.","The discrepancy between the text and Table 6 for study [60] (about 10,000 versus 2,902 significantly predicted genes) is a concrete warning that the review's transcribed numbers need independent verification before the gap analysis guides research funding or clinical planning."],"forward_implications":["If the reviewed results hold, routine H&E slides could eventually serve as a first-line molecular readout, reserving IHC, FISH, and sequencing for confirmation, which would cut cost and turnaround time.","HER2 scoring from H&E is accurate enough in validation cohorts to argue for multi-center validation and prospective testing, but currently rests on a single public benchmark.","Omic predictions, especially gene expression and mutation status, are feasible but dataset-limited; expanding paired image-omics cohorts beyond TCGA, and across ethnic groups, is the direct next step.","Domain-specific pretraining on histopathology images consistently outperforms ImageNet pretraining, indicating that foundation models trained on tissue are a natural route to better biomarker prediction.","Explainability tools such as tile-level heatmaps and attention-consistency scores will be needed to convert a statistically successful model into something a pathologist can trust and a regulator can evaluate."],"supporting_citations":[{"why":"Supplies the pan-cancer single-model evidence that hormone receptor status and actionable mutations can be inferred from H&E images.","marker":"[27]"},{"why":"Provides the Receptor Net MIL baseline for predicting ER status directly from base-level H&E stains.","marker":"[29]"},{"why":"Shows PD-L1 status can be predicted from H&E-stained tissue microarrays, supporting the non-omic biomarker category.","marker":"[33]"},{"why":"Core evidence for transcriptome-wide gene expression prediction from whole-slide images, including the gene-count claim that the review transcribes.","marker":"[60]"},{"why":"Establishes transcriptome-wide expression–morphology analysis in breast cancer, reporting thousands of significantly predicted genes.","marker":"[64]"},{"why":"Demonstrates explainable machine learning for mutations, copy-number variation, gene expression, and protein levels, linking morphology to molecular alterations.","marker":"[57]"},{"why":"Supplies a systematic pan-cancer multi-omic benchmark, a source for many of the AUCs reported across biomarker categories.","marker":"[55]"},{"why":"Provides hist2RNA as an external validation of gene expression prediction on tissue microarrays, strengthening the omic inference claim.","marker":"[73]"}],"fun_headline_variants":["AI predicts ER, PR, HER2, Ki-67 from H&E slides","Breast cancer omic profiling from H&E slides using AI","Review: AI extracts biomarkers from routine H&E histology","Data scarcity hinders AI breast cancer biomarker prediction","AI reads molecular profile from breast cancer H&E images"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the performance numbers and dataset details compiled in the tables faithfully reproduce the primary papers; the review itself contains an internal contradiction about study [60], which the text credits with roughly 10,000 significantly predicted genes while Table 6 reports 2,902, so if other rows contain similar transcription errors, the gap analysis and clinical conclusions are weakened.","fun_headline_variants_meta":{"raw":{"variants":["AI predicts ER, PR, HER2, Ki-67 from H&E slides","Breast cancer omic profiling from H&E slides using AI","Review: AI extracts biomarkers from routine H&E histology","Data scarcity hinders AI breast cancer biomarker prediction","AI reads molecular profile from breast cancer H&E images"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000445,"raw_usage":{"total_tokens":2232,"prompt_tokens":912,"completion_tokens":1320,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":1236}},"tokens_in":528,"tokens_out":1320,"duration_ms":11437,"temperature":1.0,"reasoning_tokens":1236,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:04:51.952597+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Open the original Nature Communications paper for study [60] and count the genes significantly associated with RNA-seq expression in breast cancer: if the true number is 2,902 rather than roughly 10,000, the review's text overstates the result by a factor of three, and a spot-check of the other table entries, say the AUCs claimed for ER and HER2 predictions, would be the next decisive test of whether the review's map of the field is trustworthy.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the pan-cancer single-model evidence that hormone receptor status and actionable mutations can be inferred from H&E images."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Receptor Net MIL baseline for predicting ER status directly from base-level H&E stains."},{"cited_title":"Shamai, A","cited_arxiv_id":null,"evidence_quote":"Shows PD-L1 status can be predicted from H&E-stained tissue microarrays, supporting the non-omic biomarker category."},{"cited_title":"Schmauch, A","cited_arxiv_id":null,"evidence_quote":"Core evidence for transcriptome-wide gene expression prediction from whole-slide images, including the gene-count claim that the review transcribes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes transcriptome-wide expression–morphology analysis in breast cancer, reporting thousands of significantly predicted genes."},{"cited_title":"Binder, M","cited_arxiv_id":null,"evidence_quote":"Demonstrates explainable machine learning for mutations, copy-number variation, gene expression, and protein levels, linking morphology to molecular alterations."},{"cited_title":"Arslan, J","cited_arxiv_id":null,"evidence_quote":"Supplies a systematic pan-cancer multi-omic benchmark, a source for many of the AUCs reported across biomarker categories."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides hist2RNA as an external validation of gene expression prediction on tissue microarrays, strengthening the omic inference claim."}],"review_version":1}