Pith. sign in

REVIEW 4 major objections 7 minor 24 references

A DNA language model paired with routine imaging can surface gene–phenotype associations that mutation-frequency methods miss.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 10:51 UTC pith:2PIWJEPL

load-bearing objection A clever, genuinely new hypothesis-free framework for linking cancer genomes to imaging, but the 46 novel cRCC hits are built on a statistical test whose FDR control is not established, so the discoveries should be treated as candidates, not confirmed results. the 4 major comments →

arxiv 2607.20583 v1 pith:2PIWJEPL submitted 2026-07-22 q-bio.GN cs.AI

Foundation-model-guided radiogenomic discovery linking cancer genomes to cancer scans

classification q-bio.GN cs.AI
keywords radiogenomicsgenomic language modelssomatic mutationsgene-phenotype associationstumor imagingdriver discoveryEvo 2clear cell renal cell carcinoma
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Most cancer-driver discovery ranks genes by how often they are mutated, so rarely mutated genes are invisible. This paper proposes replacing the recurrence statistic with a severity score from a DNA language model: every somatic mutation is assigned a likelihood drop, genes are summarized by these scores, and each gene's summary is correlated with tumor imaging features after adjusting for total mutation burden. In a kidney-cancer cohort, the sweep recovers known drivers and adds 46 genes outside curated cancer-gene panels, several of which are established inherited-disease genes tied to cilia and the cytoskeleton. The authors argue this turns radiogenomics from confirmation of known drivers into a hypothesis-free discovery tool, and that the same scaffold can be applied to any quantitative phenotype.

Core claim

The paper's central claim is that a zero-shot severity score from a genomic language model, correlated with routine radiology, can link rarely mutated genes to tumor phenotype without prior knowledge of cancer genes. Concretely: for every somatic single-nucleotide variant in three cancer cohorts, the model's absolute log-likelihood drop between reference and mutant sequence is taken as functional disruption; per-gene maxima, means, sums, standard deviations, and counts are built for genes with at least five carriers; and these summaries are tested against radiomic features (tumor volume, tumor-to-organ ratio, heterogeneity, necrotic fraction) by TMB-adjusted partial Spearman correlation. In

What carries the argument

The load-bearing mechanism is the Evo 2 likelihood-drop score: for each somatic variant, the absolute difference in the model's log-likelihood of the reference and mutant allele within a ±4,096 base-pair window, used as a proxy for functional disruption. Per-gene severity summaries (eight metrics, including maximum, mean, sum, standard deviation, and mutation count) are then correlated with automatically segmented radiomic features via partial Spearman correlation, residualizing both variables on total mutation burden. Benjamini–Hochberg FDR control across all gene–metric–feature tests within each cohort completes the pipeline, letting the method interrogate every mutated gene without requir

Load-bearing premise

The load-bearing premise is that the language model's absolute log-likelihood drop measures functional disruption of somatic mutations; if evolutionary constraint does not reflect oncogenic effect, the per-gene summaries contain no biological signal and the correlations are statistical artifacts.

What would settle it

Permutation test: in the kidney-cancer cohort, randomly reassign the severity values among mutations while keeping mutation counts, gene sizes, and total mutation burden fixed, then recompute the full sweep; if FDR-significant gene–imaging pairs appear at the same rate under permutation, the severity score is not carrying the signal.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Recurrence-free discovery: any mutated gene with at least five carriers can be tested, including genes too rare for conventional driver statistics.
  • Validation via recovery: the FDR hits in kidney cancer include known drivers (chromatin regulators, PI3K pathway components), supporting the idea that the severity profiles carry biological signal.
  • Novel gene candidates: 46 non-panel genes reach FDR significance, concentrated in adhesion/cytoskeleton, ciliary, and transporter biology; five of the top ten are established Mendelian disease genes.
  • The imaging modality is incidental: the same severity-score-plus-TMB-adjusted-correlation scaffold could be applied to digital pathology, lab values, electronic health records, or survival endpoints.
  • Scaling needs no new experiments: extending to larger uniformly imaged registries only requires running the language model over the mutation catalogue.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A testable extension: re-run the kidney-cancer sweep after excluding long genes or stratifying by gene length; the paper itself flags long-gene artefacts in the smaller cohorts, and several novel kidney-cancer hits (e.g., SYNE1, ALMS1) are very long, so gene size may account for part of the signal.
  • A decisive control: randomly permute the severity scores across mutations while preserving mutation counts, gene sizes, and total mutation burden; if the number of FDR hits survives, the associations depend on mutation load or gene length rather than on the severity score.
  • The thin enrichment margin (2.2-fold, p = 0.05) suggests the 46-gene list is a prioritized candidate set with unknown false-discovery proportion; replication in an independent kidney-cancer cohort would sharpen the estimate.
  • If validated, the approach supplies a functional-annotation route for the 'unknome': somatic mutations in genes of unknown function could be prioritized for experimental follow-up using only existing scans.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper proposes a hypothesis-free radiogenomic discovery framework that pairs zero-shot Evo 2 severity scores for somatic mutations with routine clinical imaging features from three TCGA cohorts (cRCC n=162, HCC n=57, BC n=121). For each gene with at least five mutation carriers, eight per-gene severity summaries are computed and correlated with six radiomic features via TMB-adjusted partial Spearman correlations; Benjamini–Hochberg correction is applied within each cohort. In cRCC, the sweep recovers known renal drivers (PBRM1, SETD2, KDM5C, etc.) and reports 46 additional genes outside OncoKB panels with FDR-significant correlations, predominantly with tumor volume. HCC and BC produce no FDR-significant hits; their nominally significant lists are described as long-gene artifacts. The paper claims this demonstrates a general discovery tool for gene–imaging associations invisible to recurrence-based methods and discusses implications for functional annotation and extension to other phenotypes.

Significance. If the statistical associations are valid, this is a genuinely novel and potentially high-impact approach: it combines an externally pretrained genomic foundation model (Evo 2) with widely available imaging to nominate low-recurrence genes whose somatic disruption may influence tumor phenotype. The study has notable strengths: Evo 2 scores are used with no task-specific fitting to imaging; mutation calls and imaging are independent public data; the pipeline is described transparently; code is promised on GitHub; and the authors explicitly acknowledge several limitations, including the evolutionary-constraint interpretation of Evo 2 scores and the correlational nature of the findings. The cRCC rediscovery of canonical drivers (PBRM1, SETD2, KDM5C, STAG2, etc.) is encouraging. However, the central statistical claim—46 novel gene–imaging associations—rests on FDR control for highly zero-inflated, sparse per-gene variables, and the manuscript provides no negative-control or permutation analysis to demonstrate that the observed hit rate (65 of 103 tested genes) is not an artifact of carrier-count structure. The borderline OncoKB enrichment (p=0.05) and the unaddressed long-gene confound

major comments (4)
  1. [§4.5, §4.6, Fig. 3f] The primary statistical test is a TMB-adjusted partial Spearman correlation between per-gene severity metrics (set to zero for all non-carriers; §4.3) and imaging features. For most genes with 5–20 carriers in n=162, these metrics are zero-inflated, near-binary variables with heavy ties. The asymptotic p-value approximation for rank correlations with heavy ties can be miscalibrated, and applying BH to 2,984 such p-values does not guarantee FDR control if the null p-values are not uniform. The claim of 46 novel discoveries is not secure without a negative-control analysis (e.g., permutation of imaging labels or Evo 2 scores) demonstrating that the FDR-significant set is not an artifact of carrier-count structure. This concern is independent of whether Evo 2 |ΔLL| is a valid functional proxy. Please add such controls and/or use a test designed for zero-inflated data (e.g., two-part model o
  2. [§2.3, §4.3] The paper acknowledges long-gene mutation-rate artifacts for the nominally significant HCC and BC lists, noting that genes with exceptionally long coding sequences inflate mutation counts and are excluded by standard driver-detection pipelines. However, the same confound is not addressed for the cRCC discoveries. Several of the top novel cRCC hits (e.g., SYNE1, ABCA13, FLG) are known to be very large genes, and gene length correlates with carrier count and therefore with the non-zero fraction of the per-gene severity variable. This could produce significant correlations with imaging features even if Evo 2 severity carries no biological signal. Please control for gene length or carrier count (e.g., as a covariate, or by stratifying on carrier count) and report whether the 46 novel hits survive such adjustment.
  3. [§2.2, Fig. 2b] The OncoKB enrichment test is used as validation that the FDR threshold isolates a biologically meaningful gene set, but the reported Fisher's exact p=0.05 (one-sided) is borderline and does not constitute strong evidence. The enrichment fold-change of 2.22 is based on small counts (19/65 vs 5/38). Please provide a confidence interval for the enrichment, report the two-sided p-value, and ideally perform a continuity-corrected or permutation-based test. As it stands, the statement in §2.2 that the threshold 'confirms' a panel-enriched subpopulation overstates the evidence.
  4. [§2.3, §3] The paper reports partial Spearman correlation coefficients but provides no confidence intervals or measures of effect-size uncertainty for the 46 novel associations. Given the small sample sizes and sparse per-gene data, the observed correlations may be unstable. Please report bootstrapped or analytic confidence intervals for the headline correlations (at least for the top novel genes in Fig. 3a), and consider showing the distribution of effect sizes to help readers judge the practical magnitude of the associations.
minor comments (7)
  1. [§2.1] The text says 'under the same family-wise control neither HCC nor BC produced a single FDR-significant hit' but the procedure is FDR control, not family-wise error rate control. Please correct the wording.
  2. [§4.2] The method states 'For each somatic single-nucleotide variant' but MAF files from TCGA typically include indels as well. Please clarify whether indels were excluded, and if so, how many variants were removed.
  3. [§4.3] The definitions of 'signed minimum ΔLL' and 'signed maximum ΔLL' are ambiguous. The signed ΔLL can be positive or negative depending on whether the mutant allele is more or less probable than the reference; 'minimum' and 'maximum' of what—across mutations in the gene, or the most negative and most positive values? Please specify.
  4. [§4.8] The statement 'No data were excluded from the analysis except as described above' should quantify the HCC quality filtering (how many scans had predicted liver volume <500 mL and were removed). This is important for reproducibility.
  5. [Fig. 3e] The ClinVar and OMIM counts are read off the figure, but the text does not state the ClinVar release date or whether counts are per gene (aggregating transcripts) or per variant. Please add these details to the Methods or figure legend.
  6. [§3] The sentence 'Nothing in the pipeline is intrinsically about imaging' is overstated; the pipeline uses imaging features as the phenotype. Better phrasing: 'The pipeline is not inherently limited to imaging.' Also, the phrase 'hypothesis-free' is used loosely—there are still choices of metrics, features, and thresholds—so consider softening.
  7. [References] Several references are dated 2026 or are preprints (e.g., refs. 12, 14, 17). Please ensure these are accurately cited and, where possible, provide published versions or DOI links.

Circularity Check

0 steps flagged

No circularity: Evo2 scores are external zero-shot inputs, imaging features are independent, and no fitted parameter is renamed as a prediction.

full rationale

The paper's derivation chain is an empirical pipeline, not a closed-form derivation. Its inputs are (i) zero-shot Evo2 |ΔLL| scores from an externally pretrained model (ref. 13, by different authors), (ii) public TCGA somatic mutation calls, and (iii) radiomic features from tumor segmentations obtained with independently pretrained nnU-Net models. None of these inputs is fitted to the target gene–imaging associations. The per-gene severity summaries (Methods 4.3) are generic aggregations, and the TMB-adjusted partial Spearman correlation (Methods 4.5) is a standard statistical operation; the Benjamini–Hochberg correction is applied to the full 2,984-test family in cRCC (Methods 4.6), not to a subset selected by the outcome. The OncoKB overlap is used as an external benchmark, not as a filter before testing: the sweep is genome-wide over all genes with at least five carriers, and the OncoKB enrichment is computed after significance is declared. The Evo2 score is explicitly described as 'a foundation-model proxy for functional disruption' (Methods 4.2), which is a stated validity assumption rather than a circular definition. The paper's own limitations (small HCC/BC cohorts, long-gene artefacts, evolutionary-constraint vs oncogenic-function gap, correlation-not-causation) are acknowledged in the Discussion and weigh on biological interpretation, not on whether the analysis reduces to its inputs. No equation in the manuscript defines a predictor in terms of the outcome, no fitted parameter is renamed as a prediction, and no load-bearing self-citation or imported uniqueness theorem is used. The skeptic's concern about zero-inflated per-gene metrics threatening FDR control is a statistical robustness issue, not circularity; it does not evidence that any result is equivalent to its inputs by construction. Therefore no circular step is identified.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

The central claim rests on four hand-chosen thresholds and five domain assumptions. None are fitted to the target result, but together they determine the 65-gene hit list; the most fragile is the Evo 2 severity proxy.

free parameters (4)
  • min_carriers_per_gene = 5
    Genes with fewer than 5 mutation carriers were excluded (Methods 4.3). This threshold is chosen by hand; at n=162 many passing genes have only 5-10 carriers, so correlations rest on very few non-zero points.
  • min_support_patients = 30
    Only gene-metric-imaging combinations with at least 30 overlapping patients were tested (Methods 4.5). Threshold choice affects the test set and multiplicity.
  • fdr_threshold = 0.05
    Benjamini-Hochberg q<0.05 used to call significance (Methods 4.6). Conventional, but the resulting '46 discoveries' depend on this cut.
  • evo2_context_window = ±4096 bp
    The likelihood drop is computed in a ±4,096 bp context (Methods 4.2); a model setting carried from Evo 2 usage, chosen before this study, not fitted here.
axioms (5)
  • domain assumption Evo 2 |ΔLL| approximates functional severity for somatic cancer mutations
    Methods 4.2: score 'used here as a foundation-model proxy for functional disruption.' Prior validation is for ClinVar/BRCA1, not imaging-linked tumor phenotypes.
  • domain assumption TMB residualization removes mutation-burden confounding
    Methods 4.5 partial Spearman residualizes per-gene metric and imaging feature on total mutation count; assumes no residual confounding by coverage, tumor size, or sample quality.
  • domain assumption Automated tumor segmentations are accurate enough for radiomic features
    Methods 4.1 relies on KiTS23, TotalSegmentator, and MAMA-MIA without per-case QC except HCC liver volume filter; segmentation errors propagate to volume and texture features.
  • domain assumption TCGA MAF somatic calls are accurate and consistently aligned to GRCh38
    Methods 4.1 verifies genome build but not variant-calling quality; false calls in long genes would inflate severity summaries.
  • standard math Benjamini-Hochberg controls FDR under the dependency of these tests
    Methods 4.6; BH validity depends on positive regression dependency, plausible but unverified for zero-inflated gene metrics and correlated radiomic features.

pith-pipeline@v1.3.0-alltime-deepseek · 8649 in / 11292 out tokens · 102371 ms · 2026-08-01T10:51:22.782346+00:00 · methodology

0 comments
read the original abstract

The function of many genes is still unknown, and conventional driver-discovery methods, which rely on how frequently a gene is mutated, cannot assess genes that are only rarely affected. Here we pair Evo~2-based genome analysis with routine clinical imaging to identify gene--phenotype associations at genome-wide scale. For every somatic mutation across three TCGA cohorts (cRCC=clear cell renal cell carcinoma, HCC=hepatocellular carcinoma, and BC=breast cancer; $n = 340$ total), Evo~2 predicts a severity score, with no task-specific training. Per-gene severity summaries are then correlated with radiomic features extracted from paired tumor segmentations, controlling for total mutation burden. In TCGA-cRCC ($n = 162$), this sweep recovers established renal-cancer drivers and identifies 46 additional genes reaching false discovery rate (FDR) significance absent from curated cancer-gene panels, several of which are Mendelian ciliopathy and cytoskeletal-disease genes. These results demonstrate that pairing a genomic language model with widely available clinical imaging can serve as a hypothesis-free discovery tool for gene--imaging associations invisible to conventional approaches.

Figures

Figures reproduced from arXiv: 2607.20583 by Christiane Kuhl, Daniel Truhn, Frederik Hauke, Ingo Kurth, Jakob Nikolas Kather, Jeremias Krause, Patrick Wienholt, Sikander Hayat, Sven Nebelung.

Figure 1
Figure 1. Figure 1: Overview of the cross-modal analysis pipeline. Somatic mutations from three [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Rediscovery of established cancer drivers through Evo 2–imaging correlation. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Novel gene discoveries, pathway context and robustness. [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

24 extracted references · 1 linked inside Pith

  1. [1]

    Pan-cancer analysis of whole genomes.Nature578, 82–93 (2020)

    The ICGC/TCGA Pan-Cancer Analysis of Whole Genomes Consortium. Pan-cancer analysis of whole genomes.Nature578, 82–93 (2020)

  2. [2]

    R., Takahashi, K., Futreal, P

    Watson, I. R., Takahashi, K., Futreal, P. A. & Chin, L. Emerging patterns of somatic mutations in cancer.Nat. Rev. Genet.14, 703–718 (2013)

  3. [3]

    Mutational landscape of cancer-driver genes across human cancers.Sci

    Sinkala, M. Mutational landscape of cancer-driver genes across human cancers.Sci. Rep. 13, 12742 (2023)

  4. [4]

    Kinnersley, B., Sud, A., Everall, A. et al. Analysis of 10,478 cancer genomes identifies candidate driver genes and opportunities for precision oncology.Nat. Genet.56, 1868– 1877 (2024)

  5. [5]

    Sherman, M. A. et al. Genome-wide mapping of somatic mutation rates uncovers drivers of cancer.Nat. Biotechnol.40, 1634–1643 (2022)

  6. [6]

    Rocha, J. J. et al. Functional unknomics: systematic screening of conserved genes of unknown function.PLoS Biol.21, e3002222 (2023). 18

  7. [7]

    Greene, D. et al. Genetic association analysis of 77,539 genomes reveals rare disease etiologies.Nat. Med.29, 679–688 (2023)

  8. [8]

    Aerts, H. J. W. L. et al. Decoding tumour phenotype by noninvasive imaging using a quantitative radiomics approach.Nat. Commun.5, 4006 (2014)

  9. [9]

    Rios Velazquez, E. et al. Somatic mutations drive distinct imaging phenotypes in lung cancer.Cancer Res.77, 3922–3930 (2017)

  10. [10]

    Zheng, D. et al. radioGW AS links radiome to genome to discover driver genes with somatic mutations for heterogeneous tumor image phenotype in pancreatic cancer.Sci. Rep.14, 12316 (2024)

  11. [11]

    Dalla-Torre, H., Gonzalez, L., Mendoza-Revilla, J. et al. Nucleotide Transformer: build- ing and evaluating robust foundation models for human genomics.Nat. Methods22, 287–297 (2025)

  12. [12]

    Avsec, ˇZ., Latysheva, N., Cheng, J. et al. Advancing regulatory variant effect prediction with AlphaGenome.Nature649, 1206–1218 (2026)

  13. [13]

    Brixi, G. et al. Genome modelling and design across all domains of life with Evo 2. Nature652, 1349–1361 (2026)

  14. [14]

    Ohno-Machado, L., Zhu, R., Zhou, X. et al. Advancing human population genomics with DNA foundation models. Preprint atResearch Square(2025)

  15. [15]

    F., Kohl, S

    Isensee, F., Jaeger, P. F., Kohl, S. A. A., Petersen, J. & Maier-Hein, K. H. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation.Nat. Methods18, 203–211 (2021)

  16. [16]

    Wasserthal, J. et al. TotalSegmentator: robust segmentation of 104 anatomic structures in CT images.Radiol. Artif. Intell.5, e230024 (2023). 19

  17. [17]

    Heller, N. et al. The KiTS21 challenge: automatic segmentation of kidneys, renal tumors, and renal cysts in corticomedullary-phase CT. Preprint atarXiv2307.01984 (2023)

  18. [18]

    Garrucho, L. et al. A large-scale multicenter breast cancer DCE-MRI benchmark dataset with expert segmentations.Sci. Data12, 453 (2025)

  19. [19]

    Chakravarty, D. et al. OncoKB: a precision oncology knowledge base.JCO Precis. Oncol. 1, PO.17.00011 (2017)

  20. [20]

    Comprehensive molecular characterization of clear cell renal cell carcinoma.Nature499, 43–49 (2013)

    Cancer Genome Atlas Research Network. Comprehensive molecular characterization of clear cell renal cell carcinoma.Nature499, 43–49 (2013)

  21. [21]

    Landrum, M. J. et al. ClinVar: improving access to variant interpretations and support- ing evidence.Nucleic Acids Res.46, D1062–D1067 (2018)

  22. [22]

    & Goto, S

    Kanehisa, M. & Goto, S. KEGG: Kyoto Encyclopedia of Genes and Genomes.Nucleic Acids Res.28, 27–30 (2000)

  23. [23]

    & Hochberg, Y

    Benjamini, Y. & Hochberg, Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing.J. R. Stat. Soc. Ser. B57, 289–300 (1995)

  24. [24]

    E., Moore, B., Amode, R

    Hunt, S. E., Moore, B., Amode, R. M. et al. Annotating and prioritizing genomic variants using the Ensembl Variant Effect Predictor—a tutorial.Hum. Mutat.43, 986–997 (2022). Author contributions F.H. conceived the study, developed the analysis pipeline, performed all experiments and wrote the manuscript. J.K. contributed to data analysis and interpretatio...