Pith. sign in

REVIEW 4 major objections 5 minor 70 references

Knowledge-guided Contextual Gene Set Analysis Using Large Language Models

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper presents cGSA, a context-aware LLM pipeline for gene set analysis, and claims it beats g:Profiler, Llama 3.1-70B, and GPT-4o by 30–37% on 102 manually curated disease gene sets while outputting 14 pathways instead of 683.

desk verdict New context-aware GSA pipeline and a useful 102-set benchmark, but the headline 30% gain is calibrated on cGSA's own outputs and needs external validation. read the letter →

arxiv 2506.04303 v1 pith:KXM5UAIO submitted 2025-06-04 q-bio.GN cs.AIcs.LG

classification q-bio.GNcs.AIcs.LG
keywords genesetanalysispathwayenrichmentlargelanguagemodelscontext-awareprioritizationdifferentialexpressionprotein-proteininteractionnetworkcommunitydetectionbiomedicalsemanticsimilarity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Gene set analysis is how researchers turn a list of disease-related genes into biological hypotheses, but standard tools treat every gene list the same, so they return hundreds of enriched pathways, many generic or redundant, forcing manual filtering. This paper introduces cGSA, which first partitions the input gene list into protein–protein interaction communities, runs ordinary enrichment analysis on each community, and then lets a large language model filter and summarize the enriched pathways against the experimental context and research objective supplied by the user. The central claim is that injecting this context into the LLM step yields pathway lists that are both more accurate and more complete: on 102 manually curated differentially expressed gene sets, cGSA achieves a harmonic mean of 0.878 at a 0.5 semantic-similarity threshold, a gain the paper reports as 30% over g:Profiler, 33% over Llama 3.1-70B, and 37% over GPT-4o. If that comparison holds, the practical payoff is that a researcher can get a short, interpretable, context-relevant pathway list instead of manually scanning a long enrichment report, making functional interpretation faster and more reproducible.

What carries the argument

The mechanism that carries the argument is a four-step pipeline: (1) gene cluster detection—build a protein–protein interaction network from STRING and partition the input genes into non-overlapping communities with the EdMot algorithm; (2) per-cluster enrichment—query Enrichr with each cluster using p<0.001 to obtain a candidate pathway list; (3) LLM pathway screening—ask GPT-4o to score each candidate pathway's relevance to the user-provided experimental conditions; and (4) LLM pathway summarization—select one representative pathway per cluster guided by the stated research objective and assign it a relevance score. The crucial design choice is that the LLM is constrained to select and condense already-significant enrichment hits rather than generating free-text functions from scratch, which the paper argues prevents both hub-gene-dominated redundancy and hallucinated generic pathway names. The relevance score is calibrated against the number of database genes and human judgment, with cutoffs of 5.5 for moderate and 7.0 for high contextual relevance.

What would settle it

An independent replication with a separate panel of annotators who re-derive ground-truth pathways from the same 102 DEG sets without seeing the original articles, combined with a similarity threshold fixed at 0.5 before any model output is scored, would settle whether the reported 30% margin persists; if cGSA's harmonic mean advantage over g:Profiler falls below the claimed margin under those conditions, the headline result is tied to the original benchmark and threshold calibration.

Watch

Extended reading notes

Core claim

The paper's core discovery, stated on its own terms, is that context-aware prioritization—not bigger models—is what fixes gene set interpretation. cGSA does not ask the LLM to invent biological functions from gene names; it feeds the LLM the statistically enriched pathways from each gene cluster and asks it to keep those that fit the stated experimental conditions and then to condense each cluster to one representative pathway aligned with the research objective. Benchmarked on 102 manually curated DEG sets drawn from 31 PubMed articles spanning 19 diseases and ten biological mechanisms, the paper reports a harmonic mean of 0.878 for accuracy and hit ratio at MedCPT similarity threshold 0.5, with a margin over g:Profiler, Llama 3.1-70B, and GPT-4o that it states as 30%, 33%, and 37%, respectively, while emitting 14 pathways per input set on average versus 683 for g:Profiler. Two independent biologists, with 90.0% inter-annotator agreement, rated cGSA's outputs at harmonic mean 0.850, close to the automatic 0.878, which the paper uses to argue that the semantic-similarity evaluation is trustworthy. Ablation results attribute 16.4% of the gain to the gene-cluster detection step; two case studies in melanoma immune checkpoint blockade and breast cancer tumor microenvironments are presented as evidence that the pipeline produces testable, context-specific hypotheses.

Load-bearing premise

The load-bearing assumption is that the 102 manually curated DEG sets—extracted from 31 PubMed articles together with their author-stated pathway annotations—are unbiased, complete, and contextually correct targets, and that the 0.5 semantic-similarity threshold calibrated on cGSA's own expert-validated outputs is not lenient in a way that inflates cGSA's advantage over the baselines.

Editorial extensions

If this is right

  • If cGSA's benchmark result is right, standard practice in functional genomics can shift from scanning long enrichment reports to reading a short ranked list of context-filtered pathways, with the 683-versus-14 reduction in output size as the headline operational gain.
  • The reported 16.4% contribution of gene-cluster detection implies that PPI-based community structure is not an optional refinement but a core part of the accuracy gain, so future GSA methods should cluster before enriching rather than enriching whole gene lists.
  • The pipeline's robustness to context lengths down to roughly 2.6-word keywords means a user who only supplies a disease or cell type can still get context-aware prioritization, not just users who write detailed study objectives.
  • Applied to the melanoma ICB case, cGSA distinguishes gene programs selected in responders (metabolism, transcriptional adaptation, cell migration, immune response) from those in non-responders (cell cycle), aligning with the invasive versus proliferative subclone phenotypes.
  • In the breast cancer anti-PD1 case, cGSA identifies FKBP5 in T cells suppressing HCAR2 in myeloid cells and TRAF1 in T cells activating NFKBIZ in endothelial cells as candidate causal links, providing concrete hypotheses for experimental follow-up.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the fairest test of the headline 30% margin would use a similarity threshold chosen before seeing any method's output, or separately calibrated on each baseline; because the 0.5 MedCPT threshold was tuned on cGSA's own human-validated outputs, an independent re-estimation could shrink the reported gap.
  • Editorial inference: since the benchmark contexts were drawn from the same articles that supplied the ground-truth pathway annotations, a stronger validation would re-annotate the same DEG sets with contexts written independently (for example, from a disease name alone) and measure how much the harmonic mean drops.
  • Editorial inference: the same pipeline could be extended beyond pathway names to output evidence-grounded claims by returning, for each summarized pathway, the supporting database genes and the LLM's rationale, which would make the expert-validation step checkable at scale.
  • Editorial inference: nothing in the method is specific to disease DEGs, so a plausible generalization is context-aware functional interpretation for single-cell cluster markers, proteomics, or species beyond human and mouse, though the current benchmark only supports human and mouse DEGs.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript introduces cGSA, a context-aware gene set analysis pipeline that integrates STRING-based gene clustering (EdMot), Enrichr enrichment analysis, and GPT-4o-based pathway screening and summarization. The authors claim that cGSA identifies pathways that are both statistically significant and biologically meaningful, reporting a harmonic mean of 0.878 at a 0.5 similarity threshold on 102 manually curated DEG sets, corresponding to a 30% improvement over g:Profiler, 33% over Llama 3.1-70B, and 37% over GPT-4o. The paper also presents ablations, two case studies, and a publicly accessible demo.

Significance. If the reported performance holds, cGSA would be a valuable contribution to practical gene set interpretation, reducing the manual effort needed to filter redundant or contextually irrelevant enrichment results. The manuscript is generally well structured, and the authors share data files, prompts, and a demo, which supports reproducibility. The central claim is, however, heavily dependent on the evaluation protocol: the 0.5 similarity threshold is calibrated on human annotations of cGSA's own outputs, the ground truths are curated from the same articles that provide the contextual inputs, and human validation of baseline methods is absent. These issues need to be addressed before the headline comparison can be accepted as reliable.

major comments (4)
  1. [Results, 'cGSA Significantly Outperforms Current Methods'; Automatic Evaluations (Methods)] The headline comparison at the 0.5 similarity threshold is post-hoc and asymmetric. The threshold is selected because it 'best aligns with the human verification result' for cGSA's outputs, and the human validation was performed only on cGSA-generated pathways (the manuscript explicitly notes this in 'High Consistency Between Automatic and Expert Annotation'). Applying a threshold calibrated on cGSA's outputs to baselines can systematically favor cGSA if, for example, MedCPT embeddings score GPT-4o-style free-text summaries as more similar to ground-truth phrases than g:Profiler's database terms, independent of biological correctness. The paper should either calibrate the threshold on all methods (or on a threshold-independent metric such as average precision), or provide human validation for the baselines. Without this, the specific 30%/33%/37% margin at threshold 0.5 is not a reliable comparative claim, even if the threshold-robustness analysis in Figure 2a is encouraging.
  2. [Data Collection (Methods) and 'Results'] The benchmark exhibits a circularity risk: the ground-truth pathways are manually extracted from the same PubMed articles that supply the 'research objectives' and 'experimental conditions' used as context for cGSA. Thus cGSA is in effect asked to reproduce the original authors' pathway selections from the same source texts, while baselines receive no such context. This does not invalidate the pipeline as a text-grounded summarization tool, but it undermines the claim that the benchmark measures generalizable biological relevance. The authors should present an evaluation where contexts and ground truths are derived from independent sources (e.g., contexts from one article and ground truths from a separate validation study), or at least quantify how much of cGSA's advantage disappears when context is paraphrased or generated independently.
  3. [Results, 'Error Analysis' and 'Relevance Score Calibration'] The asymmetry of error reporting is a load-bearing issue for the comparison. The paper concedes that cGSA misses 16.5% of ground truths and has 9.9% incorrect predictions, and it analyzes these errors in depth. No equivalent error analysis is provided for g:Profiler, Llama 3.1-70B, or GPT-4o, even though those methods are also scored with the same automatic MedCPT metric. Since ACC and HIT are computed against the same ground truths, the reported margins may reflect not only higher precision and recall but also differing failure modes that are invisible without per-method error analysis. At minimum, the paper should report the per-method counts of missed ground truths and incorrect predictions, and ideally provide a small human-annotated sample for the baselines.
  4. [Methods, Eq. (1) and Figure 4a] The contribution of gene cluster detection is not isolated from the confounding effect of context. The ablation in Figure 4a removes the entire clustering step, but the LLM still receives the same context and the same enrichment input. The 16.4% improvement could partly stem from the clustering step reducing redundancy in the candidate pathway list, which is a legitimate benefit, but the example given (PMC11350455) suggests that the clustering step also changes the pathway vocabulary (e.g., from KRAS signaling to mTORC1 signaling). The claim that clustering 'primarily' improves ground-truth coverage would be stronger if the authors also reported HIT and ACC separately for the ablation, not only the harmonic mean. Please clarify whether ACC, HIT, or both drive the 16.4% improvement.
minor comments (5)
  1. [Introduction] There is a typo in the Introduction: 'disease threptic targets' should likely be 'disease therapeutic targets'.
  2. [Results, 'cGSA Significantly Outperforms Current Methods'] The sentence 'necessitating extensive manual review of the entire result list in order not to avoid missing key pathways' is confusingly worded; it should be 'in order not to miss key pathways'.
  3. [Materials and Methods, 'Statistical analysis'] The manuscript mentions 'GPT-3.5 (version 20230613)' in the methods, but the results and figure captions only compare GPT-4o, GPT-4, Llama 3.1-70B, and g:Profiler. Clarify whether GPT-3.5 was included in the main comparison and, if not, why.
  4. [Figure 2b caption] The caption says 'GPT-4(o) indicates that both GPT-4 and GPT-4o produced the same average number of pathways per DEGs.' This is an unusual notation; consider displaying both bars or explaining the equivalence in the text.
  5. [Methods, 'Human Validations by Biologists'] The manuscript reports inter-annotator agreement as 90.0% but does not describe the calculation beyond 'absolute consistency'. Specify whether this is Cohen's kappa, percentage agreement, or another measure, and report the denominator.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the enrichment pipeline is independent, though the 0.5 similarity threshold is calibrated on cGSA's own human-validated outputs.

full rationale

The claimed derivation chain — gene clustering via STRING/EdMot, Enrichr enrichment, LLM filtering and summarization using user-provided context, and evaluation against manually curated ground truths — is self-contained. The ground truths are curated from the source articles and are not derived from cGSA's outputs, so the predicted pathways are not defined in terms of the evaluation labels. The headline 0.5 MedCPT threshold is transparently calibrated on cGSA's human-validated outputs (average similarity 0.543), which creates a fairness concern for the baseline comparison because human annotation was limited to cGSA; the paper explicitly states 'manual annotations were limited to our cGSA generated pathways (instead of predictions by all other competing methods) due to its high cost.' This is an evaluation-calibration limitation, not a circular reduction: the ACC/HIT harmonic mean is not forced by the threshold, and the paper additionally reports cGSA's average harmonic mean across thresholds 0.5–0.9 and a human-validated harmonic mean of 0.850, both of which independently support the performance ordering. The self-citations present (MedCPT as evaluation encoder, GeneAgent and Hu et al. as methodological antecedents) are either published, code-reproduced tools or non-load-bearing borrowings; no load-bearing argument reduces to an unverified self-citation. No equation-level identity or fitted-parameter-renamed-as-prediction is present, so no circular step meets the evidentiary bar.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the curated benchmark and the calibration of thresholds on that benchmark; the biological assumptions about PPI clustering and semantic similarity are standard but unverified here. No new biological entities are introduced.

free parameters (5)
  • Enrichment p-value threshold (alpha) = 0.001
    Used to filter enriched pathways per cluster in Step 2; chosen 'according to experimental results' without a reported procedure.
  • Semantic similarity threshold = 0.5
    Used to count a predicted pathway as matching a ground truth in the main benchmark; selected to align with human agreement on cGSA outputs, then applied to baselines.
  • Relevance score cutoffs = 5.5 and 7.0
    Set from Figure 3c/d based on the 1,469 cGSA outputs to define moderate and high contextual relevance; used to rank output pathways.
  • STRING confidence cutoff = 0.05
    Interactions with confidence score greater than 0.05 are kept; this very low cutoff affects clustering and is not justified beyond the reference.
  • EdMot hyperparameters (Theta)
    Hyperparameters controlling community detection are not specified in the paper and are taken from Ref. 34.
assumptions (4)
  • domain assumption Ground-truth pathway annotations from the 31 source articles are correct and contextually relevant for the DEG sets.
    The benchmark treats the original authors' selected pathways as the gold standard; this is stated in Data Collection but not independently verified.
  • domain assumption MedCPT cosine similarity is a valid proxy for biological pathway relatedness.
    The automatic evaluation uses MedCPT embeddings to match predicted and ground-truth pathway names; the 0.5 threshold is calibrated on cGSA outputs.
  • domain assumption Community structure from PPI networks (STRING + EdMot) corresponds to functional gene modules.
    The first pipeline step assumes that clustering genes by PPI interactions produces biologically meaningful groups.
  • domain assumption LLM outputs at temperature 0.0 are deterministic and faithful enough for pathway screening and scoring.
    The workflow relies on GPT-4o generating consistent relevance scores and representative pathway names, with no human verification of each intermediate step.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Knowledge-guided Contextual Gene Set Analysis Using Large Language Models." pith.science (2026). https://pith.science/paper/KXM5UAIO

@misc{pith2026250604303,
  author       = {Pith},
  title        = {Pith review of: Knowledge-guided Contextual Gene Set Analysis Using Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KXM5UAIO}},
  note         = {Machine review of arXiv:2506.04303}
}
read the original abstract

Gene set analysis (GSA) is a foundational approach for interpreting genomic data of diseases by linking genes to biological processes. However, conventional GSA methods overlook clinical context of the analyses, often generating long lists of enriched pathways with redundant, nonspecific, or irrelevant results. Interpreting these requires extensive, ad-hoc manual effort, reducing both reliability and reproducibility. To address this limitation, we introduce cGSA, a novel AI-driven framework that enhances GSA by incorporating context-aware pathway prioritization. cGSA integrates gene cluster detection, enrichment analysis, and large language models to identify pathways that are not only statistically significant but also biologically meaningful. Benchmarking on 102 manually curated gene sets across 19 diseases and ten disease-related biological mechanisms shows that cGSA outperforms baseline methods by over 30%, with expert validation confirming its increased precision and interpretability. Two independent case studies in melanoma and breast cancer further demonstrate its potential to uncover context-specific insights and support targeted hypothesis generation.

Figures

Figures reproduced from arXiv: 2506.04303 by the authors.

Figure 2
Figure 2. cGSA outperforms existing approaches by a significant margin. a, the harmonic means of accuracy (ACC) and hit ratio (HIT) achieved by cGSA and other baseline methods across different similarity score thresholds. HIT represents the proportion of ground truths covered by the cGSA’s output pathways, while ACC denotes the proportion of pathways aligned with the ground truths in all output pathways. “**” indicates a sign… view at source ↗
Figure 3
Figure 3. Annotation analysis of pathways identified by cGSA. a, the calculation of hit ratio (HIT) and accuracy (ACC) for output pathways (n=1,469) with the same functional category as the ground truths (n=1,671), i.e., function-related pathways. HIT represents the proportion of ground truths covered by function-related pathways, while ACC denotes the proportion of function-related pathways in all identified pathways. b, the… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

70 extracted references · 65 canonical work pages

  1. [1]

    Fatumo, T

    S. Fatumo, T. Chikowore, A. Choudhury, M. Ayub, A. R. Martin, K. Kuchenbaecker, A roadmap to increase diversity in genomic studies. Nature medicine, 28(2), 243-250 (2022)

  2. [2]

    Q. Wang, M. Wang, I. Choi, L. Sarrafha, M. Liang, L. Ho, K. Farrell, K.G. Beaumont, R. Sebra, C. De Sanctis, J.F. Crary, Molecular profiling of human substantia nigra identifies diverse neuron types associated with vulnerability in Parkinson’s disease. Science advances , 10(2), p.eadi8287 (2024)

  3. [3]

    Kolligundla, K.M

    L.P. Kolligundla, K.M. Sullivan, D. Mukhi, M. Andrade-Silva, H. Liu, Y. Guan, X. Gu, J. Wu, T. Doke, D. Hirohama, P. Guarnieri, Glutathione-specific gamma–glutamylcyclotransferase 1 (CHAC1) increases kidney disease risk by modulating ferroptosis. Science Translational Medicine, 17(795), p.eadn3079 (2025)

  4. [4]

    Houtekamer, M.C

    R.M. Houtekamer, M.C. van der Net, M.J. Vliem, T.E. Noordzij, L. van Uden, R.M. van Es, J.Y. Sim, E., Deguchi, K. Terai, M.A. Hopcroft, H.R. Vos, E -cadherin mechanotransduction activates EGFR-ERK signaling in epithelial monolayers by inducing ADAM-mediated ligand shedding. Science Signaling, 18(886), p.eadr7926 (2025)

  5. [5]

    Zhou, J.G

    B. Zhou, J.G. Arthur, H. Guo, T. Kim, Y. Huang, R. Pattni, T. Wang, S. Kundu, J.X. Luo, H. Lee, D.C., Nachun, Detection and analysis of complex structural variation in human genomes across populations and in brains of donors with psychiatric disorders. Cell, 187(23), pp.6687- 6706 (2024)

  6. [6]

    A. Warr, C. Robert, D. Hume, A. Archibald, N. Deeb, M. Watson, Exome sequencing: current and future perspectives. G3: Genes, Genomes, Genetics, 5(8), 1543-1550 (2015)

  7. [7]

    Hoang, C.H

    M.L. Hoang, C.H. Chen, V.S. Sidorenko, J. He, K.G. Dickman, B.H. Yun, M. Moriya, N. Niknafs, C. Douville, R. Karchin, R.J. Turesky, Mutational signature of aristolochic acid exposure as revealed by whole -exome sequencing. Science translational medicine, 5(197), pp.197ra102-197ra102 (2013)

  8. [8]

    Pickrell, J.C

    J.K. Pickrell, J.C. Marioni, A.A. Pai, J.F. Degner, B.E. Engelhardt, E. Nkadori, J.B. Veyrieras, M. Stephens, Y. Gilad, J. K. Pritchard, Understanding mechanisms underlying human gene expression variation with RNA sequencing. Nature, 464(7289), pp.768-772 (2010)

Show all 70 references
  1. [9]

    Parekh, Y

    U. Parekh, Y. Wu, D. Zhao, A. Worlikar, N. Shah, K. Zhang, P. Mali, Mapping cellular reprogramming via pooled overexpression screens with paired fitness and single -cell RNA- sequencing readout. Cell systems, 7(5), pp.548-555 (2018)

  2. [10]

    Lazzarotto, N.L

    C.R. Lazzarotto, N.L. Malinin, Y. Li, R. Zhang, Y. Yang, G. Lee, E. Cowley, Y. He, X. Lan, K. Jividen, V. Katta, CHANGE-seq reveals genetic and epigenetic effects on CRISPR –Cas9 genome-wide activity. Nature biotechnology, 38(11), pp.1317-1327 (2020)

  3. [11]

    Kumar, S

    P. Kumar, S. Kiran, S. Saha, Z. Su, T. Paulsen, A. Chatrath, Y. Shibata, E. Shibata, A. Dutta, ATAC-seq identifies thousands of extrachromosomal circular DNA in cancer and cell lines. Science advances, 6(20), p.eaba2489 (2020)

  4. [12]

    X. Zhou, P. Chen, Q. Wei, X. Shen, X. Chen, Human interactome resource and gene set linkage analysis for the functional interpretation of biologically meaningful gene sets. Bioinformatics, 29(16), pp.2024-2031 (2013)

  5. [13]

    K. Ma, S. Huang, K.K. Ng, N.J. Lake, S. Joseph, J. Xu, A. Lek, L. Ge, K.G. Woodman, K.E. Koczwara, J., Cohen, Saturation mutagenesis-reinforced functional assays for disease-related genes. Cell, 187(23), pp.6707-6724 (2024)

  6. [14]

    Chang, M

    L.Y. Chang, M. Z. Lee, Y. Wu, W.K. Lee, C.L. Ma, J.M. Chang, C.W. Chen, T.C. Huang, C. H. Lee, J.C Lee, Y.Y. Tseng, Gene set correlation enrichment analysis for interpreting and annotating gene expression profiles. Nucleic acids research, 52(3), pp.e17-e17 (2024)

  7. [15]

    Subramanian, P

    A. Subramanian, P. Tamayo, V.K. Mootha, S. Mukherjee, B.L. Ebert, M.A. Gillette, A. Paulovich, S.L. Pomeroy, T.R. Golub, E.S. Lander, J.P. Mesirov, Gene set enrichment analysis: a knowledge -based approach for interpreting genome -wide expression profiles. Proceedings of the N...

  8. [16]

    Liu, F.F

    C.J. Liu, F.F. Hu, G.Y. Xie, Y.R. Miao, X.W. Li, Y. Zeng, A.Y. Guo, GSCA: an integrated platform for gene set cancer analysis at genomic, pharmacogenomic and immunogenomic levels. Briefings in bioinformatics, 24(1), p.bbac558 (2023)

  9. [17]

    Raudvere, L

    U. Raudvere, L. Kolberg, I. Kuzmin, T. Arak, P. Adler, H. Peterson, J. Vilo, g: Profiler: a web server for functional enrichment analysis and conversions of gene lists (2019 update). Nucleic acids research, 47(W1), pp.W191-W198 (2019)

  10. [18]

    M. V. Kuleshov, M.R. Jones, A.D. Rouillard, N.F. Fernandez, Q. Duan, Z. Wang, S. Koplev, S.L. Jenkins, K.M. Jagodnik, A. Lachmann, M.G. McDermott, Enrichr: a comprehensive gene set enrichment analysis web server 2016 update. Nucleic acids research, 44(W1), pp.W90- W97 (2016)

  11. [19]

    S. Xu, E. Hu, Y. Cai, Z. Xie, X. Luo, L. Zhan, W. Tang, Q. Wang, B. Liu, R. Wang, W. Xie, Using clusterProfiler to characterize multiomics data. Nature protocols, 19(11), pp.3292-3320 (2024)

  12. [20]

    Q. Jin, Y. Yang, Q. Chen, Z. Lu, Genegpt: Augmenting large language models with domain tools for improved access to biomedical information. Bioinformatics, 40(2), p.btae075 (2024)

  13. [21]

    M. Hu, S. Alkhairy, I. Lee, R. T. Pillich, D. Fong, K. Smith, R. Bachelder, T. Ideker, D. Pratt, Evaluation of large language models for discovery of gene set function. Nature methods, 22(1), pp.82-91 (2025)

  14. [22]

    Z. Wang, Q. Jin, C. H. Wei, S. Tian, P. T. Lai, Q. Zhu, C. P. Day, C. Ross, Z. Lu, GeneAgent: self-verification language agent for gene set knowledge discovery using domain databases. arXiv preprint arXiv:2405.16205 (2024)

  15. [23]

    Aravind, Guilt by association: contextual information in genome analysis

    L. Aravind, Guilt by association: contextual information in genome analysis. Genome Research, 10(8), pp.1074-1077 (2000)

  16. [24]

    Mooney, J.T

    M.A. Mooney, J.T. Nigg, S.K. McWeeney, B. Wilmot, Functional and genomic context in pathway analysis of GWAS data. Trends in Genetics, 30(9), pp.390-400 (2014)

  17. [25]

    G. S. França, M. Baron, B. R. King, J. P. Bossowski, A. Bjornberg, M. Pour, A. Rao, A. S. Patel, S. Misirlioglu, D. Barkley, K. H. Tang, Cellular adaptation to cancer therapy along a resistance continuum. Nature, 631(8022), pp.876-883 (2024)

  18. [26]

    L. A. Baldwin, N. Bartonicek, J. Yang, S. Z. Wu, N. Deng, D. L. Roden, C. L. Chan, G. Al - Eryani, D. J. Zanker, B. S. Parker, A. Swarbrick, DNA barcoding reveals ongoing immunoediting of clonal cancer populations during metastatic progression and immunotherapy response. Natur...

  19. [27]

    J. Peng, Y. Zhou, K. Wang, Multiplex gene and phenotype network to characterize shared genetic pathways of epilepsy and autism. Scientific reports, 11(1), p.952 (2021)

  20. [28]

    Romanovsky, K

    E. Romanovsky, K. Kluck, I. Ourailidis, M. Menzel, S. Beck, M. Ball, D. Kazdal, P. Christopoulos, P. Schirmacher, T. Stiewe, A. Stenzinger, Homogenous TP53mut-associated tumor biology across mutation and cancer types revealed by transcriptome analysis. Cell Death Discovery, 9(...

  21. [29]

    Eisfeld, S

    A.K. Eisfeld, S. Schwind, R. Patel, X. Huang, R. Santhanam, C.J. Walker, J. Markowitz, K.W. Hoag, T.M. Jarvinen, B. Leffel, D. Perrotti, Intronic miR -3151 within BAALC drives leukemogenesis by deregulating the TP53 pathway. Science signaling, 7(321), pp.ra36 -ra36 (2014)

  22. [30]

    Y. Liu, C. Chen, Z. Xu, C. Scuoppo, C.D. Rillahan, J. Gao, B. Spitzer, B. Bosbach, E.R. Kastenhuber, T. Baslan, S., Ackermann, Deletions linked to TP53 loss drive cancer through p53-independent mechanisms. Nature, 531(7595), pp.471-475 (2016)

  23. [31]

    Hirsch, S

    M.G. Hirsch, S. Pal, F.R. Mehrabadi, S. Malikic, C. Gruen, A. Sassano, E. Pérez-Guijarro, G. Merlino, S. C. Sahinalp, E. K. Molloy, C. P. Day, Stochastic modeling of single -cell gene expression adaptation reveals non-genomic contribution to evolution of tumor subclones. Cell ...

  24. [32]

    Bassez, H

    A. Bassez, H. Vos, L. Van Dyck, G. Floris, I. Arijs, C. Desmedt, B. Boeckx, M. Vanden Bempt, I. Nevelsteen, K. Lambein, K. Punie, A single -cell map of intratumoral changes during anti - PD1 treatment of patients with breast cancer. Nature medicine, 27(5), pp.820-832 (2021)

  25. [33]

    Mering, M

    C.V. Mering, M. Huynen, D. Jaeggi, S. Schmidt, P Bork, B. Snel, STRING: a database of predicted functional associations between proteins. Nucleic acids research, 31(1), pp.258-261 (2003)

  26. [34]

    P.Z. Li, L. Huang, C.D . Wang, J.H. Lai, EdMot: An edge enhancement approach for motif - aware community detection. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, pp. 479-487 (2019)

  27. [35]

    Hurst, A

    A. Hurst, A. Lerer, A.P. Goucher, A. Perelman, A. Ramesh, A. Clark, A.J. Ostrow, A. Welihinda, A. Hayes, A. Radford, A. Mądry, Gpt -4o system card. arXiv preprint arXiv:2410.21276 (2024)

  28. [36]

    Jayakrishnan, M

    M. Jayakrishnan, M. Havlová, V. Veverka, C. Regnard, P.B. Becker, Genomic context - dependent histone H3K36 methylation by three Drosophila methyltransferases and implications for dedicated chromatin readers. Nucleic Acids Research, 52(13), pp.7627-7649 (2024)

  29. [37]

    Porcu, M.C

    E. Porcu, M.C. Sadler, K. Lepik, C. Auwerx, A.R. Wood, A. Weihs, M.S.B. Sleiman, D.M. Ribeiro, S. Bandinelli, T. Tanaka, M. Nauck, Differentially expressed genes reflect disease - induced rather than disease -causing changes in the transcriptome. Nature communications , 12(1),...

  30. [38]

    Achiam, S

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F.L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, Gpt -4 technical report. arXiv preprint arXiv:2303.08774 (2023)

  31. [39]

    Grattafiori, A

    A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)

  32. [40]

    Joachimiak, J.H

    M.P. Joachimiak, J.H. Caufield, N.L. Harris, H. Kim, C.J. Mungall, Gene set summarization using large language models. ArXiv, pp.arXiv-2305 (2024)

  33. [41]

    Q. Jin, W. Kim, Q. Chen, D. C. Comeau, L. Yeganova, W. J. Wilbur, Z. Lu, Medcpt: Contrastive pre -trained transformers with large -scale pubmed search logs for zero -shot biomedical information retrieval. Bioinformatics, 39(11), p.btad651 (2023)

  34. [42]

    Tanaka, M

    A. Tanaka, M. Ogawa, Y. Zhou, Y. Otani, R.C. Hendrickson, M.M. Miele, Z. Li, D.S. Klimstra, J.Y. Wang, M.H., Roehrl, Proteogenomic characterization of pancreatic neuroendocrine tumors uncovers hypoxia and immune signatures in clinically aggressive subtypes. iScience, 27(8) (2024)

  35. [43]

    Rambow, A

    F. Rambow, A. Rogiers, O. Marin-Bejar, S. Aibar, J. Femel, M. Dewaele, P. Karras, D. Brown, Y.H. Chang, M. Debiec -Rychter, C. Adriaens, Toward minimal residual disease -directed therapy in melanoma. Cell, 174(4), pp.843-855 (2018)

  36. [44]

    L. Peng, F. Wang, Z. Wang, J. Tan, L. Huang, X. Tian, G. Liu, L. Zhou, Cell –cell communication inference and analysis in the tumour microenvironments from single -cell transcriptomics: data resources and computational strategies. Briefings in bioinformatics , 23(4), p.bbac234 (2022)

  37. [45]

    J. Lv, Y. Wei, J.H. Yin, Y.P. Chen, G.Q. Zhou, C. Wei, X.Y. Liang, Y. Zhang, C.J. Zhang, S.W. He, Q.M. He, The tumor immune microenvironment of nasopharyngeal carcinoma after gemcitabine plus cisplatin treatment. Nature Medicine, 29(6), pp.1424-1436 (2023)

  38. [46]

    Griffiths, P.A

    J.I. Griffiths, P.A. Cosgrove, E.F. Medina, A. Nath, J. Chen, F.R. Adler, J.T. Chang, Q.J. Khan, A.H. Bild, Cellular interactions within the immune microenvironment underpins resistance to cell cycle inhibition in breast cancers. Nature communications, 16(1), p.2132 (2025)

  39. [47]

    Goenka, F

    A. Goenka, F. Khan, B. Verma, P. Sinha, C. C. Dmello, M. P. Jogalekar, P. Gangadaran, B. C. Ahn, Tumor microenvironment signaling and therapeutics in cancer progression. Cancer Communications, 43(5), pp.525-561 (2023)

  40. [48]

    Böttcher, E

    J.P. Böttcher, E. Bonavita, P. Chakravarty, H. Blees, M. Cabeza -Cabrerizo, S. Sammicheli, N.C. Rogers, E. Sahai, S. Zelenay, C.R. e Sousa, NK cells stimulate recruitment of cDC1 into the tumor microenvironment promoting cancer immune control. Cell, 172(5), pp.1022-1037 (2018)

  41. [49]

    L. Li, Z. Lou, L. Wang, The role of FKBP5 in cancer aetiology and chemoresistance. British journal of cancer, 104(1), pp.19-23 (2011)

  42. [50]

    Habara, Y

    M. Habara, Y. Sato, T. Goshima, M. Sakurai, H. Imai, H. Shimizu, Y. Katayama, S. Hanaki, T. Masaki, M. Morimoto, S. Nishikawa, FKBP52 and FKBP51 differentially regulate the stability of estrogen receptor in breast cancer. Proceedings of the National Academy of Sciences, 119(15...

  43. [51]

    T. Liu, L. Zhang, D. Joo, S. C. Sun, NF-κB signaling in inflammation. Signal transduction and targeted therapy, 2(1), pp.1-9 (2017)

  44. [52]

    McCulloch, G.R

    T.R. McCulloch, G.R. Rossi, L. Alim, P.Y. Lam, J.K. Wong, E. Coleborn, S. Kumari, C. Keane, A.J. Kueh, M.J. Herold, C. Wilhelm, Dichotomous outcomes of TNFR1 and TNFR2 signaling in NK cell -mediated immune responses during inflammation. Nature Communications, 15(1), p.9871 (2024)

  45. [53]

    Jie, C.J

    Z. Jie, C.J. Ko, H. Wang, X. Xie, Y. Li, M. Gu, L. Zhu, J.Y. Yang, T. Gao, W. Ru, and S.J. Tang, Microglia promote autoimmune inflammation via the noncanonical NF -κB pathway. Science advances, 7(36), p.eabh0609 (2021)

  46. [54]

    E. C. Graff, H. Fang, D. Wanders, R. L. Judd, Anti -inflammatory effects of the hydroxycarboxylic acid receptor 2. Metabolism, 65(2), pp.102-113 (2016)

  47. [55]

    C. Zhao, H. Wang, Y. Liu, L. Cheng, B. Wang, X. Tian, H. Fu, C. Wu, Z. Li, C. Shen, J. Yu, Biased allosteric activation of ketone body receptor HCAR2 suppresses inflammation. Molecular Cell, 83(17), pp.3171-3187 (2023)

  48. [56]

    Q. Li, I. M. Verma, NF-κB regulation in the immune system. Nature reviews immunology , 2(10), pp.725-734 (2002)

  49. [57]

    Bradley, TNF‐mediated inflammatory disease

    J. Bradley, TNF‐mediated inflammatory disease. The Journal of Pathology: A Journal of the Pathological Society of Great Britain and Ireland, 214(2), pp.149-160 (2008)

  50. [58]

    Beutler, A

    B. Beutler, A. Cerami, Tumor necrosis, cachexia, shock, and inflammation: a common mediator. Annual review of biochemistry, 57(1), pp.505-518 (1988)

  51. [59]

    Yamamoto, S

    M. Yamamoto, S. Yamazaki, S. Uematsu, S. Sato, H. Hemmi, K. Hoshino, T. Kaisho, H. Kuwata, O. Takeuchi, K. Takeshige, T. Saitoh, Regulation of Toll/IL -1-receptor-mediated gene expression by the inducible nuclear protein IκBζ. Nature, 430(6996), pp.218-222 (2004)

  52. [60]

    Zhu, J.M

    B. Zhu, J.M. Park, S.R. Coffey, A. Russo, I.U. Hsu, J. Wang, C. Su, R. Chang, T.T. Lam, P.P. Gopal, S.D. Ginsberg, Single -cell transcriptomic and proteomic analysis of Parkinson’s disease Brains. Science Translational Medicine, 16(771), p.eabo1997 (2024)

  53. [61]

    L. Meng, H. Wu, J. Wu, P. A. Ding, J. He, M. Sang, L. Liu, Mechanisms of immune checkpoint inhibitors: Insights into the regulation of circular RNAS involved in cancer hallmarks. Cell Death & Disease, 15(1), p.3 (2024)

  54. [62]

    C. T. Small, I. Vendrov, E. Durmus, H. Homaei, E. Barry, J. Cornebise, T. Suzman, D. Ganguli, C. Megill, Opportunities and risks of LLMs for scalable deliberation with Polis. arXiv preprint arXiv:2306.11932 (2023)

  55. [63]

    B. Wen, C. Xu, R. Wolfe, L.L. Wang, B. Howe, Mitigating overconfidence in large language models: A behavioral lens on confidence estimation and calibration. In NeurIPS 2024 Workshop on Behavioral Machine Learning (2024)

  56. [64]

    Abbasi Yadkori, I

    Y. Abbasi Yadkori, I. Kuzborskij, A. György, C. Szepesvari, To believe or not to believe your llm: Iterative prompting for estimating epistemic uncertainty. Advances in Neural Information Processing Systems, 37, pp.58077-58117 (2024)

  57. [65]

    Azure OpenAI in Azure AI Foundry Models [Internet]

    Microsoft Corporation. Azure OpenAI in Azure AI Foundry Models [Internet]. Microsoft Learn; 2025 May 20 [cited 2025 May 23]. Available from: https://learn.microsoft.com/en- us/azure/ai-services/openai/concepts/models?tabs=global-standard%2Cstandard-chat- completions

  58. [66]

    T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, Transformers: State -of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system ...

  59. [67]

    Razin, E.S

    S.V. Razin, E.S. Ioudinkova, O.L. Kantidze, O.V. Iarovaia, Co- regulated genes and gene clusters. Genes, 12(6), p.907 (2021)

  60. [68]

    D. J. Clarke, G. B. Marino, E. Z. Deng, Z. Xie, J. E. Evangelista, A. Ma’ayan, Rummagene: massive mining of gene sets from supporting materials of biomedical research publications. Communications Biology, 7(1), p.482 (2024)

  61. [69]

    R. Y. Moreno, S. B. Panina, S. Irani, H. A. Hardtke, R. Stephenson, B. M. Floyd, E. M. Marcotte, Q. Zhang, Y. J. Zhang, Thr4 phosphorylation on RNA Pol II occurs at early transcription regulating 3′-end processing. Science Advances, 10(36), p.eadq0350 (2024)

  62. [70]

    **” indicates a significant improvement with p < 0.01, while “ *

    C. H. Wei, L. Luo, R. Islamaj, P. T. Lai, Z. Lu, GNorm2: an improved gene name recognition and normalization system. Bioinformatics, 39(10), p.btad599 (2023). Acknowledgments: We would like to thank M.G. Hirsch and Teresa M. Przytycka for their helpful discussion of this work....

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.