Pith. sign in

REVIEW 17 cited by

TxGemma: Efficient and Agentic LLMs for Therapeutics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.06196 v1 pith:URTBIWRM submitted 2025-04-08 cs.AI cs.CLcs.LG

TxGemma: Efficient and Agentic LLMs for Therapeutics

classification cs.AI cs.CLcs.LG
keywords modelstxgemmatherapeutichighllmsdevelopmentgeneralisto3-mini
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Therapeutic development is a costly and high-risk endeavor that is often plagued by high failure rates. To address this, we introduce TxGemma, a suite of efficient, generalist large language models (LLMs) capable of therapeutic property prediction as well as interactive reasoning and explainability. Unlike task-specific models, TxGemma synthesizes information from diverse sources, enabling broad application across the therapeutic development pipeline. The suite includes 2B, 9B, and 27B parameter models, fine-tuned from Gemma-2 on a comprehensive dataset of small molecules, proteins, nucleic acids, diseases, and cell lines. Across 66 therapeutic development tasks, TxGemma achieved superior or comparable performance to the state-of-the-art generalist model on 64 (superior on 45), and against state-of-the-art specialist models on 50 (superior on 26). Fine-tuning TxGemma models on therapeutic downstream tasks, such as clinical trial adverse event prediction, requires less training data than fine-tuning base LLMs, making TxGemma suitable for data-limited applications. Beyond these predictive capabilities, TxGemma features conversational models that bridge the gap between general LLMs and specialized property predictors. These allow scientists to interact in natural language, provide mechanistic reasoning for predictions based on molecular structure, and engage in scientific discussions. Building on this, we further introduce Agentic-Tx, a generalist therapeutic agentic system powered by Gemini 2.5 that reasons, acts, manages diverse workflows, and acquires external domain knowledge. Agentic-Tx surpasses prior leading models on the Humanity's Last Exam benchmark (Chemistry & Biology) with 52.3% relative improvement over o3-mini (high) and 26.7% over o3-mini (high) on GPQA (Chemistry) and excels with improvements of 6.3% (ChemBench-Preference) and 2.4% (ChemBench-Mini) over o3-mini (high).

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MedPRMBench: A Fine-grained Benchmark for Process Reward Models in Medical Reasoning

    cs.CL 2026-04 unverdicted novelty 8.0

    MedPRMBench is the first fine-grained benchmark for process reward models in medical reasoning, featuring 6500 questions, 13000 chains, 113910 step labels, and a baseline that improves downstream QA accuracy by 3.2-6....

  2. SurgiQ: A Large-Scale Multi-Domain Benchmark for Evaluating Surgical Understanding in Large Language Models

    cs.CL 2026-06 unverdicted novelty 7.0

    SurgiQ is a new 13k-question surgical benchmark showing general-purpose LLMs reach 68.1% accuracy while most biomedical models lag and smaller models stay near random baseline.

  3. VibeProteinBench: An Evaluation Benchmark for Language-interfaced Vibe Protein Design

    q-bio.QM 2026-05 unverdicted novelty 7.0

    VibeProteinBench is a new benchmark evaluating LLMs on open-ended language-interfaced protein design across recognition, engineering, and generation, with no model showing strong performance in all areas.

  4. VibeProteinBench: An Evaluation Benchmark for Language-interfaced Vibe Protein Design

    q-bio.QM 2026-05 unverdicted novelty 7.0

    VibeProteinBench is a three-stage language-interfaced benchmark revealing that no current LLM performs strongly across recognition, engineering, and generation of proteins.

  5. OmicsLM: A Multimodal Large Language Model for Multi-Sample Omics Reasoning

    q-bio.GN 2026-05 unverdicted novelty 7.0

    OmicsLM integrates continuous omics embeddings into LLMs for multi-sample biological reasoning, matching specialized models on profile tasks while outperforming them and general LLMs on language-guided QA over real ex...

  6. The limits of bio-molecular modeling with large language models : a cross-scale evaluation

    cs.LG 2026-04 unverdicted novelty 7.0

    LLMs perform adequately on bio-molecular classification tasks but remain weak on regression, with hybrid architectures outperforming others on long sequences and fine-tuning hurting generalization.

  7. TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

    q-bio.QM 2026-07 conditional novelty 6.5

    A 52.6B-token multi-domain biology pretraining corpus with tool enrichment and new binding/localization instructions doubles a fixed base LLM's matched biology-eval score with little language forgetting.

  8. Divergence Decoding: Training-Free Capability Fusion

    cs.AI 2026-07 conditional novelty 6.0

    Divergence Decoding routes each token to either a domain specialist or a general reasoning LLM based on Jensen-Shannon divergence, outperforming either model alone on most tested scientific tasks.

  9. TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

    q-bio.QM 2026-07 conditional novelty 6.0

    TheBioCollection, a 52.6B-token unified biology corpus with tool-computed text and new instruction tasks, raises a fixed 16B LLM's score on its matched biology eval from 0.223 to 0.499 (2.24×).

  10. MolDeTox: Evaluating Language Model's Stepwise Fragment Editing for Molecular Detoxification

    cs.AI 2026-05 unverdicted novelty 6.0

    MolDeTox is a new benchmark that shows fragment-level stepwise editing by LLMs and VLMs improves structural validity and detoxification quality over prior toxicity-focused evaluations.

  11. An explainable hypothesis-driven approach to Drug-Induced Liver Injury with HADES

    cs.AI 2026-05 unverdicted novelty 6.0

    HADES is an agentic AI system that generates mechanistic hypotheses for drug-induced liver injury using molecular, metabolite, and pathway evidence, outperforming prior binary classifiers on the new DILER benchmark wh...

  12. Evaluating Agentic Bioinformatics through Function, Evidence, and Validation

    cs.AI 2026-07 conditional novelty 5.0

    Agentic bioinformatics systems mostly demonstrate planning and tool execution but rarely prospective empirical validation, so the paper argues evaluation should center on inspectable workflow trajectories (FEV) rather...

  13. Benchmarking open-source tools for in silico antiviral drug discovery

    q-bio.BM 2026-05 conditional novelty 5.0

    Boltz-2 and fine-tuned DrugFormDTA lead ML-based binding prediction while GNINA leads docking tools on a cleaned antiviral dataset, with performance varying by viral protein.

  14. Bolek: A Multimodal Language Model for Molecular Reasoning

    cs.LG 2026-05 unverdicted novelty 5.0

    Bolek injects Morgan fingerprint embeddings into an instruction-tuned text model, then fine-tunes on molecular alignment and synthetic chain-of-thought tasks to improve performance and grounding on 15 TDC binary class...

  15. Beyond SMILES: Evaluating Agentic Systems for Drug Discovery

    q-bio.QM 2026-02 conditional novelty 5.0

    Drug-discovery AI agents are built for small-molecule, big-pharma settings and lack peptide, in vivo, training-loop, small-lab, and multi-objective capabilities, even though LLMs themselves can reason about peptides.

  16. ToxiEval-ZKP: A Structure-Private Verification Framework for Molecular Toxicity Repair Tasks

    cs.CR 2025-08 unverdicted novelty 5.0

    ToxiEval-ZKP applies zero-knowledge proofs to enable private verification that generative AI molecules meet multidimensional toxicity repair criteria.

  17. From Knowledge to Action: Outcomes of the 2025 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry

    cond-mat.mtrl-sci 2026-05 unverdicted novelty 2.0

    Hackathon submissions indicate LLMs are moving from general assistants toward composable multi-agent systems for structuring scientific knowledge and automating tasks in materials science and chemistry.