REVIEW 2 major objections 4 minor 224 references
AI biology agents should be judged by workflow correctness, not final answers alone.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 05:43 UTC pith:64NS3EAL
load-bearing objection Useful evaluation scaffold for agentic bioinformatics; the empirical V-stage coding is looser than the stated gate, so treat the counts as indicative. the 2 major comments →
Evaluating Agentic Bioinformatics through Function, Evidence, and Validation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is normative: agentic bioinformatics should be evaluated through workflow correctness rather than final-answer correctness alone. The paper operationalizes this as three non-interchangeable properties—demonstrated workflow operations (Function F1–F6), traceable support for actions and claims (Evidence E1–E6), and use-case-specific cumulative assurance (Validation V0–V4)—and maps 109 systems plus 28 benchmarks across six biological domains. The empirical finding is lopsided progress: 94 of 109 systems reach the V2 replayability gate, but only 7 reach V4 prospective empirical testing, and closed-loop empirical refinement appears in a single mapped system. If the paper is righ
What carries the argument
The Function–Evidence–Validation (FEV) framework. Function records what the system demonstrably does (planning, coordination, tool execution, state and trace maintenance, repair, verification). Evidence records traceable sources (literature, knowledge bases, measurements, software outputs, model outputs, experimental observations). Validation is a five-stage cumulative ladder from illustrative output (V0) to demonstrated execution (V1), replayable computation (V2), scientifically evaluated computation (V3), and prospective empirical evaluation (V4). The ladder is the load-bearing instrument: it separates 'it ran' from 'it can be replayed' from 'it was scientifically tested.'
Load-bearing premise
The headline numbers (94 of 109 at V2, 7 at V4) rest on the authors applying their own stated V2 criteria—identifiable inputs, parameters, dependencies, intermediate artifacts, and execution traces—consistently across 109 heterogeneous papers, and that coding has not been checked by independent raters.
What would settle it
Re-code the 109 mapped systems from the paper's own tables using only the explicit V2 minimum (inputs, parameters, dependencies, intermediate artifacts, and execution traces all present). If a stricter coder assigns substantially fewer than 94 systems to V2—for instance, because 'public code and instructions' is treated as insufficient without traceable execution artifacts—the claimed replayability gap and the V-stage distribution would shift together.
If this is right
- Benchmarks in agentic bioinformatics should grade trajectories—tool calls, parameters, artifacts, failure recovery—alongside final answers, rather than treating endpoint accuracy as the whole score.
- Published system claims should carry an explicit V-stage and qualifiers (benchmark, expert, statistics, robustness, external, prospective, closed-loop), so a reader can see what assurance a given use case actually has.
- The replayability gap (94 of 109 at V2, 7 at V4) implies that most current systems are inspectable but not empirically tested; funding and evaluation efforts should move toward prospective, claim-aligned experiments.
- A minimum reporting standard for agentic bioinformatics papers would follow from FEV: document workflow scope, models, tools, parameters, environments, artifacts, failures, approval points, and the validation stage.
Where Pith is reading between the lines
- If FEV became a reporting norm, the field's progress measures would shift from 'can it answer' to 'can it be audited'—a change that would likely re-rank many systems and reduce the prestige of broad-but-untested agents.
- The V-ladder could be extended with a V5 for closed-loop empirical refinement at scale, since the paper counts only one such system; the distinction between one-shot prospective testing (P) and feedback-driven cycles (C) is likely to become central as wet-lab integration grows.
- A testable extension would be an inter-rater reliability study of FEV coding: if independent coders disagree widely on V-stage assignments, the framework needs tighter operational definitions before it can serve as a community standard.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the Function–Evidence–Validation (FEV) framework for evaluating agentic bioinformatics systems. Function records demonstrable workflow operations (F1–F6), Evidence records traceable sources supporting actions and claims (E1–E6), and Validation records cumulative assurance stages (V0–V4) with orthogonal qualifiers. The authors apply FEV to 109 system entries and 28 benchmark resources, representing 128 unique publications, and present a cross-domain synthesis showing that planning and tool-mediated execution are common while replayability, external validation, and prospective empirical testing are less well established. They conclude that agentic bioinformatics should be assessed through workflow correctness rather than final-answer correctness alone.
Significance. The paper makes a timely and useful conceptual contribution. The FEV framework separates operational capability, evidentiary support, and validation assurance in a way that is more explicit than most existing reviews, and the supplementary tables provide an unusually detailed audit trail of 109 systems. The distinction between using empirical data as evidence and prospectively testing an agent-generated output is particularly valuable, as is the cumulative V0–V4 gate. The accounting is internally consistent (109 + 28 − 9 = 128 unique publications), and the authors are transparent about the review's scope and limitations. The main risk is that the quantitative V-stage synthesis is not calibrated: several V2/V3 assignments appear more lenient than the stated criteria, and no inter-rater reliability or replay check is reported. This weakens the specific numeric claims but does not invalidate the normative core, which would only be strengthened if replayability is even rarer than reported.
major comments (2)
- [§S2 Tables S8, S9, S23; Table S5; §S1] The V2 gate is defined in Table S5 as requiring identifiable inputs, parameters, dependencies, intermediate artifacts, AND execution traces. Several assignments use a looser bar. AI-HOPE (Table S8) is coded V2 [S] on 'executable analyses on identifiable retrospective datasets,' with no listed scripts, parameter records, logs, or traces. HEAL-KGGen (Table S9) is coded V2 [B] because 'public code, requirements, test data, graph files, and instructions support replay'; instructions and graph files are not execution traces. SwiftDossier (Table S23) is coded V2 [H,R] from 'executable retrieval and analysis artifacts' without the required intermediate artifacts/traces. Since V3 subsumes V2, the 69 V3 entries inherit this slack. Section S1 reports no inter-rater reliability or calibration exercise and no actual replay check. This undermines quantitative statements such as 'most systems are clas
- [§2 and §S1; Figures S3–S5] The eligibility boundary is deliberately broad, including 'agent-adjacent' systems, and §S1 states that aggregate analyses include both unless otherwise stated. However, the main quantitative synthesis does not report the full-agentic subset separately. Entries such as ChatNT (Table S7), VibeGen (Table S20), and ORI (Table S21) are coded as agent-adjacent predictive or model–laboratory loops, yet they contribute to the same FEV prevalence and V-stage distribution as full multi-agent workflow systems. The claim that 'planning and tool-mediated execution have advanced' across agentic bioinformatics is therefore hard to interpret. Please report the full-agentic-only distribution or provide a sensitivity analysis showing that the qualitative conclusions are unchanged when agent-adjacent entries are excluded.
minor comments (4)
- [References] References 49 and 50 are identical (Huang et al., 'Autonomous biomedical research with an artificial intelligence agent'). Please consolidate or distinguish them. References 117 and 118 also appear to be two versions of BioMaster and should be cross-referenced explicitly.
- [Figure S3b] The zero count for V0 is partly an artifact of the eligibility filter that excludes purely conversational systems. Add a note clarifying that V0 is retained on the complete scale but that no system meeting the inclusion criteria was assigned to it.
- [Abstract and §2] The term 'workflow correctness' is used in the abstract and conclusion but never given a compact definition in the body. A brief formal definition early in Section 2 would help readers understand the exact relationship between FEV and workflow correctness.
- [Figure S5b] The 'observed gap' score (1 − share) measures absence of reporting, not absence of capability. Since the paper codes only reported capabilities, consider relabeling this as a 'reporting gap' or adding an explicit sentence that unreported capabilities were treated as not demonstrated.
Circularity Check
No significant circularity: FEV is a proposed analytical framework applied to external systems; no prediction reduces to a fit or self-citation.
full rationale
The paper introduces the Function–Evidence–Validation framework as a normative analytical lens rather than deriving it from data. The central claim — 'agentic bioinformatics should be assessed through workflow correctness rather than final-answer correctness alone' — is an argument supported by the framework's definitions and by a structured mapping of 109 external systems and 28 benchmarks. There are no equations, no fitted parameters, and no quantity predicted from a subset of the same data. The V-stage assignments are coding judgments about other groups' systems, not outputs of a model fitted to those systems; even if some assignments are lenient relative to Table S5 (e.g., AI-HOPE, HEAL-KGGen, SwiftDossier receive V2 without the table's required intermediate artifacts and execution traces), that is a measurement-reliability concern about the review's empirical summary, not a circular reduction of the paper's conclusions to its inputs. The only mildly self-referential element is Table S2, where the authors' own review is the only row marked with full checks on the feature checklist that they themselves defined; this is a positioning table rather than load-bearing evidence for the normative claim, and it does not constitute circularity in the sense of a prediction or derivation that reduces to its inputs. The manuscript also explicitly disclaims that 'Unreported or insufficiently documented capabilities were not coded' (Supplementary Section S2), a transparent limitation rather than a circular step. No load-bearing self-citations or imported uniqueness theorems appear. The paper is therefore not circular; its principal risks are external validity and coding reliability, not derivation circularity.
Axiom & Free-Parameter Ledger
free parameters (2)
- V-stage assignment thresholds (V0–V4)
- Eligibility boundary: 'agentic or agent-adjacent'
axioms (3)
- domain assumption The inspectable workflow trajectory, not architecture or final output, is the correct primary unit of analysis for agentic bioinformatics.
- domain assumption Unreported or insufficiently documented capabilities can be treated as absent for the purpose of the field-wide map.
- domain assumption The structured sample of 128 publications is adequate to support temporal claims such as 'planning and tool-mediated execution have advanced more rapidly than...'.
read the original abstract
Large language model agents increasingly plan, execute, and interpret biological analyses, yet fluent responses, successful tool calls, and benchmark performance alone do not establish scientific credibility. Existing reviews primarily organize biological agents by application, architecture, and agentic capability, but do not jointly operationalize the accountability of agent-generated workflows. We address this gap by treating the inspectable workflow trajectory, rather than architecture or final output alone, as the primary unit of analysis. We introduce the Function--Evidence--Validation (FEV) framework, which separates demonstrated workflow operations, traceable support for actions and claims, and use-case-specific validation. Using FEV, we map 109 agentic or agent-adjacent systems and 28 benchmark or evaluation resources, representing 128 unique publications across genomics, single-cell and spatial omics, protein science, drug discovery, computational pathology, and general bioinformatics automation. Across domains, planning and tool-mediated execution have advanced more rapidly than replayability, provenance, robust scientific assessment, external validation, and prospective empirical testing. We therefore argue that agentic bioinformatics should be assessed through workflow correctness rather than final-answer correctness alone. FEV provides a practical basis for comparing systems and designing transparent, auditable, and scientifically accountable bioinformatics workflows.
Reference graph
Works this paper leans on
-
[1]
The galaxy platform for accessible, reproducible, and collaborative data analyses: 2024 update.Nucleic acids research, 52(W1):W83–W94, 2024
2024
-
[2]
Uniprot: the universal protein knowledgebase in 2025.Nucleic acids research, 53(D1):D609–D617, 2025
2025
-
[3]
The gene ontology knowledgebase in 2026.Nucleic Acids Research, 54(D1):D1779–D1792, 2026
2026
-
[4]
Llm4grn: Discovering causal gene regulatory networks with llms– evaluation through synthetic data generation
Tejumade Afonja, Ivaxi Sheth, Ruta Binkyte, Waqar Hanif, Shubhi Ambast, Charles Mwangi Kaumbutha, Matthias Becker, and Mario Fritz. Llm4grn: Discovering causal gene regulatory networks with llms– evaluation through synthetic data generation. InICLR 2025 Workshop on Machine Learning for Genomics Explorations
2025
-
[5]
Multi-agent ai enables evidence-based cell annotation in single-cell transcriptomics
Gautam Ahuja, Alex Antill, Yi Su, Giovanni Marco Dall ˘2019Olio, Sukhitha Basnayake, Göran Karlsson, and Parashar Dhapola. Multi-agent ai enables evidence-based cell annotation in single-cell transcriptomics. bioRxiv, pages 2025–11, 2025
2025
-
[6]
Cellvoyager: Ai compbio agent generates new insights by autonomously analyzing biological data.Nature Methods, 23(4):749–759, 2026
Samuel Alber, Bowen Chen, Eric Sun, Alina Isakova, Aaron J Wilk, and James Zou. Cellvoyager: Ai compbio agent generates new insights by autonomously analyzing biological data.Nature Methods, 23(4):749–759, 2026
2026
-
[7]
Tactic: An explainable multi-agent architecture for classification & interpretable reasoning in spatial transcriptomics
Abdel Rahman Alsabbagh, Mahmoud Zahran, Ali Balubaid, Sumeer Ahmad Khan, Robert Lehmann, Xabier Martinez de Morentin, Vincenzo Lagani, Narsis A Kiani, David Gomez-Cabrero, and Jesper Tegnér. Tactic: An explainable multi-agent architecture for classification & interpretable reasoning in spatial transcriptomics. InICML 2025 Generative AI and Biology (GenBio...
2025
-
[8]
Retrieval augmented generation for large language models in healthcare: A systematic review.PLOS Digital Health, 4(6):e0000877, 2025
Lameck Mbangula Amugongo, Pietro Mascheroni, Steven Brooks, Stefan Doering, and Jan Seidel. Retrieval augmented generation for large language models in healthcare: A systematic review.PLOS Digital Health, 4(6):e0000877, 2025
2025
-
[9]
Jonathan Bragg, Mike D’Arcy, Nishant Balepur, Dan Bareket, Bhavana Dalvi, Sergey Feldman, Dany Haddad, Jena D Hwang, Peter Jansen, Varsha Kishore, et al. Astabench: Rigorous benchmarking of ai agents with a scientific research suite.arXiv preprint arXiv:2510.21652, 2025
Pith/arXiv arXiv 2025
-
[10]
Francesco Branda, Mohamed M Ahmed, Massimo Ciccozzi, Pietro Hiram Guzzi, and Fabio Scarpa. The next paradigm in bioinformatics: a review of multi-agent systems and foundational models for end-to-end scientific discovery.Briefings in Bioinformatics, 27(3):bbag245–bbag245, 2026
2026
-
[11]
Empowering ai data scientists using a multi-agent llm framework with self- evolving capabilities for autonomous, tool-aware biomedical data analyses
Dechao Bu, Jingbo Sun, Kun Li, Zihao He, Wei Huang, Jinlin Hu, Shanshan Zhang, Shuangshuang Lei, Peipei Huo, Zhihao Wang, et al. Empowering ai data scientists using a multi-agent llm framework with self- evolving capabilities for autonomous, tool-aware biomedical data analyses. Nature Biomedical Engineering, pages 1–16, 2026
2026
-
[12]
Mozi: Governed autonomy for drug discovery llm agents.arXiv preprint arXiv:2603.03655, 2026
He Cao, Siyu Liu, Fan Zhang, Zijing Liu, Hao Li, Bin Feng, Shengyuan Bai, Leqing Chen, Kai Xie, and Yu Li. Mozi: Governed autonomy for drug discovery llm agents.arXiv preprint arXiv:2603.03655, 2026
arXiv 2026
-
[13]
Biomics: A foundational agent for grounded and autonomous multi-omics interpretation.bioRxiv, 2026
Lei Cao, Yuntain Li, Hua Qin, Yanbang Shang, Yilin Zhang, Bogdan Jovanovic, Lazar Djokic, Tianyi Xia, Luni Hu, Haiyang Hou, Xingxing Ning, Li’ang Lin, Hao Qiu, Ziqing Deng, Yuxiang Li, Yong Zhang, and Shuangsang Fang. Biomics: A foundational agent for grounded and autonomous multi-omics interpretation.bioRxiv, 2026
2026
-
[14]
Jingyun Chen, Linghan Cai, Zhikang Wang, Yi Huang, Songhan Jiang, Shenjin Huang, Hongpeng Wang, and Yongbing Zhang. Pathagent: Toward interpretable analysis of whole-slide pathology images via large language model-based agentic reasoning.arXiv preprint arXiv:2511.17052, 2025
arXiv 2025
-
[15]
Stat: A multi-agent framework for integrated and interactive spatial transcriptomics analysis
Yuheng Chen, Shi Han, Zitong Chao, Yuyao Liu, Fan Zhang, Hao Chen, Jiguang Wang, Jiashun Xiao, and Can Yang. Stat: A multi-agent framework for integrated and interactive spatial transcriptomics analysis. bioRxiv, pages 2026–05, 2026
2026
-
[16]
Scienceagentbench: Toward rigorous assessment of language agents for data-driven scientific discovery
Ziru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang, Boshi Wang, Botao Yu, Yifei Li, Zeyi Liao, Chen Wei, Zitong Lu, et al. Scienceagentbench: Toward rigorous assessment of language agents for data-driven scientific discovery. InInternational Conference on Learning Representations, volume 2025, pages 96934–96990, 2025
2025
-
[17]
Accurate proteome-wide missense variant effect prediction with alphamissense.Science, 381(6664):eadg7492, 2023
Jun Cheng, Guido Novati, Joshua Pan, Clare Bycroft, Akvil ˙e Žemgulyt˙e, Taylor Applebaum, Alexander Pritzel, Lai Hong Wong, Michal Zielinski, Tobias Sargeant, et al. Accurate proteome-wide missense variant effect prediction with alphamissense.Science, 381(6664):eadg7492, 2023
2023
-
[18]
Jihye Choi, Nils Palumbo, Prasad Chalasani, Matthew M Engelhard, Somesh Jha, Anivarya Kumar, and David Page. Malade: Orchestration of llm-powered agents with retrieval augmented generation for pharmacovigilance.arXiv preprint arXiv:2408.01869, 2024
Pith/arXiv arXiv 2024
-
[19]
Chip-gpt: a managed large language model for robust data extraction from biomedical database records.Briefings in bioinformatics, 25(2):bbad535, 2024
Olivier Cinquin. Chip-gpt: a managed large language model for robust data extraction from biomedical database records.Briefings in bioinformatics, 25(2):bbad535, 2024
2024
-
[20]
Olivier Cinquin. Steering veridical large language model analyses by correcting and enriching generated database queries: first steps toward chatgpt bioinformatics.Briefings in Bioinformatics, 26(1):bbaf045, 2025
2025
-
[21]
Scientific workflows for computational reproducibility in the life sciences: Status, challenges and opportunities.Future Generation Computer Systems, 75:284–298, 2017
Sarah Cohen-Boulakia, Khalid Belhajjame, Olivier Collin, Jérôme Chopard, Christine Froidevaux, Alban Gaignard, Konrad Hinsen, Pierre Larmande, Yvan Le Bras, Frédéric Lemoine, et al. Scientific workflows for computational reproducibility in the life sciences: Status, challenges and opportunities.Future Generation Computer Systems, 75:284–298, 2017
2017
-
[22]
Elisa: An interpretable hybrid generative ai agent for expression-grounded discovery in single-cell genomics
Omar Coser. Elisa: An interpretable hybrid generative ai agent for expression-grounded discovery in single-cell genomics. InThe 2026 Workshop on Generative and Agentic AI for Biology
2026
-
[23]
Data normalization for addressing the challenges in the analysis of single-cell transcriptomic datasets.BMC genomics, 25(1):444, 2024
Raquel Cuevas-Diaz Duran, Haichao Wei, and Jiaqian Wu. Data normalization for addressing the challenges in the analysis of single-cell transcriptomic datasets.BMC genomics, 25(1):444, 2024
2024
-
[24]
scgpt: toward building a foundation model for single- cell multi-omics using generative ai.Nature methods, 21(8):1470–1480, 2024
Haotian Cui, Chloe Wang, Hassaan Maan, Kuan Pang, Fengning Luo, Nan Duan, and Bo Wang. scgpt: toward building a foundation model for single- cell multi-omics using generative ai.Nature methods, 21(8):1470–1480, 2024
2024
-
[25]
A multimodal conversational agent for dna, rna and protein tasks.Nature Machine Intelligence, 7(6):928–941, 2025
Bernardo P de Almeida, Guillaume Richard, Hugo Dalla-Torre, Christopher Blum, Lorenz Hexemer, Priyanka Pandey, Stefan Laurent, Chandana Rajesh, Marie Lopez, Alexandre Laterre, et al. A multimodal conversational agent for dna, rna and protein tasks.Nature Machine Intelligence, 7(6):928–941, 2025
2025
-
[26]
Nextflow enables reproducible computational workflows.Nature biotechnology, 35(4):316–319, 2017
Paolo Di Tommaso, Maria Chatzou, Evan W Floden, Pablo Prieto Barja, Emilio Palumbo, and Cedric Notredame. Nextflow enables reproducible computational workflows.Nature biotechnology, 35(4):316–319, 2017
2017
-
[27]
Automating exploratory proteomics research via language models.arXiv preprint arXiv:2411.03743, 2024
Ning Ding, Shang Qu, Linhai Xie, Yifei Li, Zaoqu Liu, Kaiyan Zhang, Yibai Xiong, Yuxin Zuo, Zhangren Chen, Ermo Hua, et al. Automating exploratory proteomics research via language models.arXiv preprint arXiv:2411.03743, 2024
Pith/arXiv arXiv 2024
-
[28]
Large language model agents for biological intelligence across genomics, proteomics, spatial biology, and biomedicine.Briefings in Bioinformatics, 27(2):bbag110, 2026
Sajib Acharjee Dip, Dipanwita Mallick, Uddip Acharjee Shuvo, Shovito Barua Soumma, Fazle Rafsani, Bikash Kumar Paul, Nazifa Ahmed Moumi, Shafayat Ahmed, and Liqing Zhang. Large language model agents for biological intelligence across genomics, proteomics, spatial biology, and biomedicine.Briefings in Bioinformatics, 27(2):bbag110, 2026
2026
-
[29]
Can lightweight llm agents improve spatial transcriptomics annotation?bioRxiv, pages 2025–11, 2025
Sajib Acharjee Dip and Liqing Zhang. Can lightweight llm agents improve spatial transcriptomics annotation?bioRxiv, pages 2025–11, 2025
2025
-
[30]
Pengfei Du. Memory for autonomous llm agents: Mechanisms, evaluation, and emerging frontiers.arXiv preprint arXiv:2603.07670, 2026
arXiv 2026
-
[31]
Dionizije Fa, Marko Culjak, Bruno Pandza, and Mateo Cupic. Bioagent bench: An ai agent evaluation suite for bioinformatics.arXiv preprint arXiv:2601.21800, 2026
Pith/arXiv arXiv 2026
-
[32]
Gabriele Fossi, Youssef Boulaimen, Leila Outemzabet, Nathalie Jeanray, Stephane Gerart, Sebastien Vachenc, Joanna Giemza, and Salvatore Raieli. Swiftdossier: tailored automatic dossier for drug discovery with llms and agents.arXiv preprint arXiv:2409.15817, 2024. Agentic Bioinformatics through Function, Evidence, and Validation, 2026, Volume , Issue13
Pith/arXiv arXiv 2024
-
[33]
Bowen Gao, Yanwen Huang, Yiqiao Liu, Wenxuan Xie, Wei-Ying Ma, Ya- Qin Zhang, and Yanyan Lan. Pharmagents: Building a virtual pharma with large language model agents.arXiv preprint arXiv:2503.22164, 2025
Pith/arXiv arXiv 2025
-
[34]
Empowering biomedical discovery with ai agents.Cell, 187(22):6125–6151, 2024
Shanghua Gao, Ada Fang, Yepeng Huang, Valentina Giunchiglia, Ayush Noori, Jonathan Richard Schwarz, Yasha Ektefaie, Jovana Kondic, and Marinka Zitnik. Empowering biomedical discovery with ai agents.Cell, 187(22):6125–6151, 2024
2024
-
[35]
Shanghua Gao, Richard Zhu, Zhenglun Kong, Ayush Noori, Xiaorui Su, Curtis Ginder, Theodoros Tsiligkaridis, and Marinka Zitnik. Txagent: an ai agent for therapeutic reasoning across a universe of tools.arXiv preprint arXiv:2503.10970, 2025
Pith/arXiv arXiv 2025
-
[36]
Fukang Ge, Jiarui Zhu, Linjie Zhang, Haowen Xiao, Xiangcheng Bao, Fangnan Xie, Danyang Chen, Yanrui Lu, Yuting Wang, Ziqian Guan, et al. Autobinder agent: An mcp-based agent for end-to-end protein binder design.arXiv preprint arXiv:2602.00019, 2026
arXiv 2026
-
[37]
Protagents: protein discovery via large language model multi-agent collaborations combining physics and machine learning.Digital Discovery, 3(7):1389–1409, 2024
Alireza Ghafarollahi and Markus J Buehler. Protagents: protein discovery via large language model multi-agent collaborations combining physics and machine learning.Digital Discovery, 3(7):1389–1409, 2024
2024
-
[38]
Alireza Ghafarollahi and Markus J Buehler. Sparks: Multi-agent artificial intelligence model discovers protein design principles.arXiv preprint arXiv:2504.19017, 2025
Pith/arXiv arXiv 2025
-
[39]
Accelerating scientific discovery with co-scientist.Nature, pages 1–3, 2026
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Artiom Myaskovsky, Grzegorz Glowaty, Felix Weissenberger, Alessio Orlandi, Dan Popovici, et al. Accelerating scientific discovery with co-scientist.Nature, pages 1–3, 2026
2026
-
[40]
Promptbio-bench: Benchmarking llm-based bioinformatics agents for end-to-end data analysis.bioRxiv, pages 2026–05, 2026
Wenbin Guo, Minzhe Zhang, Bowei Han, Youjia Ma, Yang Leng, Shishir Hebbar, Xiaoyuan Zhou, Wenhao Gu, Xiao Yang, and Shashi Dhar. Promptbio-bench: Benchmarking llm-based bioinformatics agents for end-to-end data analysis.bioRxiv, pages 2026–05, 2026
2026
-
[41]
Large- scale foundation model on single-cell transcriptomics.Nature methods, 21(8):1481–1491, 2024
Minsheng Hao, Jing Gong, Xin Zeng, Chiming Liu, Yucheng Guo, Xingyi Cheng, Taifeng Wang, Jianzhu Ma, Xuegong Zhang, and Le Song. Large- scale foundation model on single-cell transcriptomics.Nature methods, 21(8):1481–1491, 2024
2024
-
[42]
Perturboagent: An llm-based agent for designing iterative perturb-seq experiments
Minsheng Hao, Hanchen Wang, Gabriele Scalia, Aviv Regev, et al. Perturboagent: An llm-based agent for designing iterative perturb-seq experiments. InMachine Learning in Computational Biology, pages 44–64. PMLR, 2025
2025
-
[43]
Functional protein design and enhancement with ontology reinforcement iteration.Nature Communications, 17(1):4158, 2026
Bing He, Chenchen Qin, Yu Zhao, Long-Kai Huang, Zihan Wu, Fang Wang, Fandi Wu, Fan Yang, and Jianhua Yao. Functional protein design and enhancement with ontology reinforcement iteration.Nature Communications, 17(1):4158, 2026
2026
-
[44]
George Hong and Daniel Trejo Banos. Nano bio-agents (nba): Small language model agents for genomics.arXiv preprint arXiv:2509.19566, 2025
arXiv 2025
-
[45]
Biogen: evidence-grounded multi-agent reasoning framework for transcriptomic interpretation in antimicrobial resistance.Frontiers in Bioinformatics, 6:1846404, 2026
Elias Hossain, Mehrdad Shoeibi, Ivan Garibay, and Niloofar Yousefi. Biogen: evidence-grounded multi-agent reasoning framework for transcriptomic interpretation in antimicrobial resistance.Frontiers in Bioinformatics, 6:1846404, 2026
2026
-
[46]
Memory in the age of ai agents.arXiv preprint arXiv:2512.13564, 2025
Yuyang Hu, Shichun Liu, Yanwei Yue, Guibin Zhang, Boyang Liu, Fangyi Zhu, Jiahang Lin, Honglin Guo, Shihan Dou, Zhiheng Xi, et al. Memory in the age of ai agents.arXiv preprint arXiv:2512.13564, 2025
Pith/arXiv arXiv 2025
-
[47]
Chao Hui Huang. Qust-llm: Integrating large language models for comprehensive spatial transcriptomics analysis.arXiv preprint arXiv:2406.14307, 2024
Pith/arXiv arXiv 2024
-
[48]
Omnicellagent: Towards ai co-scientists for scientific discovery in precision medicine.bioRxiv, 2025
Di Huang, Hao Li, Wenyu Li, Heming Zhang, Patricia Dickson, Ming Zhan, J Philip Miller, Carlos Cruchaga, Michael Province, Yixin Chen, et al. Omnicellagent: Towards ai co-scientists for scientific discovery in precision medicine.bioRxiv, 2025
2025
-
[50]
Autonomous biomedical research with an artificial intelligence agent.Science, page eadz4351, 2026
Kexin Huang, Serena Zhang, Hanchen Wang, Yuanhao Qu, Yingzhou Lu, Ryan Li, Yusuf Roohani, Lin Qiu, Shiyi Cao, Gavin Li, et al. Autonomous biomedical research with an artificial intelligence agent.Science, page eadz4351, 2026
2026
-
[51]
Wenxuan Huang, Mingyu Tsoi, Yanhao Huang, Xinjie Mao, Xue Xia, Hao Wu, Jiaqi Wei, Yuejin Yang, Lang Yu, Cheng Tan, et al. Harmonycell: Automating single-cell perturbation modeling under semantic and distribution shifts.arXiv preprint arXiv:2603.01396, 2026
arXiv 2026
-
[52]
Molbench: A benchmark of ai models for molecular property prediction
Xiuyu Jiang, Liqin Tan, Jianhuan Cen, and Qingsong Zou. Molbench: A benchmark of ai models for molecular property prediction. InInternational Symposium on Benchmarking, Measuring and Optimization, pages 53–70. Springer, 2023
2023
-
[53]
Genegpt: augmenting large language models with domain tools for improved access to biomedical information.Bioinformatics, 40(2):btae075, 2024
Qiao Jin, Yifan Yang, Qingyu Chen, and Zhiyong Lu. Genegpt: augmenting large language models with domain tools for improved access to biomedical information.Bioinformatics, 40(2):btae075, 2024
2024
-
[54]
Biolab: End-to-end autonomous life sciences research with multi-agents system integrating biological foundation models.BioRxiv, pages 2025–09, 2025
Ruofan Jin, Yucheng Guo, Yuanhao Qu, Ming Yang, Chun Shang, Qirong Yang, Linlin Chao, Yi Zhou, Ruilai Xu, Ziyao Xu, et al. Biolab: End-to-end autonomous life sciences research with multi-agents system integrating biological foundation models.BioRxiv, pages 2025–09, 2025
2025
-
[55]
Stella: Self-evolving llm agent for biomedical research.arXiv preprint arXiv:2507.02004, 2025
Ruofan Jin, Zaixi Zhang, Mengdi Wang, and Le Cong. Stella: Self-evolving llm agent for biomedical research.arXiv preprint arXiv:2507.02004, 2025
Pith/arXiv arXiv 2025
-
[56]
Evaluating agentic ai for biological discovery in autonomous and copilot settings.bioRxiv, pages 2026–06, 2026
Shreya Johri, Erica Maria Pimenta, Josephine Yates, Jingxin Fu, Erik L Bao, Hyeji Jun, Brendan Reardon, Sasha Bacot, Maha Shady, Doris Fu, et al. Evaluating agentic ai for biological discovery in autonomous and copilot settings.bioRxiv, pages 2026–06, 2026
2026
-
[57]
Highly accurate protein structure prediction with alphafold.nature, 596(7873):583–589, 2021
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold.nature, 596(7873):583–589, 2021
2021
-
[58]
Investigating reproducibility and tracking provenance–a genomic workflow case study.BMC bioinformatics, 18(1):337, 2017
Sehrish Kanwal, Farah Zaib Khan, Andrew Lonie, and Richard O Sinnott. Investigating reproducibility and tracking provenance–a genomic workflow case study.BMC bioinformatics, 18(1):337, 2017
2017
-
[59]
Sharing interoperable workflow provenance: A review of best practices and their practical application in cwlprov.GigaScience, 8(11):giz095, 2019
Farah Zaib Khan, Stian Soiland-Reyes, Richard O Sinnott, Andrew Lonie, Carole Goble, and Michael R Crusoe. Sharing interoperable workflow provenance: A review of best practices and their practical application in cwlprov.GigaScience, 8(11):giz095, 2019
2019
-
[60]
Progressive multi-agent reasoning for biological perturbation prediction
Hyomin Kim, Sang-Yeon Hwang, Jaechang Lim, Yinhua Piao, Yunhak Oh, Woo Youn Kim, Chanyoung Park, Sungsoo Ahn, and Junhyeok Jeon. Progressive multi-agent reasoning for biological perturbation prediction. arXiv preprint arXiv:2602.07408, 2026
Pith/arXiv arXiv 2026
-
[61]
Benchmarking and behavioral characterization of llm agents for protein design.bioRxiv, pages 2026–05, 2026
Jeonghyeon Kim and Philip Romero. Benchmarking and behavioral characterization of llm agents for protein design.bioRxiv, pages 2026–05, 2026
2026
-
[62]
Empowering bioinformatics communities with nextflow and nf-core.Genome Biology, 26(1):228, 2025
Björn E Langer, Andreia Amaral, Marie-Odile Baudement, Franziska Bonath, Mathieu Charles, Praveen Krishna Chitneedi, Emily L Clark, Paolo Di Tommaso, Sarah Djebali, Philip A Ewels, et al. Empowering bioinformatics communities with nextflow and nf-core.Genome Biology, 26(1):228, 2025
2025
-
[63]
Rag-enhanced collaborative llm agents for drug discovery
Namkyeong Lee, Edward De Brouwer, Ehsan Hajiramezanali, Tommaso Biancalani, Chanyoung Park, and Gabriele Scalia. Rag-enhanced collaborative llm agents for drug discovery. InProceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 561–569, 2026
2026
-
[64]
Empowering clinical trial design with agentic intelligence and real-world data.Nature Communications, 17(1):5501, 2026
Haoyang Li, Weishen Pan, Suraj Rajendran, Chengxi Zang, and Fei Wang. Empowering clinical trial design with agentic intelligence and real-world data.Nature Communications, 17(1):5501, 2026
2026
-
[65]
Kun Li, Zhennan Wu, Shoupeng Wang, Jia Wu, Shirui Pan, and Wenbin Hu. Drugpilot: Llm-based parameterized reasoning agent for drug discovery.arXiv preprint arXiv:2505.13940, 2025. 14 Agentic Bioinformatics through Function, Evidence, and Validation, 2026, Volume , Issue
Pith/arXiv arXiv 2025
-
[66]
Loka Li, Duzhen Zhang, Xingbo Du, Leonard Song, Zixiao Wang, Assanali Aukenov, Noel Thomas, Shakhnazar Sailaukan, Yonghan Yang, Feilong Chen, et al. Bioxarena: Benchmarking llm agents on multi-modal biomedical machine learning tasks.arXiv preprint arXiv:2605.15766, 2026
Pith/arXiv arXiv 2026
-
[67]
A co-evolving agentic ai system for medical imaging analysis.arXiv preprint arXiv:2509.20279, 2025
Songhao Li, Jonathan Xu, Tiancheng Bao, Yuxuan Liu, Yuchen Liu, Yihang Liu, Lilin Wang, Wenhui Lei, Sheng Wang, Yinuo Xu, et al. A co-evolving agentic ai system for medical imaging analysis.arXiv preprint arXiv:2509.20279, 2025
arXiv 2025
-
[68]
Wsi-llava: A multimodal large language model for whole slide image
Yuci Liang, Xinheng Lyu, Wenting Chen, Meidan Ding, Jipeng Zhang, Xiangjian He, Song Wu, Xiaohan Xing, Sen Yang, Xiyue Wang, et al. Wsi-llava: A multimodal large language model for whole slide image. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 22718–22727, 2025
2025
-
[69]
Bridging artificial intelligence and biological sciences: a comprehensive review of large language models in bioinformatics.Briefings in Bioinformatics, 26(4):bbaf357, 2025
Anqi Lin, Junpu Ye, Chang Qi, Lingxuan Zhu, Weiming Mou, Wenyi Gan, Dongqiang Zeng, Bufu Tang, Mingjia Xiao, Guangdi Chu, et al. Bridging artificial intelligence and biological sciences: a comprehensive review of large language models in bioinformatics.Briefings in Bioinformatics, 26(4):bbaf357, 2025
2025
-
[70]
Evolutionary-scale prediction of atomic-level protein structure with a language model.Science, 379(6637):1123–1130, 2023
Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model.Science, 379(6637):1123–1130, 2023
2023
-
[71]
Spatial transcriptomics ai agent charts hpsc-pancreas maturation in vivo.bioRxiv, pages 2025–04, 2025
Zuwan Lin, Wenbo Wang, Arnau Marin-Llobet, Qiang Li, Samuel D Pollock, Xin Sui, Almir Aljovic, Jaeyong Lee, Jongmin Baek, Ningyue Liang, et al. Spatial transcriptomics ai agent charts hpsc-pancreas maturation in vivo.bioRxiv, pages 2025–04, 2025
2025
-
[72]
Autoct: Automating interpretable clinical trial prediction with llm agents
Fengze Liu, Haoyu Wang, Joonhyuk Cho, Dan Roth, and Andrew Lo. Autoct: Automating interpretable clinical trial prediction with llm agents. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 30933–30958, 2025
2025
-
[73]
Genotex: An llm agent benchmark for automated gene expression data analysis, 2025
Haoyang Liu, Shuyu Chen, Ye Zhang, and Haohan Wang. Genotex: An llm agent benchmark for automated gene expression data analysis, 2025
2025
-
[74]
Haoyang Liu, Yijiang Li, and Haohan Wang. Genomas: A multi- agent framework for scientific discovery via code-driven gene expression analysis.arXiv preprint arXiv:2507.21035, 2025
Pith/arXiv arXiv 2025
-
[75]
Lsm-copilot: A skill-flow agent for fluorescence microscopy analysis
Ruofan Liu, Pengcheng Chen, and Eric J Seibel. Lsm-copilot: A skill-flow agent for fluorescence microscopy analysis. InFirst Workshop on Agent Skills, 2026
2026
-
[76]
Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines
Siru Liu, Allison B McCoy, and Adam Wright. Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines. Journal of the American Medical Informatics Association, 32(4):605–615, 2025
2025
-
[77]
Sizhe Liu, Yizhou Lu, Siyu Chen, Xiyang Hu, Jieyu Zhao, Yingzhou Lu, and Yue Zhao. Drugagent: Automating ai-aided drug discovery programming through llm multi-agent collaboration.arXiv preprint arXiv:2411.15692, 2024
Pith/arXiv arXiv 2024
-
[78]
Drbioright 2.0: an llm-powered bioinformatics chatbot for large- scale cancer functional proteomics analysis.Nature communications, 16(1):2256, 2025
Wei Liu, Jun Li, Yitao Tang, Yining Zhao, Chaozhong Liu, Meiyi Song, Zhenlin Ju, Shwetha V Kumar, Yiling Lu, Rehan Akbani, et al. Drbioright 2.0: an llm-powered bioinformatics chatbot for large- scale cancer functional proteomics analysis.Nature communications, 16(1):2256, 2025
2025
-
[79]
Benchmarking llm-based agents for single-cell omics analysis.Genome Biology, 27(1):123, 2026
Yang Liu, Lu Zhou, Xiawei Du, Ruikun He, Xuguang Zhang, Rongbo Shen, and Yixue Li. Benchmarking llm-based agents for single-cell omics analysis.Genome Biology, 27(1):123, 2026
2026
-
[80]
Toursynbio- search: A large language model driven agent framework for unified search method for protein engineering
Yungeng Liu, Zan Chen, Yu Guang Wang, and Yiqing Shen. Toursynbio- search: A large language model driven agent framework for unified search method for protein engineering. In2024 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pages 5395–5400. IEEE, 2024
2024
-
[81]
Autoproteinengine: A large language model driven agent framework for multimodal automl in protein engineering
Yungeng Liu, Zan Chen, Yuguang Wang, and Yiqing Shen. Autoproteinengine: A large language model driven agent framework for multimodal automl in protein engineering. InProceedings of the 31st International Conference on Computational Linguistics: Industry Track, pages 422–430, 2025
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.