BioDSA-1K is a large, publication-grounded benchmark for evaluating AI agents on biomedical hypothesis validation, including non-verifiable cases.
BioAgents: Democratizing Bioinformatics Analysis with Multi-Agent Systems
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Creating end-to-end bioinformatics workflows requires diverse domain expertise, which poses challenges for both junior and senior researchers as it demands a deep understanding of both genomics concepts and computational techniques. While large language models (LLMs) provide some assistance, they often fall short in providing the nuanced guidance needed to execute complex bioinformatics tasks, and require expensive computing resources to achieve high performance. We thus propose a multi-agent system built on small language models, fine-tuned on bioinformatics data, and enhanced with retrieval augmented generation (RAG). Our system, BioAgents, enables local operation and personalization using proprietary data. We observe performance comparable to human experts on conceptual genomics tasks, and suggest next steps to enhance code generation capabilities.
citation-role summary
citation-polarity summary
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research
BioDSA-1K is a large, publication-grounded benchmark for evaluating AI agents on biomedical hypothesis validation, including non-verifiable cases.