REVIEW 13 cited by
STELLA: Self-Evolving LLM Agent for Biomedical Research
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
STELLA: Self-Evolving LLM Agent for Biomedical Research
read the original abstract
The rapid growth of biomedical data, tools, and literature has created a fragmented research landscape that outpaces human expertise. While AI agents offer a solution, they typically rely on static, manually curated toolsets, limiting their ability to adapt and scale. Here, we introduce STELLA, a self-evolving AI agent designed to overcome these limitations. STELLA employs a multi-agent architecture that autonomously improves its own capabilities through two core mechanisms: an evolving Template Library for reasoning strategies and a dynamic Tool Ocean that expands as a Tool Creation Agent automatically discovers and integrates new bioinformatics tools. This allows STELLA to learn from experience. We demonstrate that STELLA achieves state-of-the-art accuracy on a suite of biomedical benchmarks, scoring approximately 26\% on Humanity's Last Exam: Biomedicine, 54\% on LAB-Bench: DBQA, and 63\% on LAB-Bench: LitQA, outperforming leading models by up to 6 percentage points. More importantly, we show that its performance systematically improves with experience; for instance, its accuracy on the Humanity's Last Exam benchmark almost doubles with increased trials. STELLA represents a significant advance towards AI Agent systems that can learn and grow, dynamically scaling their expertise to accelerate the pace of biomedical discovery.
Forward citations
Cited by 13 Pith papers
-
Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolving Skill Memory
SkeMex distills agent trajectories into value-aware skills organized in general/task/action branches and evolves them via a closed-loop Read-Write-Assess-Govern process, outperforming prior memory agents on clinical tasks.
-
BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks
BioXArena benchmarks LLM agents on generating end-to-end ML pipelines for 76 multi-modal biomedical tasks, with MLEvolve plus Gemini-3.1-Pro scoring highest at 0.666.
-
ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity
ABC-Bench evaluates LLM agents on three biosecurity-relevant biology tasks and reports that agents outperformed median human experts, with wet-lab confirmation of successful DNA assembly by one model.
-
AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows
AgentCo-op retrieves and assembles existing agents and tools into interoperable workflows for open-world scientific tasks, showing effectiveness in genomics case studies and competitive benchmark results with lower costs.
-
Harnesses for Inference-Time Alignment over Execution Trajectories
Partial harnesses for LLM agents, specifying only initial execution steps, achieve higher pass rates than fully decomposed workflows, as analyzed through trajectory alignment and validated in synthetic and terminal be...
-
DrugSAGE:Self-evolving Agent Experience for Efficient State-of-the-Art Drug Discovery
DrugSAGE accumulates cross-task memory of skills, statistical evidence, and recurring errors to let LLM agents achieve top-ranked performance on molecular property prediction tasks with reduced or zero test-time search.
-
"Excuse me, may I say something..." CoLabScience, A Proactive AI Assistant for Biomedical Discovery and LLM-Expert Collaborations
The paper introduces CoLabScience with PULI, a positive-unlabeled RL framework for proactive interventions in streaming biomedical dialogues, plus the BSDD benchmark dataset, claiming superior performance over baselines.
-
ClimAgent: LLM as Agents for Autonomous Open-ended Climate Science Analysis
ClimAgent claims an autonomous LLM-agent stack plus ClimaBench yield a 40.21% gain over base LLMs on rigorousness and practicality of climate analysis solutions.
-
Evaluating Agentic Bioinformatics through Function, Evidence, and Validation
Agentic bioinformatics systems mostly demonstrate planning and tool execution but rarely prospective empirical validation, so the paper argues evaluation should center on inspectable workflow trajectories (FEV) rather...
-
Self-Improvements in Modern Agentic Systems: A Survey
Self-improving agents are classified by what they update — foundation-model weights or the surrounding scaffold — and by the signal that drives the update, under a single formal operator.
-
Beyond Prompt-Based Planning: MCP-Native Graph Planning-based Biomedical Agent System
BioManus introduces graph-scaffolded planning over MCP servers to decouple biomedical agent planning from tool inventory size, reporting improved accuracy and efficiency on BioAgentBench and LAB-Bench.
-
ClimAgent: LLM as Agents for Autonomous Open-ended Climate Science Analysis
ClimAgent is an LLM-agent system for autonomous climate research that outperforms standard LLMs by 40.21% on the new ClimaBench benchmark in solution rigorousness and practicality.
-
TusoAI: Agentic Optimization for Scientific Methods
TusoAI is an LLM-based agent that builds and iteratively optimizes domain-specific computational methods for scientific data analysis, outperforming expert baselines on RNA-seq denoising and earth monitoring while rep...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.