Pith. sign in

REVIEW 13 cited by

STELLA: Self-Evolving LLM Agent for Biomedical Research

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.02004 v1 pith:SNWMJTQW submitted 2025-07-01 cs.AI cs.CLq-bio.BM

STELLA: Self-Evolving LLM Agent for Biomedical Research

classification cs.AI cs.CLq-bio.BM
keywords stellaagentbiomedicalaccuracyexamexperienceexpertisehumanity
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The rapid growth of biomedical data, tools, and literature has created a fragmented research landscape that outpaces human expertise. While AI agents offer a solution, they typically rely on static, manually curated toolsets, limiting their ability to adapt and scale. Here, we introduce STELLA, a self-evolving AI agent designed to overcome these limitations. STELLA employs a multi-agent architecture that autonomously improves its own capabilities through two core mechanisms: an evolving Template Library for reasoning strategies and a dynamic Tool Ocean that expands as a Tool Creation Agent automatically discovers and integrates new bioinformatics tools. This allows STELLA to learn from experience. We demonstrate that STELLA achieves state-of-the-art accuracy on a suite of biomedical benchmarks, scoring approximately 26\% on Humanity's Last Exam: Biomedicine, 54\% on LAB-Bench: DBQA, and 63\% on LAB-Bench: LitQA, outperforming leading models by up to 6 percentage points. More importantly, we show that its performance systematically improves with experience; for instance, its accuracy on the Humanity's Last Exam benchmark almost doubles with increased trials. STELLA represents a significant advance towards AI Agent systems that can learn and grow, dynamically scaling their expertise to accelerate the pace of biomedical discovery.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolving Skill Memory

    cs.AI 2026-06 unverdicted novelty 7.0

    SkeMex distills agent trajectories into value-aware skills organized in general/task/action branches and evolves them via a closed-loop Read-Write-Assess-Govern process, outperforming prior memory agents on clinical tasks.

  2. BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks

    cs.CE 2026-05 unverdicted novelty 7.0

    BioXArena benchmarks LLM agents on generating end-to-end ML pipelines for 76 multi-modal biomedical tasks, with MLEvolve plus Gemini-3.1-Pro scoring highest at 0.666.

  3. ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity

    cs.AI 2026-06 unverdicted novelty 6.0

    ABC-Bench evaluates LLM agents on three biosecurity-relevant biology tasks and reports that agents outperformed median human experts, with wet-lab confirmation of successful DNA assembly by one model.

  4. AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows

    cs.AI 2026-05 unverdicted novelty 6.0

    AgentCo-op retrieves and assembles existing agents and tools into interoperable workflows for open-world scientific tasks, showing effectiveness in genomics case studies and competitive benchmark results with lower costs.

  5. Harnesses for Inference-Time Alignment over Execution Trajectories

    cs.LG 2026-05 unverdicted novelty 6.0

    Partial harnesses for LLM agents, specifying only initial execution steps, achieve higher pass rates than fully decomposed workflows, as analyzed through trajectory alignment and validated in synthetic and terminal be...

  6. DrugSAGE:Self-evolving Agent Experience for Efficient State-of-the-Art Drug Discovery

    cs.LG 2026-05 unverdicted novelty 6.0

    DrugSAGE accumulates cross-task memory of skills, statistical evidence, and recurring errors to let LLM agents achieve top-ranked performance on molecular property prediction tasks with reduced or zero test-time search.

  7. "Excuse me, may I say something..." CoLabScience, A Proactive AI Assistant for Biomedical Discovery and LLM-Expert Collaborations

    cs.CL 2026-04 unverdicted novelty 6.0

    The paper introduces CoLabScience with PULI, a positive-unlabeled RL framework for proactive interventions in streaming biomedical dialogues, plus the BSDD benchmark dataset, claiming superior performance over baselines.

  8. ClimAgent: LLM as Agents for Autonomous Open-ended Climate Science Analysis

    cs.AI 2026-04 unverdicted novelty 5.5

    ClimAgent claims an autonomous LLM-agent stack plus ClimaBench yield a 40.21% gain over base LLMs on rigorousness and practicality of climate analysis solutions.

  9. Evaluating Agentic Bioinformatics through Function, Evidence, and Validation

    cs.AI 2026-07 conditional novelty 5.0

    Agentic bioinformatics systems mostly demonstrate planning and tool execution but rarely prospective empirical validation, so the paper argues evaluation should center on inspectable workflow trajectories (FEV) rather...

  10. Self-Improvements in Modern Agentic Systems: A Survey

    cs.AI 2026-07 conditional novelty 5.0

    Self-improving agents are classified by what they update — foundation-model weights or the surrounding scaffold — and by the signal that drives the update, under a single formal operator.

  11. Beyond Prompt-Based Planning: MCP-Native Graph Planning-based Biomedical Agent System

    cs.AI 2026-06 unverdicted novelty 5.0

    BioManus introduces graph-scaffolded planning over MCP servers to decouple biomedical agent planning from tool inventory size, reporting improved accuracy and efficiency on BioAgentBench and LAB-Bench.

  12. ClimAgent: LLM as Agents for Autonomous Open-ended Climate Science Analysis

    cs.AI 2026-04 unverdicted novelty 5.0

    ClimAgent is an LLM-agent system for autonomous climate research that outperforms standard LLMs by 40.21% on the new ClimaBench benchmark in solution rigorousness and practicality.

  13. TusoAI: Agentic Optimization for Scientific Methods

    cs.AI 2025-09 unverdicted novelty 5.0

    TusoAI is an LLM-based agent that builds and iteratively optimizes domain-specific computational methods for scientific data analysis, outperforming expert baselines on RNA-seq denoising and earth monitoring while rep...