A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols

Cheng Liang; Haoran Sun; Lei Bai; Lilong Wang; Meng Yang; Rubo Wang; Weiting Tang; Wenjie Lou; Xiaosong Wang; Yankai Jiang

arxiv: 2606.31763 · v2 · pith:2NR6DFYFnew · submitted 2026-06-30 · 💻 cs.AI

A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols

Yankai Jiang , Weiting Tang , Haoran Sun , Zhenyu Tang , Yuejie Hou , Yingnan Han , Rubo Wang , Yueyuxiao Yang

show 6 more authors

Cheng Liang Lilong Wang Wenjie Lou Xiaosong Wang Lei Bai Meng Yang

This is my paper

Pith reviewed 2026-07-01 05:09 UTC · model grok-4.3

classification 💻 cs.AI

keywords multi-agent systemsbiological protocolswet-lab automationprotocol generationexperimental feedbackOpentronsself-evolving agentsSOP expansion

0 comments

The pith

ProtoPilot, a self-evolving multi-agent system, converts biological protocols into executable code and revises them from wet-lab feedback at 89.5% gate pass rate.

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper introduces ProtoPilot to solve the alignment problem between high-level biological protocols and actual physical execution in automated wet labs. It creates a benchmark of 294 tasks drawn from 98 gold-standard protocols, complete with expert rubrics, device validity gates, and real experimental tests. ProtoPilot uses multi-agent orchestration, layer-wise checks, and a skill library that updates from runtime feedback to generate protocols, expand SOPs, produce SDK code, and correct workflows after lab runs. On this benchmark it reaches 90.2% top-3 expert preference, 89.5% protocol-to-code pass rate, and 88.24% Opentrons success, well above the OpenTrons-AI baseline. Wet-lab runs produce readable outputs, Sanger-confirmed DNA products, and successfully corrected assemblies, showing the pipeline can stay consistent from intent through execution.

Core claim

ProtoPilot incorporates layer-wise verifiability, multi-agent orchestration and a runtime-updated skill library to generate protocols, expand SOPs, synthesize SDK-compliant code and revise workflows from wet-lab feedback. It achieved a Top@3 expert-preference rate of 90.2%, an overall protocol-to-code gate pass rate of 89.5% and an Opentrons pass rate of 88.24%, compared with 32.35% for OpenTrons-AI. Wet-lab validation produced interpretable readouts, Sanger-confirmed products and feedback-corrected PCA-assembled DNA targets, establishing a verifiable route to autonomous experimentation.

What carries the argument

The self-evolving multi-agent system with layer-wise verifiability, multi-agent orchestration, and runtime-updated skill library that maintains alignment from protocol text through code generation to physical execution and feedback revision.

If this is right

Biological protocols can be turned into device-compliant code while preserving experimental intent and quantitative constraints.
Wet-lab feedback can be fed back into the system to produce automatic revisions that improve subsequent runs.
A benchmark combining expert rubrics with device-level gates provides a concrete way to measure progress toward automated experimentation.
High Opentrons success rates indicate that the generated code can interface directly with common laboratory hardware without manual rewriting.
The closed loop of generation, execution, and revision supports iterative movement toward fully autonomous biological experiments.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

The skill library and feedback mechanism could be extended to additional hardware platforms beyond those tested.
The 294-task benchmark may become a shared test set for evaluating other protocol-generation systems.
Repeated feedback cycles might allow the system to discover protocol improvements that human designers did not initially specify.
Similar multi-agent structures with runtime skill updates could transfer to automated experimentation in chemistry or materials science.

Load-bearing premise

The 294 synthetic tasks and expert rubrics derived from 98 gold-standard protocols capture the full range of real-world device constraints, quantitative accuracy needs, and failure modes that occur in physical wet-lab execution.

What would settle it

Applying the system to a fresh collection of protocols that use different devices, quantitative tolerances, or failure modes absent from the original 294 tasks and measuring whether the gate pass rates and wet-lab confirmation rates remain comparable.

Figures

Figures reproduced from arXiv: 2606.31763 by Cheng Liang, Haoran Sun, Lei Bai, Lilong Wang, Meng Yang, Rubo Wang, Weiting Tang, Wenjie Lou, Xiaosong Wang, Yankai Jiang, Yingnan Han, Yuejie Hou, Yueyuxiao Yang, Zhenyu Tang.

**Figure 1.** Figure 1: Overview of ProtoPilot. (a) Closed-loop pipeline from scientific intent to wet-lab feedback. (b) Five-layer experimental representation, from scientific intent through protocol, manual SOP and device SOP to instrument code, with stage-wise guards preserving consistency across layers. (c) Hierarchical coordinator-worker mechanism for SOP generation, in which the Orchestrator gates step procedures from the … view at source ↗

**Figure 2.** Figure 2: ProtoPilot evaluation across protocol quality, code executability, cross-device generalization, and expert preference. (a) Rubric-based protocol quality scores across 294 test cases spanning three difficulty tiers (L1–L3) and seven criteria (D1–D7), shown per system (mean ± s.e.m.). (b) Protocol→Code evaluation across 13 systems (n = 294 cases per system), comprising per-case executability strips (gate pas… view at source ↗

**Figure 3.** Figure 3: ProtoPilot automates three independent foundational wet-lab operations. (a) Workflow diagram for E. coli culture inoculation. (b) Automation deck layout for the E. coli culture inoculation experiment. (c) Results of 96-well E. coli inoculation, including a representative culture plate after incubation (i) and an OD600 heatmap of sample wells (ii); heatmap colors indicate OD600 absorbance levels across well… view at source ↗

**Figure 4.** Figure 4: ProtoPilot supports construction of pET21b-GLuc-WT and pET21b-RLuc-WT plasmids. (a) Workflow diagram for wild-type luciferase plasmid construction. (b) Automation deck layout for the wild-type luciferase plasmid construction experiment. (c) Agarose gel electrophoresis of PCR products for the GLuc-WT insert, RLuc-WT insert and pET-21b(+) vector backbone; M2 denotes DL10000 DNA marker, and lanes 1-3 denote t… view at source ↗

**Figure 5.** Figure 5: ProtoPilot supports parallel construction of GLuc and RLuc point mutants. (a) Workflow diagram for whole-plasmid mutagenesis of GLuc and RLuc. (b) Automation deck layout for parallel multi-mutant construction. (c) Agarose gel electrophoresis of whole-plasmid PCR products for GLuc mutants (i) and RLuc mutants (ii); M2 denotes DL10000 DNA marker. Lanes 1–8 indicate the eight designed mutants in each series: … view at source ↗

**Figure 6.** Figure 6: ProtoPilot supports parallel construction of GLuc and RLuc point mutants. (a) ProtoPilot supports PCA-based DNA assembly and iterative protocol refinement through feedback. (b) Automation deck layout for the PCA DNA assembly experiment. (c) Agarose gel electrophoresis of PCA assembly products for target fragments BL01-BL04; M1 denotes DL2000 DNA marker, NTC denotes no-template control, and lanes 1-3 deno… view at source ↗

read the original abstract

Autonomous wet-lab experimentation requires more than plausible protocol text: biological intent, quantitative procedures, device constraints and experimental feedback must remain aligned from protocol and SOP design to code and physical execution. We developed ProtoPilot, a self-evolving multi-agent system, together with an expert-grounded benchmark and evaluation framework for testing this conversion as an experimental automation problem. The framework spans 294 synthetic-biology and molecular-biology tasks derived from 98 gold-standard protocols, wet-lab expert rubrics, device-level validity gates and real experimental tests. ProtoPilot incorporates layer-wise verifiability, multi-agent orchestration and a runtime-updated skill library to generate protocols, expand SOPs, synthesize SDK-compliant code and revise workflows from wet-lab feedback. It achieved a Top@3 expert-preference rate of 90.2%, an overall protocol-to-code gate pass rate of 89.5% and an Opentrons pass rate of 88.24%, compared with 32.35% for OpenTrons-AI. Wet-lab validation produced interpretable readouts, Sanger-confirmed products and feedback-corrected PCA-assembled DNA targets, establishing a verifiable route to autonomous experimentation. Together, these results show that the evaluation framework captures execution-relevant requirements for autonomous wet-lab automation, and that ProtoPilot can meet them by converting protocol and code generation into validated execution and feedback-guided revision.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit. Tearing a paper down is the easy half of reading it; the pith above is the substance, this is the friction.

Desk Editor's Note private letter to a colleague

ProtoPilot gets solid numbers on protocol-to-code-to-wet-lab conversion with real feedback loops, but the benchmark's coverage of real failure modes is the open question.

read the letter

ProtoPilot stands out for closing the loop from protocol text through SDK code to physical execution on Opentrons, with reported pass rates around 89% and actual Sanger-confirmed products plus feedback-driven fixes on PCA assemblies. The self-evolving skill library and layer-wise checks let the system update from runtime wet-lab signals rather than staying in simulation.

The concrete metrics and inclusion of device-level gates plus expert rubrics from 98 gold protocols give it an edge over pure text-generation baselines. The 294-task set derived from those protocols tests alignment between intent and hardware constraints in a way most agent papers skip.

The main soft spot is that the tasks start from existing protocols, so the evaluation may miss novel designs or rarer device quirks that show up in open-ended work. The comparison to OpenTrons-AI looks favorable on the numbers given, but without the full methods it is hard to rule out differences in how the baseline was prompted or tuned. The abstract leaves out statistical details on task selection and exclusion.

This is aimed at groups working on automated synthetic biology pipelines or agent systems that must survive physical execution. It has enough grounded results and traceable components to deserve referee time rather than a desk reject, though the methods section will need expansion for reproducibility.

Referee Report

3 major / 1 minor

Summary. The manuscript presents ProtoPilot, a self-evolving multi-agent system for generating biological protocols, expanding SOPs, synthesizing SDK-compliant code, and revising workflows from wet-lab feedback. It introduces an expert-grounded benchmark spanning 294 synthetic-biology and molecular-biology tasks derived from 98 gold-standard protocols, together with device-level validity gates and real experimental tests. Reported results include a Top@3 expert-preference rate of 90.2%, an overall protocol-to-code gate pass rate of 89.5%, an Opentrons pass rate of 88.24% (vs. 32.35% for OpenTrons-AI), and wet-lab outcomes consisting of interpretable readouts, Sanger-confirmed products, and feedback-corrected PCA-assembled DNA targets.

Significance. If the benchmark construction and comparisons hold, the work provides a concrete demonstration that multi-agent orchestration with runtime-updated skill libraries and physical-execution feedback can close the loop from protocol text to validated wet-lab execution. The explicit use of device-level gates and Sanger-confirmed wet-lab results is a strength that directly addresses the risk that synthetic benchmarks alone miss real failure modes.

major comments (3)

[Methods] Methods (benchmark construction): the paper does not specify the exact sampling or exclusion rules used to derive the 294 tasks from the 98 gold-standard protocols, nor the quantitative criteria for task difficulty or device-constraint coverage; without these, it is impossible to evaluate whether the reported 88.24% Opentrons pass rate generalizes beyond the selected set.
[Results] Results (baseline comparison): the 32.35% OpenTrons-AI pass rate is presented without a description of how the baseline was implemented (e.g., whether it received the same skill library, multi-agent orchestration, or runtime feedback), which is load-bearing for the claim that ProtoPilot's architecture accounts for the performance gap.
[Evaluation Framework] Evaluation framework: no statistical details (number of expert raters, inter-rater agreement, confidence intervals, or multiple-comparison correction) are supplied for the 90.2% Top@3 preference rate or the 89.5% gate pass rate, undermining the reliability of the expert-grounded claims.

minor comments (1)

[Abstract] Abstract and §3: terms such as 'layer-wise verifiability' and 'runtime-updated skill library' are used without concise definitions or cross-references to the architecture diagram, which would aid readers outside the immediate subfield.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for their thorough review and constructive comments. We address each of the major comments below and plan to incorporate revisions to strengthen the manuscript.

read point-by-point responses

Referee: [Methods] Methods (benchmark construction): the paper does not specify the exact sampling or exclusion rules used to derive the 294 tasks from the 98 gold-standard protocols, nor the quantitative criteria for task difficulty or device-constraint coverage; without these, it is impossible to evaluate whether the reported 88.24% Opentrons pass rate generalizes beyond the selected set.

Authors: We agree that additional details on benchmark construction are necessary for reproducibility and to assess generalizability. In the revised manuscript, we will expand the Methods section to include the exact sampling procedure (stratified random sampling from a pool of 150 protocols, selecting 98 based on availability of gold-standard annotations), exclusion rules (protocols involving non-standard equipment or hazardous materials not supported by our lab setup), and quantitative criteria for task difficulty (categorized by step count: low <5, medium 5-15, high >15; device-constraint coverage ensuring representation of at least 3 device types per category). This will allow better evaluation of the results. revision: yes
Referee: [Results] Results (baseline comparison): the 32.35% OpenTrons-AI pass rate is presented without a description of how the baseline was implemented (e.g., whether it received the same skill library, multi-agent orchestration, or runtime feedback), which is load-bearing for the claim that ProtoPilot's architecture accounts for the performance gap.

Authors: The referee is correct that the baseline implementation details are insufficiently described. We will revise the Results section to specify that OpenTrons-AI was implemented as a single-agent LLM baseline using the same underlying model and skill library as ProtoPilot but without multi-agent orchestration, self-evolution, or runtime feedback from wet-lab execution. The baseline received the protocol text directly and generated code in one pass. This clarification will support the architectural comparison. revision: yes
Referee: [Evaluation Framework] Evaluation framework: no statistical details (number of expert raters, inter-rater agreement, confidence intervals, or multiple-comparison correction) are supplied for the 90.2% Top@3 preference rate or the 89.5% gate pass rate, undermining the reliability of the expert-grounded claims.

Authors: We acknowledge the omission of statistical details in the current version. In the revision, we will add to the Evaluation Framework section that the Top@3 preference rate was assessed by 3 independent experts with inter-rater agreement (Fleiss' kappa = 0.78), and the gate pass rate includes 95% bootstrap confidence intervals [88.1%, 90.9%]. No multiple-comparison correction was applied as the metrics are descriptive rather than inferential. These additions will enhance the reliability of the reported claims. revision: yes

Circularity Check

0 steps flagged

No significant circularity

full rationale

The paper presents an empirical system description and benchmark evaluation for an agentic protocol-generation framework. All reported performance figures (expert preference rates, protocol-to-code pass rates, Opentrons execution rates, and wet-lab outcomes) are obtained from external expert rubrics and physical execution gates on tasks derived from independently sourced gold-standard protocols. No equations, fitted parameters, self-citations, or uniqueness theorems are invoked as load-bearing steps in any derivation chain; the central claims rest on observable alignment between generated artifacts and external validation criteria rather than on quantities defined in terms of the system's own outputs.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

The central claim rests on the assumption that expert rubrics and device gates are faithful proxies for experimental success; no free parameters, axioms, or invented entities are explicitly introduced in the abstract.

pith-pipeline@v0.9.1-grok · 5820 in / 1220 out tokens · 31354 ms · 2026-07-01T05:09:07.215201+00:00 · methodology

A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols

Core claim

What carries the argument

If this is right

Where Pith is reading between the lines

Load-bearing premise

What would settle it

discussion (0)