Pith. sign in

REVIEW 6 cited by

BioPlanner: Automatic Evaluation of LLMs on Protocol Planning in Biology

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.10632 v1 pith:YLCTQ7HY submitted 2023-10-16 cs.CL cs.AIcs.RO

classification cs.CLcs.AIcs.RO
keywords protocolspseudocodeevaluationplanningscientificautomaticevaluateexperiments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The ability to automatically generate accurate protocols for scientific experiments would represent a major step towards the automation of science. Large Language Models (LLMs) have impressive capabilities on a wide range of tasks, such as question answering and the generation of coherent text and code. However, LLMs can struggle with multi-step problems and long-term planning, which are crucial for designing scientific experiments. Moreover, evaluation of the accuracy of scientific protocols is challenging, because experiments can be described correctly in many different ways, require expert knowledge to evaluate, and cannot usually be executed automatically. Here we present an automatic evaluation framework for the task of planning experimental protocols, and we introduce BioProt: a dataset of biology protocols with corresponding pseudocode representations. To measure performance on generating scientific protocols, we use an LLM to convert a natural language protocol into pseudocode, and then evaluate an LLM's ability to reconstruct the pseudocode from a high-level description and a list of admissible pseudocode functions. We evaluate GPT-3 and GPT-4 on this task and explore their robustness. We externally validate the utility of pseudocode representations of text by generating accurate novel protocols using retrieved pseudocode, and we run a generated protocol successfully in our biological laboratory. Our framework is extensible to the evaluation and improvement of language model planning abilities in other areas of science or other areas that lack automatic evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AEGIS: Assay-Aware Protocol Validation and Runtime Monitoring for Open-Source Liquid Handling Robots

    cs.RO 2026-07 conditional novelty 6.0 of 10

    AEGIS combines a rule-guided LLM protocol validator with a PCA/VLM visual runtime monitor to catch silent liquid-handling failures on the Opentrons OT-2, reporting adjusted F1 0.97 and average precision 0.89 on small ...

  2. BioMARS: A Multi-Agent Robotic System for Autonomous Biological Experiments

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A three-agent LLM/VLM system generated and executed cell-culture protocols on a dual-arm robot, matching manual passaging in viability, with optimization results shown only in a simulated benchmark.

  3. Contemporary AI foundation models increase biological weapons risk

    cs.CY 2025-06 reject novelty 6.0 of 10

    Contemporary AI models can provide accurate technical guidance for key steps of recovering live poliovirus, challenging the assumption that tacit knowledge blocks biological weapons development.

  4. Beyond SMILES: Evaluating Agentic Systems for Drug Discovery

    q-bio.QM 2026-02 conditional novelty 5.0 of 10

    Drug-discovery AI agents are built for small-molecule, big-pharma settings and lack peptide, in vivo, training-loop, small-lab, and multi-objective capabilities, even though LLMs themselves can reason about peptides.

  5. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.

  6. From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines

    cs.DL 2026-06 unverdicted novelty 3.0 of 10

    LLMs accelerate research workflows from idea generation to writing but introduce challenges like hallucination, bias, opacity, and ten systemic risks requiring new governance frameworks.

Pith tools