Pith. sign in

REVIEW 7 cited by

BioPlanner: Automatic Evaluation of LLMs on Protocol Planning in Biology

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.10632 v1 pith:YLCTQ7HY submitted 2023-10-16 cs.CL cs.AIcs.RO

BioPlanner: Automatic Evaluation of LLMs on Protocol Planning in Biology

classification cs.CL cs.AIcs.RO
keywords protocolspseudocodeevaluationplanningscientificautomaticevaluateexperiments
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The ability to automatically generate accurate protocols for scientific experiments would represent a major step towards the automation of science. Large Language Models (LLMs) have impressive capabilities on a wide range of tasks, such as question answering and the generation of coherent text and code. However, LLMs can struggle with multi-step problems and long-term planning, which are crucial for designing scientific experiments. Moreover, evaluation of the accuracy of scientific protocols is challenging, because experiments can be described correctly in many different ways, require expert knowledge to evaluate, and cannot usually be executed automatically. Here we present an automatic evaluation framework for the task of planning experimental protocols, and we introduce BioProt: a dataset of biology protocols with corresponding pseudocode representations. To measure performance on generating scientific protocols, we use an LLM to convert a natural language protocol into pseudocode, and then evaluate an LLM's ability to reconstruct the pseudocode from a high-level description and a list of admissible pseudocode functions. We evaluate GPT-3 and GPT-4 on this task and explore their robustness. We externally validate the utility of pseudocode representations of text by generating accurate novel protocols using retrieved pseudocode, and we run a generated protocol successfully in our biological laboratory. Our framework is extensible to the evaluation and improvement of language model planning abilities in other areas of science or other areas that lack automatic evaluation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AEGIS: Assay-Aware Protocol Validation and Runtime Monitoring for Open-Source Liquid Handling Robots

    cs.RO 2026-07 conditional novelty 6.0

    AEGIS combines a rule-guided LLM protocol validator with a PCA/VLM visual runtime monitor to catch silent liquid-handling failures on the Opentrons OT-2, reporting adjusted F1 0.97 and average precision 0.89 on small ...

  2. Beyond SMILES: Evaluating Agentic Systems for Drug Discovery

    q-bio.QM 2026-02 conditional novelty 5.0

    Drug-discovery AI agents are built for small-molecule, big-pharma settings and lack peptide, in vivo, training-loop, small-lab, and multi-objective capabilities, even though LLMs themselves can reason about peptides.

  3. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 conditional novelty 4.0

    AI can generate research artifacts faster than it can verify them, so across all eight lifecycle stages the credible deployment mode is human-governed collaboration rather than full autonomy.

  4. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 unverdicted novelty 4.0

    The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.

  5. From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines

    cs.DL 2026-06 unverdicted novelty 3.0

    LLMs accelerate research workflows from idea generation to writing but introduce challenges like hallucination, bias, opacity, and ten systemic risks requiring new governance frameworks.

  6. From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines

    cs.DL 2026-06 conditional novelty 3.0

    A cross-disciplinary review of 151 studies concludes LLMs accelerate research workflows while introducing recurring technical and ethical risks, including ten it flags as underexplored.

  7. LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

    cs.CL 2024-12 accept novelty 3.0

    A survey that organizes LLMs-as-judges research into functionality, methodology, applications, meta-evaluation, and limitations.