Pith. sign in

REVIEW 8 cited by

MDCrow: Automating Molecular Dynamics Workflows with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.09565 v1 pith:GROXPWRK submitted 2025-02-13 cs.AI physics.chem-ph

MDCrow: Automating Molecular Dynamics Workflows with Large Language Models

classification cs.AI physics.chem-ph
keywords mdcrowmodelsautomatingtaskscomplexdifficultydynamicslanguage
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Molecular dynamics (MD) simulations are essential for understanding biomolecular systems but remain challenging to automate. Recent advances in large language models (LLM) have demonstrated success in automating complex scientific tasks using LLM-based agents. In this paper, we introduce MDCrow, an agentic LLM assistant capable of automating MD workflows. MDCrow uses chain-of-thought over 40 expert-designed tools for handling and processing files, setting up simulations, analyzing the simulation outputs, and retrieving relevant information from literature and databases. We assess MDCrow's performance across 25 tasks of varying required subtasks and difficulty, and we evaluate the agent's robustness to both difficulty and prompt style. \texttt{gpt-4o} is able to complete complex tasks with low variance, followed closely by \texttt{llama3-405b}, a compelling open-source model. While prompt style does not influence the best models' performance, it has significant effects on smaller models.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LLM-Guided Test-Time Discovery of Quantum-Chemical Approximation Algorithms

    physics.chem-ph 2026-06 unverdicted novelty 7.0

    LADeQ is an LLM-driven workflow that autonomously discovers and implements approximation algorithms for CCSD and CISD calculations, delivering speedups while respecting user-specified error tolerances.

  2. ChatMOSP: A Chemistry-Grounded Mobile Agent for Working-State Catalyst Simulations

    cond-mat.mtrl-sci 2026-05 unverdicted novelty 7.0

    ChatMOSP is an AI agent that maps natural-language descriptions of catalyst environments to validated multiscale simulations of working-state nanoparticle morphology and activity.

  3. El Agente Quntur: A research collaborator agent for quantum chemistry

    physics.chem-ph 2026-02 unverdicted novelty 7.0

    El Agente Quntur is a new multi-agent system that uses reasoning over literature and software documentation to autonomously handle the full workflow of quantum chemistry experiments in ORCA.

  4. MDForge: Agentic Molecular Dynamics Pipeline Design under Sparse Simulator Feedback

    cs.AI 2026-06 unverdicted novelty 6.0

    MDForge uses an LLM agent with multi-agent debate to densify sparse simulator feedback for automatic MD pipeline design, matching human experts on SAMPL benchmarks and identifying a lab-confirmed picomolar CB[7] binder.

  5. Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches

    cs.AI 2026-05 unverdicted novelty 6.0

    Survey of RLM adoption in 28 disciplines reveals maturity disparities via a new assessment framework, with focus on development, evaluation, and public resources.

  6. Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches

    cs.AI 2026-05 unverdicted novelty 6.0

    A survey of RLM use in 28 disciplines reveals uneven adoption and introduces a maturity assessment framework showing larger gaps when limited to public resources.

  7. Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches

    cs.AI 2026-05 unverdicted novelty 4.0

    A survey of reasoning language model adoption across 28 ERC scientific disciplines finds large maturity gaps, especially when only public resources are counted.

  8. From Knowledge to Action: Outcomes of the 2025 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry

    cond-mat.mtrl-sci 2026-05 unverdicted novelty 2.0

    Hackathon submissions indicate LLMs are moving from general assistants toward composable multi-agent systems for structuring scientific knowledge and automating tasks in materials science and chemistry.