{"id":"dd5c3492-29ea-4e9f-8336-b4ae0caaa863","arxiv_id":"2608.07637","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A selective-LLM workflow completed a stateful GCMC-MD clay desorption campaign with zero LLM calls during routine production, and blinded replay correctly diagnosed two preserved incidents.","lead":"Agent-MD is a workflow that keeps large language models out of routine molecular simulation steps, using them only to plan campaigns and review unexpected events. In a five-system clay desorption campaign, it completed 120 simulation cycles with zero live LLM calls during production.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Production escalation was not demonstrated: the only formal-production review was resolved by a code correction, and the two reasoning-agent replays were offline and never executed, so the central claim overstates Agent-MD's demonstrated event-triggered review.","rationale":"The reader identified the acceptance-criteria convergence risk, which is real and explicitly flagged in §2.3.2. I see an equally important but different risk that strikes closer to the central claim: the only production anomaly in the campaign was never routed through the LLM reasoning path, and the two replay exercises are retrospective interpretation-only tests. The paper is transparent about this: §3.3 says the workflow was corrected and later says the replays do not establish general diagnostic accuracy, §4 lists the absence of comparisons and the bounded scope, and the repository is a positive reproducibility signal. The counts in Tables 1 and 2 are internally consistent, and the same reasoning agent is used for planning and review, so the architecture is coherent. But coherence is not demonstration: a reader cannot tell from the paper whether a live reasoning-agent recommendation, passed through the validation layer, would actually be safe and sufficient to resume and complete the campaign. The concrete live replay would test exactly that path, using the workflow's own frozen incident and acceptance criteria. My read therefore reinforces the conditional verdict; it does not overturn it.","tokens_in":15282,"tokens_out":8006,"duration_ms":80828,"concrete_test":"Concrete check: Replay the K–LC0.4 RH=0.3 incident live in a staging campaign with the same frozen evidence bundle, but enable Codex CLI during production (§2.4) and automatically validate and execute the agent's recommendation before resuming the rule-based agent; then continue through RH=0.1. Success requires the reasoning-agent recommendation to be schema-valid, in scope, and to yield an RH=0.3 state accepted under the same configured criteria, with the RH=0.1 state inheriting from the corrected archive and completing normally. If the agent's recommendation instead requires human review, changes scientific parameters outside the allowed scope, or produces an accepted state that fails the recorded checks, the live escalation loop is not demonstrated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Agent-MD's central claim is that selective LLM reasoning plus deterministic execution can handle long-running campaigns. The paper's strongest evidence is Table 2: 15 states, 120 cycles, zero production reasoning calls, with 'one state requiring review' (§3.2). But that state (K–LC0.4 RH=0.3) was not resolved by the reasoning agent: it was stopped by an absolute-vs-RH-local timestep bug, and §3.3 reports 'The workflow was corrected…' without any production Codex invocation. The reasoning agent is therefore absent from the only production anomaly. The two 'blinded replay' cases in §3.3 are retrospective: they use frozen evidence, occur after the fact, and their recommendations were never executed or validated by the campaign agent, so the claimed 'validated control handoffs' of §2.4 are only schema/scope checks, not correctness checks. The authors explicitly limit these results to 'these specific examples' (§3.3) and disclaim general diagnostic accuracy (§4), but the abstract and strongest claim present 'one event-triggered review resolved correctly' as part of the demonstration. This is an overstatement: the production event demonstrates detection/pause and human/code repair, not selective LLM review. Without a live case where a reasoning-agent recommendation is actually applied and the campaign completes, the architecture's central escalation mechanism is untested.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Agent-MD, a workflow architecture for long-running molecular simulation campaigns in which an LLM-based reasoning agent is used only for campaign construction and event-triggered review, while a persistent rule-based campaign agent handles routine simulation, analysis, continuation, archiving, and state progression. The approach is demonstrated on a GCMC-MD water-vapor desorption campaign for five montmorillonite systems across three relative-humidity states (RH = 0.9, 0.3, 0.1). The campaign completed 120 segmented simulation cycles, accepted 15 system-RH states, and recorded zero non-interactive LLM calls during formal production. One production state (K-LC0.4 at RH=0.3) reached a review boundary due to a timestep-accounting bug, which was corrected without invoking the reasoning agent. Two preserved incidents were later replayed through the reasoning agent offline. The paper also reports composition-dependent differences in residual water content and basal spacing. The central claim is that selective reasoning, combined with deterministic execution and structured evidence, can provide reproducible and auditable agent-assisted simulation without placing every operation inside an LLM loop.","tokens_in":15566,"tokens_out":6218,"duration_ms":59147,"significance":"If the architecture is accepted as demonstrated, Agent-MD would be a useful contribution to the emerging area of LLM-assisted scientific workflows, and the public release of code, campaign specifications, provenance records, and frozen replay cases is a clear strength. The physical results on montmorillonite desorption are plausible and consistent with prior literature, but they are not the main contribution; the main significance lies in the workflow-design claim that event-triggered LLM intervention can be kept selective without sacrificing campaign completion. However, the current evidence for that claim is substantially weaker than the title and abstract imply: the only production-time anomaly was resolved by a code correction with zero LLM involvement, and the two reasoning-agent evaluations are retrospective replays whose recommendations were never executed or validated live. The paper is honest about these limitations in Sections 4 and 5, but the framing of the central contribution needs to be either softened or backed by a live escalation case.","major_comments":[{"comment":"The paper's central architectural claim is that event-triggered LLM review can resolve states that the rule-based agent cannot safely handle, but this mechanism is never exercised in production. The only production-time review event (K-LC0.4 at RH=0.3) was stopped by an absolute-versus-RH-local timestep bug and was resolved by a workflow code correction with zero non-interactive Codex CLI invocations (§3.2, Table 2). The two reasoning-agent evaluations in §3.3 are retrospective, frozen replay cases: the recommendations were produced after the fact, were never applied to a running campaign, and were checked only against the authors' own recorded interpretations. The abstract's phrasing that \"one state reached a review boundary\" and that replay cases \"identified the underlying workflow problems\" is literally accurate, but the title and the overall framing present event-driven escalation as a demonstrated capability. This is load-bearing: without a live case in which a reasoning-agent recommendation is validated, executed, and the campaign completes, the paper demonstrates a rule-based workflow plus an offline case study, not the full Agent-MD loop. I recommend either adding a live escalation run or clearly redefining the contribution as an architecture with retrospective illustration.","section":"§3.2, §3.3, Abstract, §5"},{"comment":"The acceptance criteria are configured CV thresholds (0.03, 0.05, 0.08 for total, interlayer, and external-surface water) plus trend and health checks, and the manuscript explicitly states these are \"operational indicators for state assessment rather than as proof of strict thermodynamic equilibrium.\" The physical conclusions, including Ca-LC0.4 maintaining a 14.4 Å basal spacing and Na-LC0.5 retaining more residual water, rest entirely on states accepted under these unvalidated heuristics. No block-averaging convergence check, comparison with substantially longer sampling, or independent-replica verification is provided for any state. Because the \"15 accepted states\" record is used both as evidence of workflow success and as the basis for the scientific comparisons, the adequacy of the convergence criteria is load-bearing. The paper should either validate the criteria on at least a subset of states or explicitly downgrade the physical claims to preliminary observations.","section":"§2.3.2, Figures 5-7, §3.4"},{"comment":"The reasoning-agent evaluation consists of two self-selected incidents, both authored by the same group, with the expected diagnosis withheld but the evidence bundle constructed by the authors; there are no negative controls, no blinded third-party assessment, and no comparison with a deterministic baseline or with human-only review. The authors acknowledge in §4 that the replays \"do not establish general failure-diagnosis capability,\" but the Results section still presents the agreements as evidence that the agent \"could interpret the preserved evidence for these specific examples.\" This is a legitimate but anecdotal demonstration; the paper should make the anecdotal status prominent in the abstract and conclusions rather than allowing the replay results to appear as a validation of the escalation mechanism.","section":"§3.3, §4"},{"comment":"The abstract claims that selective reasoning is combined with \"validated control handoffs,\" but the validation described in §2.4 is only a schema- and policy-scope check: consistency with the event identifier, allowed decision types, confidence value, and permitted parameter scope. It does not check whether a recommendation is scientifically correct. A recommendation that passes these checks could still be wrong and lead to a bad decision. The term \"validated control handoffs\" therefore overstates the safety property the system actually provides. I suggest using a more precise term such as \"policy-validated control handoffs\" or adding a correctness/verification layer before making this claim.","section":"§2.4, Abstract"}],"minor_comments":[{"comment":"In the final sentence of the Introduction, the phrase \"Agent-MDbothasapracticalworkflowarchitectureforreliableandauditablelong-runningsimulation and as a tool...\" contains missing spaces around \"Agent-MD\" and should be corrected to read \"Agent-MD both as a practical workflow architecture...\".","section":"§1, last paragraph"},{"comment":"The final absolute timestep for each system is slightly larger than the sum of RH-local steps (by 0.1 to 0.2 million steps), presumably reflecting the initial NVT equilibration; this offset should be stated explicitly in the text or table caption so readers do not infer an inconsistency.","section":"Table 1 and §2.3.1"},{"comment":"The error bars in Figures 5-7 are temporal sample standard deviations over the final 1.0 million steps and do not represent uncertainty across independent replicas; the text states this in §2.3.2 and §3.4, but the figure captions should repeat it for clarity.","section":"§3.4 and Figure captions"},{"comment":"References [7] and [8] are listed as preprints without DOIs or arXiv identifiers; adding stable identifiers would improve the reproducibility and verifiability of the related-work discussion.","section":"References [7] and [8]"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is more honest about its limitations than the title and abstract suggest, and the public code/data release is a genuine strength. My main concern is that the central claimed contribution—event-triggered LLM escalation—is not demonstrated in production; the only production anomaly was fixed without the LLM, and the reasoning-agent evaluation is entirely retrospective. This is fixable within the scope of a major revision by reframing the contribution as an architecture design plus retrospective case study, or by adding a live escalation run. I do not see evidence of any integrity problem; the paper's own limitation statements are accurate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Agent-MD is worth a look. The core idea—put LLM reasoning at campaign construction and at genuinely unresolved decision points, and keep all routine production under a rule-based agent with explicit state and provenance records—is a real and useful contrast to the tool-selection LLM agents that dominate the current literature. The paper shows a 15-state GCMC-MD desorption campaign running 120 cycles with zero production LLM calls, with one state pausing for review. That is a clean, internally consistent demonstration that long-running workflows do not need an LLM in every loop.\n\nCredit where due: the workflow counts line up, the provenance handling (accepted-state inheritance, RH-local vs absolute timestep) is carefully described, the incident write-up is coherent, and the authors are refreshingly explicit in Section 4 that the two replay cases do not establish general diagnostic accuracy. The repository link is a positive reproducibility signal. The separation between campaign specification and production execution is the kind of pattern that others in agent-assisted simulation will want to borrow.\n\nThe soft spots are real but proportionate. The main one is the event-triggered review path: the single production incident that reached a review boundary was resolved by a code correction, not by the reasoning agent, and the two replay cases are retrospective and frozen—they were never actually applied back into a running campaign. So the paper demonstrates detection/pause and rule-based recovery, plus an offline check that the LLM gives plausible advice on two self-selected cases. It does not demonstrate a live escalation-and-recovery cycle. The stress-test note is right about that, though I would not say the central claim overstates: the abstract and conclusion mostly stay within what was done, and the limitations section is honest. What is mildly over-claimed is the phrase \"validated control handoffs\"—validation here is schema/scope checking, not correctness checking.\n\nThe physical results are treated by the authors as secondary, which is the right call. Single trajectories per state, error bars are temporal fluctuations, and the acceptance criteria are operational indicators rather than equilibrium proof. The Ca-LC0.4 basal spacing and Na-LC0.5 water retention findings are plausible and consistent with prior work, but they are not a standalone validation. No baseline comparison with continuous-LLM or human-only alternatives; the authors flag this themselves. For a framework paper that is an acceptable limitation, not a fatal flaw. The citation pattern looks fine; self-citations point to the validated protocol they extended.\n\nWho is this for? Anyone building agent-assisted simulation workflows, especially in molecular simulation or other long-running stateful campaigns. It deserves peer review: the architectural contribution is clear, the evidence is honestly presented, and the weaknesses are fixable with broader evaluation. I would accept it for refereeing.","headline":"A useful architectural pattern—removing the LLM from routine production—with an honest but thin evaluation of the event-triggered LLM path.","tokens_in":16085,"tokens_out":3236,"would_cite":true,"duration_ms":31257,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Long-running simulation campaigns can run with LLM reasoning confined to planning and rare reviews.","keywords":["large language models","scientific workflows","molecular dynamics","grand canonical Monte Carlo","montmorillonite","provenance","agent-based automation","selective reasoning"],"falsifier":"Rerun any of the 15 accepted states with substantially longer sampling, say five to ten times the RH-local steps, and compare the resulting basal spacing and water populations to the reported accepted-state means and error bars; if the new means lie outside the reported temporal standard deviations, the acceptance criteria were too permissive. A cleaner test is to run two independent replicas from the same restart with different random seeds and check whether the inter-replica spread exceeds the intra-trajectory error bars the paper reports.","tokens_in":15074,"feed_emoji":"🧪","tokens_out":5616,"duration_ms":48388,"temperature":0.7,"pith_summary":"Long-running molecular simulation campaigns are collections of repeated, stateful operations: run a segment, analyze, continue or accept, and inherit the accepted state for the next condition. This paper tries to establish that such campaigns do not need an LLM inside every loop. It presents a framework where an LLM reasoning agent is used only to build the campaign plan and to review states that fixed rules cannot resolve, while a persistent rule-based agent performs routine simulation, analysis, archiving, and state progression. In a demonstration with five montmorillonite systems and three humidity steps, 120 simulation cycles and 15 states completed with zero live reasoning-agent calls during production, and the one ambiguous state was later handled correctly by an event-triggered review. A sympathetic reader would take the point to be that reproducible, auditable agent-assisted simulation can be achieved at lower cost and with less model variability.","feed_headline":"120 simulation cycles completed with zero live LLM calls","feed_subtitle":"An event-triggered review handled the one ambiguous state; LLM reasoning stayed at planning and replay.","key_machinery":"The central mechanism is the separation of roles enforced by a machine-readable campaign specification: a reasoning agent proposes the plan and reviews unresolved events, while a persistent rule-based campaign agent reads approved policies and explicit state records to select the next action. The load-bearing object is the 'campaign state'—a record that for each system-RH state tracks current configuration, sampling progress, acceptance status, archived restarts, and provenance lineage—together with a small set of review-event identifiers such as CONVERGENCE_ANOMALY and PROVENANCE_CONFLICT, plus an evidence bundle stored in agent_event.json and evidence_bundle.json. Any reasoning-agent recommendation is schema- and policy-validated before it can affect production, and state-local (RH-local) step accounting is what prevents the inherited-absolute-timestep misinterpretation from corrupting the campaign. This machinery carries the argument by showing that routine decisions are rule-decidable and only ambiguous states need flexible interpretation.","core_discovery":"On its own terms, the paper's central claim is that selective placement of LLM reasoning—at campaign construction and at event-triggered review—is sufficient to run a long, state-dependent simulation campaign, provided the routine layer is a rule-based campaign agent operating from an approved machine-readable specification, explicit persistent state, and provenance-aware restart inheritance. The supporting demonstration is a GCMC-MD water-vapor desorption campaign over five montmorillonite systems across RH = 0.9, 0.3, and 0.1, with state-specific sampling lengths determined by repeated reassessment. The workflow accepted all 15 system-RH states and completed 10 RH transitions without a single non-interactive reasoning-agent invocation during production. One state (K-LC0.4 at RH=0.3) hit a review boundary because a production limit was evaluated with the inherited absolute timestep rather than the state-local step count; the workflow was corrected, resumed, and accepted. Two frozen incidents—that timestep misaccounting and a stale smoke-test archive—were replayed through the reasoning agent in blind mode, and its recommendations matched the recorded resolutions. The campaign also produced physical results: Ca-LC0.4 retained more interlayer water and a roughly 14.4 Angstrom basal spacing at low RH, while Na-LC0.5 retained the most residual water, a finding the paper attributes to the combined variation of layer charge and compensating Na+ population.","pith_inferences":["The zero-invocation production result suggests that the LLM cost of a campaign can be nearly independent of campaign length, provided acceptance criteria are conservative enough—a cost ladder the paper does not directly measure.","The event-triggered design could transfer to other restartable scientific pipelines, such as adaptive free-energy calculations or materials screening, wherever successor states inherit only from authoritative archives; the paper proposes but does not test such transfers.","A natural next experiment is to inject a set of seeded production anomalies (wrong restart, timestep misaccounting, stale archive) and measure the reasoning agent's diagnostic accuracy against known ground truth, since the current two-case replay cannot establish general failure-diagnosis performance."],"forward_implications":["If selective reasoning is sufficient, long-running simulation campaigns can be automated around a deterministic core, cutting token costs and eliminating model variability in routine actions.","A reusable pattern emerges: define explicit state, acceptance criteria, restart provenance, and permitted next actions in machine-readable form, and use an LLM only at planning and at review boundaries.","Scientific software should expose structured agent-oriented interfaces (current state, valid restart, completed analyses) rather than forcing an LLM to interpret human-oriented directories and logs.","Workflow incidents can be safely contained: a state that fails rule-based resolution pauses before propagating to downstream states, and reasoning-agent proposals are validated before execution.","The physical findings on cation- and charge-dependent dehydration of montmorillonite are direct corollaries of the accepted-state records, contingent on those records being true converged states."],"supporting_citations":[{"why":"Supplies the all-LLM reasoning-action baseline that this work's selective design is meant to improve upon.","marker":"[5]"},{"why":"Establishes the pattern of persistent, stateful workflow management and data provenance for simulation campaigns.","marker":"[3]"},{"why":"Shows reproducible, traceable, composable coarse-grained simulation workflows, grounding the provenance-aware progression idea.","marker":"[4]"},{"why":"Provides the previously validated GCMC-MD clay-water protocol and the Ca-montmorillonite hysteresis behavior this campaign builds on.","marker":"[17]"},{"why":"Provides the clay-system model generation tool used to construct the montmorillonite structures from mineral composition.","marker":"[26]"},{"why":"Supplies the ClayFF force field describing clay frameworks and exchangeable ions.","marker":"[27]"},{"why":"Supplies the SPC/E water model used for water molecules.","marker":"[28]"},{"why":"Provides the molecular dynamics engine that executes the simulations.","marker":"[29]"},{"why":"Gives the saturation vapor pressure of SPC/E water at 300 K, used to set the imposed water-vapor pressures at each RH.","marker":"[31]"}],"fun_headline_variants":["LLM only at planning and event review in 120-cycle run","120 cycles, zero live LLM calls","Selective LLM: 120 cycles without a runtime reasoning call","Event-triggered LLM review: 120 cycles, one flag","LLM reasoning only at planning and escalations in 120 cycles"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The acceptance criteria—coefficients of variation of 0.03, 0.05, and 0.08 for total, interlayer, and external-surface water, plus trend and health checks—are treated as sufficient signs that a state is converged; the paper itself calls these operational indicators rather than proof of thermodynamic equilibrium.","fun_headline_variants_meta":{"raw":{"variants":["LLM only at planning and event review in 120-cycle run","120 cycles, zero live LLM calls","Selective LLM: 120 cycles without a runtime reasoning call","Event-triggered LLM review: 120 cycles, one flag","LLM reasoning only at planning and escalations in 120 cycles"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001096,"raw_usage":{"total_tokens":4676,"prompt_tokens":1148,"completion_tokens":3528,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":764,"completion_tokens_details":{"reasoning_tokens":3440}},"tokens_in":764,"tokens_out":3528,"duration_ms":28673,"temperature":1.0,"reasoning_tokens":3440,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T00:27:23.150848+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun any of the 15 accepted states with substantially longer sampling, say five to ten times the RH-local steps, and compare the resulting basal spacing and water populations to the reported accepted-state means and error bars; if the new means lie outside the reported temporal standard deviations, the acceptance criteria were too permissive. A cleaner test is to run two independent replicas from the same restart with different random seeds and check whether the inter-replica spread exceeds the intra-trajectory error bars the paper reports.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the all-LLM reasoning-action baseline that this work's selective design is meant to improve upon."},{"cited_title":"S.; Dodd, P","cited_arxiv_id":null,"evidence_quote":"Establishes the pattern of persistent, stateful workflow management and data provenance for simulation campaigns."},{"cited_title":"J.; Rudzinski, J","cited_arxiv_id":null,"evidence_quote":"Shows reproducible, traceable, composable coarse-grained simulation workflows, grounding the provenance-aware progression idea."},{"cited_title":"T.; Liang, J.-J.; Kalinichev, A","cited_arxiv_id":null,"evidence_quote":"Supplies the ClayFF force field describing clay frameworks and exchangeable ions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the SPC/E water model used for water molecules."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the saturation vapor pressure of SPC/E water at 300 K, used to set the imposed water-vapor pressures at each RH."}],"review_version":1}