{"id":"34206b08-b921-437a-90a2-1eae1672c3c4","arxiv_id":"2607.18249","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An LLM-guided multi-agent system converts natural-language questions into executable constraint-based metabolic-model workflows and generates drug-therapy hypotheses, demonstrated on RA and MS immune-cell models.","lead":"MechAInistic is a multi-agent AI system that turns plain-language questions about cell metabolism into executable, model-grounded analyses of paired metabolic models. It aims to let biologists without programming skills ask drug-target questions and receive traceable, quantitative hypotheses.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central capability claim rests on a single un-replicated run per use case, without public code or full tool-call logs; stochasticity and manual curation could account for the reported traceability.","rationale":"The central claim is a capability claim about MechAInistic, not a claim about disease biology. The reader's weakest_assumption focuses on the fidelity of the reconstructed disease models. That concern is real but secondary, because the paper explicitly frames its outputs as hypotheses and disclaims model validity. The more load-bearing issue is whether the reported workflow is reproducibly produced by the system itself. The paper provides no code, no full audit logs, and only one run per use case, while LLM temperature 0.5 implies stochasticity. Manual validation after report generation identified errors (title-author mismatch, incorrect ICDHy gene identifier, alternative drug preferred), meaning the final narrative incorporates human curation. These facts do not disprove the claim, but they weaken the evidence that the system alone produces the traceable workflow described. The proposed repeated-runs test would settle whether the demonstrated workflow is a robust characteristic of the system or an artifact of a single favorable run. The verdict remains CONDITIONAL, as the reader had it, but for a somewhat different reason.","tokens_in":24589,"tokens_out":4598,"duration_ms":57340,"concrete_test":"Run MechAInistic 10 times on each of the two use cases using the same uploaded models, the same prompt, fixed random seeds, and pinned LLM versions. Record full tool-call transcripts and final drug hypotheses. Then compare: (1) the fraction of runs that reproduce the reported core conclusions (AKGDm/Devimistat for RA; ICDHy/ivosidenib for MS), (2) run-to-run variance in flux-distance metrics and perturbation results, and (3) how many runs required manual correction to be 'traceable.' If fewer than 80% of runs reproduce the reported conclusions, the paper should narrow its claim to a proof-of-concept single demonstration rather than a general capability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims MechAInistic converts a natural-language question into an executable, model-grounded, traceable workflow. The support is two use cases, each run once with LLM temperature 0.5. The paper's own Discussion states that LLM behavior can vary across runs, and Methods describe manual validation after report generation, including corrections (title–author mismatch, ICDHy gene identifier, vorasidenib over ivosidenib). Without multiple independent runs, exact prompt/seed/version reporting, complete tool-call transcripts, and public code, the demonstrated workflow may reflect curated post-hoc reporting rather than a reproducible system output. This is load-bearing because the central claim is about what the system itself does, not about the biological validity of the reconstructed models. If reruns yield different workflows or different drug targets, the single-run demonstrations are anecdotal and the 'converts' claim is overstated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents MechAInistic, a multi-agent LLM-guided system for reasoning over paired genome-scale constraint-based metabolic models. The system uses an Architect-Reviewer-Task architecture to translate a natural-language biological question into an executable, tool-grounded workflow, then produces a structured report with flux-based quantitative evidence, literature support, and explicit limitations. The authors demonstrate the system on two therapeutic hypothesis-generation use cases: RA naive B cells (recommending Devimistat/CPI-613 targeting OGDH) and MS CD4+ Th17 cells (recommending ivosidenib targeting IDH1/IDH2). They also compare MechAInistic with general-purpose LLM chatbots using a nine-axis rubric, reporting that MechAInistic achieves better model grounding and traceability. The manuscript is transparent about several limitations, including dependence on input model quality, variability of LLM behavior across runs, the need for manual validation, and documented citation-level errors.","tokens_in":24790,"tokens_out":7344,"duration_ms":83668,"significance":"If the central claim is supported, MechAInistic would be a valuable contribution: it lowers the barrier to constraint-based metabolic modeling, provides an auditable chain from user query to model-derived results, and demonstrates a meaningful use of multi-agent LLM systems for scientific reasoning. The architecture is clearly described, the tool registry is sensible, and the authors explicitly frame their recommendations as model-grounded hypotheses rather than validated therapies. The honesty about failure modes and the inclusion of two real biological use cases are strengths. However, the current evidence is closer to a pilot demonstration than a validated system claim: each use case is run once, the outputs are manually curated, and the central claim that the system 'converts' a natural-language question into a workflow is therefore not yet established with the reproducibility expected for a methods paper.","major_comments":[{"comment":"The two use-case demonstrations are each a single execution at temperature 0.5 with no random seed or model-version pinning reported in the main text. The paper's own Discussion states that LLM behavior can vary across runs, so a single success does not establish that MechAInistic reliably 'converts' natural-language questions into workflows. Public code, full tool-call transcripts, and independent reruns are needed; without these, the central claim in the Abstract is anecdotal. I ask the authors to report at least 3-5 runs per use case, with the variance in workflow structure and recommendations, and to release logs and code.","section":"Results 'Model-grounded workflows'; Methods 'LLM configuration and prompt management'"},{"comment":"The final reports appear to have been produced after manual validation that corrected a title-author mismatch, an ICDHy annotation error, and the final drug choice (vorasidenib over ivosidenib). The manuscript does not show the raw pre-validation system outputs or a diff of what human reviewers changed. Because the central claim is that the system itself generates traceable reports, the reader must be able to distinguish system output from curated post-processing. Please include raw outputs or an explicit change log.","section":"Methods 'Manual validation of MechAInistic outputs'; Results use cases"},{"comment":"The primary evidence for improved grounding over general-purpose LLMs is a nine-axis rubric scored qualitatively by the authors, with no blinding and no inter-rater reliability assessment. Since the same group designed MechAInistic and the rubric, the comparison is vulnerable to confirmation bias. Please provide independent scoring, or at minimum raw comparator outputs and detailed per-axis justifications, to support the claim of 'most consistent workflow-level performance'.","section":"Methods 'Comparator LLM evaluation' and 'Nine-axis evaluation rubric'"},{"comment":"Reference 41 is cited as the source of the RA bulk RNA-seq data, but the reference is Wang et al. (Nat. Commun. 2018) on IL-21 in SLE, not RA. Reference 38 is cited for the healthy Naive B model but is the COMO pipeline preprint. These provenance errors prevent readers from reproducing the input models and weaken the biological use cases. Please verify all accessions/references and, if the text is correct, add the GEO accessions so the models can be reconstructed.","section":"Methods 'Metabolic model reconstruction and use-case setup'"},{"comment":"The system operationalizes 'therapeutic efficacy' as reduction of Euclidean flux distance to the healthy state (e.g., 25.247 to 23.069 cited as '8.6% improvement'). This is an unvalidated modeling assumption, not a model-derived theorem; different distance metrics or objective functions could change the ranking of targets and drugs. The paper should explicitly label this assumption and include sensitivity analysis across distance measures (Euclidean, Manhattan, cosine, etc.) to support the robustness of the recommended hypotheses.","section":"Results use cases; Methods 'Constraint-based modeling backend'"}],"minor_comments":[{"comment":"The phrase 'the all summarized tool call results' is ungrammatical; should be 'all the summarized tool call results'.","section":"Figure 2D / Results Evaluation phase"},{"comment":"Capitalization of COBRApy is inconsistent ('CobraPy' appears in the Constraint-based modeling backend section); please use the standard 'COBRApy'.","section":"Throughout"},{"comment":"Model versions such as 'Kimi K2.6, Qwen3.5, Gemma-4' should include access dates and provider endpoints; it is also not clear in the main text which agent was assigned to which model.","section":"Table 3 and Results"},{"comment":"The Abstract states that MechAInistic 'proposed ivosidenib as an FDA-approved repurposing candidate,' but the Results note that human review suggested vorasidenib may be mechanistically preferable. Please qualify the Abstract to make clear that ivosidenib is the system's output and that expert review altered the compound-level recommendation.","section":"Abstract vs. Results"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely publishable after major revision. The central architecture is promising and the authors are transparent about limitations, which is a positive sign. However, the evaluation is currently closer to a pilot demonstration than to a validated system claim, and the provenance errors in the Methods section need correction. I do not see evidence of misconduct, but the manuscript's abstract overstates the strength of the current evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper before the field floods with LLM-agent metabolic tools: MechAInistic is one of the more serious attempts, and it is not just a wrapper around ChatGPT. The Architect-Reviewer-Task architecture is sensible, the tool registry is real (COBRApy, HiGHS, flux distance, perturbation, DGIdb, PubMed), and the authors did the right thing by pitting general-purpose LLMs against it on the same models and prompt. Their audit of Claude showing plausible narrative hiding execution-level errors is a nice piece of evidence for why tool grounding matters.\n\nWhat is genuinely new: applying a reviewer-supervised agentic loop specifically to paired constraint-based models and producing two concrete, model-grounded drug hypotheses (Devimistat/CPI-613 for RA naive B cells, ivosidenib for MS Th17 cells). The paper is also refreshingly honest about its own failure modes: a citation mismatch, an annotation error (ICDHy), the need for manual correction of the final drug, and the explicit statement that the disease models require separate validation. That honesty is real and should be credited.\n\nThe soft spots are the ones the stress test names. Each use case is a single run at temperature 0.5 with no seed, no complete tool-call transcript, and no public code repository. The authors manually validated and corrected the outputs after generation, so the \"traceable\" workflow is partly a post hoc curation. The RA model comes from four RNA-seq samples from one study; the MS model from nineteen. The nine-axis rubric was scored by the same team that built the system. None of this is fatal on its own, but together it means the headline claim—\"converts natural-language questions into executable, model-grounded workflows\"—is supported more as a demonstration of concept than as a reproducible system property. The authors themselves acknowledge LLM variability across runs, which undercuts any single-run demonstration.\n\nMy bottom line: this deserves serious peer review, not desk rejection. The architecture and the comparator methodology are worth the field's attention. But the referees should push hard for public code, multiple independent runs, and raw tool logs; without those, the central claim remains anecdotal. As it stands, I would not yet cite it as a validated system, but I would tell people to watch it.\n\nSend it to review, with a request for substantial revision and reproducibility evidence.","headline":"A genuinely useful, honestly reported multi-agent system for constraint-based metabolic modeling, but the evidence for the central 'converts' claim rests on two single runs with manual correction and no public code.","tokens_in":25271,"tokens_out":1373,"would_cite":false,"duration_ms":19374,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MechAInistic claims that an LLM-based multi-agent system can turn plain-English biology questions into executable, traceable workflows over genome-scale metabolic models, yielding auditable drug-target hypotheses.","keywords":["agentic AI","workflow orchestration","metabolic modeling","constraint-based modeling","mechanistic AI","drug repurposing","flux balance analysis","therapeutic hypothesis generation"],"falsifier":"Construct two steady-state models that are identical except for a single known reaction knockout or a single known drug-target perturbation, run MechAInistic with the standard prompt, and check whether the system's highest-ranked single target is that known reaction; an exhaustive in-silico screen can then verify whether the recommended target actually minimizes flux distance to the target state.","tokens_in":24502,"feed_emoji":"🧬","tokens_out":5194,"duration_ms":49314,"temperature":0.7,"pith_summary":"MechAInistic claims that large language models become trustworthy for metabolic-model reasoning when their outputs are forced through a fixed registry of executable modeling tools and a reviewer agent that checks each step. The paper builds a three-agent system—an Architect that plans tool calls, a Reviewer that scores plans and evidence, and a Task agent that summarizes outputs and searches PubMed—and shows it converting a plain-English drug-repurposing question into a concrete workflow on paired disease/healthy metabolic models. In two immune-cell use cases the system nominated Devimistat/CPI-613 against OGDH in rheumatoid arthritis naive B cells and ivosidenib against IDH in multiple sclerosis Th17 cells, each choice backed by flux-distance metrics and simulated perturbations. The authors deliberately treat the drug identities as hypotheses, not validated answers; the headline result is that a natural-language question can be turned into an auditable chain of model-derived numbers, perturbation simulations, and literature support.","feed_headline":"LLM agents turn plain-English queries into metabolic-model screens","feed_subtitle":"Two disease use cases show the system flags OGDH and IDH as drug targets, with an auditable chain back to the models.","key_machinery":"The load-bearing mechanism is the separation of planning from verification. Three role-specialized agents—Architect, Reviewer, and Task—operate on a session state that stores the uploaded models, the user query, and a complete record of tool calls. The Architect proposes tool calls drawn from a fixed registry of COBRApy-backed functions (flux-balance simulation, flux-distance calculation, reaction/pathway comparison, gene/reaction knockout, dose-response, drug-target lookup, PubMed retrieval); the Reviewer scores plans and intermediate evidence on a numeric scale, forcing regeneration when scores are low; the Task agent summarizes raw JSON tool outputs into plain text, cutting token usage by","core_discovery":"The central claim is that workflow structure, not language-model size, is what makes LLM-based reasoning trustworthy: MechAInistic's Architect-Reviewer-Task loop keeps biological conclusions grounded in executable constraint-based modeling rather than free-form text. Concretely, the paper shows the system taking the same prompt—identify a single drug therapy that restores disease-state flux to the healthy state while minimizing off-target disruption—and, in both the RA and MS use cases, loading the JSON models, computing baseline flux distances, prioritizing reactions that carry disease-only flux, simulating inhibition, and retrieving literature before naming a therapy. The system also outpe","pith_inferences":["If the Architect-Reviewer pattern generalizes, the same grounding trick could be applied to other simulation-backed biological questions, such as kinetic models or agent-based immune models, not just constraint-based metabolism.","The flux-distance-minimization objective used here could serve as a community benchmark: a documented library of paired models with known interventions would let anyone measure whether agentic LLM tools recover the known target.","Because LLM outputs vary across runs and providers, practical adoption will likely require pinned model versions, deterministic seeds, and local caching of external database calls; the paper reports configurations but does not promise bit-for-bit reproducibility.","The RA and MS reconstructions rest on 4 and 19 bulk RNA-seq samples respectively; if those models are not faithful, the mechanism-level claims become artifacts of the reconstruction, which suggests that agent tools may push validation effort upstream into model building."],"forward_implications":["A biologist who can phrase a question in English can run a full constraint-based-model drug-target screen without writing code, provided they can supply paired current/target models.","The same Architect-Reviewer-Task workflow transfers to any pair of COBRA-compatible models, so the method is not tied to the two immune-cell use cases shown.","Comparisons of LLM-based analysis tools should be scored on traceability and execution audit, not on the plausibility of the final narrative or the identity of the recommended drug.","Because the system keeps quantitative outputs in the report, hypotheses can be independently re-run and audited—a prerequisite for using AI-generated biology in a lab setting.","The paper's own caveats imply that the limiting factor for hypothesis quality becomes the input metabolic model, not the agent software."],"fun_headline_variants":["Workflow, not model size, keeps LLM metabolic reasoning grounded","LLM agents convert natural language into executable metabolic screens","MechAInistic: structure over scale for model-grounded drug hypotheses","LLM agents find drug targets with executable metabolic-model workflows"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The disease-state metabolic models, reconstructed from very few bulk RNA-seq samples (4 for RA, 19 for MS) via COMO on Recon3D, faithfully capture the metabolic state of those immune cells, so that flux differences from the healthy reference reflect disease biology rather than reconstruction artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Workflow, not model size, keeps LLM metabolic reasoning grounded","LLM agents convert natural language into executable metabolic screens","MechAInistic: structure over scale for model-grounded drug hypotheses","LLM agents find drug targets with executable metabolic-model workflows"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000278,"raw_usage":{"total_tokens":1511,"prompt_tokens":784,"completion_tokens":727,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":655}},"tokens_in":528,"tokens_out":727,"duration_ms":7517,"temperature":1.0,"reasoning_tokens":655,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T14:33:01.450494+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct two steady-state models that are identical except for a single known reaction knockout or a single known drug-target perturbation, run MechAInistic with the standard prompt, and check whether the system's highest-ranked single target is that known reaction; an exhaustive in-silico screen can then verify whether the recommended target actually minimizes flux distance to the target state.","supporting_citations":[],"review_version":1}