{"id":"47cf35ff-14cd-4737-9027-9da84bdef689","arxiv_id":"2508.19800","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"The paper proposes, but does not implement or validate, a multi-agent AI framework for cross-scale modeling of human biology from molecules to whole body, with sketches of metastasis scoring and drug development.","lead":"This preprint proposes an architecture for a 'Full-Body AI Agent' made of seven biological-level AI agents coordinated by a supervisor, aimed at modeling disease and drug response from molecules to the whole organism. It is a vision paper: no working system, data release, or validation is provided, so its usefulness is as a roadmap, not as demonstrated capability.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Temporal alignment of heterogeneous data, admitted unsolved in §7, is load-bearing; without it the cross-scale loop cannot predict long-term efficacy or toxicity.","rationale":"The paper is transparent: it calls the framework a proposal, lists challenges, and admits temporal alignment is unsolved. That admission is exactly what makes this concern load-bearing rather than an external disagreement with consensus. The central claim is about predictive modeling of long-term outcomes; such modeling is inherently temporal and cross-scale. The architecture's reasoning mechanism relies on iterative exchange of outputs, which presupposes a shared coordinate system in time. No such method is specified. The example in Case 1 does not exercise the temporal dimension and omits most agents. Therefore, as it stands, the framework cannot realize its claimed advantage. The reader's weakest assumption correctly identifies the same issue. I do not think this changes the verdict; if anything it strengthens the rejection. A roadmap can still be useful, but the central claim is currently unsubstantiated. I recommend no change to the reader's REJECT verdict.","tokens_in":33687,"tokens_out":3339,"duration_ms":38063,"concrete_test":"Select a longitudinal publicly available cohort with repeated measurements at ≥2 biological levels (e.g., TCGA with serial transcriptomic and imaging, or MIMIC-IV with serial labs and notes). Define the temporal-alignment protocol that maps each level onto a shared timeline, and implement the Full-Body AI-Agent to predict a long-term endpoint (e.g., 5-year progression-free survival or a late-onset toxicity). Compare its time-dependent discrimination (e.g., AUC at 5 years) with a localized model that uses only one level (e.g., molecular markers). The concern is settled if the full-body model fails to beat the localized baseline, or if the protocol cannot be concretely specified from the paper's Data Commons/Common Format sections.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on the Full-Body AI Agent's ability to predict long-term efficacy and toxicity. That prediction is longitudinal, requiring temporally aligned data across biological scales. The paper's own §7 states: 'the temporal alignment of heterogeneous datasets, particularly for dynamic physiological processes, is still a largely unsolved problem.' This is not a peripheral caveat; it is the load-bearing premise of the architecture. The framework's bidirectional feedback loop (§3.4) requires outputs from one level to be iteratively combined with others, and without a common temporal reference, the loop cannot propagate information through time. The only quantitative example (Case 1) uses a static snRNA-seq snapshot; the Tissue and Organ agents are not applied, and no longitudinal endpoint is predicted. Thus the paper supplies no instantiation of the temporal-alignment mechanism and explicitly concedes it is unsolved. If this premise fails, the system cannot deliver the comparative advantage claimed in the abstract, and 'long-term efficacy and toxicity beyond localized models' is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a multi-agent AI architecture, the \"Full-Body AI Agent,\" consisting of a supervisory agent and seven biological-level agents (molecule, organelle, cell, tissue, organ, organ system, body system). The authors describe a reasoning pipeline spanning data perception, hypothesis generation, task decomposition, execution, and iterative multi-scale feedback, and they propose two applications: a Metastasis AI Agent with Initiation, Dissemination, and Colonization scores, and a Drug AI Agent intended to couple molecular discovery with organoid/organ-on-chip models under whole-body constraints. The only quantitative exercise is an Initiation Score computed from a single NSCLC snRNA-seq sample; the Tissue and Organ agents are not applied. The abstract claims that the approach \"enables the predictive modeling of long-term efficacy and toxicity beyond what localized models alone can achieve.\"","tokens_in":34010,"tokens_out":5360,"duration_ms":64601,"significance":"If the framework were implemented and validated, it could meaningfully advance multi-scale biomedical modeling by formalizing cross-level integration and providing an organizing substrate for heterogeneous data. The paper's strengths are its systematic decomposition of biological levels, its broad literature synthesis, and the reasonable three-phase metastasis scoring concept. However, the manuscript provides no implementation, no end-to-end validation, and no comparison against localized models. The one worked example is a static, single-sample analysis with a partly circular design, and the paper itself concedes in §7 that temporal alignment of heterogeneous datasets is \"still a largely unsolved problem.\" Thus the central predictive claim is not supported by the evidence presented. As a vision or perspective, the paper has heuristic value, but as a research claim it falls short of the stated scope.","major_comments":[{"comment":"The only quantitative demonstration does not support the claim to predict metastatic potential beyond localized models. It analyzes one static snRNA-seq sample (N2254) and computes an Initiation Score using curated EMT, stemness, and metabolic gene sets with hand-set thresholds (top 20% hybrid EMT; top 25% high-initiation, threshold 0.324) and PCA first-component weights. Cells classified as high-initiation are then used in differential expression analysis that reports upregulation of CD44, FN1, and VIM; these are expected components of the very EMT/stemness/invasion programs used to construct the score, so the \"validation\" is partly circular. The authors also state that the Tissue AI Agent and Organ AI Agent were not applied, so the exercise does not demonstrate cross-scale integration. No clinical endpoint, independent cohort, or comparison with existing localized predictors is provide","section":"§5, Case 1 (Initiation Score)"},{"comment":"The central abstract claim concerns \"long-term efficacy and toxicity,\" which is inherently longitudinal. Yet §7 explicitly states that \"the temporal alignment of heterogeneous datasets, particularly for dynamic physiological processes, is still a largely unsolved problem.\" The bidirectional feedback loop in §3.4 requires outputs from one biological level to be iteratively exchanged with others, but no mechanism, formalism, or data structure for temporal alignment is defined or demonstrated. The only worked example is a static snapshot. Without a temporal reference frame, the system cannot propagate information through time, and the claimed advantage over localized models is therefore unsupported. This is a load-bearing gap, not a peripheral caveat.","section":"§7 and §3.4"},{"comment":"The drug development case study is a tool enumeration rather than a demonstration. The text lists AutoDock-GPU, GROMACS, pkCSM, hERG-Block, Tox21, retrosynthesis tools, and organoid/organ-on-chip concepts, but no end-to-end agent run, no quantitative prediction of efficacy or toxicity, and no benchmark against standard preclinical models is reported. The cited successes (ISM001-055, Halicin, Baricitinib) are external discoveries made by other systems and cannot serve as evidence for the proposed framework. The conclusion that the Drug AI Agent \"can transcend conventional siloed stages\" is therefore a hope, not a result.","section":"§5, Case 2 (Drug AI Agent)"},{"comment":"The claimed comparative advantage over existing multi-agent systems rests on the reliability of LLM-based agents to perceive, reason, and execute domain-specific biological analyses, and on convergence of the iterative cross-scale loop without amplifying errors. No evidence is provided for either. Section 6.1 itself lists data integration, interpretability, scalability, and data quality as open challenges, but the paper does not show how the framework mitigates them. Until at least a pilot implementation with end-to-end results is presented, the assertion that this architecture outperforms localized models or existing multi-agent systems remains a conjecture.","section":"§3.5 and §6.1"}],"minor_comments":[{"comment":"The legend reads \"Predefined detachable biological tasks for the Organ System AI Agent,\" but Section 4.7 describes the Body System AI Agent. Please correct the mismatch.","section":"Figure 10 legend"},{"comment":"Reference [53] is cited for PharmAgents but points to a paper on biodegradable metal–organic frameworks, and reference [54] is cited for DrugAgent but points to a paper on lithium void formation in solid-state batteries. These citations appear to be unrelated to the claimed agents; please verify and correct.","section":"References [53] and [54]"},{"comment":"There is a typo: \"evaluate the Initiation core\" should presumably read \"evaluate the Initiation score.\" Also, \"Molecular AI Agent\" and \"Molecule AI Agent\" are used interchangeably; please standardize.","section":"§5, Case 1"},{"comment":"The table lists \"Biomeni\" while the text and references use \"Biomni.\" Please harmonize.","section":"Table 1"},{"comment":"Code and data processing scripts for the Case 1 snRNA-seq analysis are not provided. Given the emphasis on provenance and auditability in Section 3, the authors should make these available, ideally alongside the Supplementary Table 1 repository.","section":"Reproducibility"}],"recommendation":"reject","confidential_remarks":"I see this manuscript as a well-structured position paper rather than a demonstration of the claimed capabilities. The central predictive claim is not merely missing a benchmark; the only quantitative example is circular in its selection of high-initiation cells, and the temporal-alignment requirement for long-term prediction is explicitly acknowledged as unsolved. These are load-bearing issues. If the authors wish to pursue publication, they could reframe the work as an explicitly conceptual framework and remove unsupported claims about demonstrated superiority; however, in its current form the manuscript does not meet the evidentiary standard for the claims made."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper before reading it: it is a roadmap, not a research claim. The authors propose a seven-level hierarchy of AI agents (molecule to body) coordinated by a supervisor, and they describe two applications—metastasis scoring and drug development. There is no running system, no validation of the central framework, and the one quantitative example is illustrative. If you read it as a vision paper, it is useful; if you read it as a claimed advance, it falls short.\n\nWhat is new and good: the synthesis is genuinely careful. The paper catalogs existing multi-agent systems (Robin, OriGene, Biomni, etc.), maps data types and standard formats to each biological level, and lists concrete tools per agent. The seven-level decomposition is a plausible organizing scheme, and the bidirectional feedback loop (lower-level mechanisms constrain higher-level states and vice versa) is a sensible design principle. The two case studies are described in enough detail to show how the framework would operate, and the authors are honest about several limitations.\n\nNow the soft spots. The worked Initiation Score example is partly circular: cells are selected as \"high initiation\" using curated EMT, stemness, and metabolism gene sets, then differential expression reports those same genes as enriched. Thresholds (top 20%, top 25%, PCA first component) are hand-set with no sensitivity analysis. More importantly, the paper's abstract claims predictive modeling of long-term efficacy and toxicity \"beyond what localized models alone can achieve,\" but nothing in the paper instantiates that. The system needs temporally aligned data across scales to make longitudinal predictions, and Section 7 admits that alignment is \"a largely unsolved problem.\" That is not a peripheral caveat; it is the load-bearing premise. Without it, the cross-scale loop cannot propagate information through time, and the claimed comparative advantage is unsupported. The Drug AI Agent is also entirely conceptual—no drugs were designed or tested. Finally, the reference list has a concrete problem: refs [53] and [54] point to materials-science papers (biodegradable metals, lithium batteries) instead of PharmAgents and DrugAgent. That is sloppy and undermines confidence in the citation check.\n\nBottom line: as a roadmap for groups building multi-scale biomedical agents, it is a reasonable starting point. As a paper claiming predictive capability, it is not there. I would send it to peer review because the synthesis is valuable and the limitations are explicitly acknowledged, but I would expect heavy revision—either reposition it as a perspective piece or add real validation with baselines and external outcomes. I would not cite it as evidence of capability, but I might cite it as an architecture proposal.","headline":"A serious, well-organized roadmap for multi-scale biomedical AI agents, but the predictive claims are not backed by implementation; the temporal-alignment problem is admitted and load-bearing.","tokens_in":34434,"tokens_out":1761,"would_cite":true,"duration_ms":22305,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that a supervisory AI agent coordinating seven biological-level agents can model human physiology across scales, enabling long-term drug efficacy and toxicity prediction beyond localized models.","keywords":["Full-Body AI Agent","multi-agent systems","cross-scale reasoning","systems biology","tumor metastasis scoring","drug development","organoids and organ-on-chip","large language models"],"falsifier":"Take a retrospective cohort with long-term follow-up and compute the three-phase metastasis scores from the same multi-scale data; if the scores do not predict metastatic progression beyond standard staging or beyond a molecular-only score, the cross-scale advantage fails. Likewise, a prospective evaluation of the Drug AI Agent's predicted organ-level toxicities against clinical trial outcomes would settle the efficacy/toxicity claim.","tokens_in":33620,"feed_emoji":"🧬","tokens_out":5952,"duration_ms":60346,"temperature":0.7,"pith_summary":"The paper proposes a Full-Body AI Agent: a supervisor AI that coordinates seven specialized agents—molecule, organelle, cell, tissue, organ, organ system, and body system—to model human biology from molecular to whole-body scales. The central claim is that because the agents exchange results iteratively and bidirectionally, the system can predict long-term drug efficacy and toxicity better than models confined to one biological level. The framework is demonstrated in two cases: a metastasis AI Agent that scores tumor progression in three phases, and a drug AI Agent that guides organoid and chip-based preclinical models under whole-body physiological constraints. A worked example computes an Initiation Score from snRNA-seq data for a lung cancer sample, showing one phase in practice. A sympathetic reader would care because the proposal directly targets the translational gap where molecular findings fail to predict systemic outcomes.","feed_headline":"Seven AI agents model the body from molecule to whole organism","feed_subtitle":"A supervisor agent unites seven scale-specific agents so drug and metastasis predictions get whole-body constraints.","key_machinery":"The load-bearing mechanism is the hierarchical multi-agent loop: a supervisory Full-Body AI Agent plus seven basic agents, each bound to one biological scale, communicating through a standardized Data Commons. The loop enforces bidirectional constraints—molecular and cellular findings propagate upward to tissue, organ, and system levels, while systemic plausibility filters propagate downward to refine or reject lower-level explanations. The implemented demonstrations are the Metastasis AI Agent's initiation/dissemination/colonization scores and the Drug AI Agent's full-body physiological wrapping of preclinical organoid and chip models.","core_discovery":"The central discovery is architectural: a multi-level biological problem can be decomposed into tasks assigned to seven biology-grounded agents, whose outputs are then re-integrated through iterative bidirectional exchange. The Full-Body AI Agent acts as both supervisor and integrator, perceiving multi-modal data, generating cross-scale hypotheses, decomposing problems, and looping conclusions from one biological level back into others until a physiologically plausible whole-body model converges. On this basis the paper claims predictive power for long-term efficacy and toxicity that localized models lack, and it instantiates the claim in a three-phase metastasis scoring system and a drug-de","pith_inferences":["The three-phase metastasis scores look like a testable prognostic instrument: if they beat conventional staging in retrospective cohorts, they could become a clinical biomarker; the paper does not itself run that validation.","If temporal alignment is solved, the same supervisor architecture could evolve into a continuous digital twin of an individual patient, updating predictions as new clinical and wearable data arrive.","The 'whole-body constraints' idea suggests a new reporting standard for in vitro models: every organoid or chip result should include a statement of which systemic constraints it may violate.","The seven-agent decomposition could be transferred to non-human species or even multi-organism systems by swapping level-specific agents, a direction the paper leaves implicit."],"forward_implications":["If the framework works, cancer metastasis risk could be scored phase-by-phase (initiation, dissemination, colonization), letting clinicians target stage-specific vulnerabilities.","Drug developers could place organoid and organ-on-chip results inside whole-body constraints, catching long-term and distal toxicities before clinical trials.","Molecular discoveries would no longer be interpreted in isolation; every finding could be traced through tissue, organ, and system effects.","The same architecture could reduce the more-than-90% attrition of drug candidates by exposing system-level failures earlier in the pipeline.","Standardized data commons would make multi-omics, imaging, and clinical data interoperable across the seven biological levels."],"supporting_citations":[{"why":"Supplies the prior multi-agent virtual-cell vision that the Full-Body framework extends to whole-organism physiology.","marker":"[29]"},{"why":"Robin is the closed-loop multi-agent baseline; the paper contrasts its scale-bounded loops with full-body cross-scale reasoning.","marker":"[30]"},{"why":"AI Co-Scientist is used in the pipeline for hypothesis proposal and ranking, and compared as an epistemic-refinement system.","marker":"[35]"},{"why":"Coscientist provides the LLM-driven autonomous experimentation and protocol-parsing layer for data perception and execution.","marker":"[23]"},{"why":"ReAct supplies the planning policy for task decomposition in the reasoning pipeline.","marker":"[52]"},{"why":"Bioteque's pre-computed biomedical embeddings ground the Data Commons approach for standardizing heterogeneous data.","marker":"[38]"},{"why":"scGPT provides pretrained single-cell representations that the Cell AI Agent uses for cell-state analysis.","marker":"[41]"},{"why":"The NSCLC single-nucleus dataset (sample N2254) is the worked example where the Initiation Score is actually computed.","marker":"[113]"},{"why":"Organoids and organs-on-chips are the preclinical platforms that the Drug AI Agent wraps with whole-body constraints.","marker":"[118]"},{"why":"Documents the more-than-90% clinical attrition rate that motivates the claim that localized models cannot predict systemic outcomes.","marker":"[116, 117]"}],"fun_headline_variants":["AI agent integrates seven biological scales for whole-body modeling","Full-body AI agent links molecules to organs in one framework","Two AI agents show full-body modeling for metastasis and drugs","Supervisor AI integrates seven scale-specific agents to model body","Multi-scale AI framework predicts drug effects and metastasis across body"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The framework assumes that heterogeneous biological data at different scales can be standardized and temporally aligned into a single reasoning loop; the paper itself flags temporal alignment as largely unsolved.","fun_headline_variants_meta":{"raw":{"variants":["AI agent integrates seven biological scales for whole-body modeling","Full-body AI agent links molecules to organs in one framework","Two AI agents show full-body modeling for metastasis and drugs","Supervisor AI integrates seven scale-specific agents to model body","Multi-scale AI framework predicts drug effects and metastasis across body"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00061,"raw_usage":{"total_tokens":2689,"prompt_tokens":767,"completion_tokens":1922,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":1842}},"tokens_in":511,"tokens_out":1922,"duration_ms":13737,"temperature":1.0,"reasoning_tokens":1842,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:28:19.712556+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a retrospective cohort with long-term follow-up and compute the three-phase metastasis scores from the same multi-scale data; if the scores do not predict metastatic progression beyond standard staging or beyond a molecular-only score, the cross-scale advantage fails. Likewise, a prospective evaluation of the Drug AI Agent's predicted organ-level toxicities against clinical trial outcomes would settle the efficacy/toxicity claim.","supporting_citations":[{"cited_title":"Large language models streamline automated machine learning for clinical studies,","cited_arxiv_id":null,"evidence_quote":"Supplies the prior multi-agent virtual-cell vision that the Full-Body framework extends to whole-organism physiology."},{"cited_title":"A pathologist–AI collaboration framework for enhancing diagnostic accuracies and efficiencies,","cited_arxiv_id":null,"evidence_quote":"Robin is the closed-loop multi-agent baseline; the paper contrasts its scale-bounded loops with full-body cross-scale reasoning."},{"cited_title":"A multiscale approach for biomedical machine learning,","cited_arxiv_id":null,"evidence_quote":"Coscientist provides the LLM-driven autonomous experimentation and protocol-parsing layer for data perception and execution."},{"cited_title":"Spatialagent: An autonomous ai agent for spatial biology,","cited_arxiv_id":null,"evidence_quote":"ReAct supplies the planning policy for task decomposition in the reasoning pipeline."},{"cited_title":"Neutrophil extracellular traps produced during inflammation awaken dormant cancer cells in mice,","cited_arxiv_id":null,"evidence_quote":"The NSCLC single-nucleus dataset (sample N2254) is the worked example where the Initiation Score is actually computed."},{"cited_title":"The brain–heart axis: integrative cooperation of neural, mechanical and biochemical pathways,","cited_arxiv_id":null,"evidence_quote":"Organoids and organs-on-chips are the preclinical platforms that the Drug AI Agent wraps with whole-body constraints."}],"review_version":1}