Pith. sign in

REVIEW 4 major objections 4 minor 44 references

Shared context graphs can expand discovery breadth beyond what a single AI model achieves, making team coordination a scaling axis for science.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 05:49 UTC pith:3RISXYY3

load-bearing objection Plausible architecture and a real field deployment, but the headline breadth claim is built on a confounded comparison and a post-hoc rubric. the 4 major comments →

arxiv 2607.13220 v3 pith:3RISXYY3 submitted 2026-07-14 cs.AI cs.CEcs.HC

Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science

classification cs.AI cs.CEcs.HC
keywords networked intelligenceactive context graphhuman-AI team sciencescientific discoverycontext routingmulti-omicssparse conditional computationteam science
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper sets out to show that the bottleneck in AI-accelerated science is not only model capability but the movement of scientific context across the people and agents who can act on it. It introduces Mycelium, a runtime built on an active context graph—a shared, provenance-bearing graph of observations, hypotheses, findings, and experiment proposals—that routes relevant context to the researcher or agent whose next decision it can inform. In a week-long multi-omics campaign, the networked setup surfaced 25 of 26 mapped scientific artifacts, including 17 experiment-ready ones, versus 17 and 18 for two standalone-agent baselines with the same data. The paper also formalizes when a network is necessary: a single scaled model can subsume a network when all evidence can be merged into one context, but independent corroboration and non-mergeable contexts make a networked architecture irreducible. The significance, if true, is that coordinating distributed expertise—not just scaling models—can change the trajectory of an active investigation and compress a multi-month human-mediated iteration cycle into a sprint.

Core claim

The paper claims that scaling the connections between reasoners, not just the model, expands discovery breadth in open-ended science. Its central evidence is a controlled comparison in which Mycelium, coordinating three domain experts and autonomous investigations through a shared active context graph, produced a mechanistic model of carbon-overflow rewiring in Pseudomonas putida and a grounded experiment plan, while two standalone agents with identical data and maximum reasoning effort produced narrower coverage. The paper further claims that networked intelligence is sparse conditional computation over distributed contexts: a finding should be routed when its expected decision value for th

What carries the argument

The active context graph G=(V,E) stores typed scientific entries—datasets, observations, interpretations, hypotheses, findings, open questions, recommendations, and experiment proposals—connected by lineage and epistemic edges (derived_from, supports, contradicts) that preserve provenance. Routing is scored as a message-passing step across neighboring belief states, but belief states remain isolated so contradictions are preserved rather than silently merged. The routing value Δ(v→j)=V_j(H_j⊕v)−V_j(H_j) formalizes when propagating a claim improves the receiver's expected decision value. This framework makes the system's core operation explicit: decide which cross-context links are worth surf

Load-bearing premise

The load-bearing premise is that the 26-artifact audit is an objective, externally anchored measure of evidence-to-action value; the artifact set was assembled from the campaign outcome and scored by the authors without pre-registration or independent raters, so if the measure partly reflects what Mycelium produced, the breadth advantage is partly expected.

What would settle it

Run the same campaign with a pre-registered artifact rubric fixed before execution, scored by independent raters blind to which system produced each trace; if the breadth gap between networked and standalone execution narrows to parity, the central breadth claim is falsified. Alternatively, exhibit a real non-mergeable context in which a monolithic model given all digitally available data nevertheless fails while the network succeeds, to test the irreducibility claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Networked execution is a complementary scaling axis: it improves problem-space exploration even when localized artifact specificity is comparable to standalone agents.
  • A finding's value is defined by its effect on a downstream decision, so systems can prioritize routing by estimated decision-value gain rather than by textual similarity.
  • Preserving disagreement across isolated belief states—rather than merging them into one joint distribution—is a deliberate mechanism for independent corroboration.
  • Claims routed through the graph retain provenance and grounding, making the team's collaborative lineage physically auditable.
  • The framework outlines concrete boundary conditions: a monolithic model suffices when all evidence can be merged and error correlation is acceptable; a network is required under independent-corroboration or non-mergeable-context conditions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A testable extension: sweep the routing threshold τ or the system's proactivity level across campaigns to measure the precision–recall trade-off of surfaced cross-context links; the paper identifies this as a calibration problem but does not characterize the curve.
  • If the evaluation's artifact set is partly defined post hoc by what the system surfaced, the 25-vs-17 gap may shrink under a pre-registered rubric; this is not tested in the paper.
  • The same architecture could enable cross-institution collaboration where raw data cannot leave a firewall, since only typed, provenanced claims—not underlying datasets—need to be routed; the paper gestures at cross-project fusion but leaves this unexplored.
  • The paper's irreducibility claim depends on how often non-mergeable contexts actually occur in practice; measuring that frequency in real team-science settings would determine whether the formal boundary is a live constraint or a rare edge case.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces Mycelium, a runtime architecture for human-AI team science built around an 'active context graph' (ACG) that stores typed, provenance-aware scientific entries and routes relevant context between human experts and autonomous agents. The authors report a one-week Pseudomonas putida multi-omics campaign in which three domain experts worked asynchronously through chat interfaces while Mycelium captured observations, propagated findings across expert threads, and grounded experiment proposals. The main empirical claim is that this networked execution expands discovery breadth: Table 2 reports that Mycelium surfaced 25 of 26 mapped scientific artifacts (17 experiment-ready) versus 17 (9) and 18 (11) for two single-agent baselines. A Supplementary Note formalizes networked intelligence as sparse conditional computation over contexts, with Claim 1 describing an efficiency regime and Claim 2 asserting an irreducibility regime when independent corroboration or non-mergeable contexts are required.

Significance. If the central claim were established, the paper would make a useful architectural contribution to a timely problem: how to move scientific context among humans, agents, and instruments while preserving provenance. The paper's strengths include a concrete deployed system, a clearly described ontology and routing mechanism, an honest limitations section, and a formal framework that gives the field a shared vocabulary. The authors also make a good-faith effort to release baseline prompts and a campaign dataset. However, the empirical evaluation as presented does not support the strong causal claim that the active context graph, rather than the presence of three human experts, drove the breadth advantage. The evaluation rubric also appears to be constructed from the campaign outcomes and scored by the authors without blinding. The formal irreducibility claim is either definitional or unsupported by evidence. The significance to the human-AI team-science community would be high, but the current evidence is insufficient to justify the paper's central conclusions.

major comments (4)
  1. [§3.3, Table 2, §3.1] The headline breadth comparison confounds two variables: the Mycelium arm includes three human domain experts (User-E, User-L, User-J) plus the ACG runtime, while both baselines are single-agent LLM sessions with no human participants. The 25 vs 17–18 artifact gap can therefore be explained by human expertise participation alone, not by cross-context routing. The text claims 'To isolate the effect of networked intelligence,' but there is no control with the same three experts working independently without the shared graph and without system-initiated pollination. Without this control, the causal claim that the active context graph materially improves team science is unverified. Add a same-experts/no-routing condition and blind scoring of all conditions.
  2. [Supplementary Method S2, Table 3] The 26-artifact global set appears to have been assembled from the campaign outcomes, and Table 3 includes artifacts that only Mycelium surfaced (A11, A19–A22, A25). If artifact membership in A depends on the outputs of the system being measured, the breadth metric is partly defined by the system it evaluates, making the 25/26 advantage expected rather than demonstrated. In addition, scores were assigned by the authors without pre-registration, blinding, inter-rater reliability checks, or an externally anchored evidence-to-action gold standard. Because Table 2 is the only quantitative evidence for the central claim, provide an independent artifact taxonomy defined before scoring, with blind dual annotation, and report inter-rater agreement.
  3. [Supplementary Note, Claim 2] Condition (ii), 'non-mergeable contexts,' makes a network necessary by definition when contexts literally cannot be merged, but the note provides no measurement of how often this condition arose in the campaign or in realistic team science. Condition (i) assumes independent error profiles across experts, yet human experts often share training, literature, and databases, so the independence assumption is not argued or tested. The note itself concedes that Claim 1 does not establish a rigorous separation and that the example 'illustrates the proposed mechanism, rather than establishing a universal separation.' As written, the formal section does not substantiate the abstract's claim that the framework establishes when a networked approach is essential. Either reframe these conditions as sufficiency statements or provide empirical tests of their incidence.
  4. [§3.3, Box S1] The comparison is not matched on computational budget: the two standalone baselines were configured for 'maximum reasoning effort (xhigh)' while Mycelium 'was run with the default high setting for inference cost and consistency.' Since the measured outcome is discovery breadth, this asymmetry could explain part of the Table 2 gap. The paper should either use identical inference settings across conditions or explicitly justify why the higher baseline budget does not undermine the conclusion that the difference is architectural rather than a matter of reasoning effort.
minor comments (4)
  1. [Throughout] Several typos and editorials: 'the scientis or agent' (§3), 'succeeds' for 'succeeds' in Claim 2, 'also also' (§3.2), 'over. over.' (§4.2), 'shows an trace' (Fig. S1 caption).
  2. [Supplementary Method S1] Box S1 says the harness was 'released' and schemas were 'omitted', but no repository URL or artifact DOI is provided. To make Table 2 reproducible, release the full prompts, code, execution traces, and the complete scoring matrix.
  3. [Table 2] The 'average specificity' is computed only over surfaced artifacts. Because Mycelium surfaced almost the entire set while baselines surfaced fewer, this conditional mean is not directly comparable across conditions. Report precision-recall-style tradeoffs or also report total score separately for the shared subset of artifacts.
  4. [Figure 4] The narrative relies heavily on Figure 4, but the figure is not visible in the submitted text and the caption does not define all arrows and node types. Please provide a higher-resolution figure with a full legend.

Circularity Check

2 steps flagged

Breadth metric and Claim 2(ii) are partially self-definitional; the 26-artifact set and non-mergeability premise encode the outcome by construction.

specific steps
  1. self definitional [Section 3.3 'Impact of network scaling'; Supplementary Method S2 (breadth-weighted score); Table 3 (A11, A19–A22)]
    "We evaluated these executions by extracting a global set of 26 unique “scientific artifacts”, defined as any traceable analytical claim, mechanistic hypothesis, or actionable decision rule capable of directing the execution of experimental protocols. ... To formalize our evaluation of the evidence-to-action pipeline, we define the global set of 26 mapped scientific artifacts as A."

    The evaluation set A is extracted from the runs being compared, and Table 3 scores several artifacts only for Mycelium (A11: 4 0 0; A19–A22: 3 0 0). Breadth is then the mean over A with absences scored zero. Including Mycelium-only outputs in A therefore guarantees part of Mycelium’s 25/26 vs 17/18 margin by construction. The breadth advantage is a self-referential audit of a set that contains the system’s own outputs, rather than an independent external criterion.

  2. self definitional [Supplementary Note, Claim 2 [Irreducibility Regime], condition (ii)]
    "Claim 2 [Irreducibility Regime]: A factored network suceeds whenever either of two conditions holds: ... (ii) Non-mergeable contexts: Real-world team science frequently involves evidence that cannot physically or organizationally be placed into a single context window. ... Because this local evidence is held in the human expert or behind a physical firewall, dense model integration is impossible."

    Condition (ii) is the conclusion restated as a premise: ‘non-mergeable’ is defined as ‘cannot physically or organizationally be placed into a single context window’, so concluding that a single-window monolithic model cannot integrate such contexts is true by definition. The note supplies no independent test for when contexts are non-mergeable nor any estimate of how often this condition arises, so the formal ‘network is essential’ claim reduces to its definitional input.

full rationale

Most of the paper’s formal machinery is conditional and not circular. The routing-value definition Δ(v→j)=Vj(Hj⊕v)−Vj(Hj) is an analytic setup, and the efficiency regime explicitly disclaims a rigorous separation (“This does not establish a rigorous separation between networked and monolithic intelligence”). No load-bearing self-citation chain is present: the author-overlapping references are contextual and no uniqueness theorem is imported. The missing no-routing expert control is a serious confound (breadth could come from the three human experts rather than from the graph), but confounding is an evaluation-validity issue, not circularity, so it is not counted in the score. The circularity score is raised by two specific reductions: (1) the 26-artifact global set in Section 3.3/S2 is extracted from the executions and includes Mycelium-only artifacts, so the breadth score is partly self-referential; and (2) Claim 2(ii) makes network essentiality equivalent to the definition of non-mergeable contexts. These are partial rather than total reductions: Mycelium still had to surface 25 artifacts, and the irreducibility note is self-identified as illustrative, so the result is not fully forced. Score 6 reflects central claims with partial construction-by-definition.

Axiom & Free-Parameter Ledger

1 free parameters · 4 axioms · 0 invented entities

The paper introduces a software system (Mycelium) and a conceptual framing, not a new physical or mathematical entity. No invented_entities are listed. The ledger's central items are the uncalibrated routing threshold and the domain assumptions behind the irreducibility argument.

free parameters (1)
  • Routing utility threshold τ = not specified
    In the Supplementary Note, a candidate edge routes when Δ(v→j) > τ, with τ described as 'a system-defined threshold'; no value, calibration procedure, or sensitivity analysis is given.
axioms (4)
  • domain assumption A monolithic model's errors across domains are correlated, so independent corroboration requires a network of distinct experts/agents.
    Used in Claim 2(i) to argue network irreducibility; asserted rather than demonstrated. Two models with different data or priors could in principle provide independent error profiles.
  • domain assumption Real scientific work involves evidence that cannot physically or organizationally be placed into a single context window.
    Claim 2(ii); plausible and true in some settings, but it makes the network necessary by definition whenever it holds, and the paper does not establish how common this condition is.
  • domain assumption Routing value Δ(v→j) can be estimated in practice from belief updates and expected utilities.
    The formal framework assumes an implementation-specific model of belief updates and utility; no estimator, algorithm, or measurement is presented.
  • ad hoc to paper The 26-artifact rubric validly operationalizes the evidence-to-action pipeline.
    Supplementary Method S2 defines metrics over an artifact set that appears to be constructed from the campaign itself; validity is assumed, not independently established.

pith-pipeline@v1.3.0-alltime-deepseek · 14427 in / 16715 out tokens · 165610 ms · 2026-08-02T05:49:59.108948+00:00 · methodology

0 comments
read the original abstract

Most AI-for-science systems focus on scaling a single reasoning process by using better models, larger context windows, long-horizon agentic execution, or digital co-scientists working with one principal user. However, challenging scientific problems are rarely solved by one reasoner alone. They are solved by teams whose members carry different priors, experimental background, tacit knowledge, and domain-trained intuitions. The open problem is therefore not only how to scale models, but how to develop "networked intelligence", scaling the connections between humans and AI systems so that a result or hypothesis produced in one context reaches another person, agent, instrument or robot that can act on it. We introduce Mycelium, an active shared workspace that automatically connects researchers and AI agents. As human users and agents work, the system captures important observations and hypotheses, tracks how they relate to the team's evolving knowledge model, and routes them to the person or agent whose next decision they can inform. We evaluate Mycelium through a real-world scientific discovery use case: a biological multi-omics campaign where shared context turned a local analytical finding into a cross-expert mechanistic constraint and ultimately into an experimental design. Finally, we describe networked intelligence as sparse conditional computation over distributed scientific contexts. This framework establishes when a scaled standalone agent is sufficient, and when isolated data and specialized expertise make a networked approach essential.

Figures

Figures reproduced from arXiv: 2607.13220 by Aivett Bilbao, Alex Beliaev, Chris Oehmen, Erin Bredeweg, Jason McDermott, Jaydeep P. Bardhan, Jeffrey J. Czajka, Josh Elmore, Katherine Wolf, Kelly Stratton, Kristin Burnum Johnson, Kylee Tate, Lummy M. O. Monteiro, Paul Piehowski, Robert Rallo, Scott Baker, Sutanay Choudhury, Yuqian Gao.

Figure 1
Figure 1. Figure 1: The Mycelium runtime architecture. The system coordinates distributed scientific workflows across four distinct domains: Source domains and actors (Left): Researchers and AI agents operate within shared workspaces, reading and writing typed entries to the shared graph. Ac￾tive context graph (Center): The core routing layer managing persistent project state. It maps provenance-aware relations (e.g., generat… view at source ↗
Figure 2
Figure 2. Figure 2: Participant-facing runtime. Researchers interact with the Mycelium network using a standard, familiar AI chat interface. Rather than functioning as an isolated chatbot, MCP connects this chat window directly to the team’s shared active context graph. This allows a user to query the entire project’s history, dispatch complex autonomous workflows, and review results, all without leaving a simple conversation… view at source ↗
Figure 3
Figure 3. Figure 3: Autonomous workflow execution. The Mycelium runtime leverages dynamic code generation to construct multi-step analytical pipelines on the fly for both interactive and autonomous analysis. The figure above shows illustrative computational dataflow graphs. The execution engine supports automated fault tolerance, allowing the system to recover from runtime errors and suc￾cessfully output formalized findings t… view at source ↗
Figure 4
Figure 4. Figure 4: Context routing enables emergent collaborative discovery. The campaign is partitioned into three expert threads (rows): regulatory reasoning (User-E), proteomic/pathway analysis (User-L), and phenotype-guided design (User-J). Solid arrows represent intra-thread rea￾soning; dashed arrows represent context routed by Mycelium across participants. 3.2 Enabling team coordination via shared context propagation T… view at source ↗
Figure 5
Figure 5. Figure 5: Graph evolution reveals the diversity of scientific intent. Cumulative snapshots show isolated analyses converging into a unified model, with colors marking intents from data quality control (gray) and gluconate-overflow mechanisms (orange) to experiment design (yellow). 9 [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

44 extracted references · 1 canonical work pages

  1. [1]

    Washington, DC: The National Academies Press, 2015

    National Research Council,Enhancing the Effectiveness of Team Science. Washington, DC: The National Academies Press, 2015. N. J. Cooke and M. L. Hilton, Eds.; doi:10.17226/19007

  2. [2]

    Highly accurate protein structure prediction with alphafold,

    J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko,et al., “Highly accurate protein structure prediction with alphafold,”nature, vol. 596, no. 7873, pp. 583–589, 2021

  3. [3]

    Autonomous chemical research with large language models,

    D. A. Boiko, R. MacKnight, B. Kline, and G. Gomes, “Autonomous chemical research with large language models,”Nature, vol. 624, pp. 570–578, 2023

  4. [4]

    Chemreasoner: Heuristic search over a large language model’s knowledge space using quantum-chemical feedback,

    H. W. Sprueill, C. Edwards, K. Agarwal, M. V. Olarte, U. Sanyal, C. Johnston, H. Liu, H. Ji, and S. Choudhury, “Chemreasoner: Heuristic search over a large language model’s knowledge space using quantum-chemical feedback,” inProceedings of the 41st International Conference on Machine Learning, 2024

  5. [5]

    Augmenting large language models with chemistry tools,

    A. M. Bran, S. Cox, O. Schilter, C. Baldassari, A. D. White, and P. Schwaller, “Augmenting large language models with chemistry tools,”Nature Machine Intelligence, vol. 6, pp. 525–535, 2024

  6. [6]

    Using a gpt-5-driven autonomous lab to optimize the cost and titer of cell-free protein synthesis,

    A. A. Smith, E. L. Wong, R. C. Donovan, B. A. Chapman, R. Harry, P. Tirandazi, P. Kanigowska, E. A. Gendreau, R. H. Dahl, M. Jastrzebski,et al., “Using a gpt-5-driven autonomous lab to optimize the cost and titer of cell-free protein synthesis,”bioRxiv, pp. 2026– 02, 2026

  7. [7]

    Report for the doe office of science workshop on envisioning frontiers in ai and computing for biological research,

    D. Ushizima, C. Henry, P. Balaprakash, A. Biswas, A. Hoarfrost, K. Hofmockel, N. Kumar, A. Ramanathan, T. Northen, M. Lentz,et al., “Report for the doe office of science workshop on envisioning frontiers in ai and computing for biological research,” tech. rep., US Department of Energy (USDOE), Washington, DC (United States). Office of ..., 2026

  8. [8]

    Energy department launches ‘Genesis Mission’ to transform American science and innovation

    U.S. Department of Energy, “Energy department launches ‘Genesis Mission’ to transform American science and innovation.”https://www.energy.gov/articles/ energy-department-launches-genesis-mission-transform-american-science-and-innovation, 2025

  9. [9]

    Integrated research infrastructure architecture blueprint activity (final report 2023),

    W. L. Miller, D. Bard, A. Boehnlein, K. Fagnan, C. Guok, E. Lançon, S. J. Ramprakash, M. Shankar, N. Schwarz, and B. L. Brown, “Integrated research infrastructure architecture blueprint activity (final report 2023),” tech. rep., US Department of Energy (USDOE), Wash- ington, DC (United States). Office of ..., 2023

  10. [10]

    AutoGen: Enabling next-gen LLM applications via multi-agent conversation

    Q. Wuet al., “AutoGen: Enabling next-gen LLM applications via multi-agent conversation.” arXiv:2308.08155, 2023

  11. [11]

    Multi-agent collaboration mechanisms: A survey of LLMs

    K.-T. Tranet al., “Multi-agent collaboration mechanisms: A survey of LLMs.” arXiv:2501.06322, 2025

  12. [12]

    The virtual lab of AI agents designs new SARS-CoV-2 nanobodies,

    K. Swanson, W. Wu, N. L. Bulaong, J. E. Pak, and J. Zou, “The virtual lab of AI agents designs new SARS-CoV-2 nanobodies,”Nature, vol. 646, pp. 716–723, 2025

  13. [13]

    Empowering biomedical discovery with AI agents,

    S. Gao, A. Fang, Y. Huang,et al., “Empowering biomedical discovery with AI agents,”Cell, vol. 187, no. 22, pp. 6125–6151, 2024. 20

  14. [14]

    Agentic AI for scientific discovery: A survey of progress, challenges, and future directions,

    M. Gridach, J. Nanavati, K. Zine El Abidine, L. Mendes, and C. Mack, “Agentic AI for scientific discovery: A survey of progress, challenges, and future directions,”arXiv preprint arXiv:2503.08979, 2025

  15. [15]

    Model context protocol

    Anthropic, “Model context protocol.”https://modelcontextprotocol.io, 2024. Accessed 2026

  16. [16]

    Agent-to-agent (A2A) protocol

    Google, “Agent-to-agent (A2A) protocol.”https://a2aprotocol.ai, 2025. Accessed 2026

  17. [17]

    A survey of agent interoperability protocols: MCP, ACP, A2A, and ANP

    A. Ehtesham, A. Singh, G. K. Gupta, and S. Kumar, “A survey of agent interoperability protocols: MCP, ACP, A2A, and ANP.” arXiv:2505.02279, 2025

  18. [18]

    Mem0: Building production-ready AI agents with scalable long-term memory

    P. Chhikara, D. Khant, S. Aryan, T. Singh, and D. Yadav, “Mem0: Building production-ready AI agents with scalable long-term memory.” arXiv:2504.19413, 2025

  19. [19]

    UMP: A transport-neutral memory protocol for AI agents

    Universal Memory Protocol, “UMP: A transport-neutral memory protocol for AI agents.” https://universalmemoryprotocol.io, 2026. Accessed 2026

  20. [20]

    Memory in the age of AI agents

    Y. Huet al., “Memory in the age of AI agents.” arXiv:2512.13564, 2025

  21. [21]

    Orchestrated Platform for Autonomous Laboratories (OPAL)

    U.S. Department of Energy, “Orchestrated Platform for Autonomous Laboratories (OPAL).” https://opal-doe.org/, 2026. Accessed: 2026-06-24

  22. [22]

    Fusion, propagation, and structuring in belief networks,

    J. Pearl, “Fusion, propagation, and structuring in belief networks,” inProbabilistic and Causal Inference: The Works of Judea Pearl, pp. 139–188, 2022

  23. [23]

    Neural message passing for quantum chemistry,

    J. Gilmer, S. S. Schoenholz, P. F. Riley, O. Vinyals, and G. E. Dahl, “Neural message passing for quantum chemistry,” inProceedings of the 34th International Conference on Machine Learning, 2017

  24. [24]

    Engineering glucose metabolism for enhanced muconic acid production in pseudomonas putida kt2440,

    G. J. Bentley, N. Narayanan, R. K. Jha, D. Salvachúa, J. R. Elmore, G. L. Peabody, B. A. Black, K. Ramirez, A. De Capite, W. E. Michener,et al., “Engineering glucose metabolism for enhanced muconic acid production in pseudomonas putida kt2440,”Metabolic engineering, vol. 59, pp. 64–75, 2020

  25. [25]

    Pseudomonas putida as a functional chassis for industrial biocatalysis: from native biochemistry to trans-metabolism,

    P. I. Nikel and V. de Lorenzo, “Pseudomonas putida as a functional chassis for industrial biocatalysis: from native biochemistry to trans-metabolism,”Metabolic engineering, vol. 50, pp. 142–155, 2018

  26. [26]

    Riding the sulfur cycle–metabolism of sulfonates and sulfate esters in gram- negative bacteria,

    M. A. Kertesz, “Riding the sulfur cycle–metabolism of sulfonates and sulfate esters in gram- negative bacteria,”FEMS microbiology reviews, vol. 24, no. 2, pp. 135–175, 2000

  27. [27]

    Prompting Claude Opus 4.8

    Anthropic, “Prompting Claude Opus 4.8.”https://platform.claude.com/docs/en/ build-with-claude/prompt-engineering/prompting-claude-opus-4-8, 2026

  28. [28]

    Scienceagentbench: Toward rigorous assessment of language agents for data-driven scientific discovery,

    Z. Chen, S. Chen, Y. Ning, Q. Zhang, B. Wang, B. Yu, Y. Li, Z. Liao, C. Wei, Z. Lu,et al., “Scienceagentbench: Toward rigorous assessment of language agents for data-driven scientific discovery,” inInternational Conference on Learning Representations, vol. 2025, pp. 96934– 96990, 2025

  29. [29]

    Bixbench: a comprehensive benchmark for llm-based agents in computational biology,

    L. Mitchener, J. M. Laurent, A. Andonian, B. Tenmann, S. Narayanan, G. P. Wellawatte, A. White, L. Sani, and S. G. Rodriques, “Bixbench: a comprehensive benchmark for llm-based agents in computational biology,”arXiv preprint arXiv:2503.00096, 2025. 21

  30. [30]

    Graph of trace: Visualizing execution traces of scientific agent,

    T. Gao, H. Li, J. Li, T. Zhao, R. Shi, W. Wang, Z. Wu, and L. Mi, “Graph of trace: Visualizing execution traces of scientific agent,”arXiv preprint arXiv:2606.15116, 2026

  31. [31]

    From agent traces to trust: Evidence tracing and execution provenance in llm agents,

    Y. Wang, J. Zhang, T. Cai, Z. Liu, Q. Sun, Z. Sun, Z. Wu, M. Zhang, and Y. Zhu, “From agent traces to trust: Evidence tracing and execution provenance in llm agents,”arXiv preprint arXiv:2606.04990, 2026

  32. [32]

    Swe- bench: Can language models resolve real-world github issues?,

    C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan, “Swe- bench: Can language models resolve real-world github issues?,” inInternational Conference on Learning Representations, vol. 2024, pp. 54107–54157, 2024

  33. [33]

    Proagent: building proactive cooperative agents with large language models,

    C. Zhang, K. Yang, S. Hu, Z. Wang, G. Li, Y. Sun, C. Zhang, Z. Zhang, A. Liu, S.-C. Zhu,et al., “Proagent: building proactive cooperative agents with large language models,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 38, pp. 17591–17599, 2024

  34. [34]

    Principles of mixed-initiative user interfaces,

    E. Horvitz, “Principles of mixed-initiative user interfaces,” inProceedings of the SIGCHI con- ference on Human Factors in Computing Systems, pp. 159–166, 1999

  35. [35]

    The FAIR guiding principles for scientific data management and stewardship,

    M. D. Wilkinsonet al., “The FAIR guiding principles for scientific data management and stewardship,”Scientific Data, vol. 3, p. 160018, 2016. doi:10.1038/sdata.2016.18

  36. [36]

    Adaptive mixtures of local experts,

    R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton, “Adaptive mixtures of local experts,”Neural computation, vol. 3, no. 1, pp. 79–87, 1991

  37. [37]

    Outra- geously large neural networks: The sparsely-gated mixture-of-experts layer,

    N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outra- geously large neural networks: The sparsely-gated mixture-of-experts layer,”arXiv preprint arXiv:1701.06538, 2017

  38. [38]

    Packaging research artefacts with RO-Crate,

    S. Soiland-Reyeset al., “Packaging research artefacts with RO-Crate,”Data Science, vol. 5, no. 2, pp. 97–138, 2022. doi:10.3233/DS-210053

  39. [39]

    LLNL pushes frontier of fusion target design with AI(multi-agentdesignassistant, NNSAASC)

    Lawrence Livermore National Laboratory, “LLNL pushes frontier of fusion target design with AI(multi-agentdesignassistant, NNSAASC).”https://www.llnl.gov/article/53216, 2025. Accessed 2026

  40. [40]

    Osprey: Production-ready agentic AI for safety-critical control systems,

    T. Hellert, J. Montenegro, and A. Sulc, “Osprey: Production-ready agentic AI for safety-critical control systems,”APL Machine Learning, vol. 4, no. 1, p. 016103, 2026. doi:10.1063/5.0306302

  41. [41]

    A multi-agent system for automating scientific discovery,

    A. E. Ghareeb, B. Chang, L. Mitchener, A. Yiu, C. J. Szostkiewicz, D. Shved, G. J. Gyimesi, J. M. Laurent, S. M. Wright, M. T. Razzak,et al., “A multi-agent system for automating scientific discovery,”Nature, pp. 1–3, 2026

  42. [42]

    Empowering scientific workflows with federated agents (Academy)

    A. Kamatar, J. G. Pauloski, Y. Babuji, R. Chard, M. Sakarvadia, K. Chard, and I. Foster, “Empowering scientific workflows with federated agents (Academy).” arXiv:2505.05428, 2025

  43. [43]

    Uniprot: the universal protein knowledgebase in 2025,

    “Uniprot: the universal protein knowledgebase in 2025,”Nucleic acids research, vol. 53, no. D1, pp. D609–D617, 2025

  44. [44]

    Kegg: bio- logical systems database as a model of the real world,

    M. Kanehisa, M. Furumichi, Y. Sato, Y. Matsuura, and M. Ishiguro-Watanabe, “Kegg: bio- logical systems database as a model of the real world,”Nucleic acids research, vol. 53, no. D1, pp. D672–D677, 2025. 22