Pith. sign in

REVIEW 2 major objections 4 minor 1 cited by

Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research

T0 review · 2 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This perspective argues that foundation models—versatile AI agents and robotic foundation models—are the key technology for flexible laboratory automation, able to adapt to diverse samples, devices, and data formats without full…

desk verdict A useful, well-hedged perspective on foundation models for lab automation that earns its place as an overview; the central 'key technology' claim is a considered prediction, not a proven result, and the paper mostly says so itself. read the letter →

arxiv 2506.12312 v1 pith:QZCZZLD6 submitted 2025-06-14 cs.RO cs.CLphysics.chem-ph

classification cs.ROcs.CLphysics.chem-ph
keywords laboratoryautomationfoundationmodelslargelanguageroboticsmaterialsscienceAIagentsself-drivinglaboratories
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review argues that the next generation of laboratory automation in materials research will be driven not by stricter standardization of hardware but by foundation models: large language models, multimodal AI, and robotic foundation models. These models play two roles—cognitive work such as planning and data analysis, and physical work such as controlling instruments and robots. The paper identifies highly versatile AI agents and robotic foundation models as the key technology for adapting to varied sample types, devices, communication standards, and databases, and lays out a roadmap based on these dual roles. It acknowledges serious limits, especially the transformer architecture's data hunger and the gap between web-trained knowledge and real lab practice, and argues these can be addressed through synthetic data, benchmarks, and external cognitive systems. If the roadmap holds, open-ended experiments could be automated without waiting for industry-wide hardware standardization.

What carries the argument

The machinery is the foundation model itself, defined as a large pretrained AI model, such as a large language model, multimodal model, or robotic foundation model, that can be adapted across tasks. The paper divides its roles into cognitive (experiment planning, data analysis, report writing) and physical (device control, sensing, orchestrating instruments), and uses a two-axis scheme—'brain' autonomy versus 'body' automation—to place existing systems and future steps on a roadmap. The argument's internal mechanism is that general-purpose pretrained models transfer knowledge across tasks via zero-shot learning, replacing task-specific, rule-based control, while the limiting mechanism is the transformer scaling law, which requires far more data than a single experiment or lab note provides.

What would settle it

Run the current best LLM-based lab agents on a held-out set of real laboratory protocols described only in single experimental notes, with no additional training, and compare their success rate against simple rule-based automation. If the agents do not clearly outperform on novel, varied tasks, the central claim that foundation models are the key to flexible open-ended laboratory automation would be contradicted.

Watch

Extended reading notes

Core claim

The central claim is that foundation models, rather than modular hardware standards, are the pivotal enabling technology for flexible laboratory automation. The paper proposes that AI agents and robot foundation models can act as a 'brain' that plans experiments and interprets data and a 'body' that physically operates equipment, allowing robots to work in open-ended lab environments without uniform modularization of instruments. Even before humanoid robots mature, foundation models can automate the standardization of data from analytical instruments and coordinate specialized modules. The paper presents a two-axis roadmap—autonomy of decision-making versus automation of operational mechanisms—and concludes that reaching fully autonomous laboratories requires openly licensed datasets, quantitative benchmarks for lab tasks, and human-AI integration, including a 'human actuator' stage in which people carry out detailed AI instructions.

Load-bearing premise

The roadmap assumes that the data inefficiency and limited scientific knowledge of today's transformer-based models can be overcome with synthetic data, specialized external cognitive systems, and better benchmarks; if the data bottleneck is a fundamental limit rather than a fixable engineering problem, fully autonomous laboratories would not arrive no matter how well the robots improve.

Editorial extensions

If this is right

  • Laboratory automation would no longer depend on universal hardware and communication standards; robots with foundation models could handle varied instruments directly.
  • AI agents could turn experimental plans into working control programs and fine-tune them, reducing the cost and rigidity of bespoke automation.
  • Robots trained partly on cooking and human demonstration videos could learn lab skills such as pipetting, powder handling, and glassware manipulation, and combine them into new protocols.
  • A 'human actuator' stage—humans following detailed AI instructions—could produce standardized, multimodal process data while AI tools remain ahead of dexterity limits.
  • The main competitive bottleneck would shift from hardware to data: whoever curates, shares, and benchmarks lab datasets would determine how quickly autonomous labs advance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If cheap AI agents can draft full research pipelines, the economics of early-stage materials research may shift toward many parallel, low-cost automated attempts rather than a few carefully chosen human experiments.
  • Cooking is a natural test bed: a robot that can reliably execute a novel recipe from video and language instruction would be strong evidence that the same approach can run laboratory experiments.
  • The 'human actuator' model, if adopted widely, could generate large corpora of expert-labelled process data, but also raises the risk that tacit skills are encoded incompletely or with systematic bias.
  • One testable extension of the roadmap is that performance on a standardized open benchmark of lab-perception and manipulation tasks should predict a system's ability to run unseen experimental protocols; if it does not, the foundation-model-centered strategy would need revision.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This perspective argues that foundation models—large language models, multimodal models, and robotic foundation models—are the key enabling technology for flexible laboratory automation in materials research, allowing open-ended experiments without full hardware standardization. The paper categorizes the roles of foundation models into cognitive (planning, data analysis, writing) and physical (hardware control, manipulation) functions, reviews current specialized automation and modular standardization, surveys recent progress in AI agents and robot foundation models, and discusses adjacent progress in cooking robotics as a proxy for lab manipulation. It concludes with a roadmap that positions AI agents and robotic foundation models at the center of future laboratory automation while acknowledging persistent challenges in data efficiency, tacit knowledge, precision manipulation, and safety.

Significance. If the central thesis is correct, the paper outlines a plausible path toward a paradigm shift in laboratory automation from rigid, standardized systems to flexible, general-purpose robots guided by foundation models. The paper is a valuable synthesis for a broad community: it gathers scattered literature from robotics, AI, chemistry, and materials science, and it is notably honest about the limits of current evidence—for example, the glassware misidentification is explicitly presented as a single anecdote, and the data-efficiency limitation of transformers is acknowledged with reference to scaling-law studies. The roadmap is falsifiable in principle, and the explicit calls for laboratory-specific benchmarks, datasets, and human-AI integration are concrete and actionable. The main weakness is that the central claim rests on extrapolation from non-laboratory domains, which the authors themselves partially concede.

major comments (2)
  1. [Section 3.1, Section 3.4] Section 3.1 asserts that 'the key technology for flexible adaptation to various sample types, devices, communication standards, and databases lies in highly versatile AI agents and foundation models for robots,' yet the supporting evidence in Sections 3.2-3.4 is drawn almost entirely from programming, web-based agents, cooking, and tabletop manipulation. The manuscript itself notes in Section 3.4 that ChatGPT misidentifies a common piece of laboratory glassware and that lab-specific tacit knowledge (size specifications, nuanced techniques, undocumented procedures) is largely missing from web-scale training data. Because this assertion is the central thesis, it should be explicitly framed as a research hypothesis rather than an established fact, with a falsifiable validation path (e.g., quantitative benchmarks on lab-specific perception, planning, and manipulation tasks) and a discussion of what evidence would strengthen or weaken the claim. Without this framing, the roadmap risks being read as an extrapolation from adjacent domains rather than a grounded assessment.
  2. [Section 3.3] The proposal to overcome transformer data inefficiency using synthetic data augmentation is not supported by a discussion of a critical circularity risk: if synthetic data are generated by the same foundation models that lack the target tacit knowledge, the augmentation may propagate or amplify existing errors. The paper states that 'a naive approach of merely training models on scientific data may not effectively capture user-expected information' and then suggests expanding methodologies with synthetic data, but it does not specify a source of synthetic data that is independent of the deficient model (e.g., physics-based simulators with ground-truth state or expert-curated data). Adding a concrete discussion of how the synthetic data would be generated, validated, and shown to improve real-world lab performance would make the roadmap more credible.
minor comments (4)
  1. [References] Reference [52] cites arXiv:2501.05789, which is the same identifier as reference [49]; the intended arXiv number for the survey on large language model based agents should be corrected.
  2. [Reference [85]] The DOI for reference [85] is malformed ('10.1146/((please'), likely a placeholder artifact; it should be completed or removed.
  3. [Figure 5 caption] The caption says 'using an Unrealistic engine,' which should read 'using the Unreal Engine.'
  4. [Section 4.2] In the description of RoboCat, 'from as few as around 100 demonstrations' is slightly vague; a precise number or range would be clearer, though this is a minor issue.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a perspective/review with no derivations, fitted parameters, or predictions that reduce to its own inputs.

full rationale

This is a perspective and literature-review paper, not a derivation. There are no equations, no fitted parameters, and no quantitative predictions that could reduce to the paper's own inputs by construction. The central claim, that foundation models are the key technology for flexible laboratory automation, is presented as an interpretive position supported by a broad set of external references (e.g., RT-1, RoboCat, AI Scientist, scaling-law analyses, and robotics reviews) rather than by any computation performed in this paper. The few self-citations by the authors (refs. 20, 69, 84, 97) are used as supporting examples or contextual background: for instance, ref. 69 is cited to illustrate the general point that naive training on scientific data may not capture user-expected information, alongside an independent scaling-law reference (ref. 68). None of these self-citations is invoked as a uniqueness theorem, a forced choice, or the exclusive evidence for a load-bearing premise. The paper explicitly identifies unresolved limitations, such as data inefficiency and the absence of lab-specific tacit knowledge in web-scale datasets, which further indicates that the authors are not disguising assumptions as results. Whether the roadmap's extrapolation from cooking and tabletop robotics to materials laboratories is sound is a scientific-correctness question, not a circularity question. The analysis therefore finds no circular step and assigns a score of 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper is a review, so it does not introduce new parameters or entities. Its roadmap depends on domain assumptions about the future capabilities of AI models and the solvability of data bottlenecks.

assumptions (3)
  • domain assumption Foundation models (LLMs) will continue to improve and can be adapted to laboratory domains.
    The entire review relies on the expectation that general-purpose AI models can be trained or prompted to handle chemistry/materials lab tasks. This is not proven and is a central theme of Sections 3.2 and 3.4.
  • domain assumption The main bottleneck is data and benchmarks, not fundamental hardware or algorithmic limitations.
    In Section 5, the conclusion states 'conditions for effective data collection and usage in real-world laboratory automation remain insufficient, creating a bottleneck.' This assumes that more and better data will yield substantial progress.
  • domain assumption Laboratory tasks are similar enough to cooking and other domains for transfer learning to work.
    Section 4.3 draws parallels between cooking and lab experiments, using this to justify approaches from robotics research. The similarity is plausible but not quantitatively established.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research." pith.science (2026). https://pith.science/paper/QZCZZLD6

@misc{pith2026250612312,
  author       = {Pith},
  title        = {Pith review of: Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QZCZZLD6}},
  note         = {Machine review of arXiv:2506.12312}
}
read the original abstract

This review explores the potential of foundation models to advance laboratory automation in the materials and chemical sciences. It emphasizes the dual roles of these models: cognitive functions for experimental planning and data analysis, and physical functions for hardware operations. While traditional laboratory automation has relied heavily on specialized, rigid systems, foundation models offer adaptability through their general-purpose intelligence and multimodal capabilities. Recent advancements have demonstrated the feasibility of using large language models (LLMs) and multimodal robotic systems to handle complex and dynamic laboratory tasks. However, significant challenges remain, including precision manipulation of hardware, integration of multimodal data, and ensuring operational safety. This paper outlines a roadmap highlighting future directions, advocating for close interdisciplinary collaboration, benchmark establishment, and strategic human-AI integration to realize fully autonomous experimental laboratories.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AI4Research: A Survey of Artificial Intelligence for Scientific Research

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A survey that organizes AI-for-research work into five tasks, comprehension, survey, discovery, writing, and peer review, and compiles associated tools and benchmarks.

Reference graph

Works this paper leans on

115 extracted references · 33 canonical work pages · cited by 1 Pith paper

  1. [52]

    Z. Xi, W. Chen, X. Guo, et al. The Rise and Potential of Large Language Model Based Agents: A Survey. 2023;arXiv:2501.05789

  2. [49]

    H. Bar, R. Hochstrasser, B. Papenfub. SiLA: Basic standards for rapid integration in laboratory automation. J Lab Autom. 2012;17(2):86. doi: 10.1177/2211068211424550

  3. [1]

    G. Tom, S. P. Schmid, S. G. Baird, et al. Self -Driving Laboratories for Chemistry and Materials Science. Chem. Rev. 2024;124(16):9633. doi: 10.1021/acs.chemrev.4c00055

  4. [2]

    M. C. Ramos, C. J. Collison, A. D. White. A review of large language models and autonomous agents in chemistry. Chem. Sci. 2025;16(6):2514. doi: 10.1039/d4sc03921a

  5. [3]

    Yoshikawa, Y

    N. Yoshikawa, Y. Asano, D. N. Futaba, et al. Self-driving laboratories in Japan. Digital Discovery

  6. [4]

    N. J. Szymanski, B. Rendy, Y. Fei, et al. An autonomous laboratory for the accelerated synthesis of novel materials. Nature. 2023;624(7990):86. doi: 10.1038/s41586-023-06734-w

  7. [5]

    Williams, E

    K. Williams, E. Bilsland, A. Sparkes, et al. Cheaper faster drug development validated by the repositioning of drugs against neglected tropical diseases. J R Soc Interface. 2015;12(104):20141289. doi: 10.1098/rsif.2014.1289

  8. [6]

    OpenAI, GPT-4 Technical Report, https://cdn.openai.com/papers/gpt-4.pdf, 2023

Show all 115 references
  1. [7]

    K. M. Jablonka, P. Schwaller, A. Ortega -Guerrero, et al. Leveraging large language models for predictive chemistry. Nature Machine Intelligence. 2024;6(2):161. doi: 10.1038/s42256 -023-00788-1

  2. [8]

    Gomez-Bombarelli, J

    R. Gomez-Bombarelli, J. N. Wei, D. Duvenaud, et al. Automatic Chemical Design Using a Data - Driven Continuous Representation of Molecules. ACS Cent. Sci. 2018;4(2):268. doi: 10.1021/acscentsci.7b00572

  3. [10]

    Z. Zhou, X. Li, R. N. Zare. Optimizing Chemical Reactions with Deep Reinforcement Learning. ACS Cent. Sci. 2017;3(12):1337. doi: 10.1021/acscentsci.7b00492

  4. [12]

    Y. Su, X. Wang, Y. Ye, et al. Automation and machine learning augmented by large language models in a catalysis study. Chem. Sci. 2024;15(31):12200. doi: 10.1039/d3sc07012c

  5. [13]

    D. A. Boiko, R. MacKnight, B. Kline, et al. Autonomous chemical research with large language models. Nature. 2023;624(7992):570. doi: 10.1038/s41586-023-06792-0

  6. [14]

    M. Ahn, A. Brohan, N. Brown, et al. Do As I Can, Not As I Say: Grounding Language in Robotic Affordances. 2022;arXiv:2204.0169

  7. [15]

    Brohan, N

    A. Brohan, N. Brown, J. Carbajal, et al. RT -1: Robotics Transformer for Real -World Control at Scale. 2022;arXiv:2212.06817

  8. [16]

    Zhang, H

    P. Zhang, H. Zhang, H. Xu, et al. Scaling Laws in Scientific Discovery with AI and Robot Scientists. 2025;arXiv:2503.22444

  9. [17]

    Taylor, M

    R. Taylor, M. Kardas, G. Cucurull, et al. Galactica: A Large Language Model for Science. 2022;arXiv:2211.09085

  10. [18]

    Jumper, R

    J. Jumper, R. Evans, A. Pritzel, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596(7873):583. doi: 10.1038/s41586-021-03819-2

  11. [19]

    Varadi, S

    M. Varadi, S. Anyango, M. Deshpande, et al. AlphaFold Protein Structure Database: massively expanding the structural coverage of protein -sequence space with high -accuracy models. Nucleic Acids Res. 2022;50(D1):D439. doi: 10.1093/nar/gkab1061

  12. [20]

    Hatakeyama-Sato, N

    K. Hatakeyama-Sato, N. Yamane, Y. Igarashi, et al. Prompt engineering of GPT -4 for chemical research: what can/cannot be done? Sci. Technol. Adv. Mater.: Methods. 2023;3(1):2260300. doi: 10.1080/27660400.2023.2260300

  13. [21]

    Ramprasad, R

    R. Ramprasad, R. Batra, G. Pilania, et al. Machine learning in materials informatics: recent applications and prospects. Npj Comput. Mater. 2017;3(1):54. doi: 10.1038/s41524 -017-0056-5

  14. [22]

    Y. Chen, J. Kirchmair. Cheminformatics in Natural Product -based Drug Discovery. Mol. Inform. 2020;39(12):e2000171. doi: 10.1002/minf.202000171

  15. [23]

    D. A. Boiko, R. MacKnight, G. Gomes. Emergent autonomous scientific research capabilities of large language models. 2023;arXiv:2304.05332

  16. [24]

    C. Lu, C. Lu, R. T. Lange, et al. The AI Scientist: Towards Fully Automated Open -Ended Scientific Discovery. 2024;arXiv:2408.06292. doi: arXiv:2408.06292

  17. [25]

    J. Choi, G. Nam, J. Choi, et al. A Perspective on Foundation Models in Chemistry. JACS Au

  18. [26]

    frugal twin

    S. Lo, S. G. Baird, J. Schrier, et al. Review of low -cost self-driving laboratories in chemistry and materials science: the “frugal twin” concept. Digital Discovery. 2024;3(5):842. doi: 10.1039/d3dd00223c

  19. [27]

    doi: 10.1021/jacsau.4c01160

  20. [28]

    K. Olsen. The first 110 years of laboratory automation: technologies, applications, and the creative scientist. J Lab Autom. 2012;17(6):469. doi: 10.1177/2211068212455631

  21. [29]

    Schneider

    G. Schneider. Automating drug discovery. Nature Reviews Drug Discovery. 2018;17(2):97. doi: 10.1038/nrd.2017.232

  22. [30]

    Q. Jia, H. Chu, Z. Jin, et al. High -throughput single -сell sequencing in cancer research. Signal Transduction and Targeted Therapy. 2022;7(1):145. doi: 10.1038/s41392-022-00990-4

  23. [31]

    C. W. Coley, D. A. Thomas, 3rd, J. A. M. Lummiss, et al. A robotic platform for flow synthesis of organic compounds informed by AI planning. Science. 2019;365(6453). doi: 10.1126/science.aax1566

  24. [32]

    Akiyama, H

    S. Akiyama, H. Akitsu, R. Tamura, et al. Autonomous Optimization of Discrete Reaction Parameters: Mono -Functionalization of a Bifunctional Substrate via a Suzuki -Miyaura Reaction. 2024;chemrxiv-2024-bnj6p-v2. doi: 10.26434/chemrxiv-2024-bnj6p-v2

  25. [33]

    Burger, P

    B. Burger, P. M. Maffettone, V. V. Gusev, et al. A mobile robotic chemist. Nature. 2020;583(7815):237. doi: 10.1038/s41586-020-2442-2

  26. [34]

    Matsuda, K

    S. Matsuda, K. Nishioka, S. Nakanishi. High -throughput combinatorial screening of multi - component electrolyte additives to improve the performance of Li metal secondary batteries. Sci Rep. 2019;9(1):6211. doi: 10.1038/s41598-019-42766-x

  27. [35]

    Shimizu, S

    R. Shimizu, S. Kobayashi, Y. Watanabe, et al. Autonomous materials synthesis by machine learning and robotics. APL Materials. 2020;8(11):111110. doi: 10.1063/5.0020370

  28. [36]

    Terada, Y

    M. Terada, Y. Kogawa, Y. Shibata, et al. Robotic cell processing facility for clinical research of retinal cell therapy. SLAS Technol. 2023;28(6):449. doi: 10.1016/j.slast.2023.10.004

  29. [37]

    Nishio, A

    K. Nishio, A. Aiba, K. Takihara, et al. A digital laboratory with a modular measurement system and standardized data format. Digital Discovery. 2025. doi: 10.1039/d4dd00326h

  30. [38]

    A. W. Liu, A. Villar-Briones, N. M. Luscombe, et al. Automated phenol-chloroform extraction of high molecular weight genomic DNA for use in long -read single -molecule sequencing. F1000Res. 2022;11:240. doi: 10.12688/f1000research.109251.1

  31. [39]

    Holland, J

    I. Holland, J. A. Davies. Automation in the Life Science Research Laboratory. Front Bioeng Biotechnol. 2020;8:571777. doi: 10.3389/fbioe.2020.571777

  32. [40]

    R. D. King, J. Rowland, S. G. Oliver, et al. The automation of science. Science. 2009;324(5923):85. doi: 10.1126/science.1165620

  33. [41]

    R. D. King, K. E. Whelan, F. M. Jones, et al. Functional genomic hypothesis generation and experimentation by a robot scientist. Nature. 2004;427(6971):247. doi: 10.1038/nature02236

  34. [42]

    Sasamata, D

    M. Sasamata, D. Shimojo, H. Fuse, et al. Establishment of a Robust Platform for Induced Pluripotent Stem Cell Research Using Maholo LabDroid. SLAS Technol. 2021;26(5):441. doi: 10.1177/24726303211000690

  35. [43]

    Yachie, K

    N. Yachie, K. Takahashi, T. Katayama, et al. Robotic crowd biology with Maholo LabDroids. Nat. Biotechnol. 2017;35(4):310. doi: 10.1038/nbt.3758

  36. [44]

    N. Rupp, K. Peschke, M. Koppl, et al. Establishment of low-cost laboratory automation processes using AutoIt and 4-axis robots. SLAS Technol. 2022;27(5):312. doi: 10.1016/j.slast.2022.07.001

  37. [45]

    G. N. Kanda, T. Tsuzuki, M. Terada, et al. Robotic search for optimal cell culture in regenerative medicine. Elife. 2022;11:e77007. doi: 10.7554/eLife.77007

  38. [46]

    K. G. Webber, O. Clemens, V. Buscaglia, et al. Review of the opportunities and limitations for powder-based high-throughput solid-state processing of advanced functional ceramics. J. Eur. Ceram. Soc. 2024;44(15). doi: 10.1016/j.jeurceramsoc.2024.116780

  39. [47]

    L. M. Roch, F. Hase, C. Kreisbeck, et al. ChemOS: An orchestration software to democratize autonomous discovery. PLoS One. 2020;15(4):e0229862. doi: 10.1371/journal.pone.0229862

  40. [48]

    J. Zhou, M. Luo, L. Chen, et al. A multi -robot–multi-task scheduling system for autonomous chemistry laboratories. Digital Discovery. 2025;4(3):636. doi: 10.1039/d4dd00313f

  41. [50]

    M. Moor, O. Banerjee, Z. S. H. Abad, et al. Foundation models for generalist medical artificial intelligence. Nature. 2023;616(7956):259. doi: 10.1038/s41586-023-05881-4

  42. [54]

    Y. Jin, Q. Zhao, Y. Wang, et al. AgentReview: Exploring Peer Review Dynamics with LLM Agents. 2024. doi: arXiv:2406.12708

  43. [55]

    L. Wang, C. Ma, X. Feng, et al. A survey on large language model based autonomous agents. Frontiers of Computer Science. 2024;18(6). doi: 10.1007/s11704-024-40231-1

  44. [56]

    Huang, F

    Y. Huang, F. Cheng, F. Zhou, et al. ROMAS: A Role -Based Multi-Agent System for Database monitoring and Planning. 2024;arXiv:2412.13520

  45. [57]

    Rasheed, M

    Z. Rasheed, M. Waseem, K. K. Kemell, et al. Large Language Models for Code Generation: The Practitioners Perspective. 2025;arXiv:2501.16998

  46. [58]

    Jansen, M

    P. Jansen, M. -A. Côté, T. Khot, et al. DiscoveryWorld: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents. 2024;arXiv:2406.06769

  47. [59]

    Z. Chen, S. Chen, Y. Ning, et al. ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery. 2024;arXiv:2410.05080

  48. [60]

    S. Gao, A. Fang, Y. Huang, et al. Empowering biomedical discovery with AI agents. Cell. 2024;187(22):6125. doi: 10.1016/j.cell.2024.09.022

  49. [61]

    J. Baek, S. K. Jauhar, S. Cucerzan, et al. ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models. 2025;arXiv:2404.07738

  50. [62]

    Huang, J

    Q. Huang, J. Vora, P. Liang, et al. MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation. 2023;arXiv:2310.03302

  51. [63]

    J. S. Chan, N. Chowdhury, O. Jaffe, et al. MLE -bench: Evaluating Machine Learning Agents on Machine Learning Engineering. 2024;arXiv:2410.07095

  52. [64]

    Behrouz, P

    A. Behrouz, P. Zhong, V. Mirrokni. Titans: Learning to Memorize at Test Time. 2024;arXiv:2501.00663

  53. [65]

    https://sakana.ai/ai-scientist-first-publication/

  54. [66]

    Y. Yao, J. Duan, K. Xu, et al. A survey on large language model (LLM) security and privacy: The Good, The Bad, and The Ugly. High-Confidence Computing. 2024;4(2). doi: 10.1016/j.hcc.2024.100211

  55. [67]

    B. J. Shields, J. Stevens, J. Li, et al. Bayesian reaction optimization as a tool for chemical synthesis. Nature. 2021;590(7844):89. doi: 10.1038/s41586-021-03213-y

  56. [68]

    Allen-Zhu, Y

    Z. Allen-Zhu, Y. Li. Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws. 2024;arXiv:2404.05405

  57. [69]

    Kaplan, S

    J. Kaplan, S. McCandlish, T. Henighan, et al. Scaling Laws for Neural Language Models. 2020;arXiv:2001.08361

  58. [70]

    K. G. Yager. Towards a science exocortex. Digital Discovery. 2024;3(10):1933. doi: 10.1039/d4dd00178h

  59. [71]

    Hatakeyama-Sato, Y

    K. Hatakeyama-Sato, Y. Igarashi, S. Katakami, et al. Teaching Specific Scientific Knowledge into Large Language Models through Additional Training. 2023;arXiv:2312.03360

  60. [72]

    Z. Song, M. Ju, C. Ren, et al. LLM -Feynman: Leveraging Large Language Models for Universal Scientific Formula and Theory Discovery. 2025;arXiv:2503.06512

  61. [73]

    Y. Ruan, C. Lu, N. Xu, et al. An automatic end -to-end chemical synthesis development platform powered by large language models. Nat. Commun. 2024;15(1):10160. doi: 10.1038/s41467 -024-54457-x

  62. [74]

    Eppel, H

    S. Eppel, H. Xu, M. Bismuth, et al. Computer Vision for Recognition of Materials and Vessels in Chemistry Lab Settings and the Vector -LabPics Data Set. ACS Cent. Sci. 2020;6(10):1743. doi: 10.1021/acscentsci.0c00460

  63. [75]

    Cheng, S

    X. Cheng, S. Zhu, Z. Wang, et al. Intelligent vision for the detection of chemistry glassware toward AI robotic chemists. Artificial Intelligence Chemistry. 2023;1(2). doi: 10.1016/j.aichem.2023.100016

  64. [76]

    X. Chen, J. Yang. X-Fi: A Modality-Invariant Foundation Model for Multimodal Human Sensing. 2025;arXiv:2410.10167

  65. [77]

    T. T. Brunye, T. Drew, D. L. Weaver, et al. A review of eye tracking for understanding and improving diagnostic interpretation. Cogn Res Princ Implic. 2019;4(1):7. doi: 10.1186/s41235 -019-0159- 2

  66. [78]

    S. M. Kargar, B. Yordanov, C. Harvey, et al. Emerging Trends in Realistic Robotic Simulations: A Comprehensive Systematic Literature Review. IEEE Access. 2024:1. doi: 10.1109/access.2024.3404881

  67. [79]

    Xiao, P.-F

    H. Xiao, P.-F. Wang. LLM A*: Human in the Loop Large Language Models Enabled A* Search for Robotics. 2024;arXiv:2312.01797

  68. [80]

    M. Kaup, C. Wolff, H. Hwang, et al. A Review of Nine Physics Engines for Reinforcement Learning Research. 2024;arXiv:2407.08590

  69. [81]

    Learning Robotic Powder Weighing from Simulation for Laboratory Automation

    Y. Kadokawa, M. Hamaya, K. Tanaka, "Learning Robotic Powder Weighing from Simulation for Laboratory Automation", presented at 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 1-5 Oct. 2023, 2023

  70. [82]

    M. V. Taylor, Z. Muwaffak, M. R. Penny, et al. Optimising digital twin laboratories with conversational AIs: enhancing immersive training and simulation through virtual reality. Digital Discovery

  71. [83]

    Agarwal, A

    N. Agarwal, A. Ali, M. Bala, et al. Cosmos World Foundation Model Platform for Physical AI. 2025;arXiv:2501.03575

  72. [84]

    Hatakeyama-Sato, K

    K. Hatakeyama-Sato, K. Oyaizu. Integrating multiple materials science projects in a single neural network. Commun. Mater. 2020;1:article number: 49. doi: 10.1038/s43246 -020-00052-8

  73. [85]

    doi: 10.1039/d4dd00330f

  74. [86]

    Firoozi, J

    R. Firoozi, J. Tucker, S. Tian, et al. Foundation Models in Robotics: Applications, Challenges, and the Future. 2023;arXiv:2312.07843

  75. [87]

    Bousmalis, G

    K. Bousmalis, G. Vezzani, D. Rao, et al. RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation. 2023;arXiv:2306.11706

  76. [88]

    C. Tang, B. Abbatematteo, J. Hu, et al. Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes. 2024;arXiv:2408.03539. doi: 10.1146/((please

  77. [89]

    S. Reed, K. Zolna, E. Parisotto, et al. A Generalist Agent. 2022;arXiv:2205.06175

  78. [90]

    J. Sun, Q. Zhang, G. Han, et al. Trinity: A Modular Humanoid Robot AI System. 2025;arXiv:2503.08338

  79. [91]

    T. Z. Zhao, J. Tompson, D. Driess, et al. ALOHA Unleashed: A Simple Recipe for Robot Dexterity. 2024;arXiv:2410.13126

  80. [92]

    Z. Gu, J. Li, W. Shen, et al. Humanoid Locomotion and Manipulation: Current Progress and Challenges in Control, Planning, and Learning. 2025;arXiv:2501.02116

  81. [93]

    Darvish, M

    K. Darvish, M. Skreta, Y. Zhao, et al. ORGANA: A Robotic Assistant for Automated Chemistry Experimentation and Characterization. 2024;arXiv:2401.06949

  82. [94]

    Black, N

    K. Black, N. Brown, D. Driess, et al. π0: A Vision -Language-Action Flow Model for General Robot Control. 2024;arXiv:2410.24164

  83. [95]

    T. Song, M. Luo, X. Zhang, et al. A Multiagent-Driven Robotic AI Chemist Enabling Autonomous Chemical Research On Demand. J. Am. Chem. Soc. 2025. doi: 10.1021/jacs.4c17738

  84. [96]

    K. Jang, J. Park, H. J. Min. Developing a Cooking Robot System for Raw Food Processing Based on Instance Segmentation. IEEE Access. 2024;12:106857. doi: 10.1109/access.2024.3436849

  85. [97]

    Sochacki, X

    G. Sochacki, X. Zhang, A. Abdulali, et al. Towards practical robotic chef: Review of relevant work and future challenges. Journal of Field Robotics. 2024;41(5):1596. doi: 10.1002/rob.22321

  86. [98]

    M. S. Sakib, Y. Sun. From Cooking Recipes to Robot Task Trees -- Improving Planning Correctness and Task Efficiency by Leveraging LLMs with a Knowledge Network. 2023;arXiv:2309.09181

  87. [99]

    D. M. Le, R. Guo, W. Xu, et al. Improved Instruction Ordering in Recipe-Grounded Conversation. 2023;arXiv:2305.17280

  88. [100]

    Hatakeyama-Sato, H

    K. Hatakeyama-Sato, H. Ishikawa, S. Takaishi, et al. Semiautomated experiment with a robotic system and data generation by foundation models for synthesis of polyamic acid particles. Polym. J. 2024;56(11):977. doi: 10.1038/s41428-024-00930-9

  89. [101]

    Perrett, A

    T. Perrett, A. Darkhalil, S. Sinha, et al. HD -EPIC: A Highly-Detailed Egocentric Video Dataset. 2025;arXiv:2502.04144

  90. [102]

    Kanazawa, K

    N. Kanazawa, K. Kawaharazuka, Y. Obinata, et al. Real-world cooking robot system from recipes based on food state recognition using foundation models and PDDL. Advanced Robotics. 2024;38(18):1318. doi: 10.1080/01691864.2024.2407136

  91. [103]

    Sochacki, A

    G. Sochacki, A. Abdulali, N. K. Hosseini, et al. Recognition of Human Chef’s Intentions for Incremental Learning of Cookbook by Robotic Salad Chef. IEEE Access. 2023;11:57006. doi: 10.1109/access.2023.3276234

  92. [104]

    W. Gu, S. Kondepudi, L. Huang, et al. Continual Skill and Task Learning via Dialogue. 2024;arXiv:2409.03166

  93. [105]

    Kadalagere Sampath, N

    S. Kadalagere Sampath, N. Wang, H. Wu, et al. Review on human‐like robot manipulation using dexterous hands. Cognitive Computation and Systems. 2023;5(1):14. doi: 10.1049/ccs2.12073

  94. [106]

    Yamane, Y

    K. Yamane, Y. Saigusa, S. Sakaino, et al. Soft and Rigid Object Grasping With Cross -Structure Hand Using Bilateral Control -Based Imitation Learning. IEEE Robotics and Automation Letters. 2024;9(2):1198. doi: 10.1109/lra.2023.3335768

  95. [107]

    D. Noh, H. Nam, K. Gillespie, et al. YORI: Autonomous Cooking System Utilizing a Modular Robotic Kitchen and a Dual-Arm Proprioceptive Manipulator. 2024;arXiv:2405.11094

  96. [108]

    U. Kim, D. Jung, H. Jeong, et al. Integrated linkage -driven dexterous anthropomorphic robotic hand. Nat. Commun. 2021;12(1):7177. doi: 10.1038/s41467-021-27261-0

  97. [109]

    Billard, D

    A. Billard, D. Kragic. Trends and challenges in robot manipulation. Science. 2019;364(6446). doi: 10.1126/science.aat8414

  98. [110]

    S. Park, S. Lee, M. Choi, et al. Learning to Transfer Human Hand Skills for Robot Manipulations. 2025;arXiv:2501.04169

  99. [111]

    F. Iida, C. Laschi. Soft Robotics: Challenges and Perspectives. Procedia Computer Science. 2011;7:99. doi: 10.1016/j.procs.2011.12.030

  100. [112]

    Hegde, J

    C. Hegde, J. Su, J. M. R. Tan, et al. Sensing in Soft Robotics. ACS Nano. 2023;17(16):15277. doi: 10.1021/acsnano.3c04089

  101. [113]

    Zhang, J

    N. Zhang, J. Ren, Y. Dong, et al. Soft robotic hand with tactile palm-finger coordination. Nat. Commun. 2025;16(1):2395. doi: 10.1038/s41467-025-57741-6

  102. [114]

    Koroyasu, T

    Y. Koroyasu, T. V. Nguyen, S. Sasaguri, et al. Microfluidic platform using focused ultrasound passing through hydrophobic meshes with jump availability. PNAS Nexus. 2023;2(7):pgad207. doi: 10.1093/pnasnexus/pgad207

  103. [115]

    D. Rus, M. T. Tolley. Design, fabrication and control of soft robots. Nature. 2015;521(7553):467. doi: 10.1038/nature14543

  104. [116]

    Polygerinos, N

    P. Polygerinos, N. Correll, S. A. Morin, et al. Soft Robotics: Review of Fluid‐Driven Intrinsically Soft Devices; Manufacturing, Sensing, Control, and Applications in Human‐Robot Interaction Adv. Eng. Mater. 2017;19(12). doi: 10.1002/adem.201700016

  105. [118]

    H. Sato, C. W. Berry, Y. Peeri, et al. Remote radio control of insect flight. Front. Integr. Neurosci. 2009;3:24. doi: 10.3389/neuro.07.024.2009

  106. [119]

    Mazumder, C

    M. Mazumder, C. Banbury, X. Yao, et al. DataPerf: Benchmarks for Data -Centric AI Development. 2022;arXiv:2207.10062

  107. [2025]

    doi: 10.1039/d4dd00387j

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.