Pith. sign in

REVIEW 4 major objections 5 minor 4 cited by

LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A unified taxonomy and formal language can describe every LLM-based agentic reasoning framework as a modification of one general reasoning loop.

desk verdict A useful, well-organized survey of agentic reasoning frameworks whose taxonomy and notation are the real contribution; the scenario-level comparisons are illustrative, not demonstrated, and the framework-vs-model confound is acknowledged but unresolved. read the letter →

arxiv 2508.17692 v1 pith:36TAFVDB submitted 2025-08-25 cs.AI cs.CL

classification cs.AIcs.CL
keywords agenticreasoningLLM-basedagentsframeworksmulti-agentsystemstoolusepromptengineeringapplicationscenariossurveytaxonomy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey tries to establish that the exploding field of LLM-based agentic reasoning can be organized by a single conceptual key: a general reasoning loop in which an agent maintains persistent context, selects actions from a small action space, and updates goals and context until a termination condition is met. It claims that every current framework is a modification of that loop, and it sorts those modifications into three progressive levels: single-agent methods, tool-based methods, and multi-agent methods. It then uses the taxonomy to compare frameworks across scientific discovery, healthcare, software engineering, and social and economic simulation, and to collect the evaluation strategies used in each scenario. A sympathetic reader would care because, if the taxonomy is right, researchers gain a common language for describing, comparing, and reusing agent designs that previously resisted horizontal comparison.

What carries the argument

The load-bearing object is Algorithm 1, a general reasoning loop formalized with a persistent context $C_k$, an action space $\mathcal{A} = \{a_{\text{reason}}, a_{\text{tool}}, a_{\text{reflect}}\}$, goals $g$, external tools $t$, and a termination condition $Q$. The loop runs initialization, action selection, context update, and step counting until $Q$ is met, and each family of methods is presented as a modification of one specific line of this algorithm. This lets the survey claim a unified description of otherwise heterogeneous frameworks, and it is the mechanism that makes cross-scenario comparison possible.

What would settle it

Find a published agent framework whose reasoning cannot be written as an instance of Algorithm 1, or run an ablation where the organizational architecture is exchanged while the same LLM and tools are kept and performance is unchanged; either result would show the taxonomy is not universal.

Watch

Extended reading notes

Core claim

The paper's central claim is that framework-level agent reasoning can be decomposed into a single formal loop, and that every agentic reasoning framework is best understood as a modification of that loop. The loop begins with a user query and a goal, initializes a persistent context, and then repeatedly selects an action from a three-element action space, produces an output, updates the context, and increments a step counter until a termination condition is satisfied. On top of this loop, the survey places three progressive levels of methodology: single-agent methods that enrich the initial context or let the agent improve through reflection and iteration; tool-based methods that expand a single tool into a toolkit and address integration, selection, and utilization; and multi-agent methods that organize multiple agents through centralized, decentralized, or hierarchical architectures and let them interact through cooperation, competition, or negotiation. The survey further claims to be the first unified methodological taxonomy of this kind, and it uses the taxonomy to review representative frameworks and evaluation practices in several major application domains.

Load-bearing premise

The survey assumes that the reasoning an agent displays can be analyzed separately from the capabilities of the underlying language model, even though it concedes that separating framework design from model-level improvements is often hard.

Editorial extensions

If this is right

  • Any existing or future agent framework can be described as annotations on Algorithm 1, so method papers could report exactly which reasoning steps they modify.
  • Cross-domain comparisons become meaningful: a medical multi-agent debate and a chemistry tool-chaining system can be compared by which lines of the loop they change.
  • Evaluation practices can be tied to framework categories, allowing benchmark designers to target specific steps such as context update or tool selection.
  • The three-level progression offers a design recipe: start with prompting, add tools, then organize multiple agents, because each level builds on the previous one.
  • The scenario reviews provide a map of which framework types currently dominate each application area, which can guide new work toward underused designs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A controlled experiment the survey does not run would be to hold the LLM fixed, vary only the line of Algorithm 1 that a framework modifies, and measure performance; the taxonomy predicts each line's contribution will be separable. This is an editorial extension, not a paper claim.
  • The same formal loop could index benchmarks by which reasoning step they stress, giving a finer-grained evaluation matrix than the per-scenario lists the survey provides. This is an editorial suggestion.
  • A natural formal extension would add a composition operator so that single-agent, tool-based, and multi-agent modifications can be explicitly combined, a step the survey leaves implicit in its discussion of hybrid frameworks.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper surveys LLM-based agentic reasoning frameworks. It proposes a three-level taxonomy—single-agent methods, tool-based methods, and multi-agent methods—and introduces a formal language (Table 1 and Algorithm 1) intended to describe the multi-step reasoning loop of any such framework. It then reviews applications in scientific research, healthcare, software engineering, social and economic simulation, and a short 'others' section, with tables summarizing evaluation strategies and benchmark results. The central claims are that this is the first unified methodological taxonomy for framework-level agentic reasoning, and that comparing framework types across scenarios reveals how they 'dominate framework-level reasoning.'

Significance. The paper's organizational contribution is genuinely useful: mapping a wide range of recent works onto a shared reasoning loop, and collecting evaluation setups across domains, could serve as a convenient reference for researchers entering the area. The breadth of coverage is a strength, as are the consolidated tables (Tables 2–8) that gather methods, inspirations, datasets, and benchmark numbers in one place. However, the paper's central comparative claim depends on two assumptions that are asserted rather than tested: that framework design can be separated from model-level capability, and that an informal selection of 'representative works' supports scenario-level generalizations. If the authors can either control for the confound or substantially soften the dominance claims, the survey will be a solid reference; in its current form, the comparative conclusions outrun the evidence presented.

major comments (4)
  1. [Section 1 (page 3); Section 3.5; Table 6] The load-bearing assumption that framework-level reasoning can be analyzed separately from model-level capability is explicitly conceded in Section 1: 'it is often hard to clearly separate whether enhanced capabilities of an agent come from careful framework design, model-level improvements, or technological advancements.' Section 3.5 then excludes fine-tuning and RL from the classification, yet Section 4.5 and Figure 9 describe systems whose gains depend on fine-tuning (e.g., AIME's outer loop), and Table 6 compares frameworks running on different base LLMs (GPT-3.5 vs. GPT-4) as if the differences were attributable to framework design. This confound undermines the central comparative claim. Please either control for base model and inference budget in the comparative tables, or explicitly re-frame the scenario discussions as qualitative mappings of design patterns and clearly list the model-level confounds for each system.
  2. [Section 3.1, Algorithm 1, Eq. (2)] The formal language is a stated contribution, but it is not consistently defined. Table 1 defines the action space as A = {a_reason, a_tool, a_reflect}, while Algorithm 1 line 5 uses a'_k ∈ A for a context-updating action and Eq. (2) writes C_{k+1} = a'_reflect(...). The set of context-update actions a' is never formally defined, and reflection appears both as an element of A and as a member of a'. Without a precise definition of A and of the a/a' distinction, the formal language cannot serve as a 'unified formal language' for comparing frameworks. Please clarify the notation, for example by defining A_output and A_update explicitly.
  3. [Abstract; Section 4] The Abstract claims the survey will 'analyze how these frameworks dominate framework-level reasoning by comparing their applications across different scenarios,' but Section 4 is organized as a scenario-by-scenario enumeration without a cross-scenario comparative analysis. The comparative tables (e.g., Tables 6 and 7) mix base models and do not test whether single-agent, tool-based, or multi-agent design is associated with performance differences; Table 8 reports simulation scale rather than comparative outcomes. Please add an explicit cross-scenario synthesis (for example, a matrix mapping framework types to scenario requirements and observed design patterns) and state the evidentiary status of any dominance claims.
  4. [Figure 1 caption; Section 4 introduction] The survey does not report a reproducible paper-selection protocol. The caption of Figure 1 and the Section 4 introduction describe the included papers as 'a diverse set of representative works,' but no inclusion/exclusion criteria, search strings, screening procedure, or coverage statistics are provided. Because the scenario-level observations are the basis for the central comparative claim, the informal selection leaves the generalizability of those observations unclear. Please add a methods paragraph specifying how the corpus was collected, filtered, and updated, and state the limitations of the resulting coverage.
minor comments (5)
  1. [Throughout] There are numerous typos and grammatical slips, including 'porvides' (Section 3.2.1), 'Ru-Based Selection' (Section 3.3.2, should be 'Rule-Based'), 'Pandy et al.' (Section 4.2.2, should be 'Pandey'), 'AI-Reseracher' (Section 4.1.5), and 'an unified' (Abstract). A careful proofreading pass is needed.
  2. [Figure 1] The sentence 'For 2025, we predict the overall amount of papers linearly based on data accessed at 14th August' should specify the extrapolation method and the exact data cutoff, or the prediction should be removed.
  3. [Table 2] The abbreviation legend in Table 2 lists Role and Task under Prompt Engineering but omits environment simulation and in-context learning, even though those sub-methods are discussed in Section 3.2.1 and some entries may rely on them. Please either add the missing abbreviations or explain why they are not needed.
  4. [Section 3.2.2 and Section 3.4.2] The 'Interactive Learning' mechanism in the single-agent section and the 'Cooperation' mechanism in the multi-agent section both involve goal updates driven by environmental or partner feedback; the paper would benefit from an explicit discussion of how these categories are distinguished at the boundary.
  5. [Section 4.5] The 'Others' section is considerably thinner than the main scenario sections and relies on review-level references (e.g., [175], [271]) rather than the detailed system-level treatment used elsewhere. This is acceptable as a scoping choice, but it should be stated explicitly so readers do not expect the same depth of coverage.

Circularity Check

0 steps flagged · score 2.0 of 10

No circularity: the survey's taxonomy and scenario comparisons are self-contained; the lone self-citation is not load-bearing.

full rationale

The central claim is an organizational one: a three-level taxonomy (single-agent, tool-based, multi-agent) with a generic reasoning algorithm. Nothing in Sections 3 or 4 is derived from the authors' own prior results. The only self-citation is [217], which appears in the introduction as one of two examples of LLM-based agentic reasoning frameworks ('researchers have been actively exploring the use of LLMs as a core engine to build LLM-based agentic reasoning frameworks capable of executing complex, multi-step reasoning tasks [217, 266]'); it is not used to justify the taxonomy, the formal language, or any scenario-level comparison. The paper's comparative tables (e.g., Tables 6 and 7) report external benchmark numbers from the cited works, not fitted or predicted values. The framework-versus-model confound identified by the skeptic is a legitimate validity concern about whether framework-level comparisons are meaningful across different base LLMs, but it is not circularity: the paper's classification is not equivalent to its inputs by construction, and there is no hidden derivation or self-citation chain that forces the conclusions. Accordingly, the appropriate finding is minor self-citation without load-bearing circularity, consistent with a non-circularity score of 2.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No numbers are fitted; the linear extrapolation of 2025 publication counts in Fig. 1 is a minor forecast, not load-bearing. The taxonomy and notation are descriptive constructs, not entities with independent evidence. All load-bearing premises are domain assumptions about coverage, expressiveness, and separability; none are proven.

assumptions (4)
  • domain assumption The three-level taxonomy (single-agent, tool-based, multi-agent) is jointly exhaustive for LLM-based agentic reasoning frameworks.
    The survey asserts this organization without a formal proof or a systematic test; it is the frame on which all scenario analyses hang. Section 3 presents it as the top-level split (Fig. 2).
  • domain assumption Selected papers are representative of the field.
    Inclusion criteria are informal: 'mainly select technical papers published at top computer science conferences' and 'a diverse set of representative works' (Fig. 1 caption, Section 1). If the selection is biased, the scenario-level conclusions would not generalize.
  • domain assumption The unified formal language (Alg. 1, action space A = {a_reason, a_tool, a_reflect}) is expressive enough to describe all frameworks.
    The survey abstracts memory, retrieval, sandboxed environments, and human interruption into a single tool t (Section 3.1). If this abstraction loses a dimension that matters for reasoning, equations such as Eq. 6-15 would misrepresent how frameworks differ.
  • domain assumption Framework-level reasoning can be analyzed separately from model-level capabilities.
    The paper itself flags this as hard: 'it is often hard to clearly separate whether enhanced capabilities of an agent come from careful framework design, model-level improvements, or technological advancements' (Section 1, page 3). The entire survey proceeds from the assumption that such separation is meaningful.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios." pith.science (2026). https://pith.science/paper/36TAFVDB

@misc{pith2026250817692,
  author       = {Pith},
  title        = {Pith review of: LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/36TAFVDB}},
  note         = {Machine review of arXiv:2508.17692}
}
read the original abstract

Recent advances in the intrinsic reasoning capabilities of large language models (LLMs) have given rise to LLM-based agent systems that exhibit near-human performance on a variety of automated tasks. However, although these systems share similarities in terms of their use of LLMs, different reasoning frameworks of the agent system steer and organize the reasoning process in different ways. In this survey, we propose a systematic taxonomy that decomposes agentic reasoning frameworks and analyze how these frameworks dominate framework-level reasoning by comparing their applications across different scenarios. Specifically, we propose an unified formal language to further classify agentic reasoning systems into single-agent methods, tool-based methods, and multi-agent methods. After that, we provide a comprehensive review of their key application scenarios in scientific discovery, healthcare, software engineering, social simulation, and economics. We also analyze the characteristic features of each framework and summarize different evaluation strategies. Our survey aims to provide the research community with a panoramic view to facilitate understanding of the strengths, suitable scenarios, and evaluation practices of different agentic reasoning frameworks.

Figures

Figures reproduced from arXiv: 2508.17692 by the authors.

Figure 1
Figure 1. Number of publications regarding LLM-based Agentic Frameworks from 2020 to 2025 in journals and [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Taxonomy of our proposed agentic reasoning frameworks. We decompose agentic reasoning methods [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Prompt engineerng for agentic reasoning framework. We summarize four types of methods: [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Paradigms of self-improvement for an LLM-based agent. We introduce three core mechanisms. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Tool-based reasoning frameworks of LLM-based agent. [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 6
Figure 6. Figure 6: A taxonomy of Multi-agent reasoning frameworks, categorized in [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]
Figure 7
Figure 7. Figure 7: The overview of our selected paper of agentic reasoning frameworks across different application [PITH_FULL_IMAGE:figures/full_fig_p017_7.png]
Figure 8
Figure 8. Figure 8: A summarization of pipeline proposed by BioDiscovery-Agent [ [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]
Figure 9
Figure 9. Figure 9: Pipeline of AIME [260]. The framework is built upon two self-play loops. a) Inner Loop: a doctor agent continuously optimize its behavior based on real-time feedback from a Critic module during simulated dialogues. b) Outer Loop: the optimized simulated dialogues and o…
Figure 10
Figure 10. Figure 10: The fundamental paradigm for social simulation based on an Agentic Reasoning Framework. In this [PITH_FULL_IMAGE:figures/full_fig_p030_10.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AXE: Grey-Box Exploitability Confirmation for Localized Vulnerability Reports

    cs.CR 2026-02 conditional novelty 7.0 of 10

    Grey-box metadata (CWE + code location) plus a multi-agent LLM workflow raises automated web-exploit confirmation from ~10% to 30% on CVE-Bench, with actionable PoC output.

  2. The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Reasoning messages between heterogeneous VLMs can be routed through the image-token span: a distilled universal codec plus affine alignment transmits latent traces across model families, cutting wall-clock time in sma...

  3. APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning

    cs.CV 2026-08 conditional novelty 5.0 of 10

    A profiling-guided, LLM-driven framework combines structured pruning and mixed-precision quantization-aware training, reporting 13-18x bit-operation reductions with modest accuracy loss on ImageNet and CIFAR-10.

  4. Agents in the Wild: Where Research Meets Deployment

    cs.AI 2026-07 unverdicted

    A tutorial description reviewing the state of LLM agent deployment, with no new research findings.

Reference graph

Works this paper leans on

300 extracted references · 3 canonical work pages · cited by 4 Pith papers

  1. [1]

    Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans, Tim Green, Alexander Pritzel, Olaf Ronneberger, Lindsay Willmore, Andrew J Ballard, Joshua Bambrick, et al. 2024. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 8016 (2024), 493–500

  2. [2]

    Vaibhav Adlakha, Parishad BehnamGhader, Xing Han Lu, Nicholas Meade, and Siva Reddy. 2024. Evaluating correctness and faithfulness of instruction-following models for question answering. Transactions of the Association for Computational Linguistics 12 (2024), 681–699

  3. [3]

    Ali AhmadiTeshnizi, Wenzhi Gao, and Madeleine Udell. 2024. OptiMUS: scalable optimization modeling with (MI) LP solvers and large language models. In Proceedings of the 41st International Conference on Machine Learning . 577–596

  4. [4]

    Zaid Al-Ars, Obinna Agba, Zhuoran Guo, Christiaan Boerkamp, Ziyaad Jaber, and Tareq Jaber. 2023. Nlice: Synthetic medical record generation for effective primary healthcare differential diagnosis. In 2023 IEEE 23rd International Conference on Bioinformatics and Bioengineering (BIBE) . IEEE, 397–402

  5. [5]

    Mohammad Almansoori, Komal Kumar, and Hisham Cholakkal. 2025. Self-Evolving Multi-Agent Simulations for Realistic Clinical Interactions. arXiv preprint arXiv:2503.22678 (2025)

  6. [6]

    Amr Almorsi, Mohanned Ahmed, and Walid Gomaa. 2024. Guided code generation with llms: A multi-agent framework for complex code tasks. In 2024 12th International Japan-Africa Conference on Electronics, Communications, and Computations (JAC-ECC) . IEEE, 215–218

  7. [7]

    Vinicius M Alves, Eugene Muratov, Denis Fourches, Judy Strickland, Nicole Kleinstreuer, Carolina H Andrade, and Alexander Tropsha. 2015. Predicting chemically-induced skin reactions. Part I: QSAR models of skin sensitization and their application to identify potentially hazardous compounds. Toxicology and applied pharmacology 284, 2 (2015), 262–272

  8. [8]

    Theonie Anastassiadis, Sean W Deacon, Karthik Devarajan, Haiching Ma, and Jeffrey R Peterson. 2011. Comprehensive assay of kinase catalytic activity reveals features of kinase inhibitor selectivity. Nature biotechnology 29, 11 (2011), 1039–1045

Show all 300 references
  1. [9]

    introduction to Computational Chemistry

    D Armstrong. 2024. Exercises from “introduction to Computational Chemistry”(CHM 323), University of toronto. Personal communication (2024)

  2. [10]

    Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021. Program synthesis with large language models.arXiv preprint arXiv:2108.07732 (2021)

  3. [11]

    Reza Averly, Frazier N Baker, and Xia Ning. 2025. Liddia: Language-based intelligent drug discovery agent. arXiv preprint arXiv:2502.13959 (2025)

  4. [12]

    Kaito Baba, Chaoran Liu, Shuhei Kurita, and Akiyoshi Sannai. 2025. Prover Agent: An Agent-based Framework for Formal Mathematical Proofs. arXiv preprint arXiv:2506.19923 (2025)

  5. [13]

    Seongsu Bae, Daeun Kyung, Jaehee Ryu, Eunbyeol Cho, Gyubok Lee, Sunjun Kweon, Jungwoo Oh, Lei Ji, Eric Chang, Tackeun Kim, et al. 2023. Ehrxqa: A multi-modal question answering dataset for electronic health records with chest x-ray images. Advances in Neural Information Proces...

  6. [14]

    Jinheon Baek, Sujay Kumar Jauhar, Silviu Cucerzan, and Sung Ju Hwang. 2025. ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association ...

  7. [15]

    Kinjal Basu, Ibrahim Abdelaziz, Kiran Kate, Mayank Agarwal, Maxwell Crouse, Yara Rizk, Kelsey Bradford, Asim Munawar, Sadhana Kumaravel, Saurabh Goyal, et al . 2024. Nestful: A benchmark for evaluating llms on nested sequences of api calls. arXiv preprint arXiv:2409.03797 (2024)

  8. [16]

    Suhana Bedi, Iddah Mlauzi, Daniel Shin, Sanmi Koyejo, and Nigam H Shah. 2025. The Optimization Paradox in Clinical AI Multi-Agent Systems. arXiv preprint arXiv:2506.06574 (2025)

  9. [17]

    RM Belbin and V Brown. 2012. Team roles at work. Routledge (2012)

  10. [18]

    Peter Belcak, Greg Heinrich, Shizhe Diao, Yonggan Fu, Xin Dong, Saurav Muralidharan, Yingyan Celine Lin, and Pavlo Molchanov. 2025. Small Language Models are the Future of Agentic AI. arXiv preprint arXiv:2506.02153 (2025)

  11. [19]

    A Patrícia Bento, Anna Gaulton, Anne Hersey, Louisa J Bellis, Jon Chambers, Mark Davies, Felix A Krüger, Yvonne Light, Lora Mak, Shaun McGlinchey, et al. 2014. The ChEMBL bioactivity database: an update. Nucleic acids research 42, D1 (2014), D1083–D1090

  12. [20]

    Helen M Berman, John Westbrook, Zukang Feng, Gary Gilliland, Talapady N Bhat, Helge Weissig, Ilya N Shindyalov, and Philip E Bourne. 2000. The protein data bank. Nucleic acids research 28, 1 (2000), 235–242

  13. [21]

    G Richard Bickerton, Gaia V Paolini, Jérémy Besnard, Sorel Muresan, and Andrew L Hopkins. 2012. Quantifying the chemical beauty of drugs. Nature chemistry 4, 2 (2012), 90–98

  14. [22]

    Lisa Bortolotti. 2011. Does reflection lead to wise choices? Philosophical Explorations 14, 3 (2011), 297–313

  15. [23]

    Islem Bouzenia, Premkumar Devanbu, and Michael Pradel. 2025. RepairAgent: An Autonomous, LLM-Based Agent for Program Repair. In 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE) . IEEE Computer Society, 694–694

  16. [24]

    Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. 2023. Chemcrow: Augmenting large-language models with chemistry tools. arXiv preprint arXiv:2304.05376 (2023)

  17. [25]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901

  18. [26]

    Bokai Cao, Saizhuo Wang, Xinyi Lin, Xiaojun Wu, Haohan Zhang, Lionel M Ni, and Jian Guo. 2025. From deep learning to LLMs: a survey of AI in quantitative investment. arXiv preprint arXiv:2503.21422 (2025)

  19. [27]

    Julia Carnevale, Eric Shifrut, Nupura Kale, William A Nyberg, Franziska Blaeschke, Yan Yi Chen, Zhongmei Li, Sagar P Bapat, Morgan E Diolaiti, Patrick O’Leary, et al . 2022. RASA2 ablation in T cells boosts antigen sensitivity and long-term function. Nature 609, 7925 (2022), 174–182

  20. [28]

    Mert Cemri, Melissa Z Pan, Shuyi Yang, Lakshya A Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, et al. 2025. Why do multi-agent llm systems fail? arXiv preprint arXiv:2503.13657 (2025)

  21. [29]

    Jingyi Chai, Shuo Tang, Rui Ye, Yuwen Du, Xinyu Zhu, Mengcheng Zhou, Yanfeng Wang, Siheng Chen, et al. 2025. SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity’s Last Exam? arXiv preprint arXiv:2507.05241 (2025)

  22. [30]

    Jiajun Chai, Zijie Zhao, Yuanheng Zhu, and Dongbin Zhao. 2025. A Survey of Cooperative Multi-Agent Reinforcement Learning for Multi-Task Scenarios. Artificial Intelligence Science and Engineering 1, 2 (2025), 98–121

  23. [31]

    Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2024. ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate. InThe Twelfth International Conference on Learning Representations

  24. [32]

    Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, et al. 2025. MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering. In The Thirteenth International Conference ...

  25. [33]

    Crystal T Chang, Hodan Farah, Haiwen Gui, Shawheen Justin Rezaei, Charbel Bou-Khalil, Ye-Jean Park, Akshay Swaminathan, Jesutofunmi A Omiye, Akaash Kolluri, Akash Chaurasia, et al . 2024. Red teaming large language models in medicine: real-world insights on model behavior. med...

  26. [34]

    Guangyao Chen, Siwei Dong, Yu Shu, Ge Zhang, Jaward Sesay, Börje Karlsson, Jie Fu, and Yemin Shi. 2024. AutoAgents: a framework for automatic agent generation. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence. 22–30

  27. [35]

    Hanjie Chen, Zhouxiang Fang, Yash Singla, and Mark Dredze. 2025. Benchmarking large language models on answering and explaining challenging medical questions. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Ling...

  28. [36]

    Kexin Chen, Junyou Li, Kunyi Wang, Yuyang Du, Jiahui Yu, Jiamin Lu, Lanqing Li, Jiezhong Qiu, Jianzhang Pan, Yi Huang, et al. 2023. Chemist-X: Large language model-empowered agent for reaction condition recommendation in chemical synthesis. arXiv preprint arXiv:2311.10776 (202...

  29. [37]

    Kai Chen, Xinfeng Li, Tianpei Yang, Hewei Wang, Wei Dong, and Yang Gao. 2025. MDTeamGPT: A Self-Evolving LLM- based Multi-Agent Framework for Multi-Disciplinary Team Medical Consultation. arXiv preprint arXiv:2503.13856 (2025)

  30. [38]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)

  31. [39]

    Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Chen Qian, et al. 2024. AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors. In ICLR

  32. [40]

    Weize Chen, Ziming You, Ran Li, Chen Qian, Chenyang Zhao, Cheng Yang, Ruobing Xie, Zhiyuan Liu, Maosong Sun, et al. 2025. Internet of Agents: Weaving a Web of Heterogeneous Agents for Collaborative Intelligence. In The Thirteenth International Conference on Learning Representations

  33. [41]

    Xuanzhong Chen, Ye Jin, Xiaohao Mao, Lun Wang, Shuyang Zhang, and Ting Chen. 2024. RareAgents: Autonomous Multi-disciplinary Team for Rare Disease Diagnosis and Treatment. arXiv e-prints (2024), arXiv–2412

  34. [42]

    Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2024. Teaching Large Language Models to Self-Debug. In The Twelfth International Conference on Learning Representations

  35. [43]

    Xuanzhong Chen, Xiaohao Mao, Qihan Guo, Lun Wang, Shuyang Zhang, and Ting Chen. 2024. RareBench: can LLMs serve as rare diseases specialists?. In Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining. 4850–4861

  36. [44]

    Xi Chen, Huahui Yi, Mingke You, WeiZhi Liu, Li Wang, Hairui Li, Xue Zhang, Yingman Guo, Lei Fan, Gang Chen, et al. 2025. Enhancing diagnostic capability with multi-agents conversational large language models. NPJ digital medicine 8, 1 (2025), 159

  37. [45]

    Yuxuan Chen, Xu Zhu, Hua Zhou, and Zhuyin Ren. 2024. MetaOpenFOAM: an LLM-based multi-agent framework for CFD. arXiv preprint arXiv:2407.21320 (2024)

  38. [46]

    Yuxuan Chen, Xu Zhu, Hua Zhou, and Zhuyin Ren. 2025. Metaopenfoam 2.0: Large language model driven chain of thought for automating cfd simulation and post-processing. arXiv preprint arXiv:2502.00498 (2025)

  39. [47]

    Zehui Chen, Kuikun Liu, Qiuchen Wang, Jiangning Liu, Wenwei Zhang, Kai Chen, and Feng Zhao. 2025. Mind- Search: Mimicking Human Minds Elicits Deep AI Searcher. In The Thirteenth International Conference on Learning Representations

  40. [48]

    Zhaoling Chen, Xiangru Tang, Gangda Deng, Fang Wu, Jialong Wu, Zhiwei Jiang, Viktor Prasanna, Arman Cohan, and Xingyao Wang. 2025. Locagent: Graph-guided llm agents for code localization. arXiv preprint arXiv:2503.09089 (2025)

  41. [49]

    Yuan Chiang, Elvis Hsieh, Chia-Hong Chou, and Janosh Riebesell. 2025. LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval. In AI for Accelerated Materials Design-ICLR

  42. [50]

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168 (2021)

  43. [51]

    1000 Genomes Project Consortium et al. 2012. An integrated map of genetic variation from 1,092 human genomes. Nature 491, 7422 (2012), 56

  44. [52]

    Steven M Corsello, Joshua A Bittker, Zihan Liu, Joshua Gould, Patrick McCarren, Jodi E Hirschman, Stephen E Johnston, Anita Vrcic, Bang Wong, Mariya Khan, et al. 2017. The Drug Repurposing Hub: a next-generation drug library and information resource. Nature medicine 23, 4 (201...

  45. [53]

    Debrup Das, Debopriyo Banerjee, Somak Aditya, and Ashish Kulkarni. 2024. MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human La...

  46. [54]

    Tom DeMarco and Tim Lister. 2013. Peopleware: productive projects and teams . Addison-Wesley

  47. [55]

    Han Ding, Yinheng Li, Junhao Wang, and Hang Chen. 2024. Large language model agent in financial trading: A survey. arXiv preprint arXiv:2408.06361 (2024)

  48. [56]

    Yihong Dong, Jiazheng Ding, Xue Jiang, Ge Li, Zhuo Li, and Zhi Jin. 2025. Codescore: Evaluating code generation by learning code execution. ACM Transactions on Software Engineering and Methodology 34, 3 (2025), 1–22

  49. [57]

    Yihong Dong, Xue Jiang, Zhi Jin, and Ge Li. 2024. Self-collaboration code generation via chatgpt. ACM Transactions on Software Engineering and Methodology 33, 7 (2024), 1–38

  50. [58]

    Mengge Du, Yuntian Chen, Zhongzheng Wang, Longfeng Nie, and Dongxiao Zhang. 2024. LLM4ED: Large Language Models for Automatic Equation Discovery. CoRR (2024)

  51. [59]

    Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch. 2024. Improving factuality and reasoning in language models through multiagent debate. In Proceedings of the 41st International Conference on Machine Learning. 11733–11763. J. ACM, Vol. 37, No. 4, Ar...

  52. [60]

    Zane Durante, Qiuyuan Huang, Naoki Wake, Ran Gong, Jae Sung Park, Bidipta Sarkar, Rohan Taori, Yusuke Noda, Demetri Terzopoulos, Yejin Choi, et al. 2024. Agent ai: Surveying the horizons of multimodal interaction. arXiv preprint arXiv:2401.03568 (2024)

  53. [61]

    Abul Ehtesham, Aditi Singh, Gaurav Kumar Gupta, and Saket Kumar. 2025. A survey of agent interoperability protocols: Model context protocol (mcp), agent communication protocol (acp), agent-to-agent protocol (a2a), and agent network protocol (anp). arXiv preprint arXiv:2505.022...

  54. [62]

    Peter Ertl and Ansgar Schuffenhauer. 2009. Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions. Journal of cheminformatics 1, 1 (2009), 8

  55. [63]

    Caoyun Fan, Jindou Chen, Yaohui Jin, and Hao He. 2024. Can large language models serve as rational players in game theory? a systematic analysis. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 17960–17967

  56. [64]

    Zhihao Fan, Lai Wei, Jialong Tang, Wei Chen, Wang Siyuan, Zhongyu Wei, and Fei Huang. 2025. AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator. In Proceedings of the 31st International Conference on Computational Linguistics . 10183–10213

  57. [65]

    Huihui Fang, Fei Li, Junde Wu, Huazhu Fu, Xu Sun, Jaemin Son, Shuang Yu, Menglu Zhang, Chenglang Yuan, Cheng Bian, et al. 2022. Refuge2 challenge: A treasure trove for multi-dimension analysis and evaluation in glaucoma screening. arXiv preprint arXiv:2202.08994 (2022)

  58. [66]

    Arsene Fansi Tchango, Rishab Goel, Zhi Wen, Julien Martel, and Joumana Ghosn. 2022. Ddxplus: A new dataset for automatic medical diagnosis. Advances in neural information processing systems 35 (2022), 31306–31318

  59. [67]

    Sorouralsadat Fatemi and Yuheng Hu. 2024. FinVision: A multi-agent framework for stock market prediction. In Proceedings of the 5th ACM International Conference on AI in Finance . 582–590

  60. [68]

    Xiang Fei, Xiawu Zheng, and Hao Feng. 2025. MCP-Zero: Proactive Toolchain Construction for LLM Agents from Scratch. arXiv preprint arXiv:2506.01056 (2025)

  61. [69]

    Helen V Firth, Shola M Richards, A Paul Bevan, Stephen Clayton, Manuel Corpas, Diana Rajan, Steven Van Vooren, Yves Moreau, Roger M Pettett, and Nigel P Carter. 2009. DECIPHER: database of chromosomal imbalance and phenotype in humans using ensembl resources. The American Jour...

  62. [70]

    Paul G Francoeur, Tomohide Masuda, Jocelyn Sunseri, Andrew Jia, Richard B Iovanisci, Ian Snyder, and David R Koes

  63. [71]

    Bowen Gao, Yanwen Huang, Yiqiao Liu, Wenxuan Xie, Wei-Ying Ma, Ya-Qin Zhang, and Yanyan Lan. 2025. Pharma- gents: Building a virtual pharma with large language model agents. arXiv preprint arXiv:2503.22164 (2025)

  64. [72]

    Chen Gao, Xiaochong Lan, Zhihong Lu, Jinzhu Mao, Jinghua Piao, Huandong Wang, Depeng Jin, and Yong Li. 2023. S3: Social-network Simulation System with Large Language Model-Empowered Agents. CoRR (2023)

  65. [73]

    Huan-ang Gao, Jiayi Geng, Wenyue Hua, Mengkang Hu, Xinzhe Juan, Hongzhang Liu, Shilong Liu, Jiahao Qiu, Xuan Qi, Yiran Wu, et al. 2025. A survey of self-evolving agents: On path to artificial super intelligence. arXiv preprint arXiv:2507.21046 (2025)

  66. [74]

    Shen Gao, Yuntao Wen, Minghang Zhu, Jianing Wei, Yuhan Cheng, Qunzi Zhang, and Shuo Shang. 2024. Simulating financial market via large language model based agents. arXiv preprint arXiv:2406.19966 (2024)

  67. [75]

    Shanghua Gao, Richard Zhu, Zhenglun Kong, Ayush Noori, Xiao-Rui Su, Curtis Ginder, Theodoros Tsiligkaridis, and Marinka Zitnik. 2025. TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools. CoRR (2025)

  68. [76]

    Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023. Enabling Large Language Models to Generate Text with Citations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 6465–6488

  69. [77]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 2, 1 (2023)

  70. [78]

    Alireza Ghafarollahi and Markus Buehler. 2024. ProtAgents: Protein discovery via large language model multi-agent collaborations combining physics and machine learning. In ICLR Workshop on Large Language Model (LLM) Agents

  71. [79]

    Alireza Ghafarollahi and Markus J Buehler. 2025. Automating alloy design and discovery with physics-aware multimodal multiagent AI. Proceedings of the National Academy of Sciences 122, 4 (2025), e2414074122

  72. [80]

    Alireza Ghafarollahi and Markus J Buehler. 2025. SciAgents: automating scientific discovery through bioinspired multi-agent intelligent graph reasoning. Advanced Materials 37, 22 (2025), 2413523

  73. [81]

    Ali Essam Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Caralyn J Szostkiewicz, Jon M Laurent, Muhammed T Razzak, Andrew D White, Michaela M Hinks, and Samuel G Rodriques. 2025. Robin: A multi-agent system for automating scientific discovery. arXiv preprint arXiv:25...

  74. [82]

    Majid Ghasemi, Amir Hossein Moosavi, and Dariush Ebrahimi. 2024. A comprehensive survey of reinforcement learning: From algorithms to practical challenges. arXiv preprint arXiv:2411.18892 (2024)

  75. [83]

    Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, et al. 2025. Towards an AI co-scientist. arXiv preprint arXiv:2502.18864 J. ACM, Vol. 37, No. 4, Article 111. Publication da...

  76. [84]

    Zhiming Gou, Zili Li, Zili Wang, Ming Li, Zhen Wang, and Enhong Chen. 2024. OLVERA: A Framework for Open- ended Code Snippet Verification and Rectification using LLMs. InInternational Conference on Learning Representations (ICLR)

  77. [85]

    Zhibin Gou, Zhihong Shao, Yeyun Gong, Yujiu Yang, Nan Duan, Weizhu Chen, et al. 2024. CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing. In The Twelfth International Conference on Learning Representations

  78. [86]

    Mourad Gridach, Jay Nanavati, Khaldoun Zine El Abidine, Lenon Mendes, and Christina Mack. 2025. Agentic ai for scientific discovery: A survey of progress, challenges, and future directions. arXiv preprint arXiv:2503.08979 (2025)

  79. [87]

    Sven Gronauer and Klaus Diepold. 2022. Multi-agent deep reinforcement learning: a survey. Artificial Intelligence Review 55, 2 (2022), 895–943

  80. [88]

    Yu Gu, Yiheng Shu, Hao Yu, Xiao Liu, Yuxiao Dong, Jie Tang, Jayanth Srinivasa, Hugo Latapie, and Yu Su. 2024. Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language...

  81. [89]

    T Guo, X Chen, Y Wang, R Chang, S Pei, NV Chawla, O Wiest, and X Zhang. 2024. Large Language Model based Multi-Agents: A Survey of Progress and Challenges.. In 33rd International Joint Conference on Artificial Intelligence (IJCAI 2024). IJCAI; Cornell arxiv

  82. [90]

    Taicheng Guo, Bozhao Nan, Zhenwen Liang, Zhichun Guo, Nitesh Chawla, Olaf Wiest, Xiangliang Zhang, et al. 2023. What can large language models do in chemistry? a comprehensive benchmark on eight tasks. Advances in Neural Information Processing Systems 36 (2023), 59662–59688

  83. [91]

    Xuehang Guo, Xingyao Wang, Yangyi Chen, Sha Li, Chi Han, Manling Li, and Heng Ji. 2025. SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering. In Forty-second International Conference on Machine Learning

  84. [92]

    Zikang Guo, Benfeng Xu, Xiaorui Wang, and Zhendong Mao. 2025. MIRROR: Multi-agent Intra-and Inter-Reflection for Optimized Reasoning in Tool Learning. arXiv preprint arXiv:2505.20670 (2025)

  85. [93]

    Deepak Gupta, Kush Attal, and Dina Demner-Fushman. 2023. A dataset for medical instructional video classification and question answering. Scientific Data 10, 1 (2023), 158

  86. [94]

    Jiuzhou Han, Wray Buntine, and Ehsan Shareghi. 2024. Towards Uncertainty-Aware Language Agent. In Findings of the Association for Computational Linguistics ACL 2024 . 6662–6685

  87. [95]

    Tessa Han, Aounon Kumar, Chirag Agarwal, and Himabindu Lakkaraju. 2024. Medsafetybench: Evaluating and improving the medical safety of large language models. Advances in Neural Information Processing Systems 37 (2024), 33423–33454

  88. [96]

    Xuehai He, Yichen Zhang, Luntian Mou, Eric Xing, and Pengtao Xie. 2020. Pathvqa: 30000+ questions for medical visual question answering. arXiv preprint arXiv:2003.10286 (2020)

  89. [97]

    Xinyi He, Jiaru Zou, Yun Lin, Mengyu Zhou, Shi Han, Zejian Yuan, and Dongmei Zhang. 2024. CoCoST: Automatic Complex Code Generation with Online Searching and Correctness Testing. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing . 19433–19451

  90. [98]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring Massive Multitask Language Understanding. In International Conference on Learning Representations

  91. [99]

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring Mathematical Problem Solving With the MATH Dataset. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks ...

  92. [100]

    Daniel Scott Himmelstein, Antoine Lizee, Christine Hessler, Leo Brueggeman, Sabrina L Chen, Dexter Hadley, Ari Green, Pouya Khankhanian, and Sergio E Baranzini. 2017. Systematic integration of biomedical knowledge prioritizes drugs for repurposing. elife 6 (2017), e26726

  93. [101]

    Sebastian Hofstätter, Jiecao Chen, Karthik Raman, and Hamed Zamani. 2023. Fid-light: Efficient and effective retrieval- augmented text generation. InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1437–1447

  94. [102]

    Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, et al. 2024. MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. In The Twelfth International Conference on Learning Rep...

  95. [103]

    SU Hongjin, Ruoxi Sun, Jinsung Yoon, Pengcheng Yin, Tao Yu, and Sercan O Arik. 2025. Learn-by-interact: A Data- Centric Framework For Self-Adaptive Agents in Realistic Environments. In The Thirteenth International Conference on Learning Representations

  96. [104]

    Max A Horlbeck, Albert Xu, Min Wang, Neal K Bennett, Chong Y Park, Derek Bogdanoff, Britt Adamson, Eric D Chow, Martin Kampmann, Tim R Peterson, et al. 2018. Mapping the genetic landscape of human cells. Cell 174, 4 (2018), 953–967. J. ACM, Vol. 37, No. 4, Article 111. Publica...

  97. [105]

    Xinyi Hou, Yanjie Zhao, Shenao Wang, and Haoyu Wang. 2025. Model context protocol (mcp): Landscape, security threats, and future research directions. arXiv preprint arXiv:2503.23278 (2025)

  98. [106]

    Brian Hu, Bill Ray, Alice Leung, Amy Summerville, David Joy, Christopher Funk, and Arslan Basharat. 2024. Language Models are Alignable Decision-Makers: Dataset and Application to the Medical Triage Domain. In Proceedings of the 2024 Conference of the North American Chapter of...

  99. [107]

    Zhiyuan Hu, Chumin Liu, Xidong Feng, Yilun Zhao, See-Kiong Ng, Anh Tuan Luu, Junxian He, Pang Wei Koh, and Bryan Hooi. 2024. Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in Large Language Models. In ICLR 2024 Workshop on Large Language Model...

  100. [108]

    Chenghua Huang, Shisong Chen, Zhixu Li, Jianfeng Qu, Yanghua Xiao, Jiaxin Liu, and Zhigang Chen. 2024. Geoagent: To empower llms using geospatial tools for address standardization. In Findings of the Association for Computational Linguistics ACL 2024. 6048–6063

  101. [109]

    Daoyi Huang, Jianping Jiang, Tingting Zhao, Shengnan Wu, Pin Li, Yongfen Lyu, Jincai Feng, Mingyue Wei, Zhixing Zhu, Jianlei Gu, et al. 2023. diseaseGPS: auxiliary diagnostic system for genetic disorders based on genotype and phenotype. Bioinformatics 39, 9 (2023), btad517

  102. [110]

    Dong Huang, Jie M Zhang, Michael Luck, Qingwen Bu, Yuhao Qing, and Heming Cui. 2023. Agentcoder: Multi-agent- based code generation with iterative testing and optimisation. arXiv preprint arXiv:2312.13010 (2023)

  103. [111]

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2025. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Info...

  104. [112]

    Xu Huang, Weiwen Liu, Xiaolong Chen, Xingmei Wang, Hao Wang, Defu Lian, Yasheng Wang, Ruiming Tang, and Enhong Chen. 2024. Understanding the planning of LLM agents: A survey. arXiv preprint arXiv:2402.02716 (2024)

  105. [113]

    Yangyu Huang, Tianyi Gao, Haoran Xu, Qihao Zhao, Yang Song, Zhipeng Gui, Tengchao Lv, Hao Chen, Lei Cui, Scarlett Li, et al. 2025. Peace: Empowering geologic map holistic understanding with mllms. In Proceedings of the Computer Vision and Pattern Recognition Conference . 3899–3908

  106. [114]

    Yoshitaka Inoue, Tianci Song, Xinling Wang, Augustin Luna, and Tianfan Fu. 2025. DrugAgent: Multi-Agent Large Language Model-Based Reasoning for Drug-Target Interaction Prediction. In ICLR Workshop on Machine Learning for Genomics Explorations

  107. [115]

    John J Irwin, Teague Sterling, Michael M Mysinger, Erin S Bolstad, and Ryan G Coleman. 2012. ZINC: a free tool to discover chemistry for biology. Journal of chemical information and modeling 52, 7 (2012), 1757–1768

  108. [116]

    Shoichi Ishida, Tomohiro Sato, Teruki Honma, and Kei Terayama. 2025. Large language models open new way of AI-assisted molecule design for chemists. Journal of Cheminformatics 17, 1 (2025), 36

  109. [117]

    Md Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez. 2024. MapCoder: Multi-Agent Code Generation for Competitive Problem Solving. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 4912–4944

  110. [118]

    Juyong Jiang, Fan Wang, Jiasi Shen, Sungju Kim, and Sunghun Kim. 2024. A survey on large language models for code generation. arXiv preprint arXiv:2406.00515 (2024)

  111. [119]

    Xue Jiang, Yihong Dong, Lecheng Wang, Zheng Fang, Qiwei Shang, Ge Li, Zhi Jin, and Wenpin Jiao. 2024. Self- planning code generation with large language models. ACM Transactions on Software Engineering and Methodology 33, 7 (2024), 1–30

  112. [120]

    Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. 2024. SWE-bench: Can Language Models Resolve Real-world Github Issues?. In The Twelfth International Conference on Learning Representations

  113. [121]

    Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences 11, 14 (2021), 6421

  114. [122]

    Haolin Jin, Linghan Huang, Haipeng Cai, Jun Yan, Bo Li, and Huaming Chen. 2024. From llms to llm-based agents for software engineering: A survey of current, challenges and future. arXiv preprint arXiv:2408.02479 (2024)

  115. [123]

    Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. 2019. PubMedQA: A Dataset for Biomedical Research Question Answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on...

  116. [124]

    Wengong Jin, Connor Coley, Regina Barzilay, and Tommi Jaakkola. 2017. Predicting organic reaction outcomes with weisfeiler-lehman network. Advances in neural information processing systems 30 (2017)

  117. [125]

    Alistair EW Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J Pollard, Sicheng Hao, Benjamin Moody, Brian Gow, et al. 2023. MIMIC-IV, a freely accessible electronic health record dataset. Scientific data 10, 1 (2023), 1. J. ACM, Vol. 37, No. ...

  118. [126]

    Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . ...

  119. [127]

    René Just, Darioush Jalali, and Michael D Ernst. 2014. Defects4J: A database of existing faults to enable controlled testing studies for Java programs. In Proceedings of the 2014 international symposium on software testing and analysis . 437–440

  120. [128]

    Yeonghun Kang and Jihan Kim. 2024. ChatMOF: an artificial intelligence system for predicting and generating metal-organic frameworks using large language models. Nature communications 15, 1 (2024), 4705

  121. [129]

    Zixuan Ke, Fangkai Jiao, Yifei Ming, Xuan-Phi Nguyen, Austin Xu, Do Xuan Long, Minzhi Li, Chengwei Qin, Peifeng Wang, Silvio Savarese, et al. 2025. A survey of frontiers in llm reasoning: Inference scaling, learning to reason, and agentic systems. arXiv preprint arXiv:2504.090...

  122. [130]

    M Keestra et al. 2017. Metacognition and Reflection by Interdisciplinary Experts: Insights from Cognitive Science and Philosophy. Issues in Interdisciplinary Studies 35 (2017)

  123. [131]

    Jaechang Kim, Jinmin Goh, Inseok Hwang, Jaewoong Cho, and Jungseul Ok. 2025. Bridging the Gap between Expert and Language Models: Concept-guided Chess Commentary Generation and Evaluation. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Asso...

  124. [132]

    Junseok Kim, Nakyeong Yang, and Kyomin Jung. 2024. Persona is a Double-edged Sword: Mitigating the Negative Impact of Role-playing Prompts in Zero-shot Reasoning Tasks. arXiv preprint arxiv:2408.08631 (2024)

  125. [133]

    Sehoon Kim, Suhong Moon, Ryan Tabrizi, Nicholas Lee, Michael W Mahoney, Kurt Keutzer, and Amir Gholami. 2024. An LLM compiler for parallel function calling. In Proceedings of the 41st International Conference on Machine Learning . 24370–24391

  126. [134]

    Yubin Kim, Hyewon Jeong, Chanwoo Park, Eugene Park, Haipeng Zhang, Xin Liu, Hyeonhoon Lee, Daniel McDuff, Marzyeh Ghassemi, Cynthia Breazeal, et al. 2025. Tiered Agentic Oversight: A Hierarchical Multi-Agent System for AI Safety in Healthcare. arXiv preprint arXiv:2506.12482 (2025)

  127. [135]

    Yubin Kim, Chanwoo Park, Hyewon Jeong, Yik S Chan, Xuhai Xu, Daniel McDuff, Hyeonhoon Lee, Marzyeh Ghassemi, Cynthia Breazeal, and Hae W Park. 2024. Mdagents: An adaptive collaboration of llms for medical decision-making. Advances in Neural Information Processing Systems 37 (2...

  128. [136]

    Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, Xin Zhou, Enzhi Wang, and Xiaohang Dong

  129. [137]

    Yuqing Kong, Yunqi Li, Yubo Zhang, Zhihuan Huang, and Jinzhao Wu. 2022. Eliciting thinking hierarchy without a prior. Advances in Neural Information Processing Systems 35 (2022), 13329–13341

  130. [138]

    Adarsh Kumarappan, Mo Tiwari, Peiyang Song, Robert Joseph George, Chaowei Xiao, and Anima Anandkumar. 2025. LeanAgent: Lifelong Learning for Formal Theorem Proving. In The Thirteenth International Conference on Learning Representations

  131. [139]

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for C...

  132. [140]

    Alexey Lagunin, Dmitrii Filimonov, Alexey Zakharov, Wei Xie, Ying Huang, Fucheng Zhu, Tianxiang Shen, Jianhua Yao, and Vladimir Poroikov. 2009. Computer-aided prediction of rodent carcinogenicity by PASS and CISOC-PSCT. QSAR & Combinatorial Science 28, 8 (2009), 806–810

  133. [141]

    Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Wen-tau Yih, Daniel Fried, Sida Wang, and Tao Yu. 2023. DS-1000: A natural and reliable benchmark for data science code generation. InInternational Conference on Machine Learning . PMLR, 18319–18345

  134. [142]

    Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. 2018. A dataset of clinically generated visual questions and answers about radiology images. Scientific data 5, 1 (2018), 1–10

  135. [143]

    Andrew Laverick, Kristen Surrao, Inigo Zubeldia, Boris Bolliet, Miles Cranmer, Antony Lewis, Blake Sherwin, and Julien Lesgourgues. 2024. Multi-Agent System for Cosmological Parameter Analysis. arXiv preprint arXiv:2412.00431 (2024)

  136. [144]

    Chaehong Lee, Varatheepan Paramanayakam, Andreas Karatzas, Yanan Jian, Michael Fore, Heming Liao, Fuxun Yu, Ruopu Li, Iraklis Anagnostopoulos, and Dimitrios Stamoulis. 2025. Multi-Agent Geospatial Copilots for Remote Sensing Workflows. arXiv preprint arXiv:2501.16254 (2025)

  137. [145]

    Namkyeong Lee, Edward De Brouwer, Ehsan Hajiramezanali, Tommaso Biancalani, Chanyoung Park, and Gabriele Scalia. 2025. RAG-Enhanced Collaborative LLM Agents for Drug Discovery. In ICLR Workshop on Machine Learning for Genomics Explorations. J. ACM, Vol. 37, No. 4, Article 111....

  138. [146]

    Sunjae Lee, Junyoung Choi, Jungjae Lee, Munim Hasan Wasi, Hojun Choi, Steve Ko, Sangeun Oh, and Insik Shin

  139. [147]

    Mosh Levy, Alon Jacoby, and Yoav Goldberg. 2024. Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 15339–15353

  140. [148]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing s...

  141. [149]

    In Proceedings of the 30th Annual International Conference on Mobile Computing and Networking

    Mobilegpt: Augmenting llm with human-like app memory for mobile task automation. In Proceedings of the 30th Annual International Conference on Mobile Computing and Networking . 1119–1133

  142. [150]

    Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. 2023. Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems 36 (20...

  143. [151]

    Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023. Camel: Communicative agents for" mind" exploration of large language model society. Advances in Neural Information Processing Systems 36 (2023), 51991–52008

  144. [152]

    Binxu Li, Tiankai Yan, Yuanting Pan, Jie Luo, Ruiyang Ji, Jiayuan Ding, Zhe Xu, Shilong Liu, Haoyu Dong, Zihao Lin, et al. 2024. MMedAgent: Learning to Use Medical Tools with Multi-modal Agent. In Findings of the Association for Computational Linguistics: EMNLP 2024 . 8745–8760

  145. [153]

    Nian Li, Chen Gao, Mingyu Li, Yong Li, and Qingmin Liao. 2024. EconAgent: Large Language Model-Empowered Agents for Simulating Macroeconomic Activities. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 15523–15536

  146. [154]

    Xinzhe Li. 2025. A review of prominent paradigms for llm-based agents: Tool use, planning (including rag), and feedback learning. In Proceedings of the 31st International Conference on Computational Linguistics . 9760–9779

  147. [155]

    Junkai Li, Yunghwei Lai, Weitao Li, Jingyi Ren, Meng Zhang, Xinhui Kang, Siyu Wang, Peng Li, Ya-Qin Zhang, Weizhi Ma, et al. 2024. Agent hospital: A simulacrum of hospital with evolvable medical agents. arXiv preprint arXiv:2405.02957 (2024)

  148. [156]

    Xiaonan Li and Xipeng Qiu. 2023. Finding Support Examples for In-Context Learning. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 6219–6235

  149. [157]

    Yuan Li, Yixuan Zhang, and Lichao Sun. 2023. Metaagents: Simulating interactions of human behaviors for llm-based task-oriented coordination via collaborative generative agents. arXiv preprint arXiv:2310.06500 (2023)

  150. [158]

    Zhang, Yiling Lou, Tianlin Li, Weisong Sun, Yang Liu, and Xuanzhe Liu

    Xinyue Li, Zhenpeng Chen, Jie M. Zhang, Yiling Lou, Tianlin Li, Weisong Sun, Yang Liu, and Xuanzhe Liu. 2024. Benchmarking Bias in Large Language Models during Role-Playing. arXiv preprint arxiv:2411.00585 (2024)

  151. [159]

    Zhenlong Li, Huan Ning, Song Gao, Krzysztof Janowicz, Wenwen Li, Samantha T Arundel, Chaowei Yang, Budhendra Bhaduri, Shaowen Wang, A Zhu, et al. 2025. Giscience in the era of artificial intelligence: A research agenda towards autonomous gis. arXiv preprint arXiv:2503.23633 (2025)

  152. [160]

    Zhucong Li, Jin Xiao, Bowei Zhang, Zhijian Zhou, Qianyu He, Fenglei Cao, Jiaqing Liang, and Yuan Qi. 2025. ChemHTS: Hierarchical Tool Stacking for Enhancing Chemical Agents. arXiv preprint arXiv:2502.14327 (2025)

  153. [161]

    Zhenlong Li and Huan Ning. 2023. Autonomous GIS: the next-generation AI-powered GIS. International Journal of Digital Earth 16, 2 (2023), 4668–4686

  154. [162]

    Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, and Zhaopeng Tu. 2024. Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate. In Proceedings of the 2024 Conference on Empirical Methods in Natural Langua...

  155. [163]

    Xun Liang, Jiawei Yang, Yezhaohui Wang, Chen Tang, Zifan Zheng, Shichao Song, Zehao Lin, Yebin Yang, Simin Niu, Hanyu Wang, et al. 2025. Surveyx: Academic survey automation via large language models. arXiv preprint arXiv:2502.14776 (2025)

  156. [164]

    Kaiqu Liang, Zixu Zhang, and Jaime F Fisac. 2024. Introspective Planning: Aligning Robots’ Uncertainty with Inherent Task Ambiguity. Advances in Neural Information Processing Systems 37 (2024), 71998–72031

  157. [165]

    Christopher A Lipinski, Franco Lombardo, Beryl W Dominy, and Paul J Feeney. 1997. Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings. Advanced drug delivery reviews 23, 1-3 (1997), 3–25

  158. [166]

    Bang Liu, Xinfeng Li, Jiayi Zhang, Jinlin Wang, Tanjin He, Sirui Hong, Hongzhang Liu, Shaokun Zhang, Kaitao Song, Kunlun Zhu, et al. 2025. Advances and challenges in foundation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems. arXiv pre...

  159. [167]

    Zhehui Liao, Maria Antoniak, Inyoung Cheong, Evie Yu-Yen Cheng, Ai-Heng Lee, Kyle Lo, Joseph Chee Chang, and Amy X Zhang. 2024. LLMs as Research Tools: A Large Scale Survey of Researchers’ Usage and Perceptions. arXiv preprint arXiv:2411.05025 (2024)

  160. [168]

    Hao Liu, Zi-Yi Dou, Yixin Wang, Nanyun Peng, and Yisong Yue. 2024. Uncertainty Calibration for Tool-Using Language Agents. In Findings of the Association for Computational Linguistics: EMNLP 2024 . 16781–16805

  161. [169]

    Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023. Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation. Advances in Neural Information Processing Systems 36 (2023), 21558–21572

  162. [170]

    Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma, Yan Yang, and Xiao-Ming Wu. 2021. Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering. In 2021 IEEE 18th international symposium on biomedical imaging (ISBI). IEEE, 1650–1654

  163. [171]

    Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. Comput. Surveys 55, 9 (2023), 1–35

  164. [172]

    Shengchao Liu, Jiongxiao Wang, Yijin Yang, Chengpeng Wang, Ling Liu, Hongyu Guo, and Chaowei Xiao. 2024. Conversational drug editing using retrieval and domain feedback. In The twelfth international conference on learning representations

  165. [173]

    Pengfei Liu, Jun Tao, and Zhixiang Ren. 2025. A quantitative analysis of knowledge-learning preferences in large language models in molecular science. Nature Machine Intelligence 7, 2 (2025), 315–327

  166. [174]

    Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al. 2024. AgentBench: Evaluating LLMs as Agents. In ICLR

  167. [175]

    Yang Liu, Weixing Chen, Yongjie Bai, Xiaodan Liang, Guanbin Li, Wen Gao, and Liang Lin. 2025. Aligning cyber space with physical world: A comprehensive survey on embodied ai. IEEE/ASME Transactions on Mechatronics (2025)

  168. [176]

    Wei Liu, Jun Li, Yitao Tang, Yining Zhao, Chaozhong Liu, Meiyi Song, Zhenlin Ju, Shwetha V Kumar, Yiling Lu, Rehan Akbani, et al. 2025. DrBioRight 2.0: an LLM-powered bioinformatics chatbot for large-scale cancer functional proteomics analysis. Nature communications 16, 1 (2025), 2256

  169. [177]

    Zhengyao Liu, Yunlong Ma, Jingxuan Xu, Junchen Ai, Xiang Gao, Hailong Sun, and Abhik Roychoudhury. 2025. Agent That Debugs: Dynamic State-Guided Vulnerability Repair. arXiv preprint arXiv:2504.07634 (2025)

  170. [178]

    Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. 2024. The ai scientist: Towards fully automated open-ended scientific discovery. arXiv preprint arXiv:2408.06292 (2024)

  171. [179]

    Yungeng Liu, Zan Chen, Yu Guang Wang, and Yiqing Shen. 2024. Toursynbio-search: A large language model driven agent framework for unified search method for protein engineering. In 2024 IEEE International Conference on Bioinformatics and Biomedicine (BIBM) . IEEE, 5395–5400

  172. [180]

    Yi Luo, Linghang Shi, Yihao Li, Aobo Zhuang, Yeyun Gong, Ling Liu, and Chen Lin. 2025. From intention to implementation: automating biomedical research via LLMs. Science China Information Sciences 68, 7 (2025), 1–18

  173. [181]

    Bohan Lyu, Xin Cong, Heyang Yu, Pan Yang, Yujia Qin, Yining Ye, Yaxi Lu, Zhong Zhang, Yukun Yan, Yankai Lin, Zhiyuan Liu, and Maosong Sun. 2023. GitAgent: Facilitating Autonomous Agent with GitHub by Tool Extension. arXiv preprint arxiv:2312.17294 (2023)

  174. [182]

    Qinyu Luo, Yining Ye, Shihao Liang, Zhong Zhang, Yujia Qin, Yaxi Lu, Yesai Wu, Xin Cong, Yankai Lin, Yingli Zhang, et al. 2024. RepoAgent: An LLM-Powered Open-Source Framework for Repository-level Code Documentation Generation. In Proceedings of the 2024 Conference on Empirica...

  175. [183]

    Aman Madaan, Niket Tandon, Prakhar Gupta, Kevin Hall, Luyu Gao, Rohan Majumder, Julian McAuley, Srijan Narayan, and Sean Welleck. 2023. Self-refine: Iterative refinement with self-feedback. In Advances in Neural Information Processing Systems, Vol. 36

  176. [184]

    Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2023. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems 36 (2023), 46534–46594

  177. [185]

    Weiyu Ma, Qirui Mi, Yongcheng Zeng, Xue Yan, Runji Lin, Yuqiao Wu, Jun Wang, and Haifeng Zhang. 2024. Large language models play starcraft ii: Benchmarks and a chain of summarization approach.Advances in Neural Information Processing Systems 37 (2024), 133386–133442

  178. [186]

    Tula Masterman, Sandi Besen, Mason Sawtell, and Alex Chao. 2024. The landscape of emerging ai agent architectures for reasoning, planning, and tool calling: A survey. arXiv preprint arXiv:2404.11584 (2024)

  179. [187]

    Andrew D McNaughton, Gautham Krishna Sankar Ramalaxmi, Agustin Kruel, Carter R Knutson, Rohith A Varikoti, and Neeraj Kumar. 2024. Cactus: Chemistry agent connecting tool usage to science. ACS omega 9, 46 (2024), 46563–46573

  180. [188]

    Agile Manifesto. 2001. Manifesto for Agile Software Development. http://www. agilemanifesto. org/ (2001)

  181. [189]

    Shikib Mehri and Maxine Eskenazi. 2020. Unsupervised Evaluation of Interactive Dialog with DialoGPT. InProceedings of the 21th Annual Meeting of the Special Interest Group on Discourse and Dialogue . 225–235

  182. [190]

    Lingrui Mei, Jiayu Yao, Yuyao Ge, Yiwei Wang, Baolong Bi, Yujun Cai, Jiazhi Liu, Mingyu Li, Zhong-Zhi Li, Duzhen Zhang, et al. 2025. A Survey of Context Engineering for Large Language Models. arXiv preprint arXiv:2507.13334 (2025)

  183. [191]

    Nikita Mehandru, Amanda K Hall, Olesya Melnichenko, Yulia Dubinina, Daniel Tsirulnikov, David Bamman, Ahmed Alaa, Scott Saponas, and Venkat S Malladi. 2025. BioAgents: Democratizing bioinformatics analysis with multi-agent systems. arXiv preprint arXiv:2501.06314 (2025). J. AC...

  184. [192]

    Marvin Minsky. 1986. Society of mind. Simon and Schuster

  185. [193]

    Lluis Morey, Luigi Aloia, Luca Cozzuto, Salvador Aznar Benitah, and Luciano Di Croce. 2013. RYBP and Cbx7 define specific biological functions of polycomb complexes in mouse embryonic stem cells. Cell reports 3, 1 (2013), 60–69

  186. [194]

    Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing...

  187. [195]

    Xinyi Mou, Zhongyu Wei, and Xuan-Jing Huang. 2024. Unveiling the Truth and Facilitating Change: Towards Agent-based Large-scale Social Movement Simulation. In Findings of the Association for Computational Linguistics ACL 2024. 4789–4809

  188. [196]

    Chunyan Mu, Muhammad Najib, and Nir Oren. 2025. Responsibility-aware Strategic Reasoning in Probabilistic Multi-Agent Systems. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. 23258–23266

  189. [197]

    Adam Moss. 2025. The AI Cosmologist I: An Agentic System for Automated Data Analysis. arXiv preprint arXiv:2504.03424 (2025)

  190. [198]

    Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian. 2023. A comprehensive overview of large language models. ACM Transactions on Intelligent Systems and Technology (2023)

  191. [199]

    Subash Neupane, Sudip Mittal, and Shahram Rahimi. 2025. Towards a hipaa compliant agentic ai system in healthcare. arXiv preprint arXiv:2504.17669 (2025)

  192. [200]

    Sriraam Natarajan, Saurabh Mathur, Sahil Sidheekh, Wolfgang Stammer, and Kristian Kersting. 2025. Human-in-the- loop or AI-in-the-loop? Automate or Collaborate?. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. 28594–28600

  193. [201]

    David Ochoa, Andrew Hercules, Miguel Carmona, Daniel Suveges, Jarrod Baker, Cinzia Malangone, Irene Lopez, Alfredo Miranda, Carlos Cruz-Castillo, Luca Fumis, et al. 2023. The next-generation Open Targets Platform: reimagined, redesigned, rebuilt. Nucleic acids research 51, D1 ...

  194. [202]

    Timothy J O’Donnell, Alex Rubinsteyn, and Uri Laserson. 2020. MHCflurry 2.0: improved pan-allele prediction of MHC class I-presented peptides by incorporating antigen processing. Cell systems 11, 1 (2020), 42–48

  195. [203]

    Huan Ning, Zhenlong Li, Temitope Akinboyewa, and M Naser Lessani. 2025. An autonomous GIS agent framework for geospatial data retrieval. International Journal of Digital Earth 18, 1 (2025), 2458688

  196. [204]

    Aosong Pan, Sameen Al-Azani, Yifei An, Zhipeng Jiang, Wen-Bin Wang, Xipeng Wan, and Man Lan. 2023. LogicLM: Empowering Large Language Models with Tool-Enhanced Logic-Evolving Reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 8500–8518

  197. [205]

    Melissa Z Pan, Mert Cemri, Lakshya A Agrawal, Shuyi Yang, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Kannan Ramchandran, Dan Klein, et al. 2025. Why do multiagent systems fail?. InICLR 2025 Workshop on Building Trust in Language Models and Applications

  198. [206]

    Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Conference on health, inference, and learning . PMLR, 248–260

  199. [207]

    Dmitrii Pantiukhin, Boris Shapkin, Ivan Kuznetsov, Antonia Anna Jost, and Nikolay Koldunov. 2025. Accelerating Earth Science Discovery via Multi-Agent LLM Systems. arXiv preprint arXiv:2503.05854 (2025)

  200. [208]

    Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology . 1–22

  201. [209]

    Himanshu Gautam Pandey, Akhil Amod, and Shivang Kumar. 2024. Advancing Healthcare Automation: Multi-Agent System for Medical Necessity Justification. In Proceedings of the 23rd Workshop on Biomedical Natural Language Processing. 39–49

  202. [210]

    Jean Piaget. 2013. The construction of reality in the child . Routledge

  203. [211]

    Giorgio Piatti, Zhijing Jin, Max Kleiman-Weiner, Bernhard Schölkopf, Mrinmaya Sachan, and Rada Mihalcea. 2024. Cooperate or collapse: Emergence of sustainable cooperation in a society of llm agents.Advances in Neural Information Processing Systems 37 (2024), 111715–111759

  204. [212]

    Anthony D Pellegrini. 2009. The role of play in human development . Oxford University Press

  205. [213]

    Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, et al. 2024. ChatDev: Communicative Agents for Software Development. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vo...

  206. [214]

    Boyu Qiao, Kun Li, Wei Zhou, Shilong Li, Qianqian Lu, and Songlin Hu. 2025. BotSim: LLM-Powered Malicious Social Botnet Simulation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. 14377–14385

  207. [215]

    Kevin Pu, KJ Kevin Feng, Tovi Grossman, Tom Hope, Bhavana Dalvi Mishra, Matt Latzke, Jonathan Bragg, Joseph Chee Chang, and Pao Siangliulue. 2025. Ideasynth: Iterative research idea development through evolving and composing J. ACM, Vol. 37, No. 4, Article 111. Publication dat...

  208. [216]

    Yihao Qin, Shangwen Wang, Yiling Lou, Jinhao Dong, Kaixin Wang, Xiaoling Li, and Xiaoguang Mao. 2024. AgentFL: Scaling LLM-based Fault Localization to Project-Level Context. CoRR (2024)

  209. [217]

    Haoxuan Qu, Xiaofei Hui, Yujun Cai, and Jun Liu. 2023. LMC: large model collaboration with cross-assessment for training-free open-set object recognition. In Proceedings of the 37th International Conference on Neural Information Processing Systems. Red Hook, NY, USA, Article 2...

  210. [218]

    Yujia Qin, Shengding Hu, Yankai Lin, Weize Chen, Ning Ding, Ganqu Cui, Zheni Zeng, Xuanhe Zhou, Yufei Huang, Chaojun Xiao, et al. 2024. Tool learning with foundation models. Comput. Surveys 57, 4 (2024), 1–40

  211. [219]

    Roshan Rao, Nicholas Bhattacharya, Neil Thomas, Yan Duan, Peter Chen, John Canny, Pieter Abbeel, and Yun Song

  212. [220]

    Suhas SP Rao, Miriam H Huntley, Neva C Durand, Elena K Stamenova, Ivan D Bochkov, James T Robinson, Adrian L Sanborn, Ido Machol, Arina D Omer, Eric S Lander, et al. 2014. A 3D map of the human genome at kilobase resolution reveals principles of chromatin looping. Cell 159, 7 ...

  213. [221]

    Yuanhao Qu, Kaixuan Huang, Ming Yin, Kanghong Zhan, Dyllan Liu, Di Yin, Henry C Cousins, William A Johnson, Xiaotong Wang, Mihir Shah, et al. 2025. CRISPR-GPT for agentic automation of gene-editing experiments. Nature Biomedical Engineering (2025), 1–14

  214. [222]

    Shuo Ren, Pu Jian, Zhenjiang Ren, Chunlin Leng, Can Xie, and Jiajun Zhang. 2025. Towards scientific intelligence: A survey of llm-based scientific agents. arXiv preprint arXiv:2503.24047 (2025)

  215. [223]

    Pat Rondon, Renyao Wei, José Cambronero, Jürgen Cito, Aaron Sun, Siddhant Sanyam, Michele Tufano, and Satish Chandra. 2025. Evaluating Agent-based Program Repair at Google. CoRR (2025)

  216. [224]

    Yusuf H Roohani, Andrew H Lee, Qian Huang, Jian Vora, Zachary Steinhart, Kexin Huang, Alexander Marson, Percy Liang, and Jure Leskovec. 2025. BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments. In The Thirteenth International Conference on Learning R...

  217. [225]

    Ruiyang Ren, Peng Qiu, Yingqi Qu, Jing Liu, Wayne Xin Zhao, Hua Wu, Ji-Rong Wen, and Haifeng Wang. 2024. BASES: Large-scale Web Search User Simulation with Large Language Model based Agents. In Findings of the Association for Computational Linguistics: EMNLP 2024 . 902–917

  218. [226]

    Yixiang Ruan, Chenyin Lu, Ning Xu, Yuchen He, Yixin Chen, Jian Zhang, Jun Xuan, Jianzhang Pan, Qun Fang, Hanyu Gao, et al. 2024. An automatic end-to-end chemical synthesis development platform powered by large language models. Nature communications 15, 1 (2024), 10160

  219. [227]

    Stuart J Russell and Peter Norvig. 2016. Artificial intelligence: a modern approach . pearson

  220. [228]

    Daniel Saeedi, Denise K Buckner, Jose C Aponte, and Amirali Aghazadeh. 2025. AstroAgents: A Multi-Agent AI for Hypothesis Generation from Mass Spectrometry Data. In Towards Agentic AI for Science: Hypothesis Generation, Comprehension, Quantification, and Validation

  221. [229]

    Giulio Rossetti, Massimo Stella, Rémy Cazabet, Katherine Abramski, Erica Cau, Salvatore Citraro, Andrea Failla, Riccardo Improta, Virginia Morini, and Valentina Pansanella. 2024. Y social: an llm-powered social media digital twin. arXiv preprint arXiv:2408.00818 (2024)

  222. [230]

    Carlos G Sanchez, Christopher M Acker, Audrey Gray, Malini Varadarajan, Cheng Song, Nadire R Cochran, Steven Paula, Alicia Lindeman, Shaojian An, Gregory McAllister, et al. 2021. Genome-wide CRISPR screen identifies protein pathways modulating tau protein levels in neurons. Co...

  223. [231]

    Samantha G Scharenberg, Wentao Dong, Ali Ghoochani, Kwamina Nyame, Roni Levin-Konigsberg, Aswini R Krishnan, Eshaan S Rawat, Kaitlyn Spees, Michael C Bassik, and Monther Abu-Remaileh. 2023. An SPNS1-dependent lysosomal lipid transport pathway that enables cell survival under c...

  224. [232]

    Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems 36 (2023), 68...

  225. [233]

    Liane Salewski, Arian Safavi, and R. Groh. 2024. Can LLMs Learn to Reason from Role-Playing?. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies

  226. [234]

    Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Michael Moor, Zicheng Liu, and Emad Barsoum. 2025. Agent laboratory: Using llm agents as research assistants. arXiv preprint arXiv:2501.04227 (2025)

  227. [235]

    Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Pontes Reis, Jeffrey Jopling, and Michael Moor. 2024. AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments. CoRR (2024)

  228. [236]

    Ralf Schmidt, Zachary Steinhart, Madeline Layeghi, Jacob W Freimer, Raymund Bueno, Vinh Q Nguyen, Franziska Blaeschke, Chun Jimmie Ye, and Alexander Marson. 2022. CRISPR activation and interference screens decode stimulation responses in primary human T cells. Science 375, 658...

  229. [237]

    Samuel Schmidgall and Michael Moor. 2025. Agentrxiv: Towards collaborative autonomous research. arXiv preprint arXiv:2503.18102 (2025). J. ACM, Vol. 37, No. 4, Article 111. Publication date: August 2025. 111:46 Zhao et al

  230. [238]

    Minju Seo, Jinheon Baek, Seongyun Lee, and Sung Ju Hwang. 2025. Paper2code: Automating code generation from scientific papers in machine learning. arXiv preprint arXiv:2504.17192 (2025)

  231. [239]

    Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023. Role play with large language models. Nature 623, 7987 (2023), 493–498

  232. [240]

    Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, et al. 2023. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Learning Representations

  233. [241]

    Sander Schulhoff, Michael Ilie, Nishant Balepur, Konstantine Kahadze, Amanda Liu, Chenglei Si, Yinheng Li, Aayush Gupta, HyoJung Han, Sevien Schulhoff, et al. 2024. The prompt report: a systematic survey of prompt engineering techniques. arXiv preprint arXiv:2406.06608 (2024)

  234. [242]

    Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Cote, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. 2021. ALFWorld: Aligning Text and Embodied Environments for Interactive Learning. In International Conference on Learning Representations

  235. [243]

    Simranjit Singh, Michael Fore, and Dimitrios Stamoulis. 2024. Geollm-engine: A realistic environment for building geospatial copilots. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 585–594

  236. [244]

    Simranjit Singh, Andreas Karatzas, Michael Fore, Iraklis Anagnostopoulos, and Dimitrios Stamoulis. 2024. An llm-tool compiler for fused parallel function calling. arXiv preprint arXiv:2405.17438 (2024)

  237. [245]

    Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems 36 (2023), 8634–8652

  238. [246]

    Isabella Stewart and Markus J Buehler. 2025. Molecular analysis and design using generative artificial intelligence via multi-agent modeling. Molecular Systems Design & Engineering 10, 4 (2025), 314–337

  239. [247]

    Buxin Su, Jiayao Zhang, Natalie Collina, Yuling Yan, Didong Li, Kyunghyun Cho, Jianqing Fan, Aaron Roth, and Weijie Su. 2025. The ICML 2023 ranking experiment: Examining author self-assessment in ML/AI peer review. J. Amer. Statist. Assoc. just-accepted (2025), 1–16

  240. [248]

    Haoyang Su, Renqi Chen, Shixiang Tang, Zhenfei Yin, Xinzhe Zheng, Jinzhe Li, Biqing Qi, Qi Wu, Hui Li, Wanli Ouyang, Philip Torr, Bowen Zhou, and Nanqing Dong. 2025. Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System. In Proce...

  241. [249]

    Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adri Garriga-Alonso, et al . 2023. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions ...

  242. [250]

    Jiankai Sun, Chuanyang Zheng, Enze Xie, Zhengying Liu, Ruihang Chu, Jianing Qiu, Jiaqi Xu, Mingyu Ding, Hongyang Li, Mengzhe Geng, et al. 2025. A survey of reasoning with foundation models: Concepts, methodologies, and outlook. Comput. Surveys 57, 11 (2025), 1–43

  243. [251]

    Damian Szklarczyk, Alberto Santos, Christian Von Mering, Lars Juhl Jensen, Peer Bork, and Michael Kuhn. 2016. STITCH 5: augmenting protein–chemical interaction networks with tissue and affinity data. Nucleic acids research 44, D1 (2016), D380–D384

  244. [252]

    Jiakai Tang, Heyang Gao, Xuchen Pan, Lei Wang, Haoran Tan, Dawei Gao, Yushuo Chen, Xu Chen, Yankai Lin, Yaliang Li, et al. 2025. GenSim: A General Social Simulation Platform with Large Language Model based Agents. In Proceedings of the 2025 Conference of the Nations of the Ame...

  245. [253]

    Houcheng Su, Weicai Long, and Yanlin Zhang. 2025. BioMaster: Multi-agent System for Automated Bioinformatics Analysis Workflow. bioRxiv (2025), 2025–01

  246. [254]

    Xiangru Tang, Anni Zou, Zhuosheng Zhang, Ziming Li, Yilun Zhao, Xingyao Zhang, Arman Cohan, and Mark Gerstein. 2024. MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning. In Findings of the Association for Computational Linguistics ACL 2024 . 599–6...

  247. [255]

    Yang Tang, Chaoqiang Zhao, Jianrui Wang, Chongzhen Zhang, Qiyu Sun, Wei Xing Zheng, Wenli Du, Feng Qian, and Juergen Kurths. 2022. Perception and navigation in autonomous systems in the era of learning: A survey. IEEE Transactions on Neural Networks and Learning Systems 34, 12...

  248. [256]

    Wei Tao, Yucheng Zhou, Yanlin Wang, Wenqiang Zhang, Hongyu Zhang, and Yu Cheng. 2024. Magis: Llm-based multi-agent framework for github issue resolution. Advances in Neural Information Processing Systems 37 (2024), 51963–51993

  249. [257]

    Jiabin Tang, Lianghao Xia, Zhonghang Li, and Chao Huang. 2025. AI-Researcher: Autonomous Scientific Innovation. arXiv preprint arXiv:2505.18705 (2025)

  250. [258]

    Khanh-Tung Tran, Dung Dao, Minh-Duong Nguyen, Quoc-Viet Pham, Barry O’Sullivan, and Hoang D Nguyen. 2025. Multi-agent collaboration mechanisms: A survey of llms. arXiv preprint arXiv:2501.06322 (2025)

  251. [259]

    Oleg Trott and Arthur J Olson. 2010. AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. Journal of computational chemistry 31, 2 (2010), 455–461

  252. [260]

    Tao Tu, Mike Schaekermann, Anil Palepu, Khaled Saab, Jan Freyberg, Ryutaro Tanno, Amy Wang, Brenna Li, Mohamed Amin, Yong Cheng, et al. 2025. Towards conversational diagnostic artificial intelligence. Nature (2025), 1–9

  253. [261]

    Raghav Thind, Youran Sun, Ling Liang, and Haizhao Yang. 2025. OptimAI: Optimization from Natural Language Using LLM-Powered AI Agents. arXiv preprint arXiv:2504.16918 (2025)

  254. [262]

    Han Wang, An Zhang, Nguyen Duy Tai, Jun Sun, Tat-Seng Chua, et al. 2024. Ali-agent: Assessing llms’ alignment with human values via agent-based evaluation. Advances in Neural Information Processing Systems 37 (2024), 99040–99088

  255. [263]

    Junyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang

  256. [264]

    Kun Wang, Guibin Zhang, Zhenhong Zhou, Jiahao Wu, Miao Yu, Shiqian Zhao, Chenlong Yin, Jinhu Fu, Yibo Yan, Hanjun Luo, et al. 2025. A comprehensive survey in llm (-agent) full stack safety: Data, training and deployment. arXiv preprint arXiv:2504.15585 (2025)

  257. [265]

    Hanbin Wang, Zhenghao Liu, Shuo Wang, Ganqu Cui, Ning Ding, Zhiyuan Liu, and Ge Yu. 2024. INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair. In Findings of the Association for Computational Linguistics ACL 2024 . 2081–2107

  258. [266]

    Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. 2024. A survey on large language model based autonomous agents. Frontiers of Computer Science 18, 6 (2024), 186345

  259. [267]

    Lei Wang, Jingsen Zhang, Hao Yang, Zhi-Yuan Chen, Jiakai Tang, Zeyu Zhang, Xu Chen, Yankai Lin, Hao Sun, Ruihua Song, et al. 2025. User behavior simulation with large language model-based agents. ACM Transactions on Information Systems 43, 2 (2025), 1–37

  260. [268]

    Advances in Neural Information Processing Systems 37 (2024), 2686–2710

    Mobile-agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration. Advances in Neural Information Processing Systems 37 (2024), 2686–2710

  261. [269]

    Ruida Wang, Rui Pan, Yuxin Li, Jipeng Zhang, Yizhen Jia, Shizhe Diao, Renjie Pi, Junjie Hu, and Tong Zhang. 2025. MA-LoT: Multi-Agent Lean-based Long Chain-of-Thought Reasoning enhances Formal Theorem Proving. arXiv e-prints (2025), arXiv–2503

  262. [270]

    Luoqi Wang, Haipeng Li, Linshu Hu, Jiarui Cai, and Zhenhong Du. 2024. Mitigating Interpretation Bias in Rock Records with Large Language Models: Insights from Paleoenvironmental Analysis. arXiv preprint arXiv:2407.09977 (2024)

  263. [271]

    Shuai Wang, Weiwen Liu, Jingxuan Chen, Yuqi Zhou, Weinan Gan, Xingshan Zeng, Yuhan Che, Shuai Yu, Xinlong Hao, Kun Shao, et al. 2024. Gui agents with foundation models: A comprehensive survey. arXiv preprint arXiv:2411.04890 (2024)

  264. [272]

    Shuangquan Wang, Huiyong Sun, Hui Liu, Dan Li, Youyong Li, and Tingjun Hou. 2016. ADMET evaluation in drug discovery. 16. Predicting hERG blockers by combining multiple pharmacophores and machine learning approaches. Molecular pharmaceutics 13, 8 (2016), 2855–2866

  265. [273]

    Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, et al. 2024. Large Language Models are not Fair Evaluators. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: L...

  266. [274]

    Wenxuan Wang, Zizhan Ma, Zheng Wang, Chenghan Wu, Jiaming Ji, Wenting Chen, Xiang Li, and Yixuan Yuan. 2025. A survey of llm-based agents in medicine: How far are we from baymax? arXiv preprint arXiv:2502.11211 (2025)

  267. [275]

    Sihan Wang, Suiyang Jiang, Yibo Gao, Boming Wang, Shangqi Gao, and Xiahai Zhuang. 2025. Empowering Medical Multi-Agents with Clinical Consultation Flow for Dynamic Diagnosis. arXiv preprint arXiv:2503.16547 (2025)

  268. [276]

    Yidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang, Xin Zhang, Zhen Wu, Meishan Zhang, Xinyu Dai, Qingsong Wen, Wei Ye, et al. 2024. Autosurvey: Large language models can automatically write surveys. Advances in neural information processing systems 37 (2024), 115119–115145

  269. [277]

    Yuqi Xie Yunfan Jiang Ajay Mandlekar Chaowei Xiao Yuke Zhu Linxi Fan Wang, Guanzhi and Anima Anandkumar

  270. [278]

    Wenxuan Wang, Xiaoyuan Liu, Kuiyi Gao, Jen-tse Huang, Youliang Yuan, Pinjia He, Shuai Wang, and Zhaopeng Tu

  271. [279]

    Ziyue Wang, Junde Wu, Linghan Cai, Chang Han Low, Xihong Yang, Qiaxuan Li, and Yueming Jin. 2025. MedAgent- Pro: Towards Evidence-Based Multi-Modal Medical Diagnosis via Reasoning Agentic Workflow. arXiv preprint arXiv:2503.18968 (2025)

  272. [280]

    Zhenhailong Wang, Haiyang Xu, Junyang Wang, Xi Zhang, Ming Yan, Ji Zhang, Fei Huang, and Heng Ji. 2025. Mobile-agent-e: Self-evolving mobile assistant for complex tasks. arXiv preprint arXiv:2501.11733 (2025)

  273. [281]

    Xingyao Wang, Boxuan Li, Yufan Song, Frank F Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, et al. 2025. OpenHands: An Open Platform for AI Software Developers as Generalist Agents. In The Thirteenth International Conference on Learning Repr...

  274. [282]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837

  275. [283]

    Jinjie Wei, Dingkang Yang, Yanshu Li, Qingyao Xu, Zhaoyu Chen, Mingcheng Li, Yue Jiang, Xiaolu Hou, and Lihua Zhang. 2024. Medaide: Towards an omni medical aide via specialized llm-based multi-agent collaboration. arXiv preprint arXiv:2410.12532 (2024)

  276. [284]

    Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang. 2025. CycleResearcher: Improving Automated Research via Automated Review. InThe Thirteenth International Conference on Learning Representations

  277. [285]

    Zhenhailong Wang, Shaoguang Mao, Wenshan Wu, Tao Ge, Furu Wei, and Heng Ji. 2024. Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration. In Proceedings of the 2024 Conference of the North American Chap...

  278. [286]

    Michael Wooldridge and Nicholas R Jennings. 1998. Pitfalls of agent-oriented development. In Proceedings of the second international conference on Autonomous agents . 385–391

  279. [287]

    Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, et al. 2024. Autogen: Enabling next-gen LLM applications via multi-agent conversations. In First Conference on Language Modeling

  280. [288]

    Hao Wei, Jianing Qiu, Haibao Yu, and Wu Yuan. 2024. Medco: Medical education copilots based on a multi-agent framework. In European Conference on Computer Vision . Springer, 119–135

  281. [289]

    Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang. 2024. Agentless: Demystifying LLM-based Software Engineering Agents. CoRR (2024)

  282. [290]

    Yijia Xiao, Edward Sun, Di Luo, and Wei Wang. 2025. TradingAgents: Multi-Agents LLM Financial Trading Framework. In The First MARW: Multi-Agent AI in the Real World Workshop at AAAI

  283. [291]

    Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh J Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. 2024. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Proces...

  284. [292]

    David S Wishart, Yannick D Feunang, An C Guo, Elvis J Lo, Ana Marcu, Jason R Grant, Tanvir Sajed, Daniel Johnson, Carin Li, Zinat Sayeeda, et al. 2018. DrugBank 5.0: a major update to the DrugBank database for 2018. Nucleic acids research 46, D1 (2018), D1074–D1082

  285. [293]

    Youjun Xu, Ziwei Dai, Fangjin Chen, Shuaishi Gao, Jianfeng Pei, and Luhua Lai. 2015. Deep learning for drug-induced liver injury. Journal of chemical information and modeling 55, 10 (2015), 2085–2093

  286. [295]

    Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande. 2018. MoleculeNet: a benchmark for molecular machine learning. Chemical science 9, 2 (2018), 513–530

  287. [299]

    Fengli Xu, Qianyue Hao, Zefang Zong, Jingwei Wang, Yunke Zhang, Jingyi Wang, Xiaochong Lan, Jiahui Gong, Tianjian Ouyang, Fanjin Meng, et al. 2025. Towards large reasoning models: A survey of reinforced reasoning with large language models. arXiv preprint arXiv:2501.09686 (2025)

  288. [2019]

    Advances in neural information processing systems 32 (2019)

    Evaluating protein transfer learning with TAPE. Advances in neural information processing systems 32 (2019)

  289. [2020]

    Journal of chemical information and modeling 60, 9 (2020), 4200–4215

    Three-dimensional convolutional neural networks and a cross-docked data set for structure-based drug design. Journal of chemical information and modeling 60, 9 (2020), 4200–4215

  290. [2023]

    Voyager: An Open-Ended Embodied Agent with Large Language Models.Intrinsically-Motivated and Open-Ended Learning Workshop @NeurIPS2023 (2023)

  291. [2024]

    In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)

    Better Zero-Shot Reasoning with Role-Play Prompting. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 4099–4113

  292. [2025]

    CoRR (2025)

    Can’t See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs. CoRR (2025)

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.