Pith. sign in

REVIEW 3 major objections 4 minor 69 references

Agentic systems built on large language models are moving from research prototypes to production-scale deployment, and this tutorial argues that robustness, safety, and reliability — not benchmark scores — are now the binding constraints.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 12:41 UTC pith:77IBTURD

load-bearing objection Solid tutorial map, but the 'production-scale deployment' hook leans on a few shaky citations — treat it as orientation, not evidence. the 3 major comments →

arxiv 2607.19336 v1 pith:77IBTURD submitted 2026-07-21 cs.AI cs.CL

Agents in the Wild: Where Research Meets Deployment

classification cs.AI cs.CL
keywords Agentic systemsLarge language modelsMulti-agent coordinationAutonomous workflowsAgent evaluationDeployment safetyHuman-in-the-loopApplied agents
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

LLM-based agentic systems — architectures that reason, plan, act, and coordinate with tools and other agents — are moving from research prototypes into production-scale use across software engineering, science, and finance, and this tutorial treats that transition as the field's central fact. It argues that deployment raises challenges that static benchmarks miss: hallucination, deadlocks, drift, and cascading errors, which call for design patterns like verification pipelines, fallback mechanisms, and human-in-the-loop supervision. The tutorial grounds these claims in case studies from pharmaceutical discovery and financial systems, and pushes evaluation toward interactive, scenario-driven tests of robustness, recovery, and security. A sympathetic reader takes away a map of what agentic AI needs in the wild: not just better models, but engineering discipline around failure and control.

Core claim

The paper's central claim is that agentic systems have crossed into production-scale deployment across domains, with vendor-reported cases such as a pharmaceutical company running hundreds of agentic applications in R&D and financial systems adopting planner–executor–verifier architectures for trading and analysis. To make such systems work in the wild, the authors argue, evaluation must move from static benchmarks toward dynamic, behavior-centric tests that probe adaptability, recovery from failure, and human-intervention behavior, and deployment must build in verification pipelines, fallback mechanisms, and human-in-the-loop supervision as standard safeguards. The tutorial synthesizes curr

What carries the argument

The tutorial's organizing device is the planner–executor–verifier architecture presented as a common pattern in deployed financial agents, together with a mitigation stack — verification pipelines, fallback mechanisms, and human-in-the-loop supervision — that the authors claim underlies successful agentic deployments. The paper also leans on modular multi-agent orchestration (peer-to-peer and hierarchical topologies) and on a family of scenario-driven, behavior-centric evaluation frameworks that test adaptability, recovery, and intervention under distribution shift and adversarial perturbations. These mechanisms carry the argument that production success is an engineering problem with identi

Load-bearing premise

The tutorial's premise that agentic systems are already running at production scale rests largely on vendor-reported deployment figures — for instance, a pharmaceutical partner's claim of over 750 agentic applications, cited from a commentary source — and if those reports are inflated, the case-study foundation and the 'field in transition' narrative lose their support.

What would settle it

An independent census of enterprise LLM-agent deployments across the cited industries that found most 'agentic applications' to be small-scale pilots requiring human approval for every consequential action would falsify the production-scale claim at the heart of the tutorial.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Evaluation of agentic systems will shift toward interactive, scenario-driven benchmarks that measure adaptability, recovery, and human-intervention behavior under distribution shifts and adversarial perturbations.
  • Production agent architectures will standardize around modular designs with independent verification and fallback paths, with human-in-the-loop supervision as a safety net.
  • Failure modes such as hallucination, deadlocks, drift, and cascading errors will be treated as first-class engineering risks, with explicit mitigation pipelines built into system design.
  • Cross-industry design patterns — planner–executor–verifier in finance, multi-agent literature and experiment pipelines in pharma — will spread as reusable templates for other high-stakes domains.
  • Academic research on agents will increasingly be judged by deployment reliability and safety rather than by benchmark performance alone.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the deployment-centric framing is correct, static benchmark scores will increasingly overstate agent readiness, and we should expect a shift toward continuous in-operation monitoring and red-teaming as the primary assurance tools.
  • The recurring emphasis on verification and fallback suggests that the realistic end-state for high-stakes agents is 'human-on-the-loop' supervision rather than full autonomy — a testable prediction for how regulated industries will actually adopt these systems.
  • The paper presents verification pipelines, fallbacks, and human-in-the-loop oversight as mitigation strategies, but does not quantify their relative contribution; a controlled field experiment varying each safeguard would turn the design-pattern advice into engineering guidance.
  • The deployment numbers cited come largely from vendor commentary; an independent audit of how many reported 'agentic applications' are production-critical versus pilot-stage would separate signal from marketing.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper is a proposal for a half-day KDD 2026 tutorial on agentic systems, authored by researchers from academia and industry. It argues that LLM-based agentic systems are rapidly moving from research prototypes to production-scale deployments, and therefore require attention to robustness, safety, and reliability beyond academic benchmarks. The tutorial outline covers foundations (reasoning, planning, multi-agent coordination, retrieval pipelines, evaluation), applied case studies in pharmaceutical/life sciences and finance, and practical design patterns and mitigation strategies. The submission is a 5-page tutorial description, including author bios and references, rather than a technical research paper.

Significance. If the tutorial is delivered as outlined, it addresses a timely and relevant topic at the intersection of AI research and industrial deployment. The authors are well-positioned to present this material, with strong credentials in both academic research and enterprise AI. The emphasis on evaluation beyond static benchmarks, failure modes (hallucination, deadlocks, drift, cascading errors), and human-in-the-loop mechanisms is valuable for practitioners. However, the tutorial's central premise—that agentic systems are 'rapidly transitioning from research prototypes to production-scale deployments'—currently relies on unverified or secondary sources for its most striking evidence. Strengthening the evidentiary basis of this claim would materially increase the credibility and impact of the tutorial.

major comments (3)
  1. [§4.1 (first paragraph)] The claims that FutureHouse's ether0 and Robin are 'capable of superhuman molecular design' and that agentic systems are 'matching or exceeding human performance' are not supported by the cited references. [29] is a NeurIPS paper on training a chemistry reasoning model; it does not establish 'superhuman' capability in a deployed product. [16] describes a multi-agent research system, not a production deployment. Please either provide primary evidence for these superlative claims or qualify them as 'reported'/'claimed' by the developers. Conflating research-level results with production capability weakens the tutorial's evidentiary foundation.
  2. [§4.1, Moderna statistic] The statement that Moderna's partnership with OpenAI 'report[s] over 750 agentic applications across R&D operations' is cited to [23], a Cureus commentary. A commentary is not a primary deployment report or an audited case study. This statistic is the flagship example of 'production-scale deployments' in the abstract; please trace it to a primary source (e.g., a formal company announcement or third-party audit) or clearly label it as a vendor/press claim. Without such sourcing, the tutorial's motivation rests on a secondary, promotional source.
  3. [Abstract and §1] The opening assertion that agentic systems are 'rapidly transitioning from research prototypes to production-scale deployments' is stated as fact. The tutorial's own examples blur this distinction: many cited systems (e.g., [2], [4], [36], [17]) are research prototypes or limited pilots, and the only explicit count of production-scale use (Moderna) is from a commentary. The tutorial description should either (a) provide at least one independently verifiable production deployment with primary sourcing, or (b) temper the claim to 'reported industrial adoption' and explicitly discuss the gap between research demonstrations and production readiness. As written, the motivation risks presenting vendor claims as established fact.
minor comments (4)
  1. [§4.1, last paragraph] The sentence 'Biomedical researchers envision AI scientists as collaborative and critically reasoning partners...' is a near-duplicate of the previous sentence citing Harvard Medical School [13]. Remove one occurrence.
  2. [Section heading] The heading '5 Tutors Bios' should be 'Tutor Bios' or 'Tutors' Biographies'.
  3. [Abstract] The abstract promises 'concrete design patterns, evaluation checklists, and templates for safe and reliable deployment,' but the tutorial description does not enumerate these deliverables. Adding a short list in §2 or §3 would clarify the expected takeaways.
  4. [References] Reference [30] contains a stray '11 pages.' in the citation text; check formatting. Also, some references (e.g., [29]) are cited in contexts that overstate their claims; see major comments.

Circularity Check

0 steps flagged

No circular reasoning found: the tutorial makes no predictive derivation and its self-citations are background references, not load-bearing evidence.

full rationale

This document is a tutorial proposal, not a research derivation. It contains no equations, fitted parameters, or constructed predictions whose outputs are equivalent to their inputs by definition. The central claims are descriptive: agentic systems are said to be 'rapidly transitioning from research prototypes to production-scale deployments' and the tutorial plans to survey reasoning, planning, multi-agent coordination, and evaluation. These claims are supported by citations to external work (e.g., [23], [29], [16]), including some self-citations ([8], [11], [28], [40], [41]), but those self-citations are used as examples of prior systems or evaluation efforts, not as the sole justification for any derived result. The concern raised in the skeptical analysis — that some deployment statistics trace to a commentary rather than a primary audit — is an evidentiary limitation of the motivating premise, not a circularity. It does not involve redefining a term in terms of its conclusion, fitting a parameter and calling it a prediction, or importing an author-specific uniqueness theorem. Under the stated rules, absence of circular derivation should be scored 0, and no circular steps are identified.

Axiom & Free-Parameter Ledger

0 free parameters · 2 axioms · 0 invented entities

The document introduces no free parameters or invented entities. Its central assumptions are the accuracy of external reports and the representativeness of the cited literature.

axioms (2)
  • domain assumption The cited deployments and benchmarks are accurately described by their sources.
    The tutorial's framing rests on external reports such as Moderna's 750 agentic applications [23] and FutureHouse's ether0 [29]; if these are inaccurate, the tutorial's depiction of the field is unsupported.
  • domain assumption The curated topic selection is representative of the field.
    The tutorial claims to offer 'a comprehensive view of the field,' which assumes the cited works are the relevant ones; no systematic selection methodology is given.

pith-pipeline@v1.3.0-alltime-deepseek · 10956 in / 11043 out tokens · 100213 ms · 2026-08-01T12:41:37.255794+00:00 · methodology

0 comments
read the original abstract

Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments across domains such as software engineering, scientific discovery, and finance. While academic work has emphasized benchmarks and algorithmic innovation, deployment raises new challenges around robustness, safety, and reliability. This tutorial brings together researchers and practitioners to explore advances in reasoning and planning, multi agent coordination, and evaluation, highlighting open challenges arising from deployment experience. Through applied case studies in pharmaceutical discovery and financial systems, we analyze common design patterns that make agentic systems successful, and discuss practical mitigation strategies for failure modes, such as verification pipelines, fallback mechanisms, and human in the loop supervision. Attendees will gain a comprehensive view of the field along with concrete design patterns, evaluation checklists, and templates for safe and reliable deployment across industries.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

69 extracted references · 2 canonical work pages

  1. [1]

    Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi

  2. [2]

    Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes

    Daniil A. Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes. 2023. Au- tonomous chemical research with large language models.Nature624, 7992 (2023), 570–578. doi:10.1038/s41586-023-06792-0

  3. [3]

    Léo Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault Le Sell- ier de Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, and Alexandre Drouin. 2024. WorkArena++: Towards Compositional Planning and Reasoning-Based Common Knowledge Work Tasks. InAdvances in Neural Infor- mation Processing Systems, Vol. 37. 5996–6051. doi:10.52202/...

  4. [4]

    Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D

    Andres M. Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D. White, and Philippe Schwaller. 2024. Augmenting large language models with chemistry tools.Nature Machine Intelligence6 (2024), 525–535. doi:10.1038/s42256-024- 00832-8

  5. [5]

    Baker, Benjamin Burns, Daniel Adu-Ampratwum, Xuhui Huang, Xia Ning, Song Gao, Yu Su, and Huan Sun

    Ziru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang, Boshi Wang, Botao Yu, Yifei Li, Zeyi Liao, Chen Wei, Zitong Lu, Vishal Dey, Mingyi Xue, Frazier N. Baker, Benjamin Burns, Daniel Adu-Ampratwum, Xuhui Huang, Xia Ning, Song Gao, Yu Su, and Huan Sun. 2025. ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discover...

  6. [6]

    Yufan Dang, Chen Qian, Xueheng Luo, Jingru Fan, Zihao Xie, Ruijie Shi, Weize Chen, Cheng Yang, Xiaoyin Che, Ye Tian, Xuantang Xiong, Lei Han, Zhiyuan Liu, and Maosong Sun. 2025. Multi-Agent Collaboration via KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea Grace Hui Yang et al. Evolving Orchestration. InAdvances in Neural Information Processing...

  7. [7]

    Xu, Siva Reddy, Graham Neubig, Quentin Cappart, Russ Salakhutdinov, and Nicolas Chapados

    Thibault Le Sellier de Chezelles, Maxime Gasse, Alexandre Lacoste, Massimo Caccia, Alexandre Drouin, Léo Boisvert, Megh Thakkar, Tom Marty, Rim Assouel, Sahar Omidi Shayegan, Lawrence Keunho Jang, Xing Han Lù, Ori Yoran, De- han Kong, Frank F. Xu, Siva Reddy, Graham Neubig, Quentin Cappart, Russ Salakhutdinov, and Nicolas Chapados. 2025. The BrowserGym Ec...

  8. [8]

    Victor Dibia, Jingya Chen, Gagan Bansal, Suff Syed, Adam Fourney, Erkang Zhu, Chi Wang, and Saleema Amershi. 2024. AUTOGEN STUDIO: A No-Code Developer Tool for Building and Debugging Multi-Agent Systems. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Delia Irazu Hernandez Farias, Tom Hope, ...

  9. [9]

    Laradji, Manuel Del Verme, Tom Marty, David Vazquez, Nicolas Chapados, and Alexandre Lacoste

    Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, David Vazquez, Nicolas Chapados, and Alexandre Lacoste

  10. [10]

    Zhuoyun Du, Chen Qian, Wei Liu, Zihao Xie, YiFei Wang, Rennai Qiu, Yufan Dang, Weize Chen, Cheng Yang, Ye Tian, Xuantang Xiong, and Lei Han. 2025. Multi- Agent Collaboration via Cross-Team Orchestration. InFindings of the Association for Computational Linguistics: ACL 2025, Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.)...

  11. [11]

    InProceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol

    WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?. InProceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235). PMLR, 11642– 11662. https://proceedings.mlr.press/v235/drouin24a.html

  12. [12]

    Bowen Gao, Yanwen Huang, Yiqiao Liu, Wenxuan Xie, Wei-Ying Ma, Ya-Qin Zhang, and Yanyan Lan. 2025. PharmAgents: Building a Virtual Pharma with Large Language Model Agents. arXiv:2503.22164 [cs.AI] doi:10.48550/arXiv.2503. 22164

  13. [13]

    Adam Fourney, Gagan Bansal, Hussein Mozannar, Cheng Tan, Eduardo Salinas, Erkang Zhu, Friederike Niedtner, Grace Proebsting, Griffin Bassman, Jack Gerrits, Jacob Alber, Peter Chang, Ricky Loynd, Robert West, Victor Dibia, Ahmed Awadal- lah, Ece Kamar, Rafah Hosn, and Saleema Amershi. 2024. Magentic-One: A Gen- eralist Multi-Agent System for Solving Comple...

  14. [14]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2023. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997 [cs.CL] doi:10.48550/arXiv.2312.10997

  15. [15]

    Shanghua Gao, Ada Fang, Yepeng Huang, Valentina Giunchiglia, Ayush Noori, Jonathan Richard Schwarz, Yasha Ektefaie, Jovana Kondic, and Marinka Zitnik

  16. [16]

    doi:10.1016/j.cell.2024.09.022

    Empowering Biomedical Discovery with AI Agents.Cell187, 22 (2024), 6125–6151. doi:10.1016/j.cell.2024.09.022

  17. [17]

    Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Ar- tiom Myaskovsky, Grzegorz Glowaty, Felix Weissenberger, Alessio Orlandi, Dan Popovici, Anil Palepu, Keran Rong, Ryutaro Tanno, Khaled Saab, Fan Zhang, Jacob Blum, Andrew Carroll, Kavita Kulkarni, Nenad Tomašev, Dina Zverinski, Ivor Rendulic, Elahe Vedadi, Florian Hasler, Luka Rim...

  18. [18]

    Alireza Ghafarollahi and Markus J. Buehler. 2025. SciAgents: Automating Scien- tific Discovery Through Bioinspired Multi-Agent Intelligent Graph Reasoning. Advanced Materials37, 22 (2025), 2413523. doi:10.1002/adma.202413523

  19. [19]

    Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Cara- lyn J

    Ali E. Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Cara- lyn J. Szostkiewicz, Dmytro Shved, Gavin J. Gyimesi, Jon M. Laurent, Saman- tha M. Wright, Muhammed T. Razzak, Andrew D. White, Silvia C. Finnemann, Michaela M. Hinks, and Samuel G. Rodriques. 2026. A Multi-Agent System for Automating Scientific Discovery.Nature655 (2026), 497–505. doi:...

  20. [20]

    Xu Huang, Weiwen Liu, Xiaolong Chen, Xingmei Wang, Hao Wang, Defu Lian, Yasheng Wang, Ruiming Tang, and Enhong Chen. 2024. Understanding the Planning of LLM Agents: A Survey. arXiv:2402.02716 [cs.AI] doi:10.48550/arXiv. 2402.02716

  21. [21]

    Kung-Hsiang Huang, Akshara Prabhakar, Sidharth Dhawan, Yixin Mao, Huan Wang, Silvio Savarese, Caiming Xiong, Philippe Laban, and Chien-Sheng Wu

  22. [22]

    Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

    Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. InICLR

  23. [23]

    Kung-Hsiang Huang, Akshara Prabhakar, Onkar Thorat, Divyansh Agarwal, Prafulla Kumar Choubey, Yixin Mao, Silvio Savarese, Caiming Xiong, and Chien- Sheng Wu. 2026. CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions.Transactions on Machine Learning Research(2026). https://openreview.net/forum?id=EPlpe3Fx1x

  24. [24]

    Bradley Knox, and Kimin Lee

    Juyong Lee, Dongyoon Hahm, June Suk Choi, W. Bradley Knox, and Kimin Lee

  25. [25]

    Shankar Kumar Jeyakumar, Alaa Alameer Ahmad, and Adrian Garret Gabriel

  26. [26]

    InNeurIPS 2024 Work- shops: OW A

    Advancing Agentic Systems: Dynamic Task Decomposition, Tool Integra- tion and Evaluation Using Novel Metrics and Dataset. InNeurIPS 2024 Work- shops: OW A. https://mlanthology.org/neuripsw/2024/jeyakumar2024neuripsw- advancing/

  27. [28]

    Shaheen E. Lakhan. 2025. The Agentic Era: Why Biopharma Must Embrace Artificial Intelligence That Acts, Not Just Informs.Cureus17, 5 (2025), e83390. doi:10.7759/cureus.83390

  28. [29]

    Siddharth Narayanan, James Braza, Ryan-Rhys Griffiths, Albert Bou, Geemi Wellawatte, Mayk Caldas Ramos, Ludovico Mitchener, Michael Pieler, Sam Rodriques, and Andrew White. 2025. Training a Scientific Reasoning Model for Chemistry. InAdvances in Neural Information Processing Sys- tems, Vol. 38. https://proceedings.neurips.cc/paper_files/paper/2025/hash/ e...

  29. [30]

    Yu, and Qing Li

    Liangbo Ning, Ziran Liang, Zhuohang Jiang, Haohao Qu, Yujuan Ding, Wenqi Fan, Xiao-Yong Wei, Shanru Lin, Hui Liu, Philip S. Yu, and Qing Li. 2025. A Survey of WebAgents: Towards Next-Generation AI Agents for Web Automation with Large Foundation Models. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2. Association ...

  30. [31]

    Ido Levy, Ben Wiesel, Sami Marreed, Alon Oved, Avi Yaeli, and Segev Shlomov

  31. [32]

    InThe Fourteenth International Conference on Learning Representations

    ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustwor- thiness in Web Agents. InThe Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=fAmhr96SUw

  32. [33]

    Yanming Liu, Xinyue Peng, Jiannan Cao, Shi Bo, Yuwei Zhang, Xuhong Zhang, Sheng Cheng, Xun Wang, Jianwei Yin, and Tianyu Du. 2025. Tool-Planner: Task Planning with Clusters across Multiple Tools. InThe Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net. https://openreview.net/forum?id=d...

  33. [34]

    Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language Models Can Teach Themselves to Use Tools. InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Infor- mation Processing Systems 2023, NeurIPS 2023, New Or...

  34. [35]

    Hussein Mozannar, Gagan Bansal, Cheng Tan, Adam Fourney, Victor Dibia, Jingya Chen, Jack Gerrits, Tyler Payne, Matheus Kunzler Maldaner, Madeleine Grunde-McLaughlin, Eric Zhu, Griffin Bassman, Jacob Alber, Peter Chang, Ricky Loynd, Friederike Niedtner, Ece Kamar, Maya Murad, Rafah Hosn, and Saleema Amershi. 2025. Magentic-UI: Towards Human-in-the-Loop Age...

  35. [36]

    Bulaong, John E

    Kyle Swanson, Wesley Wu, Nash L. Bulaong, John E. Pak, and James Zou. 2025. The Virtual Lab of AI Agents Designs New SARS-CoV-2 Nanobodies.Nature646 (2025), 716–723. doi:10.1038/s41586-025-09442-9

  36. [37]

    Aaron Xuxiang Tian, Ruofan Zhang, Jiayao Tang, Young Min Cho, Xueqian Li, Qiang Yi, Ji Wang, Zhunping Zhang, Danrui Qi, Zekun Li, Xingyu Xiang, Sharath Chandra Guntuku, Lyle Ungar, Tianyu Shi, and Chi Wang. 2025. Beyond the Strongest LLM: Multi-Turn Multi-Agent Orchestration vs. Single LLMs on Benchmarks. arXiv:2509.23537 [cs.AI] https://arxiv.org/abs/2509.23537

  37. [38]

    Joshua Owotogbe. 2025. Assessing and Enhancing the Robustness of LLM-Based Multi-Agent Systems Through Chaos Engineering. In4th IEEE/ACM International Conference on AI Engineering - Software Engineering for AI, CAIN 2025, Ottawa, ON, Canada, April 27-28, 2025. IEEE, 250–252. doi:10.1109/CAIN66642.2025.00039

  38. [39]

    Mihir Parmar, Xin Liu, Palash Goyal, Yanfei Chen, Long Le, Swaroop Mishra, Hos- sein Mobahi, Jindong Gu, Zifeng Wang, Hootan Nakhost, Chitta Baral, Chen-Yu Lee, Tomas Pfister, and Hamid Palangi. 2025. PlanGEN: A Multi-Agent Frame- work for Generating Planning and Reasoning Trajectories for Complex Problem Solving. InProceedings of the 2025 Conference on E...

  39. [40]

    Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, Jun Xu, and Ji-Rong Wen. 2025. Tool Learning with Large Language Models: A Survey.Frontiers of Computer Science19, 8 (2025), 198343. doi:10.1007/s11704- 024-40678-2

  40. [41]

    Pranav Narayanan Venkit, Philippe Laban, Yilun Zhou, Yixin Mao, and Chien- Sheng Wu. 2025. Search Engines in the AI Era: A Qualitative Understanding to the False Promise of Factual and Verifiable Source-Cited Responses in LLM-based Search. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. Association for Computing Mac...

  41. [42]

    Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language Agents with Verbal Rein- forcement Learning. InAdvances in Neural Information Processing Systems, Agents in the Wild: Where Research Meets Deployment KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea Vol. 36. 8634–8652. https://proceeding...

  42. [43]

    Yaoxiang Wang, Zhiyong Wu, Junfeng Yao, and Jinsong Su. 2025. TDAG: A multi- agent framework based on dynamic Task Decomposition and Agent Generation. Neural Networks185 (2025), 107200. doi:10.1016/J.NEUNET.2025.107200

  43. [44]

    Hui Wei, Zihao Zhang, Shenghua He, Tian Xia, Shijia Pan, and Fei Liu. 2025. PlanGenLLMs: A Modern Survey of LLM Planning Capabilities. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Moham- mad Taher Pilehvar (Eds.). Association for Compu...

  44. [45]

    Khanh-Tung Tran, Dung Dao, Minh-Duong Nguyen, Quoc-Viet Pham, Barry O’Sullivan, and Hoang D. Nguyen. 2025. Multi-Agent Collaboration Mechanisms: A Survey of LLMs. arXiv:2501.06322 [cs.AI] doi:10.48550/arXiv.2501.06322

  45. [46]

    Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal

  46. [47]

    Griffiths, Yuan Cao, and Karthik Narasimhan

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: deliberate problem solving with large language models. InProceedings of the 37th International Conference on Neural Information Processing Systems

  47. [48]

    Pranav Narayanan Venkit, Philippe Laban, Yilun Zhou, Kung-Hsiang Huang, Yixin Mao, and Chien-Sheng Wu. 2026. DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence. InInterna- tional Conference on Learning Representations. https://openreview.net/forum?id= QkaeTea16Y

  48. [49]

    Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel, Yilun Zhao, Roy Bar-Haim, Arman Cohan, and Michal Shmueli-Scheuer. 2026. A Survey on Evaluation of LLM-Based Agents. InFindings of the Association for Computational Linguistics: ACL 2026. Association for Computational Linguistics, San Diego, California, United States, 26690–26714. doi:10.18653/v1/2026.finding...

  49. [50]

    Le, Ed H

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. https://openreview.net/forum?id=1PL1NIMMrw

  50. [51]

    Suchow, Denghui Zhang, and Khaldoun Khashanah

    Yangyang Yu, Haohang Li, Zhi Chen, Yuechen Jiang, Yang Li, Jordan W. Suchow, Denghui Zhang, and Khaldoun Khashanah. 2025. FinMem: A Performance- Enhanced LLM Trading Agent With Layered Memory and Character Design. IEEE Transactions on Big Data11, 6 (2025), 3443–3459. doi:10.1109/TBDATA.2025. 3593370

  51. [52]

    Siyu Yuan, Zehui Chen, Zhiheng Xi, Junjie Ye, Zhengyin Du, and Jiecao Chen

  52. [53]

    Chi, Quoc V

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain- of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems 35 (NeurIPS 2022). New Orleans, LA, USA. https://proceedings.neurips.cc/paper/2022/hash/ 9d5609613524ecf4f15...

  53. [54]

    White, Doug Burger, and Chi Wang

    Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W. White, Doug Burger, and Chi Wang. 2024. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversations. InProceedings of the First Con- ference on Language Modeling. https://openreview.net/fo...

  54. [55]

    Wentao Zhang, Lingxuan Zhao, Haochong Xia, Shuo Sun, Jiaze Sun, Molei Qin, Xinyi Li, Yuqing Zhao, Yilei Zhao, Xinyu Cai, Longtao Zheng, Xinrun Wang, and Bo An. 2024. A Multimodal Foundation Agent for Financial Trading: Tool- Augmented, Diversified, and Generalist. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Asso...

  55. [56]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. InICLR

  56. [57]

    Bingxi Zhao, Lin Geng Foo, Ping Hu, Christian Theobalt, Hossein Rahmani, and Jun Liu. 2025. LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios. arXiv:2508.17692 [cs.AI] https://arxiv.org/abs/2508.17692

  57. [58]

    Zonghao Ying, Le Wang, Yisong Xiao, Jiakai Wang, Yuqing Ma, Jinyang Guo, Zhenfei Yin, Mingchuan Zhang, Aishan Liu, and Xianglong Liu. 2026. AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions. InProceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition. 37664–37673. https://openaccess.thecvf.com/ conte...

  58. [59]

    Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

    Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. 2024. WebArena: A Realistic Web Environment for Building Autonomous Agents. InThe Twelfth International Conference on Learning Representations. https: //openreview.net/forum?id=zlsj9akpaa

  59. [60]

    Tianyu Zhou, Pinqiao Wang, Yilin Wu, and Hongyang Yang. 2024. FinRobot: AI Agent for Equity Research and Valuation with Large Language Models. arXiv:2411.08804 [q-fin.CP] https://arxiv.org/abs/2411.08804

  60. [61]

    arXiv:2501.11425 [cs.AI] doi:10.48550/arXiv.2501.11425

    Agent-R: Training Language Model Agents to Reflect via Iterative Self- Training. arXiv:2501.11425 [cs.AI] doi:10.48550/arXiv.2501.11425

  61. [62]

    Cong Zhang, Xin Deik Goh, Dexun Li, Hao Zhang, and Yong Liu. 2025. Planning with Multi-Constraints via Collaborative Language Agents. InProceedings of the 31st International Conference on Computational Linguistics, Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert (Eds.). Association for Computational...

  62. [63]

    Hanrong Zhang, Jingyuan Huang, Kai Mei, Yifei Yao, Zhenting Wang, Chenlu Zhan, Hongwei Wang, and Yongfeng Zhang. 2025. Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-Based Agents. InThe Thirteenth International Conference on Learning Representations. https: //openreview.net/forum?id=V4y0CpX4hK

  63. [65]

    Zhexin Zhang, Shiyao Cui, Yida Lu, Jingzhuo Zhou, Junxiao Yang, Hongning Wang, and Minlie Huang. 2024. Agent-SafetyBench: Evaluating the Safety of LLM Agents. arXiv:2412.14470 [cs.AI] doi:10.48550/arXiv.2412.14470

  64. [67]

    Junhao Zheng, Shengjie Qiu, Chengming Shi, and Qianli Ma. 2025. Towards Lifelong Learning of Large Language Models: A Survey.Comput. Surveys57, 8 (2025), 1–35. doi:10.1145/3716629 Article 193

  65. [70]

    Kunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang, Shuyi Guo, Zhe Wang, Zhenhailong Wang, Cheng Qian, Robert Tang, Heng Ji, and Jiaxuan You. 2025. MultiAgentBench: Evaluating the Collaboration and Competition of LLM Agents. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for...

  66. [2023]

    InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge- Intensive Multi-Step Questions. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Toronto, Canada, 10014–10037. doi:10.18653/v1/ 2023.acl-long.557

  67. [2024]

    InThe Twelfth International Conference on Learning Representations

    Self-RAG: Learning to Retrieve, Generate, and Critique through Self- Reflection. InThe Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=hSyW5go0v8

  68. [2025]

    CRMArena: Understanding the Capacity of LLM Agents to Perform Profes- sional CRM Tasks in Realistic Environments. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Lin- guistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 20...

  69. [2026]

    doi:10.1609/aaai.v40i44.41090

    MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control.Proceedings of the AAAI Conference on Artificial Intelligence40, 44 (2026), 37565–37573. doi:10.1609/aaai.v40i44.41090