REVIEW 3 major objections 4 minor 69 references
Agentic systems built on large language models are moving from research prototypes to production-scale deployment, and this tutorial argues that robustness, safety, and reliability — not benchmark scores — are now the binding constraints.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 12:41 UTC pith:77IBTURD
load-bearing objection Solid tutorial map, but the 'production-scale deployment' hook leans on a few shaky citations — treat it as orientation, not evidence. the 3 major comments →
Agents in the Wild: Where Research Meets Deployment
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that agentic systems have crossed into production-scale deployment across domains, with vendor-reported cases such as a pharmaceutical company running hundreds of agentic applications in R&D and financial systems adopting planner–executor–verifier architectures for trading and analysis. To make such systems work in the wild, the authors argue, evaluation must move from static benchmarks toward dynamic, behavior-centric tests that probe adaptability, recovery from failure, and human-intervention behavior, and deployment must build in verification pipelines, fallback mechanisms, and human-in-the-loop supervision as standard safeguards. The tutorial synthesizes curr
What carries the argument
The tutorial's organizing device is the planner–executor–verifier architecture presented as a common pattern in deployed financial agents, together with a mitigation stack — verification pipelines, fallback mechanisms, and human-in-the-loop supervision — that the authors claim underlies successful agentic deployments. The paper also leans on modular multi-agent orchestration (peer-to-peer and hierarchical topologies) and on a family of scenario-driven, behavior-centric evaluation frameworks that test adaptability, recovery, and intervention under distribution shift and adversarial perturbations. These mechanisms carry the argument that production success is an engineering problem with identi
Load-bearing premise
The tutorial's premise that agentic systems are already running at production scale rests largely on vendor-reported deployment figures — for instance, a pharmaceutical partner's claim of over 750 agentic applications, cited from a commentary source — and if those reports are inflated, the case-study foundation and the 'field in transition' narrative lose their support.
What would settle it
An independent census of enterprise LLM-agent deployments across the cited industries that found most 'agentic applications' to be small-scale pilots requiring human approval for every consequential action would falsify the production-scale claim at the heart of the tutorial.
If this is right
- Evaluation of agentic systems will shift toward interactive, scenario-driven benchmarks that measure adaptability, recovery, and human-intervention behavior under distribution shifts and adversarial perturbations.
- Production agent architectures will standardize around modular designs with independent verification and fallback paths, with human-in-the-loop supervision as a safety net.
- Failure modes such as hallucination, deadlocks, drift, and cascading errors will be treated as first-class engineering risks, with explicit mitigation pipelines built into system design.
- Cross-industry design patterns — planner–executor–verifier in finance, multi-agent literature and experiment pipelines in pharma — will spread as reusable templates for other high-stakes domains.
- Academic research on agents will increasingly be judged by deployment reliability and safety rather than by benchmark performance alone.
Where Pith is reading between the lines
- If the deployment-centric framing is correct, static benchmark scores will increasingly overstate agent readiness, and we should expect a shift toward continuous in-operation monitoring and red-teaming as the primary assurance tools.
- The recurring emphasis on verification and fallback suggests that the realistic end-state for high-stakes agents is 'human-on-the-loop' supervision rather than full autonomy — a testable prediction for how regulated industries will actually adopt these systems.
- The paper presents verification pipelines, fallbacks, and human-in-the-loop oversight as mitigation strategies, but does not quantify their relative contribution; a controlled field experiment varying each safeguard would turn the design-pattern advice into engineering guidance.
- The deployment numbers cited come largely from vendor commentary; an independent audit of how many reported 'agentic applications' are production-critical versus pilot-stage would separate signal from marketing.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a proposal for a half-day KDD 2026 tutorial on agentic systems, authored by researchers from academia and industry. It argues that LLM-based agentic systems are rapidly moving from research prototypes to production-scale deployments, and therefore require attention to robustness, safety, and reliability beyond academic benchmarks. The tutorial outline covers foundations (reasoning, planning, multi-agent coordination, retrieval pipelines, evaluation), applied case studies in pharmaceutical/life sciences and finance, and practical design patterns and mitigation strategies. The submission is a 5-page tutorial description, including author bios and references, rather than a technical research paper.
Significance. If the tutorial is delivered as outlined, it addresses a timely and relevant topic at the intersection of AI research and industrial deployment. The authors are well-positioned to present this material, with strong credentials in both academic research and enterprise AI. The emphasis on evaluation beyond static benchmarks, failure modes (hallucination, deadlocks, drift, cascading errors), and human-in-the-loop mechanisms is valuable for practitioners. However, the tutorial's central premise—that agentic systems are 'rapidly transitioning from research prototypes to production-scale deployments'—currently relies on unverified or secondary sources for its most striking evidence. Strengthening the evidentiary basis of this claim would materially increase the credibility and impact of the tutorial.
major comments (3)
- [§4.1 (first paragraph)] The claims that FutureHouse's ether0 and Robin are 'capable of superhuman molecular design' and that agentic systems are 'matching or exceeding human performance' are not supported by the cited references. [29] is a NeurIPS paper on training a chemistry reasoning model; it does not establish 'superhuman' capability in a deployed product. [16] describes a multi-agent research system, not a production deployment. Please either provide primary evidence for these superlative claims or qualify them as 'reported'/'claimed' by the developers. Conflating research-level results with production capability weakens the tutorial's evidentiary foundation.
- [§4.1, Moderna statistic] The statement that Moderna's partnership with OpenAI 'report[s] over 750 agentic applications across R&D operations' is cited to [23], a Cureus commentary. A commentary is not a primary deployment report or an audited case study. This statistic is the flagship example of 'production-scale deployments' in the abstract; please trace it to a primary source (e.g., a formal company announcement or third-party audit) or clearly label it as a vendor/press claim. Without such sourcing, the tutorial's motivation rests on a secondary, promotional source.
- [Abstract and §1] The opening assertion that agentic systems are 'rapidly transitioning from research prototypes to production-scale deployments' is stated as fact. The tutorial's own examples blur this distinction: many cited systems (e.g., [2], [4], [36], [17]) are research prototypes or limited pilots, and the only explicit count of production-scale use (Moderna) is from a commentary. The tutorial description should either (a) provide at least one independently verifiable production deployment with primary sourcing, or (b) temper the claim to 'reported industrial adoption' and explicitly discuss the gap between research demonstrations and production readiness. As written, the motivation risks presenting vendor claims as established fact.
minor comments (4)
- [§4.1, last paragraph] The sentence 'Biomedical researchers envision AI scientists as collaborative and critically reasoning partners...' is a near-duplicate of the previous sentence citing Harvard Medical School [13]. Remove one occurrence.
- [Section heading] The heading '5 Tutors Bios' should be 'Tutor Bios' or 'Tutors' Biographies'.
- [Abstract] The abstract promises 'concrete design patterns, evaluation checklists, and templates for safe and reliable deployment,' but the tutorial description does not enumerate these deliverables. Adding a short list in §2 or §3 would clarify the expected takeaways.
- [References] Reference [30] contains a stray '11 pages.' in the citation text; check formatting. Also, some references (e.g., [29]) are cited in contexts that overstate their claims; see major comments.
Circularity Check
No circular reasoning found: the tutorial makes no predictive derivation and its self-citations are background references, not load-bearing evidence.
full rationale
This document is a tutorial proposal, not a research derivation. It contains no equations, fitted parameters, or constructed predictions whose outputs are equivalent to their inputs by definition. The central claims are descriptive: agentic systems are said to be 'rapidly transitioning from research prototypes to production-scale deployments' and the tutorial plans to survey reasoning, planning, multi-agent coordination, and evaluation. These claims are supported by citations to external work (e.g., [23], [29], [16]), including some self-citations ([8], [11], [28], [40], [41]), but those self-citations are used as examples of prior systems or evaluation efforts, not as the sole justification for any derived result. The concern raised in the skeptical analysis — that some deployment statistics trace to a commentary rather than a primary audit — is an evidentiary limitation of the motivating premise, not a circularity. It does not involve redefining a term in terms of its conclusion, fitting a parameter and calling it a prediction, or importing an author-specific uniqueness theorem. Under the stated rules, absence of circular derivation should be scored 0, and no circular steps are identified.
Axiom & Free-Parameter Ledger
axioms (2)
- domain assumption The cited deployments and benchmarks are accurately described by their sources.
- domain assumption The curated topic selection is representative of the field.
read the original abstract
Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments across domains such as software engineering, scientific discovery, and finance. While academic work has emphasized benchmarks and algorithmic innovation, deployment raises new challenges around robustness, safety, and reliability. This tutorial brings together researchers and practitioners to explore advances in reasoning and planning, multi agent coordination, and evaluation, highlighting open challenges arising from deployment experience. Through applied case studies in pharmaceutical discovery and financial systems, we analyze common design patterns that make agentic systems successful, and discuss practical mitigation strategies for failure modes, such as verification pipelines, fallback mechanisms, and human in the loop supervision. Attendees will gain a comprehensive view of the field along with concrete design patterns, evaluation checklists, and templates for safe and reliable deployment across industries.
Reference graph
Works this paper leans on
-
[1]
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi
-
[2]
Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes
Daniil A. Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes. 2023. Au- tonomous chemical research with large language models.Nature624, 7992 (2023), 570–578. doi:10.1038/s41586-023-06792-0
-
[3]
Léo Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault Le Sell- ier de Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, and Alexandre Drouin. 2024. WorkArena++: Towards Compositional Planning and Reasoning-Based Common Knowledge Work Tasks. InAdvances in Neural Infor- mation Processing Systems, Vol. 37. 5996–6051. doi:10.52202/...
-
[4]
Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D
Andres M. Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D. White, and Philippe Schwaller. 2024. Augmenting large language models with chemistry tools.Nature Machine Intelligence6 (2024), 525–535. doi:10.1038/s42256-024- 00832-8
-
[5]
Baker, Benjamin Burns, Daniel Adu-Ampratwum, Xuhui Huang, Xia Ning, Song Gao, Yu Su, and Huan Sun
Ziru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang, Boshi Wang, Botao Yu, Yifei Li, Zeyi Liao, Chen Wei, Zitong Lu, Vishal Dey, Mingyi Xue, Frazier N. Baker, Benjamin Burns, Daniel Adu-Ampratwum, Xuhui Huang, Xia Ning, Song Gao, Yu Su, and Huan Sun. 2025. ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discover...
2025
-
[6]
Yufan Dang, Chen Qian, Xueheng Luo, Jingru Fan, Zihao Xie, Ruijie Shi, Weize Chen, Cheng Yang, Xiaoyin Che, Ye Tian, Xuantang Xiong, Lei Han, Zhiyuan Liu, and Maosong Sun. 2025. Multi-Agent Collaboration via KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea Grace Hui Yang et al. Evolving Orchestration. InAdvances in Neural Information Processing...
2025
-
[7]
Xu, Siva Reddy, Graham Neubig, Quentin Cappart, Russ Salakhutdinov, and Nicolas Chapados
Thibault Le Sellier de Chezelles, Maxime Gasse, Alexandre Lacoste, Massimo Caccia, Alexandre Drouin, Léo Boisvert, Megh Thakkar, Tom Marty, Rim Assouel, Sahar Omidi Shayegan, Lawrence Keunho Jang, Xing Han Lù, Ori Yoran, De- han Kong, Frank F. Xu, Siva Reddy, Graham Neubig, Quentin Cappart, Russ Salakhutdinov, and Nicolas Chapados. 2025. The BrowserGym Ec...
2025
-
[8]
Victor Dibia, Jingya Chen, Gagan Bansal, Suff Syed, Adam Fourney, Erkang Zhu, Chi Wang, and Saleema Amershi. 2024. AUTOGEN STUDIO: A No-Code Developer Tool for Building and Debugging Multi-Agent Systems. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Delia Irazu Hernandez Farias, Tom Hope, ...
-
[9]
Laradji, Manuel Del Verme, Tom Marty, David Vazquez, Nicolas Chapados, and Alexandre Lacoste
Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, David Vazquez, Nicolas Chapados, and Alexandre Lacoste
-
[10]
Zhuoyun Du, Chen Qian, Wei Liu, Zihao Xie, YiFei Wang, Rennai Qiu, Yufan Dang, Weize Chen, Cheng Yang, Ye Tian, Xuantang Xiong, and Lei Han. 2025. Multi- Agent Collaboration via Cross-Team Orchestration. InFindings of the Association for Computational Linguistics: ACL 2025, Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.)...
-
[11]
InProceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?. InProceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235). PMLR, 11642– 11662. https://proceedings.mlr.press/v235/drouin24a.html
-
[12]
Bowen Gao, Yanwen Huang, Yiqiao Liu, Wenxuan Xie, Wei-Ying Ma, Ya-Qin Zhang, and Yanyan Lan. 2025. PharmAgents: Building a Virtual Pharma with Large Language Model Agents. arXiv:2503.22164 [cs.AI] doi:10.48550/arXiv.2503. 22164
-
[13]
Adam Fourney, Gagan Bansal, Hussein Mozannar, Cheng Tan, Eduardo Salinas, Erkang Zhu, Friederike Niedtner, Grace Proebsting, Griffin Bassman, Jack Gerrits, Jacob Alber, Peter Chang, Ricky Loynd, Robert West, Victor Dibia, Ahmed Awadal- lah, Ece Kamar, Rafah Hosn, and Saleema Amershi. 2024. Magentic-One: A Gen- eralist Multi-Agent System for Solving Comple...
-
[14]
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2023. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997 [cs.CL] doi:10.48550/arXiv.2312.10997
-
[15]
Shanghua Gao, Ada Fang, Yepeng Huang, Valentina Giunchiglia, Ayush Noori, Jonathan Richard Schwarz, Yasha Ektefaie, Jovana Kondic, and Marinka Zitnik
-
[16]
doi:10.1016/j.cell.2024.09.022
Empowering Biomedical Discovery with AI Agents.Cell187, 22 (2024), 6125–6151. doi:10.1016/j.cell.2024.09.022
-
[17]
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Ar- tiom Myaskovsky, Grzegorz Glowaty, Felix Weissenberger, Alessio Orlandi, Dan Popovici, Anil Palepu, Keran Rong, Ryutaro Tanno, Khaled Saab, Fan Zhang, Jacob Blum, Andrew Carroll, Kavita Kulkarni, Nenad Tomašev, Dina Zverinski, Ivor Rendulic, Elahe Vedadi, Florian Hasler, Luka Rim...
2026
-
[18]
Alireza Ghafarollahi and Markus J. Buehler. 2025. SciAgents: Automating Scien- tific Discovery Through Bioinspired Multi-Agent Intelligent Graph Reasoning. Advanced Materials37, 22 (2025), 2413523. doi:10.1002/adma.202413523
-
[19]
Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Cara- lyn J
Ali E. Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Cara- lyn J. Szostkiewicz, Dmytro Shved, Gavin J. Gyimesi, Jon M. Laurent, Saman- tha M. Wright, Muhammed T. Razzak, Andrew D. White, Silvia C. Finnemann, Michaela M. Hinks, and Samuel G. Rodriques. 2026. A Multi-Agent System for Automating Scientific Discovery.Nature655 (2026), 497–505. doi:...
doi:10.1038/s41586- 2026
-
[20]
Xu Huang, Weiwen Liu, Xiaolong Chen, Xingmei Wang, Hao Wang, Defu Lian, Yasheng Wang, Ruiming Tang, and Enhong Chen. 2024. Understanding the Planning of LLM Agents: A Survey. arXiv:2402.02716 [cs.AI] doi:10.48550/arXiv. 2402.02716
-
[21]
Kung-Hsiang Huang, Akshara Prabhakar, Sidharth Dhawan, Yixin Mao, Huan Wang, Silvio Savarese, Caiming Xiong, Philippe Laban, and Chien-Sheng Wu
-
[22]
Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. InICLR
2024
-
[23]
Kung-Hsiang Huang, Akshara Prabhakar, Onkar Thorat, Divyansh Agarwal, Prafulla Kumar Choubey, Yixin Mao, Silvio Savarese, Caiming Xiong, and Chien- Sheng Wu. 2026. CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions.Transactions on Machine Learning Research(2026). https://openreview.net/forum?id=EPlpe3Fx1x
2026
-
[24]
Bradley Knox, and Kimin Lee
Juyong Lee, Dongyoon Hahm, June Suk Choi, W. Bradley Knox, and Kimin Lee
-
[25]
Shankar Kumar Jeyakumar, Alaa Alameer Ahmad, and Adrian Garret Gabriel
-
[26]
InNeurIPS 2024 Work- shops: OW A
Advancing Agentic Systems: Dynamic Task Decomposition, Tool Integra- tion and Evaluation Using Novel Metrics and Dataset. InNeurIPS 2024 Work- shops: OW A. https://mlanthology.org/neuripsw/2024/jeyakumar2024neuripsw- advancing/
2024
-
[28]
Shaheen E. Lakhan. 2025. The Agentic Era: Why Biopharma Must Embrace Artificial Intelligence That Acts, Not Just Informs.Cureus17, 5 (2025), e83390. doi:10.7759/cureus.83390
-
[29]
Siddharth Narayanan, James Braza, Ryan-Rhys Griffiths, Albert Bou, Geemi Wellawatte, Mayk Caldas Ramos, Ludovico Mitchener, Michael Pieler, Sam Rodriques, and Andrew White. 2025. Training a Scientific Reasoning Model for Chemistry. InAdvances in Neural Information Processing Sys- tems, Vol. 38. https://proceedings.neurips.cc/paper_files/paper/2025/hash/ e...
2025
-
[30]
Liangbo Ning, Ziran Liang, Zhuohang Jiang, Haohao Qu, Yujuan Ding, Wenqi Fan, Xiao-Yong Wei, Shanru Lin, Hui Liu, Philip S. Yu, and Qing Li. 2025. A Survey of WebAgents: Towards Next-Generation AI Agents for Web Automation with Large Foundation Models. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2. Association ...
arXiv 2025
-
[31]
Ido Levy, Ben Wiesel, Sami Marreed, Alon Oved, Avi Yaeli, and Segev Shlomov
-
[32]
InThe Fourteenth International Conference on Learning Representations
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustwor- thiness in Web Agents. InThe Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=fAmhr96SUw
-
[33]
Yanming Liu, Xinyue Peng, Jiannan Cao, Shi Bo, Yuwei Zhang, Xuhong Zhang, Sheng Cheng, Xun Wang, Jianwei Yin, and Tianyu Du. 2025. Tool-Planner: Task Planning with Clusters across Multiple Tools. InThe Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net. https://openreview.net/forum?id=d...
2025
-
[34]
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language Models Can Teach Themselves to Use Tools. InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Infor- mation Processing Systems 2023, NeurIPS 2023, New Or...
2023
-
[35]
Hussein Mozannar, Gagan Bansal, Cheng Tan, Adam Fourney, Victor Dibia, Jingya Chen, Jack Gerrits, Tyler Payne, Matheus Kunzler Maldaner, Madeleine Grunde-McLaughlin, Eric Zhu, Griffin Bassman, Jacob Alber, Peter Chang, Ricky Loynd, Friederike Niedtner, Ece Kamar, Maya Murad, Rafah Hosn, and Saleema Amershi. 2025. Magentic-UI: Towards Human-in-the-Loop Age...
-
[36]
Kyle Swanson, Wesley Wu, Nash L. Bulaong, John E. Pak, and James Zou. 2025. The Virtual Lab of AI Agents Designs New SARS-CoV-2 Nanobodies.Nature646 (2025), 716–723. doi:10.1038/s41586-025-09442-9
-
[37]
Aaron Xuxiang Tian, Ruofan Zhang, Jiayao Tang, Young Min Cho, Xueqian Li, Qiang Yi, Ji Wang, Zhunping Zhang, Danrui Qi, Zekun Li, Xingyu Xiang, Sharath Chandra Guntuku, Lyle Ungar, Tianyu Shi, and Chi Wang. 2025. Beyond the Strongest LLM: Multi-Turn Multi-Agent Orchestration vs. Single LLMs on Benchmarks. arXiv:2509.23537 [cs.AI] https://arxiv.org/abs/2509.23537
arXiv 2025
-
[38]
Joshua Owotogbe. 2025. Assessing and Enhancing the Robustness of LLM-Based Multi-Agent Systems Through Chaos Engineering. In4th IEEE/ACM International Conference on AI Engineering - Software Engineering for AI, CAIN 2025, Ottawa, ON, Canada, April 27-28, 2025. IEEE, 250–252. doi:10.1109/CAIN66642.2025.00039
arXiv 2025
-
[39]
Mihir Parmar, Xin Liu, Palash Goyal, Yanfei Chen, Long Le, Swaroop Mishra, Hos- sein Mobahi, Jindong Gu, Zifeng Wang, Hootan Nakhost, Chitta Baral, Chen-Yu Lee, Tomas Pfister, and Hamid Palangi. 2025. PlanGEN: A Multi-Agent Frame- work for Generating Planning and Reasoning Trajectories for Complex Problem Solving. InProceedings of the 2025 Conference on E...
-
[40]
Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, Jun Xu, and Ji-Rong Wen. 2025. Tool Learning with Large Language Models: A Survey.Frontiers of Computer Science19, 8 (2025), 198343. doi:10.1007/s11704- 024-40678-2
doi:10.1007/s11704- 2025
-
[41]
Pranav Narayanan Venkit, Philippe Laban, Yilun Zhou, Yixin Mao, and Chien- Sheng Wu. 2025. Search Engines in the AI Era: A Qualitative Understanding to the False Promise of Factual and Verifiable Source-Cited Responses in LLM-based Search. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. Association for Computing Mac...
arXiv 2025
-
[42]
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language Agents with Verbal Rein- forcement Learning. InAdvances in Neural Information Processing Systems, Agents in the Wild: Where Research Meets Deployment KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea Vol. 36. 8634–8652. https://proceeding...
2023
-
[43]
Yaoxiang Wang, Zhiyong Wu, Junfeng Yao, and Jinsong Su. 2025. TDAG: A multi- agent framework based on dynamic Task Decomposition and Agent Generation. Neural Networks185 (2025), 107200. doi:10.1016/J.NEUNET.2025.107200
arXiv 2025
-
[44]
Hui Wei, Zihao Zhang, Shenghua He, Tian Xia, Shijia Pan, and Fei Liu. 2025. PlanGenLLMs: A Modern Survey of LLM Planning Capabilities. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Moham- mad Taher Pilehvar (Eds.). Association for Compu...
-
[45]
Khanh-Tung Tran, Dung Dao, Minh-Duong Nguyen, Quoc-Viet Pham, Barry O’Sullivan, and Hoang D. Nguyen. 2025. Multi-Agent Collaboration Mechanisms: A Survey of LLMs. arXiv:2501.06322 [cs.AI] doi:10.48550/arXiv.2501.06322
-
[46]
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal
-
[47]
Griffiths, Yuan Cao, and Karthik Narasimhan
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: deliberate problem solving with large language models. InProceedings of the 37th International Conference on Neural Information Processing Systems
2023
-
[48]
Pranav Narayanan Venkit, Philippe Laban, Yilun Zhou, Kung-Hsiang Huang, Yixin Mao, and Chien-Sheng Wu. 2026. DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence. InInterna- tional Conference on Learning Representations. https://openreview.net/forum?id= QkaeTea16Y
2026
-
[49]
Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel, Yilun Zhao, Roy Bar-Haim, Arman Cohan, and Michal Shmueli-Scheuer. 2026. A Survey on Evaluation of LLM-Based Agents. InFindings of the Association for Computational Linguistics: ACL 2026. Association for Computational Linguistics, San Diego, California, United States, 26690–26714. doi:10.18653/v1/2026.finding...
-
[50]
Le, Ed H
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. https://openreview.net/forum?id=1PL1NIMMrw
2023
-
[51]
Suchow, Denghui Zhang, and Khaldoun Khashanah
Yangyang Yu, Haohang Li, Zhi Chen, Yuechen Jiang, Yang Li, Jordan W. Suchow, Denghui Zhang, and Khaldoun Khashanah. 2025. FinMem: A Performance- Enhanced LLM Trading Agent With Layered Memory and Character Design. IEEE Transactions on Big Data11, 6 (2025), 3443–3459. doi:10.1109/TBDATA.2025. 3593370
-
[52]
Siyu Yuan, Zehui Chen, Zhiheng Xi, Junjie Ye, Zhengyin Du, and Jiecao Chen
-
[53]
Chi, Quoc V
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain- of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems 35 (NeurIPS 2022). New Orleans, LA, USA. https://proceedings.neurips.cc/paper/2022/hash/ 9d5609613524ecf4f15...
2022
-
[54]
White, Doug Burger, and Chi Wang
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W. White, Doug Burger, and Chi Wang. 2024. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversations. InProceedings of the First Con- ference on Language Modeling. https://openreview.net/fo...
2024
-
[55]
Wentao Zhang, Lingxuan Zhao, Haochong Xia, Shuo Sun, Jiaze Sun, Molei Qin, Xinyi Li, Yuqing Zhao, Yilei Zhao, Xinyu Cai, Longtao Zheng, Xinrun Wang, and Bo An. 2024. A Multimodal Foundation Agent for Financial Trading: Tool- Augmented, Diversified, and Generalist. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Asso...
arXiv 2024
-
[56]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. InICLR
2023
-
[57]
Bingxi Zhao, Lin Geng Foo, Ping Hu, Christian Theobalt, Hossein Rahmani, and Jun Liu. 2025. LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios. arXiv:2508.17692 [cs.AI] https://arxiv.org/abs/2508.17692
Pith/arXiv arXiv 2025
-
[58]
Zonghao Ying, Le Wang, Yisong Xiao, Jiakai Wang, Yuqing Ma, Jinyang Guo, Zhenfei Yin, Mingchuan Zhang, Aishan Liu, and Xianglong Liu. 2026. AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions. InProceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition. 37664–37673. https://openaccess.thecvf.com/ conte...
2026
-
[59]
Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig
Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. 2024. WebArena: A Realistic Web Environment for Building Autonomous Agents. InThe Twelfth International Conference on Learning Representations. https: //openreview.net/forum?id=zlsj9akpaa
2024
-
[60]
Tianyu Zhou, Pinqiao Wang, Yilin Wu, and Hongyang Yang. 2024. FinRobot: AI Agent for Equity Research and Valuation with Large Language Models. arXiv:2411.08804 [q-fin.CP] https://arxiv.org/abs/2411.08804
Pith/arXiv arXiv 2024
-
[61]
arXiv:2501.11425 [cs.AI] doi:10.48550/arXiv.2501.11425
Agent-R: Training Language Model Agents to Reflect via Iterative Self- Training. arXiv:2501.11425 [cs.AI] doi:10.48550/arXiv.2501.11425
-
[62]
Cong Zhang, Xin Deik Goh, Dexun Li, Hao Zhang, and Yong Liu. 2025. Planning with Multi-Constraints via Collaborative Language Agents. InProceedings of the 31st International Conference on Computational Linguistics, Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert (Eds.). Association for Computational...
2025
-
[63]
Hanrong Zhang, Jingyuan Huang, Kai Mei, Yifei Yao, Zhenting Wang, Chenlu Zhan, Hongwei Wang, and Yongfeng Zhang. 2025. Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-Based Agents. InThe Thirteenth International Conference on Learning Representations. https: //openreview.net/forum?id=V4y0CpX4hK
2025
-
[65]
Zhexin Zhang, Shiyao Cui, Yida Lu, Jingzhuo Zhou, Junxiao Yang, Hongning Wang, and Minlie Huang. 2024. Agent-SafetyBench: Evaluating the Safety of LLM Agents. arXiv:2412.14470 [cs.AI] doi:10.48550/arXiv.2412.14470
-
[67]
Junhao Zheng, Shengjie Qiu, Chengming Shi, and Qianli Ma. 2025. Towards Lifelong Learning of Large Language Models: A Survey.Comput. Surveys57, 8 (2025), 1–35. doi:10.1145/3716629 Article 193
doi:10.1145/3716629 2025
-
[70]
Kunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang, Shuyi Guo, Zhe Wang, Zhenhailong Wang, Cheng Qian, Robert Tang, Heng Ji, and Jiaxuan You. 2025. MultiAgentBench: Evaluating the Collaboration and Competition of LLM Agents. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for...
-
[2023]
Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge- Intensive Multi-Step Questions. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Toronto, Canada, 10014–10037. doi:10.18653/v1/ 2023.acl-long.557
doi:10.18653/v1/ 2023
-
[2024]
InThe Twelfth International Conference on Learning Representations
Self-RAG: Learning to Retrieve, Generate, and Critique through Self- Reflection. InThe Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=hSyW5go0v8
-
[2025]
CRMArena: Understanding the Capacity of LLM Agents to Perform Profes- sional CRM Tasks in Realistic Environments. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Lin- guistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 20...
-
[2026]
MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control.Proceedings of the AAAI Conference on Artificial Intelligence40, 44 (2026), 37565–37573. doi:10.1609/aaai.v40i44.41090
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.