REVIEW 2 minor 23 cited by
Agentic Reasoning for Large Language Models
T0 review · 0 major / 2 minor · reviewed 2026-05-17 · grok-4.3
Pith's one-line read Agentic reasoning turns large language models into autonomous agents that plan, act, and adapt through interaction.
desk verdict This survey organizes agentic reasoning into three layers and splits in-context from post-training methods, giving a usable map of the literature but no new techniques or results. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The three complementary dimensions—foundational agentic reasoning for core single-agent capabilities, self-evolving agentic reasoning for refinement through feedback and adaptation, and collective multi-agent reasoning for coordination and shared goals—organize the field and bridge thought with action.
What would settle it
Discovery of a major agentic reasoning method or framework that requires a fourth distinct category or shows substantial overlap across the proposed dimensions would challenge the survey's organizational structure.
Extended reading notes
Core claim
Agentic reasoning reframes large language models as autonomous agents that plan, act, and learn through continual interaction with their environments. The survey organizes this capability along three complementary dimensions: foundational agentic reasoning that establishes core single-agent skills including planning, tool use, and search in stable environments; self-evolving agentic reasoning that studies refinement through feedback, memory, and adaptation; and collective multi-agent reasoning that extends intelligence to collaborative coordination, knowledge sharing, and shared goals. These layers are further split into in-context reasoning that scales test-time interaction through structed
Load-bearing premise
The three dimensions of foundational, self-evolving, and collective agentic reasoning cover the entire field comprehensively without significant overlap or omission.
Editorial extensions
If this is right
- Foundational methods support reliable planning and tool use by single agents in stable environments.
- Self-evolving techniques enable agents to improve their own performance using memory and feedback over repeated interactions.
- Collective reasoning allows multiple agents to coordinate actions and share knowledge toward common objectives.
- Applications in robotics, healthcare, and autonomous research follow directly from applying the organized roadmap.
- Open challenges in long-horizon interaction and scalable multi-agent training must be resolved for broader deployment.
Reading between the lines
- The three-dimension roadmap could guide researchers in systematically identifying gaps for personalization of agent behaviors.
- Integrating explicit world modeling may emerge naturally as an extension of the foundational and self-evolving layers.
- Governance requirements for real-world agents might be derived from the coordination mechanisms in the collective dimension.
- Testable extensions could involve applying the in-context versus post-training split to new benchmarks in mathematics or science.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey organizes agentic reasoning for LLMs along three complementary dimensions: foundational agentic reasoning establishing core single-agent capabilities (planning, tool use, search) in stable environments; self-evolving agentic reasoning focusing on refinement via feedback, memory, and adaptation; and collective multi-agent reasoning addressing coordination, knowledge sharing, and shared goals. It distinguishes in-context orchestration from post-training optimization, reviews frameworks in applications like science, robotics, healthcare, autonomous research, and mathematics, and outlines open challenges such as personalization, long-horizon interaction, world modeling, scalable multi-agent training, and governance.
Significance. If the taxonomy holds, this survey makes a useful contribution by synthesizing a rapidly growing literature into a unified roadmap that connects reasoning processes with agentic action. The explicit separation of in-context scaling from post-training optimization provides a practical lens for comparing approaches, and the enumeration of concrete open challenges (personalization, world modeling, governance) supplies clear signposts for future work. The review of domain-specific frameworks adds concrete grounding to the high-level structure.
minor comments (2)
- [Introduction] Introduction: the positioning of the three dimensions as complementary and comprehensive would be clearer if the manuscript briefly noted selection criteria for the taxonomy and acknowledged possible boundary overlaps (e.g., adaptive multi-agent systems) rather than treating the partition as self-evident.
- [Applications and benchmarks] Applications and benchmarks section: a compact summary table mapping representative frameworks to the three dimensions, listing primary techniques and benchmark results, would make the review more scannable and allow readers to assess coverage at a glance.
Simulated Author's Rebuttal
We thank the referee for their positive assessment of our survey and for the recommendation of minor revision. The referee's summary accurately captures the three-layer taxonomy (foundational, self-evolving, and collective), the distinction between in-context orchestration and post-training optimization, and the enumerated open challenges. We appreciate the recognition that this structure provides a useful roadmap connecting reasoning processes with agentic action.
Circularity Check
No significant circularity in this literature survey
full rationale
This paper is a literature survey synthesizing existing agentic reasoning methods for LLMs into a high-level roadmap. It organizes the field along three complementary dimensions (foundational, self-evolving, and collective) and distinguishes in-context from post-training approaches, but presents this taxonomy explicitly as an editorial organizing lens drawn from external references rather than a derived result. No mathematical derivations, equations, fitted parameters, predictions, or uniqueness theorems appear in the manuscript. All claims are positioned as reviews of prior work, with no self-citation chains or self-definitional reductions that would make the central synthesis equivalent to its inputs by construction. The paper is therefore self-contained against external benchmarks and receives a score of zero.
Assumptions & free parameters
assumptions (1)
- domain assumption LLMs can be reframed as autonomous agents capable of planning, acting, and learning through interaction in dynamic environments
Cite this review
Pith. "Pith review of Agentic Reasoning for Large Language Models." pith.science (2026). https://pith.science/paper/UJUDHYFT
@misc{pith2026260112538,
author = {Pith},
title = {Pith review of: Agentic Reasoning for Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/UJUDHYFT}},
note = {Machine review of arXiv:2601.12538}
}
read the original abstract
Reasoning is a fundamental cognitive process underlying inference, problem-solving, and decision-making. While large language models (LLMs) demonstrate strong reasoning capabilities in closed-world settings, they struggle in open-ended and dynamic environments. Agentic reasoning marks a paradigm shift by reframing LLMs as autonomous agents that plan, act, and learn through continual interaction. In this survey, we organize agentic reasoning along three complementary dimensions. First, we characterize environmental dynamics through three layers: foundational agentic reasoning, which establishes core single-agent capabilities including planning, tool use, and search in stable environments; self-evolving agentic reasoning, which studies how agents refine these capabilities through feedback, memory, and adaptation; and collective multi-agent reasoning, which extends intelligence to collaborative settings involving coordination, knowledge sharing, and shared goals. Across these layers, we distinguish in-context reasoning, which scales test-time interaction through structured orchestration, from post-training reasoning, which optimizes behaviors via reinforcement learning and supervised fine-tuning. We further review representative agentic reasoning frameworks across real-world applications and benchmarks, including science, robotics, healthcare, autonomous research, and mathematics. This survey synthesizes agentic reasoning methods into a unified roadmap bridging thought and action, and outlines open challenges and future directions, including personalization, long-horizon interaction, world modeling, scalable multi-agent training, and governance for real-world deployment.
Forward citations
Cited by 23 Pith papers
-
Co-Evolving Skill Generation and Policy Optimization
Framework estimates context-dependent marginal utility of candidate skills via reward gaps in matched base vs. skill-augmented rollouts to filter skills and co-train policy as generator.
-
AVTrack: Audio-Visual Tracking in Human-centric Complex Scenes
Introduces AVTrack dataset for audio-visual tracking in challenging human-centric scenes, demonstrating performance drops in existing methods.
-
ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents
ClawForge supplies a generator that turns scenario templates into reproducible command-line tasks testing state conflict handling, where the strongest frontier model scores only 45.3 percent strict accuracy.
-
Certifying Collective Reasoning in Multi-Agent Systems via Koopman Spectral Analysis
Koopman spectral analysis of interaction traces produces convergence deadline, faction attribution, and message compression certificates that hold on a synthetic attention-consensus model of LLM debate.
-
Towards a Risk Assessment of Malicious Skill Files in Coding Agents
In sandboxed runs, Gemini CLI declared intent to run malicious preflight commands in 96.1% of cases and Qwen Code in 74.0%, while only one run showed verified execution.
-
VC-Tooler: Learning Compositional and Adaptive Visual Tool Use
VC-Tooler trains Qwen3-VL-8B on hierarchically synthesized tool trajectories via SFT then GRPO with a judge-based tool reward, reaching open-source SOTA on V* (95.8) and VTC-Bench (35.3) and transferring to 11 unseen ...
-
From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents
Outcome-verified teacher continuations from student failure prefixes, plus divergence-local comparison and suffix distillation, raise skill-free agent success over scoring-based self-distillation.
-
PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning
Adaptively expanding and compressing prompt-scaffold guidance during RL lets LLM agents train to competitive performance without any skill library at test time.
-
C-PTQ: Fisher-weighted Channel-wise Sensitivity for Post-training Quantization of MLLMs
C-PTQ weights quantization error by per-channel Fisher information of the task loss, improving low-bit accuracy of multimodal LLMs by small margins over existing channel-wise scaling methods.
-
Towards Self-Evolving Agents: A Human-Inspired Adaptive Exploration-Exploitation Framework for Genetic Network Programming
HGNP improves GNP and its variants via adaptive crossover protecting high-in-degree nodes later, early-favoring mutation of judgment-to-judgment links, and cycle elimination, with HGNP-SBGNP best on Tileworld.
-
TSQAgent: Rating Time Series Data Quality via Dedicated Agentic Reasoning
TSQAgent uses three collaborative LLM agents with analytical tools to identify relevant quality dimensions and enable quantitative comparisons for time series data, improving on standard LLM methods and leading to bet...
-
Adaptive Latent Agentic Reasoning
ALAR trains LLM agents to perform most reasoning in a latent space supervised by actions and escalates to explicit CoT only when needed, cutting tokens by up to 84.6% while preserving accuracy on search and tool-use b...
-
SEAL: Synergistic Co-Evolution of Agents and Learning Environments
SEAL co-evolves LLM agents and environments via shared turn-level failure diagnoses, yielding +8.25 to +26.25 point gains on tool-use tasks with only 400 samples.
-
Exploring Robust Multi-Agent Workflows for Environmental Data Management
A role-separated multi-agent workflow with deterministic validation gates blocked a coordinate-transformation error before publication and completed a 2,452-station dataset release in two days, based on two non-contro...
-
Orchestrating Power Grid Studies with Multi-Agent AI and MCP Servers
The authors propose MCP-based AI orchestration for TSO grid studies and present pypowsybl-mcp, but provide only qualitative evidence.
-
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application
This survey categorizes agentic environments for LLMs by eight attributes and domains, introduces symbolic and neural synthesis paradigms with evaluation, and outlines four agent evolution pathways plus three environm...
-
Rethinking Continual Experience Internalization for Self-Evolving LLM Agents
Existing methods for turning LLM interaction experience into parametric skills collapse over multiple iterations; principle-level experience, step-wise injection, and off-policy teacher distillation yield more stable ...
-
The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?
LLM-as-a-judge evaluators are claimed to reward hollow but formally elaborate reasoning under simulated consensus pressure, with a logistic detector transferring across three benchmarks.
-
A Dual-Helix Governance Approach Towards Reliable Agentic Artificial Intelligence for WebGIS Development
Externalizing domain knowledge and constraints into a persistent knowledge graph reduced output variance of an LLM coding agent on a WebGIS refactoring workflow by about half, but the evidence is thin and the abstract...
-
ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning
An episode-level variant of GRPO (ESPO) improves personalized GUI-agent reasoning on the 102-episode SmartSpot benchmark, outperforming step-wise and outcome-only training baselines.
-
Critique of Agent Model
Distinguishes agentic (externally scaffolded) from agentive (internally structured) AI systems and proposes the Goal-Identity-Configurator architecture for endogenous autonomy.
-
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
Autonomous AI becomes dependable when tool use is embedded in persistent workspaces with reusable skills, shifting evaluation from answers to task closure.
-
OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems
A review-plus-demo claiming agentic capabilities emerge from system integration, backed by a 15-task benchmark whose C1→C3 performance gap is largely built into the test design.
Reference graph
Works this paper leans on
-
[1]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022
work page 2022
-
[2]
Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, et al. Least-to-most prompting enables complex reasoning in large language models.arXiv preprint arXiv:2205.10625, 2022
work page Pith review arXiv 2022
-
[3]
Pal: Program-aided language models
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. Pal: Program-aided language models. InInternational Conference on Machine Learning, pages 10764–10799. PMLR, 2023
work page 2023
-
[4]
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023
work page 2023
-
[5]
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. InInternational Conference on Learning Representations (ICLR), 2023
work page 2023
-
[6]
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools.Advances in Neural Information Processing Systems, 36:68539–68551, 2023
work page 2023
-
[7]
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.Advances in Neural Information Processing Systems, 36:38154–38180, 2023
work page 2023
-
[8]
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6):186345, 2024
work page 2024
Show all 300 references
-
[9]
Agentic retrieval-augmented generation: A survey on agentic rag.arXiv preprint arXiv:2501.09136, 2025
Aditi Singh, Abul Ehtesham, Saket Kumar, and Tala Talaei Khoei. Agentic retrieval-augmented generation: A survey on agentic rag.arXiv preprint arXiv:2501.09136, 2025
2025 arXiv
-
[10]
A survey on retrieval-augmented text generation for large language models.arXiv preprint arXiv:2404.10981, 2024
Yizheng Huang and Jimmy Huang. A survey on retrieval-augmented text generation for large language models.arXiv preprint arXiv:2404.10981, 2024
2024 arXiv
-
[11]
Openhands: An open platform for ai software developers as generalist agents.arXiv preprint arXiv:2407.16741, 2024
Xingyao Wang, Boxuan Li, Yufan Song, Frank F Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, et al. Openhands: An open platform for ai software developers as generalist agents.arXiv preprint arXiv:2407.16741, 2024. 74 Agentic Reasoning for La...
2024 arXiv
-
[12]
Mem0: Building production-ready ai agents with scalable long-term memory.arXiv preprint arXiv:2504.19413, 2025
Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. Mem0: Building production-ready ai agents with scalable long-term memory.arXiv preprint arXiv:2504.19413, 2025
2025 arXiv
-
[13]
Memos: An operating system for memory-augmented generation (mag) in large language models.arXiv preprint arXiv:2505.22101, 2025
Zhiyu Li, Shichao Song, Hanyu Wang, Simin Niu, Ding Chen, Jiawei Yang, Chenyang Xi, Huayi Lai, Jihao Zhao, Yezhaohui Wang, et al. Memos: An operating system for memory-augmented generation (mag) in large language models.arXiv preprint arXiv:2505.22101, 2025
2025
-
[14]
Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023
2023
-
[15]
Pan, Hinrich Schütze, Volker Tresp, and Yunpu Ma
Sikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma, Kristian Kersting, Jeff Z. Pan, Hinrich Schütze, Volker Tresp, and Yunpu Ma. Memory-r1: Enhancing large language model agents to manage and utilize memories via reinforcement learning.arXi...
2025 arXiv
-
[16]
Autoagents: A framework for automatic agent generation.arXiv preprint arXiv:2309.17288, 2023
Guangyao Chen, Siwei Dong, Yu Shu, Ge Zhang, Jaward Sesay, Börje F Karlsson, Jie Fu, and Yemin Shi. Autoagents: A framework for automatic agent generation.arXiv preprint arXiv:2309.17288, 2023
2023
-
[17]
MetaGPT: Meta programming for a multi-agent collaborative framework
Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber. MetaGPT: Meta programming for a multi-agent collaborative fr...
2024
-
[18]
Unleashing cogni- tive synergy in large language models: A task-solving agent through multi-persona self-collaboration
Zhenhailong Wang, Shaoguang Mao, Wenshan Wu, Tao Ge, Furu Wei, and Heng Ji. Unleashing cogni- tive synergy in large language models: A task-solving agent through multi-persona self-collaboration. InProc. 2024 Annual Conference of the North American Chapter of the Association f...
2024
-
[19]
Battleagentbench: A benchmark for evaluating cooperation and competition capabilities of language models in multi-agent systems.arXiv preprint arXiv:2408.15971, 2024
Wei Wang, Dan Zhang, Tao Feng, Boyan Wang, and Jie Tang. Battleagentbench: A benchmark for evaluating cooperation and competition capabilities of language models in multi-agent systems.arXiv preprint arXiv:2408.15971, 2024
2024
-
[20]
Agentbench: Evaluating llms as agents.arXiv preprint arXiv:2308.03688, 2023
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang. Agentbench:...
2023 arXiv
-
[21]
Multiagentbench: Evaluating the collaboration and competition of llm agents.arXiv preprint arXiv:2503.01935, 2025
Kunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang, Shuyi Guo, Zhe Wang, Zhenhailong Wang, Cheng Qian, Xiangru Tang, Heng Ji, et al. Multiagentbench: Evaluating the collaboration and competition of llm agents.arXiv preprint arXiv:2503.01935, 2025
2025
-
[22]
Tree-of-code: A self-growing tree framework for end-to-end code generation and execution in complex tasks
Ziyi Ni, Yifan Li, Ning Yang, Dou Shen, Pin Lyu, and Daxiang Dong. Tree-of-code: A self-growing tree framework for end-to-end code generation and execution in complex tasks. InFindings of the Association for Computational Linguistics: ACL 2025, pages 9804–9819, 2025
2025
-
[23]
Search-o1: Agentic search-enhanced large reasoning models.arXiv preprint arXiv:2501.05366, 2025
XiaoxiLi, GuantingDong, JiajieJin, YuyaoZhang, YujiaZhou, YutaoZhu, PeitianZhang, andZhicheng Dou. Search-o1: Agentic search-enhanced large reasoning models.arXiv preprint arXiv:2501.05366, 2025. 75 Agentic Reasoning for Large Language Models
2025 arXiv
-
[24]
A-mem: Agentic memory for llm agents.arXiv preprint arXiv:2502.12110, 2025
Wujiang Xu, Kai Mei, Hang Gao, Juntao Tan, Zujie Liang, and Yongfeng Zhang. A-mem: Agentic memory for llm agents.arXiv preprint arXiv:2502.12110, 2025
2025 arXiv
-
[25]
Evo-memory: Benchmarking llm agent test-time learning with self-evolving memory.arXiv preprint arXiv:2511.20857, 2025
TianxinWei,NoveenSachdeva,BenjaminColeman,ZhankuiHe,YuanchenBei,XuyingNing,Mengting Ai, Yunzhe Li, Jingrui He, Ed H Chi, et al. Evo-memory: Benchmarking llm agent test-time learning with self-evolving memory.arXiv preprint arXiv:2511.20857, 2025
2025 arXiv
-
[26]
Coevolving with the other you: Fine-tuning llm with sequential cooperative multi-agent reinforcement learning
Hao Ma, Tianyi Hu, Zhiqiang Pu, Liu Boyin, Xiaolin Ai, Yanyan Liang, and Min Chen. Coevolving with the other you: Fine-tuning llm with sequential cooperative multi-agent reinforcement learning. Advances in Neural Information Processing Systems, 37:15497–15525, 2024
2024
-
[27]
Search-r1: Training llms to reason and leverage search engines with reinforcement learning.arXiv preprint arXiv:2503.09516, 2025
Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan Arik, Dong Wang, Hamed Zamani, and Jiawei Han. Search-r1: Training llms to reason and leverage search engines with reinforcement learning.arXiv preprint arXiv:2503.09516, 2025
2025 arXiv
-
[28]
Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning.arXiv preprint arXiv:2505.16421, 2025
Zhepei Wei, Wenlin Yao, Yao Liu, Weizhi Zhang, Qin Lu, Liang Qiu, Changlong Yu, Puyang Xu, Chao Zhang, Bing Yin, et al. Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning.arXiv preprint arXiv:2505.16421, 2025
2025
-
[29]
Solving olympiad geometry without human demonstrations.Nature, 625(7995):476–482, 2024
Trieu H Trinh, Yuhuai Wu, Quoc V Le, He He, and Thang Luong. Solving olympiad geometry without human demonstrations.Nature, 625(7995):476–482, 2024
2024
-
[30]
Mathematical discoveries from program search with large language models.Nature, 625(7995): 468–475, 2024
Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M Pawan Kumar, Emilien Dupont, Francisco JR Ruiz, Jordan S Ellenberg, Pengming Wang, Omar Fawzi, et al. Mathematical discoveries from program search with large language models.Nature, 625(7995)...
2024
-
[31]
Vibe coding vs
Ranjan Sapkota, Konstantinos I Roumeliotis, and Manoj Karkee. Vibe coding vs. agentic coding: Fundamentals and practical implications of agentic AI, 2025
2025
-
[32]
Vibe coding — wikipedia.https://en.wikipedia.org/wiki/Vibe_coding, 2025
Andrej Karpathy. Vibe coding — wikipedia.https://en.wikipedia.org/wiki/Vibe_coding, 2025
2025
-
[33]
Chemcrow: Augmenting large-language models with chemistry tools.arXiv preprint arXiv:2304.05376, 2023
Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. Chemcrow: Augmenting large-language models with chemistry tools.arXiv preprint arXiv:2304.05376, 2023
2023 arXiv
-
[34]
Physical ai agents: Integrating cognitive intelligence with real-world action
Fouad Bousetouane. Physical ai agents: Integrating cognitive intelligence with real-world action. arXiv preprint arXiv:2501.08944, 2025
2025
-
[35]
Matexpert: Decomposing materials discovery by mimicking human experts.arXiv preprint arXiv:2410.21317, 2024
Qianggang Ding, Santiago Miret, and Bang Liu. Matexpert: Decomposing materials discovery by mimicking human experts.arXiv preprint arXiv:2410.21317, 2024
2024
-
[36]
Voyager: An open-ended embodied agent with large language models.arXiv preprint arXiv:2305.16291, 2023
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models.arXiv preprint arXiv:2305.16291, 2023
2023 arXiv
-
[37]
Embodiedrag: Dynamic 3d scene graph retrieval for efficient and scalable robot task planning.arXiv preprint arXiv:2410.23968, 2024
Booker Meghan, Byrd Grayson, Kemp Bethany, Schmidt Aurora, and Rivera Corban. Embodiedrag: Dynamic 3d scene graph retrieval for efficient and scalable robot task planning.arXiv preprint arXiv:2410.23968, 2024. URLhttps://www.arxiv.org/abs/2410.23968. 76 Agentic Reasoning for L...
2024
-
[38]
Embodied-r: Collaborative framework for activating embodied spatial reasoning in foundation models via reinforcement learning.arXiv preprint arXiv:2504.12680, 2025
Baining Zhao, Ziyou Wang, Jianjie Fang, Chen Gao, Fanhang Man, Jinqiang Cui, Xin Wang, Xinlei Chen, Yong Li, and Wenwu Zhu. Embodied-r: Collaborative framework for activating embodied spatial reasoning in foundation models via reinforcement learning.arXiv preprint arXiv:2504.1...
2025
-
[39]
Mmedagent: Learning to use medical tools with multi-modal agent.arXiv preprint arXiv:2407.02483, 2024
Binxu Li, Tiankai Yan, Yuanting Pan, Jie Luo, Ruiyang Ji, Jiayuan Ding, Zhe Xu, Shilong Liu, Haoyu Dong, Zihao Lin, et al. Mmedagent: Learning to use medical tools with multi-modal agent.arXiv preprint arXiv:2407.02483, 2024
2024
-
[40]
Biomni: A general-purpose biomedical ai agent.biorxiv, 2025
Kexin Huang, Serena Zhang, Hanchen Wang, Yuanhao Qu, Yingzhou Lu, Yusuf Roohani, Ryan Li, Lin Qiu, Gavin Li, Junze Zhang, et al. Biomni: A general-purpose biomedical ai agent.biorxiv, 2025
2025
-
[41]
Websailor: Navigating super-human reasoning for web agent
Kuan Li, Zhongwang Zhang, Huifeng Yin, Liwen Zhang, Litu Ou, Jialong Wu, Wenbiao Yin, Baixuan Li, Zhengwei Tao, Xinyu Wang, et al. Websailor: Navigating super-human reasoning for web agent. arXiv preprint arXiv:2507.02592, 2025
2025
-
[42]
Skillweaver: Web agents can self-improve by discovering and honing skills.arXiv preprint arXiv:2504.07079, 2025
Boyuan Zheng, Michael Y Fatemi, Xiaolong Jin, Zora Zhiruo Wang, Apurva Gandhi, Yueqi Song, Yu Gu, Jayanth Srinivasa, Gaowen Liu, Graham Neubig, et al. Skillweaver: Web agents can self-improve by discovering and honing skills.arXiv preprint arXiv:2504.07079, 2025
2025 arXiv
-
[43]
Ai agents vs
Ranjan Sapkota, Konstantinos I Roumeliotis, and Manoj Karkee. Ai agents vs. agentic ai: A conceptual taxonomy, applications and challenges.arXiv preprint arXiv:2505.10468, 2025
2025
-
[44]
A dynamic llm-powered agent network for task-oriented agent collaboration
Zijun Liu, Yanzhe Zhang, Peng Li, Yang Liu, and Diyi Yang. A dynamic llm-powered agent network for task-oriented agent collaboration. InFirst Conference on Language Modeling, 2024
2024
-
[45]
Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig
Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. Webarena: A realistic web environment for building autonomous agents.arXiv preprint arXiv:2307.13854, 2023. URLhttps://...
2023 arXiv
-
[46]
Visualwebarena: Evaluat- ing multimodal agents on realistic visual web tasks.arXiv preprint arXiv:2401.13649, 2024
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Gra- ham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried. Visualwebarena: Evaluat- ing multimodal agents on realistic visual web tasks.arXiv preprint arXiv:2401.13649, 2024. URL http...
2024
-
[47]
Videowebarena: Evaluating long context multimodal agents with video understanding web tasks.arXiv preprint arXiv:2410.19100, 2024
Lawrence Jang, Yinheng Li, Dan Zhao, Charles Ding, Justin Lin, Paul Pu Liang, Rogerio Bonatti, and Kazuhito Koishida. Videowebarena: Evaluating long context multimodal agents with video understanding web tasks.arXiv preprint arXiv:2410.19100, 2024
2024
-
[48]
Alfworld: Aligning text and embodied environments for interactive learning.arXiv preprint arXiv:2010.03768, 2020
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. Alfworld: Aligning text and embodied environments for interactive learning.arXiv preprint arXiv:2010.03768, 2020
2010 arXiv
-
[49]
Mind2web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36:28091–28114, 2023
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36:28091–28114, 2023
2023
-
[50]
Mind2web 2: Evaluating agentic search with agent-as-a-judge.arXiv preprint arXiv:2506.21506, 2025
Boyu Gou, Zanming Huang, Yuting Ning, Yu Gu, Michael Lin, Weijian Qi, Andrei Kopanev, Botao Yu, Bernal Jiménez Gutiérrez, Yiheng Shu, et al. Mind2web 2: Evaluating agentic search with agent-as-a-judge.arXiv preprint arXiv:2506.21506, 2025. 77 Agentic Reasoning for Large Langua...
2025
-
[51]
Towards reasoning in large language models: A survey
Jie Huang and Kevin Chen-Chuan Chang. Towards reasoning in large language models: A survey. arXiv preprint arXiv:2212.10403, 2022
2022
-
[52]
Towards reasoning era: A survey of long chain-of-thought for reasoning large language models.arXiv preprint arXiv:2503.09567, 2025
Qiguang Chen, Libo Qin, Jinhao Liu, Dengyun Peng, Jiannan Guan, Peng Wang, Mengkang Hu, Yuhang Zhou, Te Gao, and Wanxiang Che. Towards reasoning era: A survey of long chain-of-thought for reasoning large language models.arXiv preprint arXiv:2503.09567, 2025
2025 arXiv
-
[53]
Towards large reasoning models: A survey of reinforced reasoning with large language models.arXiv preprint arXiv:2501.09686, 2025
Fengli Xu, Qianyue Hao, Zefang Zong, Jingwei Wang, Yunke Zhang, Jingyi Wang, Xiaochong Lan, Jiahui Gong, Tianjian Ouyang, Fanjin Meng, et al. Towards large reasoning models: A survey of reinforced reasoning with large language models.arXiv preprint arXiv:2501.09686, 2025
2025 arXiv
-
[54]
A survey of frontiers in llm reasoning: Inference scaling, learning to reason, and agentic systems.arXiv preprint arXiv:2504.09037, 2025
Zixuan Ke, Fangkai Jiao, Yifei Ming, Xuan-Phi Nguyen, Austin Xu, Do Xuan Long, Minzhi Li, Chengwei Qin, Peifeng Wang, Silvio Savarese, et al. A survey of frontiers in llm reasoning: Inference scaling, learning to reason, and agentic systems.arXiv preprint arXiv:2504.09037, 2025
2025
-
[55]
A survey of reinforcement learning for large reasoning models.arXiv preprint arXiv:2509.08827, 2025
Kaiyan Zhang, Yuxin Zuo, Bingxiang He, Youbang Sun, Runze Liu, Che Jiang, Yuchen Fan, Kai Tian, Guoli Jia, Pengfei Li, et al. A survey of reinforcement learning for large reasoning models.arXiv preprint arXiv:2509.08827, 2025
2025
-
[56]
The landscape of agentic reinforcement learning for llms: A survey
Guibin Zhang, Hejia Geng, Xiaohang Yu, Zhenfei Yin, Zaibin Zhang, Zelin Tan, Heng Zhou, Zhongzhi Li, Xiangyuan Xue, Yijiang Li, et al. The landscape of agentic reinforcement learning for llms: A survey. arXiv preprint arXiv:2509.02547, 2025
2025 arXiv
-
[57]
A comprehensive survey on reinforcement learning-based agentic search: Foundations, roles, optimizations, evaluations, and applications.arXiv preprint arXiv:2510.16724, 2025
Minhua Lin, Zongyu Wu, Zhichao Xu, Hui Liu, Xianfeng Tang, Qi He, Charu Aggarwal, Xiang Zhang, and Suhang Wang. A comprehensive survey on reinforcement learning-based agentic search: Foundations, roles, optimizations, evaluations, and applications.arXiv preprint arXiv:2510.16724, 2025
2025
-
[58]
A comprehensive survey of self-evolving ai agents: A new paradigm bridging foundation models and lifelong agentic systems.arXiv preprint arXiv:2508.07407, 2025
Jinyuan Fang, Yanwen Peng, Xi Zhang, Yingxu Wang, Xinhao Yi, Guibin Zhang, Yi Xu, Bin Wu, Siwei Liu, Zihao Li, et al. A comprehensive survey of self-evolving ai agents: A new paradigm bridging foundation models and lifelong agentic systems.arXiv preprint arXiv:2508.07407, 2025
2025 arXiv
-
[59]
A survey of self-evolving agents: On path to artificial super intelligence.arXiv preprint arXiv:2507.21046, 2025
Huan-ang Gao, Jiayi Geng, Wenyue Hua, Mengkang Hu, Xinzhe Juan, Hongzhang Liu, Shilong Liu, Jiahao Qiu, Xuan Qi, Yiran Wu, et al. A survey of self-evolving agents: On path to artificial super intelligence.arXiv preprint arXiv:2507.21046, 2025
2025 arXiv
-
[60]
Deepseek-r1: Incentivizingreasoningcapabilityinllmsviareinforcement learning.arXiv preprint arXiv:2501.12948, 2025
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma,PeiyiWang,XiaoBi,etal. Deepseek-r1: Incentivizingreasoningcapabilityinllmsviareinforcement learning.arXiv preprint arXiv:2501.12948, 2025
2025 arXiv
-
[61]
Deepretrieval: Hacking real search engines and retrievers with large language models via reinforcement learning.arXiv preprint arXiv:2503.00223, 2025
Pengcheng Jiang, Jiacheng Lin, Lang Cao, Runchu Tian, SeongKu Kang, Zifeng Wang, Jimeng Sun, and Jiawei Han. Deepretrieval: Hacking real search engines and retrievers with large language models via reinforcement learning.arXiv preprint arXiv:2503.00223, 2025
2025
-
[62]
Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347, 2017
2017 arXiv
-
[63]
Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024. 78 Agentic Reasoning for Large...
2024 arXiv
-
[64]
Arpo: End-to-end policy optimization for gui agents with experience replay.arXiv preprint arXiv:2505.16282, 2025
Fanbin Lu, Zhisheng Zhong, Shu Liu, Chi-Wing Fu, and Jiaya Jia. Arpo: End-to-end policy optimization for gui agents with experience replay.arXiv preprint arXiv:2505.16282, 2025
2025
-
[65]
Dapo: An open-source llm reinforcement learning system at scale
Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, et al. Dapo: An open-source llm reinforcement learning system at scale. arXiv preprint arXiv:2503.14476, 2025
2025 arXiv
-
[66]
Autogen: Enabling next-gen LLM applications via multi-agent conversations
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang. Autogen: Enabling next-gen LLM applications via multi-agent conversations. InFirst Confer...
2024
-
[67]
Camel: Commu- nicative agents for" mind" exploration of large language model society.Advances in Neural Information Processing Systems, 36:51991–52008, 2023
Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. Camel: Commu- nicative agents for" mind" exploration of large language model society.Advances in Neural Information Processing Systems, 36:51991–52008, 2023
2023
-
[68]
Gptswarm: Language agents as optimizable graphs
Mingchen Zhuge, Wenyi Wang, Louis Kirsch, Francesco Faccio, Dmitrii Khizbullin, and Jürgen Schmid- huber. Gptswarm: Language agents as optimizable graphs. InForty-first International Conference on Machine Learning, 2024
2024
-
[69]
Multi-agent deep research: Training multi-agent systems with m-grpo.arXiv preprint arXiv:2511.13288, 2025
Haoyang Hong, Jiajun Yin, Yuan Wang, Jingnan Liu, Zhe Chen, Ailing Yu, Ji Li, Zhiling Ye, Hansong Xiao, Yefei Chen, et al. Multi-agent deep research: Training multi-agent systems with m-grpo.arXiv preprint arXiv:2511.13288, 2025
2025
-
[70]
Alphaevolve: A coding agent for scientific and algorithmic discovery.arXiv preprint arXiv:2506.13131, 2025
Alexander Novikov, Ngân V˜u, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco JR Ruiz, Abbas Mehrabian, et al. Alphaevolve: A coding agent for scientific and algorithmic discovery.arXiv preprint arXiv:2506.1...
2025 arXiv
-
[71]
REWOO: Decoupling reasoning from observations for efficient augmented language models.arXiv preprint arXiv:2305.18323, 2023
BinfengXu, ZhiyuanPeng, BowenLei, SubhabrataMukherjee, YuchenLiu, andDongkuanXu. REWOO: Decoupling reasoning from observations for efficient augmented language models.arXiv preprint arXiv:2305.18323, 2023
2023 arXiv
-
[72]
LLM+P: Empowering large language models with optimal planning proficiency.arXiv preprint arXiv:2304.11477, 2023
Bo Liu, Yuqian Jiang, Xiaohan Zhang, Qiang Liu, Shiqi Zhang, Joydeep Biswas, and Peter Stone. LLM+P: Empowering large language models with optimal planning proficiency.arXiv preprint arXiv:2304.11477, 2023
2023 arXiv
-
[73]
On the planning abilities of large language models: A critical investigation.Advances in Neural Information Processing Systems, 36:75993–76005, 2023
Karthik Valmeekam, Matthew Marquez, Sarath Sreedharan, and Subbarao Kambhampati. On the planning abilities of large language models: A critical investigation.Advances in Neural Information Processing Systems, 36:75993–76005, 2023
2023
-
[74]
Graph of thoughts: Solving elaborate problems with large language models
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al. Graph of thoughts: Solving elaborate problems with large language models. InProceedings of the AAAI confer...
2024
-
[75]
Algorithm of thoughts: Enhancing exploration of ideas in large language models.arXiv preprint arXiv:2308.10379, 2023
Bilgehan Sel, Ahmad Al-Tawaha, Vanshaj Khattar, Ruoxi Jia, and Ming Jin. Algorithm of thoughts: Enhancing exploration of ideas in large language models.arXiv preprint arXiv:2308.10379, 2023
2023
-
[76]
Hypertree planning: Enhancing llm reasoning via hierarchical thinking
Runquan Gui, Zhihai Wang, Jie Wang, Chi Ma, Huiling Zhen, Mingxuan Yuan, Jianye Hao, Defu Lian, Enhong Chen, and Feng Wu. Hypertree planning: Enhancing llm reasoning via hierarchical thinking. arXiv preprint arXiv:2505.02322, 2025. 79 Agentic Reasoning for Large Language Models
2025
-
[77]
Reflect-then-plan: Offline model-based planning through a doubly bayesian lens.arXiv preprint arXiv:2506.06261, 2025
Jihwan Jeong, Xiaoyu Wang, Jingmin Wang, Scott Sanner, and Pascal Poupart. Reflect-then-plan: Offline model-based planning through a doubly bayesian lens.arXiv preprint arXiv:2506.06261, 2025
2025
-
[78]
Gorilla: Large language model connected with massive apis.Advances in Neural Information Processing Systems, 37:126544–126565, 2024
Shishir G Patil, Tianjun Zhang, Xin Wang, and Joseph E Gonzalez. Gorilla: Large language model connected with massive apis.Advances in Neural Information Processing Systems, 37:126544–126565, 2024
2024
-
[79]
Codenav: Beyond tool-use to using real-world codebases with llm agents.arXiv preprint arXiv:2406.12276, 2024
Tanmay Gupta, Luca Weihs, and Aniruddha Kembhavi. Codenav: Beyond tool-use to using real-world codebases with llm agents.arXiv preprint arXiv:2406.12276, 2024
2024
-
[80]
Plan-on-graph: Self-correcting adaptive planning of large language model on knowledge graphs.Advances in Neural Information Processing Systems, 37:37665–37691, 2024
Liyi Chen, Panrong Tong, Zhongming Jin, Ying Sun, Jieping Ye, and Hui Xiong. Plan-on-graph: Self-correcting adaptive planning of large language model on knowledge graphs.Advances in Neural Information Processing Systems, 37:37665–37691, 2024
2024
-
[81]
Tool-planner: Task planning with clusters across multiple tools
Yanming Liu, Xinyue Peng, Jiannan Cao, Yuwei Zhang, Xuhong Zhang, Sheng Cheng, Xun Wang, Jianwei Yin, and Tianyu Du. Tool-planner: Task planning with clusters across multiple tools. InThe Thirteenth International Conference on Learning Representations, 2025. URLhttps://openrev...
2025
-
[82]
Visualpredicator: Learning abstract world models with neuro-symbolic predicates for robot planning.arXiv preprint arXiv:2410.23156, 2024
Yichao Liang, Nishanth Kumar, Hao Tang, Adrian Weller, Joshua B Tenenbaum, Tom Silver, João F Henriques, and Kevin Ellis. Visualpredicator: Learning abstract world models with neuro-symbolic predicates for robot planning.arXiv preprint arXiv:2410.23156, 2024
2024
-
[83]
Llm- planner: Few-shot grounded planning for embodied agents with large language models
Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M Sadler, Wei-Lun Chao, and Yu Su. Llm- planner: Few-shot grounded planning for embodied agents with large language models. InProceedings of the IEEE/CVF international conference on computer vision, pages 2998–3009, 2023
2023
-
[84]
Agent-e: From autonomous web navigation to foundational design principles in agentic systems.arXiv preprint arXiv:2407.13032, 2024
Tamer Abuelsaad, Deepak Akkil, Prasenjit Dey, Ashish Jagmohan, Aditya Vempaty, and Ravi Kokku. Agent-e: From autonomous web navigation to foundational design principles in agentic systems.arXiv preprint arXiv:2407.13032, 2024
2024
-
[85]
Agent s: An open agentic framework that uses computers like a human.arXiv preprint arXiv:2410.08164, 2024
Saaket Agashe, Jiuzhou Han, Shuyu Gan, Jiachen Yang, Ang Li, and Xin Eric Wang. Agent s: An open agentic framework that uses computers like a human.arXiv preprint arXiv:2410.08164, 2024
2024
-
[86]
Exploratory retrieval-augmented planning for continual embodied instruction following.Advances in Neural Information Processing Systems, 37: 67034–67060, 2024
Minjong Yoo, Jinwoo Jang, Wei-Jin Park, and Honguk Woo. Exploratory retrieval-augmented planning for continual embodied instruction following.Advances in Neural Information Processing Systems, 37: 67034–67060, 2024
2024
-
[87]
Real-time anomaly detection and reactive planning with large language models.arXiv preprint arXiv:2407.08735, 2024
Rohan Sinha, Amine Elhafsi, Christopher Agia, Matthew Foutter, Edward Schmerling, and Marco Pavone. Real-time anomaly detection and reactive planning with large language models.arXiv preprint arXiv:2407.08735, 2024
2024
-
[88]
Hierarchical planning for complex tasks with knowledge graph-rag and symbolic verification.arXiv preprint arXiv:2504.04578, 2025
Cristina Cornelio, Flavio Petruzzellis, and Pietro Lio. Hierarchical planning for complex tasks with knowledge graph-rag and symbolic verification.arXiv preprint arXiv:2504.04578, 2025
2025
-
[89]
Behaviorgpt: Smart agent simulation for autonomous driving with next-patch prediction.Advances in Neural Information Processing Systems, 37:79597–79617, 2024
Zikang Zhou, HU Haibo, Xinhong Chen, Jianping Wang, Nan Guan, Kui Wu, Yung-Hui Li, Yu-Kai Huang, and Chun Jason Xue. Behaviorgpt: Smart agent simulation for autonomous driving with next-patch prediction.Advances in Neural Information Processing Systems, 37:79597–79617, 2024
2024
-
[90]
Dino-wm: World models on pre-trained visual features enable zero-shot planning.arXiv preprint arXiv:2411.04983, 2024
Gaoyue Zhou, Hengkai Pan, Yann LeCun, and Lerrel Pinto. Dino-wm: World models on pre-trained visual features enable zero-shot planning.arXiv preprint arXiv:2411.04983, 2024. 80 Agentic Reasoning for Large Language Models
2024
-
[91]
Flip: Flow-centric generative planning as general-purpose manipulation world model.arXiv preprint arXiv:2412.08261, 2024
Chongkai Gao, Haozhuo Zhang, Zhixuan Xu, Zhehao Cai, and Lin Shao. Flip: Flow-centric generative planning as general-purpose manipulation world model.arXiv preprint arXiv:2412.08261, 2024
2024
-
[92]
LLM reasoners: New evaluation, library, and analysis of step-by-step reasoning with large language models.arXiv preprint arXiv:2404.05221, 2024
Shibo Hao, Yi Gu, Haotian Luo, Tianyang Liu, Xiyan Shao, Xinyuan Wang, Shuhua Xie, Haodi Ma, Adithya Samavedhi, Qiyue Gao, et al. LLM reasoners: New evaluation, library, and analysis of step-by-step reasoning with large language models.arXiv preprint arXiv:2404.05221, 2024
2024
-
[93]
Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models
Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistic...
2023
-
[94]
Plan, verify and switch: Integrated reasoning with diverse x-of-thoughts.arXiv preprint arXiv:2310.14628, 2023
TengxiaoLiu,QipengGuo,YuqingYang, XiangkunHu,YueZhang,XipengQiu,andZhengZhang. Plan, verify and switch: Integrated reasoning with diverse x-of-thoughts.arXiv preprint arXiv:2310.14628, 2023
2023
-
[95]
Peria: Perceive, reason, imagine, act via holistic language and vision planning for manipulation.Advances in Neural Information Processing Systems, 37:17541–17571, 2024
Fei Ni, Jianye Hao, Shiguang Wu, Longxin Kou, Yifu Yuan, Zibin Dong, Jinyi Liu, MingZhi Li, Yuzheng Zhuang, and Yan Zheng. Peria: Perceive, reason, imagine, act via holistic language and vision planning for manipulation.Advances in Neural Information Processing Systems, 37:175...
2024
-
[96]
Plan-and-act: Improving planning of agents for long-horizon tasks
Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon, Hiroki Furuta, Gopala Anumanchipalli, Kurt Keutzer, and Amir Gholami. Plan-and-act: Improving planning of agents for long-horizon tasks. arXiv preprint arXiv:2503.09572, 2025
2025
-
[97]
Codeplan: Unlocking reasoning potential in large language models by scaling code-form planning
Jiaxin Wen, Jian Guan, Hongning Wang, Wei Wu, and Minlie Huang. Codeplan: Unlocking reasoning potential in large language models by scaling code-form planning. InThe Thirteenth International Conference on Learning Representations, 2024
2024
-
[98]
Wilbur: Adaptive in-context learning for robust and accurate web agents.arXiv preprint arXiv:2404.05902, 2024
Michael Lutz, Arth Bohra, Manvel Saroyan, Artem Harutyunyan, and Giovanni Campagna. Wilbur: Adaptive in-context learning for robust and accurate web agents.arXiv preprint arXiv:2404.05902, 2024
2024
-
[99]
Executable code actions elicit better llm agents
Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, and Heng Ji. Executable code actions elicit better llm agents. InForty-first International Conference on Machine Learning, 2024
2024
-
[100]
Marco: Multi-agent code optimization with real-time knowledge integration for high-performance computing.arXiv preprint arXiv:2505.03906, 2025
Asif Rahman, Veljko Cvetkovic, Kathleen Reece, Aidan Walters, Yasir Hassan, Aneesh Tummeti, Bryan Torres, Denise Cooney, Margaret Ellis, and Dimitrios S Nikolopoulos. Marco: Multi-agent code optimization with real-time knowledge integration for high-performance computing.arXiv...
2025
-
[101]
Enhancing llm reasoningwithmulti-pathcollaborativereactiveandreflectionagents.arXivpreprintarXiv:2501.00430, 2024
Chengbo He, Bochao Zou, Xin Li, Jiansheng Chen, Junliang Xing, and Huimin Ma. Enhancing llm reasoningwithmulti-pathcollaborativereactiveandreflectionagents.arXivpreprintarXiv:2501.00430, 2024
2024
-
[102]
Pre-act: Multi-step planning and reasoning improves acting in llm agents.arXiv preprint arXiv:2505.09970, 2025
Mrinal Rawat, Ambuje Gupta, Rushil Goomer, Alessandro Di Bari, Neha Gupta, and Roberto Pier- accini. Pre-act: Multi-step planning and reasoning improves acting in llm agents.arXiv preprint arXiv:2505.09970, 2025
2025
-
[103]
Rest meets react: Self-improvement for multi-step reasoning llm agent.arXiv preprint arXiv:2312.10003, 2023
Renat Aksitov, Sobhan Miryoosefi, Zonglin Li, Daliang Li, Sheila Babayan, Kavya Kopparapu, Zachary Fisher, Ruiqi Guo, Sushant Prakash, Pranesh Srinivasan, et al. Rest meets react: Self-improvement for multi-step reasoning llm agent.arXiv preprint arXiv:2312.10003, 2023. 81 Age...
2023
-
[104]
Self-planning code generation with large language models.ACM Transactions on Software Engineering and Methodology, 33(7):1–30, 2024
Xue Jiang, Yihong Dong, Lecheng Wang, Zheng Fang, Qiwei Shang, Ge Li, Zhi Jin, and Wenpin Jiao. Self-planning code generation with large language models.ACM Transactions on Software Engineering and Methodology, 33(7):1–30, 2024
2024
-
[105]
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
Dhruv Shah, Błażej Osiński, Sergey Levine, et al. Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action. InConference on robot learning, pages 492–504. PMLR, 2023
2023
-
[106]
Tree-of-traversals: A zero-shot reasoning algorithm for augmenting black-box language models with knowledge graphs.arXiv preprint arXiv:2407.21358, 2024
Elan Markowitz, Anil Ramakrishna, Jwala Dhamala, Ninareh Mehrabi, Charith Peris, Rahul Gupta, Kai- Wei Chang, and Aram Galstyan. Tree-of-traversals: A zero-shot reasoning algorithm for augmenting black-box language models with knowledge graphs.arXiv preprint arXiv:2407.21358, 2024
2024
-
[107]
Large language model guided tree-of-thought.arXiv preprint arXiv:2305.08291, 2023
Jieyi Long. Large language model guided tree-of-thought.arXiv preprint arXiv:2305.08291, 2023
2023
-
[108]
Tree search for language model agents.arXiv preprint arXiv:2407.01476, 2024
Jing Yu Koh, Stephen McAleer, Daniel Fried, and Ruslan Salakhutdinov. Tree search for language model agents.arXiv preprint arXiv:2407.01476, 2024
2024
-
[109]
Q*: Improving multi-step reasoning for llms with deliberative planning.arXiv preprint arXiv:2406.14283, 2024
Chaojie Wang, Yanchen Deng, Zhiyi Lyu, Liang Zeng, Jujie He, Shuicheng Yan, and Bo An. Q*: Improving multi-step reasoning for llms with deliberative planning.arXiv preprint arXiv:2406.14283, 2024
2024
-
[110]
Llm-a*: Large language model enhanced incremental heuristic search on path planning.arXiv preprint arXiv:2407.02511, 2024
Silin Meng, Yiwei Wang, Cheng-Fu Yang, Nanyun Peng, and Kai-Wei Chang. Llm-a*: Large language model enhanced incremental heuristic search on path planning.arXiv preprint arXiv:2407.02511, 2024
2024
-
[111]
Multimodal large language models for inverse molecular design with retrosynthetic planning.arXiv preprint arXiv:2410.04223, 2024
Gang Liu, Michael Sun, Wojciech Matusik, Meng Jiang, and Jie Chen. Multimodal large language models for inverse molecular design with retrosynthetic planning.arXiv preprint arXiv:2410.04223, 2024
2024
-
[112]
Reasoning with language model is planning with world model.arXiv preprint arXiv:2305.14992, 2023
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. Reasoning with language model is planning with world model.arXiv preprint arXiv:2305.14992, 2023
2023 arXiv
-
[113]
Agent q: Advanced reasoning and learning for autonomous ai agents.arXiv preprint arXiv:2408.07199, 2024
Pranav Putta, Edmund Mills, Naman Garg, Sumeet Motwani, Chelsea Finn, Divyansh Garg, and Rafael Rafailov. Agent q: Advanced reasoning and learning for autonomous ai agents.arXiv preprint arXiv:2408.07199, 2024
2024
-
[114]
Monte carlo thought search: Large language model querying for complex scientific reasoning in catalyst design.arXiv preprint arXiv:2310.14420, 2023
Henry W Sprueill, Carl Edwards, Mariefel V Olarte, Udishnu Sanyal, Heng Ji, and Sutanay Choudhury. Monte carlo thought search: Large language model querying for complex scientific reasoning in catalyst design.arXiv preprint arXiv:2310.14420, 2023
2023
-
[115]
Prompt-based monte-carlo tree search for goal-oriented dialogue policy planning.arXiv preprint arXiv:2305.13660, 2023
Xiao Yu, Maximillian Chen, and Zhou Yu. Prompt-based monte-carlo tree search for goal-oriented dialogue policy planning.arXiv preprint arXiv:2305.13660, 2023
2023
-
[116]
Large language models as commonsense knowledge for large-scale task planning.Advances in neural information processing systems, 36:31967–31987, 2023
Zirui Zhao, Wee Sun Lee, and David Hsu. Large language models as commonsense knowledge for large-scale task planning.Advances in neural information processing systems, 36:31967–31987, 2023
2023
-
[117]
Everything of thoughts: Defying the law of penrose triangle for thought generation.arXiv preprint arXiv:2311.04254, 2023
Ruomeng Ding, Chaoyun Zhang, Lu Wang, Yong Xu, Minghua Ma, Wei Zhang, Si Qin, Saravan Rajmohan, Qingwei Lin, and Dongmei Zhang. Everything of thoughts: Defying the law of penrose triangle for thought generation.arXiv preprint arXiv:2311.04254, 2023
2023
-
[118]
When is tree search useful for llm planning? it depends on the discriminator.arXiv preprint arXiv:2402.10890, 2024
Ziru Chen, Michael White, Raymond Mooney, Ali Payani, Yu Su, and Huan Sun. When is tree search useful for llm planning? it depends on the discriminator.arXiv preprint arXiv:2402.10890, 2024. 82 Agentic Reasoning for Large Language Models
2024
-
[119]
Latent plan transformer for trajectory abstraction: Planning as latent space inference.Advances in Neural Information Processing Systems, 37:123379–123401, 2024
Deqian Kong, Dehong Xu, Minglu Zhao, Bo Pang, Jianwen Xie, Andrew Lizarraga, Yuhao Huang, Sirui Xie, and Ying Nian Wu. Latent plan transformer for trajectory abstraction: Planning as latent space inference.Advances in Neural Information Processing Systems, 37:123379–123401, 2024
2024
-
[120]
Alphazero-like tree-search can guide large language model decoding and training.arXiv preprint arXiv:2309.17179, 2023
Xidong Feng, Ziyu Wan, Muning Wen, Stephen Marcus McAleer, Ying Wen, Weinan Zhang, and Jun Wang. Alphazero-like tree-search can guide large language model decoding and training.arXiv preprint arXiv:2309.17179, 2023
2023
-
[121]
Monte carlo tree diffusion for system 2 planning.arXiv preprint arXiv:2502.07202, 2025
Jaesik Yoon, Hyeonseo Cho, Doojin Baek, Yoshua Bengio, and Sungjin Ahn. Monte carlo tree diffusion for system 2 planning.arXiv preprint arXiv:2502.07202, 2025
2025
-
[122]
Mastering board games by external and internal planning with language models.arXiv preprint arXiv:2412.12119, 2024
John Schultz, Jakub Adamek, Matej Jusup, Marc Lanctot, Michael Kaisers, Sarah Perrin, Daniel Hennes, Jeremy Shar, Cannada Lewis, Anian Ruoss, et al. Mastering board games by external and internal planning with language models.arXiv preprint arXiv:2412.12119, 2024
2024
-
[123]
Broaden your scope! effi- cient multi-turn conversation planning for llms with semantic space.arXiv preprint arXiv:2503.11586, 2025
Zhiliang Chen, Xinyuan Niu, Chuan-Sheng Foo, and Bryan Kian Hsiang Low. Broaden your scope! effi- cient multi-turn conversation planning for llms with semantic space.arXiv preprint arXiv:2503.11586, 2025
2025
-
[124]
Self-evaluation guided beam search for reasoning.Advances in Neural Information Processing Systems, 36:41618–41650, 2023
Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, James Xu Zhao, Min-Yen Kan, Junxian He, and Michael Xie. Self-evaluation guided beam search for reasoning.Advances in Neural Information Processing Systems, 36:41618–41650, 2023
2023
-
[125]
Pathfinder: Guided search over multi-step reasoning paths.arXiv preprint arXiv:2312.05180, 2023
Olga Golovneva, Sean O’Brien, Ramakanth Pasunuru, Tianlu Wang, Luke Zettlemoyer, Maryam Fazel- Zarandi, and Asli Celikyilmaz. Pathfinder: Guided search over multi-step reasoning paths.arXiv preprint arXiv:2312.05180, 2023
2023
-
[126]
Discriminator-guided em- bodied planning for llm agent
Haofu Qian, Chenjia Bai, Jiatao Zhang, Fei Wu, Wei Song, and Xuelong Li. Discriminator-guided em- bodied planning for llm agent. InThe Thirteenth International Conference on Learning Representations, 2025
2025
-
[127]
Stream of search (sos): Learning to search in language.arXiv preprint arXiv:2404.03683, 2024
Kanishk Gandhi, Denise Lee, Gabriel Grand, Muxin Liu, Winson Cheng, Archit Sharma, and Noah D Goodman. Stream of search (sos): Learning to search in language.arXiv preprint arXiv:2404.03683, 2024
2024
-
[128]
System-1
Swarnadeep Saha, Archiki Prasad, Justin Chih-Yao Chen, Peter Hase, Elias Stengel-Eskin, and Mohit Bansal. System-1. x: Learning to balance fast and slow planning with language models.arXiv preprint arXiv:2407.14414, 2024
2024
-
[129]
Intelligent virtual assistants with llm-based process automation.arXiv preprint arXiv:2312.06677, 2023
Yanchu Guan, Dong Wang, Zhixuan Chu, Shiyu Wang, Feiyue Ni, Ruihua Song, Longfei Li, Jinjie Gu, and Chenyi Zhuang. Intelligent virtual assistants with llm-based process automation.arXiv preprint arXiv:2312.06677, 2023
2023
-
[130]
Enhancing llm-based agents via global planning and hierarchical execution.arXiv preprint arXiv:2504.16563, 2025
Junjie Chen, Haitao Li, Jingli Yang, Yiqun Liu, and Qingyao Ai. Enhancing llm-based agents via global planning and hierarchical execution.arXiv preprint arXiv:2504.16563, 2025
2025
-
[131]
Divide and conquer: Grounding llms as efficient decision-making agents via offline hierarchical reinforcement learning.arXiv preprint arXiv:2505.19761, 2025
Zican Hu, Wei Liu, Xiaoye Qu, Xiangyu Yue, Chunlin Chen, Zhi Wang, and Yu Cheng. Divide and conquer: Grounding llms as efficient decision-making agents via offline hierarchical reinforcement learning.arXiv preprint arXiv:2505.19761, 2025. 83 Agentic Reasoning for Large Language Models
2025
-
[132]
Swe- search: Enhancing software agents with monte carlo tree search and iterative refinement.arXiv preprint arXiv:2410.20285, 2024
Antonis Antoniades, Albert Örwall, Kexun Zhang, Yuxi Xie, Anirudh Goyal, and William Wang. Swe- search: Enhancing software agents with monte carlo tree search and iterative refinement.arXiv preprint arXiv:2410.20285, 2024
2024
-
[133]
Llm-brain: Ai-driven fast generation of robot behaviour tree based on large language model
Artem Lykov and Dzmitry Tsetserukou. Llm-brain: Ai-driven fast generation of robot behaviour tree based on large language model. In2024 2nd International Conference on Foundation and Large Language Models (FLLM), pages 392–397. IEEE, 2024
2024
-
[134]
Robot behavior-tree-based task generation with large language models.arXiv preprint arXiv:2302.12927, 2023
Yue Cao and CS Lee. Robot behavior-tree-based task generation with large language models.arXiv preprint arXiv:2302.12927, 2023
2023
-
[135]
Btgenbot: Behavior tree generation for robotic tasks with lightweight llms
Riccardo Andrea Izzo, Gianluca Bardaro, and Matteo Matteucci. Btgenbot: Behavior tree generation for robotic tasks with lightweight llms. In2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 9684–9690. IEEE, 2024
2024
-
[136]
Do as i can, not as i say: Grounding language in robotic affordances.arXiv preprint arXiv:2204.01691, 2022
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. Do as i can, not as i say: Grounding language in robotic affordances.arXiv preprint arXiv:2204.01691, 2022
2022 arXiv
-
[137]
Inner monologue: Embodied reasoning through planning with language models.arXiv preprint arXiv:2207.05608, 2022
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al. Inner monologue: Embodied reasoning through planning with language models.arXiv preprint arXiv:2207.05608, 2022
2022 arXiv
-
[138]
Leveraging pre- trained large language models to construct and utilize world models for model-based task planning
Lin Guan, Karthik Valmeekam, Sarath Sreedharan, and Subbarao Kambhampati. Leveraging pre- trained large language models to construct and utilize world models for model-based task planning. Advances in Neural Information Processing Systems, 36:79081–79094, 2023
2023
-
[139]
Leveraging environment interaction for automated pddl translation and planning with large language models.Advances in Neural Information Processing Systems, 37:38960–39008, 2024
Sadegh Mahdavi, Raquel Aoki, Keyi Tang, and Yanshuai Cao. Leveraging environment interaction for automated pddl translation and planning with large language models.Advances in Neural Information Processing Systems, 37:38960–39008, 2024
2024
-
[140]
Thought of search: Planning with language models through the lens of efficiency.Advances in Neural Information Processing Systems, 37:138491–138568, 2024
Michael Katz, Harsha Kokel, Kavitha Srinivas, and Shirin Sohrabi Araghi. Thought of search: Planning with language models through the lens of efficiency.Advances in Neural Information Processing Systems, 37:138491–138568, 2024
2024
-
[141]
Planning anything with rigor: General-purpose zero-shot planning with llm-based formalized programming.arXiv preprint arXiv:2410.12112, 2024
Yilun Hao, Yang Zhang, and Chuchu Fan. Planning anything with rigor: General-purpose zero-shot planning with llm-based formalized programming.arXiv preprint arXiv:2410.12112, 2024
2024
-
[142]
From an llm swarm to a pddl-empowered hive: Planning self-executed instructions in a multi-modal jungle.arXiv preprint arXiv:2412.12839, 2024
Kaustubh Vyas, Damien Graux, Yijun Yang, Sébastien Montella, Chenxin Diao, Wendi Zhou, Pavlos Vougiouklis, Ruofei Lai, Yang Ren, Keshuang Li, et al. From an llm swarm to a pddl-empowered hive: Planning self-executed instructions in a multi-modal jungle.arXiv preprint arXiv:241...
2024
-
[143]
Atomic reasoning for scientific table claim verification
Yuji Zhang, Qingyun Wang, Cheng Qian, Jiateng Liu, Chenkai Sun, Denghui Zhang, Tarek Abdelzaher, Chengxiang Zhai, Preslav Nakov, and Heng Ji. Atomic reasoning for scientific table claim verification. arXiv preprint arXiv:2506.06972, 2025
2025
-
[144]
Diffuserlite: Towards real-time diffusion planning.Advances in Neural Information Processing Systems, 37:122556– 122583, 2024
Zibin Dong, Jianye Hao, Yifu Yuan, Fei Ni, Yitian Wang, Pengyi Li, and Yan Zheng. Diffuserlite: Towards real-time diffusion planning.Advances in Neural Information Processing Systems, 37:122556– 122583, 2024. 84 Agentic Reasoning for Large Language Models
2024
-
[145]
Goal-space planning with subgoal models.Journal of Machine Learning Research, 25(330):1–57, 2024
Chunlok Lo, Kevin Roice, Parham Mohammad Panahi, Scott M Jordan, Adam White, Gabor Mihucz, Farzane Aminmansour, and Martha White. Goal-space planning with subgoal models.Journal of Machine Learning Research, 25(330):1–57, 2024
2024
-
[146]
Agent-oriented planning in multi-agent systems.arXiv preprint arXiv:2410.02189, 2024
Ao Li, Yuexiang Xie, Songze Li, Fugee Tsung, Bolin Ding, and Yaliang Li. Agent-oriented planning in multi-agent systems.arXiv preprint arXiv:2410.02189, 2024
2024
-
[147]
Goplan: Goal-conditioned offline reinforcement learning by planning with learned models.arXiv preprint arXiv:2310.20025, 2023
Mianchu Wang, Rui Yang, Xi Chen, Hao Sun, Meng Fang, and Giovanni Montana. Goplan: Goal-conditioned offline reinforcement learning by planning with learned models.arXiv preprint arXiv:2310.20025, 2023
2023
-
[148]
Retrointext: A multimodal large language model enhanced framework for retrosynthetic planning via in-context representation learning
Chenglong Kang, Xiaoyi Liu, and Fei Guo. Retrointext: A multimodal large language model enhanced framework for retrosynthetic planning via in-context representation learning. InThe Thirteenth International Conference on Learning Representations, 2025
2025
-
[149]
Beyond autoregression: Discrete diffusion for complex reasoning and planning.arXiv preprint arXiv:2410.14157, 2024
Jiacheng Ye, Jiahui Gao, Shansan Gong, Lin Zheng, Xin Jiang, Zhenguo Li, and Lingpeng Kong. Beyond autoregression: Discrete diffusion for complex reasoning and planning.arXiv preprint arXiv:2410.14157, 2024
2024
-
[150]
Planagent: A multi-modal large language agent for closed-loop vehicle motion planning.arXiv preprint arXiv:2406.01587, 2024
Yupeng Zheng, Zebin Xing, Qichao Zhang, Bu Jin, Pengfei Li, Yuhang Zheng, Zhongpu Xia, Kun Zhan, Xianpeng Lang, Yaran Chen, et al. Planagent: A multi-modal large language agent for closed-loop vehicle motion planning.arXiv preprint arXiv:2406.01587, 2024
2024
-
[151]
Long-horizon planning for multi-agent robots in partially observable environments.Advances in Neural Information Processing Systems, 37:67929–67967, 2024
Sid Nayak, Adelmo Morrison Orozco, Marina Have, Jackson Zhang, Vittal Thirumalai, Darren Chen, Aditya Kapoor, Eric Robinson, Karthik Gopalakrishnan, James Harrison, et al. Long-horizon planning for multi-agent robots in partially observable environments.Advances in Neural Info...
2024
-
[152]
Robust watermarking for diffusion models: A unified multi-dimensional recipe, 2024
Tianxin Wei, Ruizhong Qiu, Yifan Chen, Yunzhe Qi, Jiacheng Lin, Wenju Xu, Sreyashi Nag, Ruirui Li, Hanqing Lu, Zhengyang Wang, Chen Luo, Hui Liu, Suhang Wang, Jingrui He, Qi He, and Xianfeng Tang. Robust watermarking for diffusion models: A unified multi-dimensional recipe, 2024
2024
-
[153]
Latte: Collaborative test-time adaptation of vision-language models in federated learning
Wenxuan Bao, Ruxi Deng, Ruizhong Qiu, Tianxin Wei, Hanghang Tong, and Jingrui He. Latte: Collaborative test-time adaptation of vision-language models in federated learning. InProceedings of the IEEE/CVF International Conference on Computer Vision, 2025
2025
-
[154]
WAPITI: A watermark for finetuned open-source LLMs, 2024
Lingjie Chen, Ruizhong Qiu, Siyu Yuan, Zhining Liu, Tianxin Wei, Hyunsik Yoo, Zhichen Zeng, Deqing Yang, and Hanghang Tong. WAPITI: A watermark for finetuned open-source LLMs, 2024
2024
-
[155]
Breaking silos: Adaptive model fusion unlocks better time series forecasting
Zhining Liu, Ze Yang, Xiao Lin, Ruizhong Qiu, Tianxin Wei, Yada Zhu, Hendrik Hamann, Jingrui He, and Hanghang Tong. Breaking silos: Adaptive model fusion unlocks better time series forecasting. In Proceedings of the 42nd International Conference on Machine Learning, 2025
2025
-
[156]
Logic query of thoughts: Guiding large language models to answer complex logic queries with knowledge graphs, 2024
Lihui Liu, Zihao Wang, Ruizhong Qiu, Yikun Ban, Eunice Chan, Yangqiu Song, Jingrui He, and Hanghang Tong. Logic query of thoughts: Guiding large language models to answer complex logic queries with knowledge graphs, 2024
2024
-
[157]
Class-imbalanced graph learning without class rebalancing
Zhining Liu, Ruizhong Qiu, Zhichen Zeng, Hyunsik Yoo, David Zhou, Zhe Xu, Yada Zhu, Kommy Weldemariam, Jingrui He, and Hanghang Tong. Class-imbalanced graph learning without class rebalancing. InProceedings of the 41st International Conference on Machine Learning, 2024. 85 Age...
2024
-
[158]
AIM: Attributing, interpreting, mitigatingdataunfairness
Zhining Liu, Ruizhong Qiu, Zhichen Zeng, Yada Zhu, Hendrik Hamann, and Hanghang Tong. AIM: Attributing, interpreting, mitigatingdataunfairness. InProceedingsofthe30thACMSIGKDDConference on Knowledge Discovery and Data Mining, pages 2014–2025, 2024
2014
-
[159]
Topological augmentation for class-imbalanced node classification, 2023
Zhining Liu, Zhichen Zeng, Ruizhong Qiu, Hyunsik Yoo, David Zhou, Zhe Xu, Yada Zhu, Kommy Weldemariam, Jingrui He, and Hanghang Tong. Topological augmentation for class-imbalanced node classification, 2023
2023
-
[160]
Abdelzaher, Jiawei Han, and Hanghang Tong
Zhichen Zeng, Ruizhong Qiu, Wenxuan Bao, Tianxin Wei, Xiao Lin, Yuchen Yan, Tarek F. Abdelzaher, Jiawei Han, and Hanghang Tong. Pave your own path: Graph gradual domain adaptation on fused Gromov–Wasserstein geodesics, 2025
2025
-
[161]
Graph mixup on approximate Gromov–Wasserstein geodesics
Zhichen Zeng, Ruizhong Qiu, Zhe Xu, Zhining Liu, Yuchen Yan, Tianxin Wei, Lei Ying, Jingrui He, and Hanghang Tong. Graph mixup on approximate Gromov–Wasserstein geodesics. InProceedings of the 41st International Conference on Machine Learning, 2024
2024
-
[162]
Moralise: A structured benchmark for moral alignment in visual language models, 2025
Xiao Lin, Zhining Liu, Ze Yang, Gaotang Li, Ruizhong Qiu, Shuke Wang, Hui Liu, Haotian Li, Sumit Keswani, Vishwa Pardeshi, et al. Moralise: A structured benchmark for moral alignment in visual language models, 2025
2025
-
[163]
BackTime: Backdoor attacks on multivariate time series forecasting
Xiao Lin, Zhining Liu, Dongqi Fu, Ruizhong Qiu, and Hanghang Tong. BackTime: Backdoor attacks on multivariate time series forecasting. InAdvances in Neural Information Processing Systems, volume 37, 2024
2024
-
[164]
Saffron-1: Safety inference scaling, 2025
Ruizhong Qiu, Gaotang Li, Tianxin Wei, Jingrui He, and Hanghang Tong. Saffron-1: Safety inference scaling, 2025
2025
-
[165]
Ask, and it shall be given: On the Turing completeness of prompting
Ruizhong Qiu, Zhe Xu, Wenxuan Bao, and Hanghang Tong. Ask, and it shall be given: On the Turing completeness of prompting. In13th International Conference on Learning Representations, 2025
2025
-
[166]
How efficient is LLM-generated code? A rigorous & high-standard benchmark
Ruizhong Qiu, Weiliang Will Zeng, Hanghang Tong, James Ezick, and Christopher Lott. How efficient is LLM-generated code? A rigorous & high-standard benchmark. In13th International Conference on Learning Representations, 2025
2025
-
[167]
TUCKET: A tensor time series data structure for efficient and accurate factor analysis over time ranges.Proceedings of the VLDB Endowment, 17(13), 2024
Ruizhong Qiu, Jun-Gi Jang, Xiao Lin, Lihui Liu, and Hanghang Tong. TUCKET: A tensor time series data structure for efficient and accurate factor analysis over time ranges.Proceedings of the VLDB Endowment, 17(13), 2024
2024
-
[168]
Recon- structing graph diffusion history from a single snapshot
Ruizhong Qiu, Dingsu Wang, Lei Ying, H Vincent Poor, Yifang Zhang, and Hanghang Tong. Recon- structing graph diffusion history from a single snapshot. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 1978–1988, 2023
1978
-
[169]
DIMES: A differentiable meta solver for combinatorial optimization problems
Ruizhong Qiu, Zhiqing Sun, and Yiming Yang. DIMES: A differentiable meta solver for combinatorial optimization problems. InAdvances in Neural Information Processing Systems, volume 35, pages 25531–25546, 2022
2022
-
[170]
Discrete-state continuous-time diffusion for graph generation
Zhe Xu, Ruizhong Qiu, Yuzhong Chen, Huiyuan Chen, Xiran Fan, Menghai Pan, Zhichen Zeng, Mahashweta Das, and Hanghang Tong. Discrete-state continuous-time diffusion for graph generation. InAdvances in Neural Information Processing Systems, volume 37, 2024
2024
-
[171]
Model-free graph data selection under distribution shift, 2025
Ting-Wei Li, Ruizhong Qiu, and Hanghang Tong. Model-free graph data selection under distribution shift, 2025. 86 Agentic Reasoning for Large Language Models
2025
-
[172]
Transformer copilot: Learning from the mistake log in llm fine-tuning, 2025
Jiaru Zou, Yikun Ban, Zihao Li, Yunzhe Qi, Ruizhong Qiu, Ling Yang, and Jingrui He. Transformer copilot: Learning from the mistake log in llm fine-tuning, 2025. URLhttps://arxiv.org/abs/ 2505.16270
2025
-
[173]
Gradient compressed sensing: A query-efficient gradient estimator for high-dimensional zeroth-order optimization
Ruizhong Qiu and Hanghang Tong. Gradient compressed sensing: A query-efficient gradient estimator for high-dimensional zeroth-order optimization. InProceedings of the 41st International Conference on Machine Learning, 2024
2024
-
[174]
Embracing plasticity: Balancing stability and plasticity in continual recommender systems
Hyunsik Yoo, SeongKu Kang, Ruizhong Qiu, Charlie Xu, Fei Wang, and Hanghang Tong. Embracing plasticity: Balancing stability and plasticity in continual recommender systems. InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information ...
2025
-
[175]
Generalizable recommender systemduringtemporalpopularitydistributionshifts
Hyunsik Yoo, Ruizhong Qiu, Charlie Xu, Fei Wang, and Hanghang Tong. Generalizable recommender systemduringtemporalpopularitydistributionshifts. InProceedingsofthe31stACMSIGKDDConference on Knowledge Discovery and Data Mining, 2025
2025
-
[176]
Ensuring user-side fairness in dynamic recommender systems
Hyunsik Yoo, Zhichen Zeng, Jian Kang, Ruizhong Qiu, David Zhou, Zhining Liu, Fei Wang, Charlie Xu, Eunice Chan, and Hanghang Tong. Ensuring user-side fairness in dynamic recommender systems. InProceedings of the ACM on Web Conference 2024, pages 3667–3678, 2024
2024
-
[177]
Group fairness via group consensus
Eunice Chan, Zhining Liu, Ruizhong Qiu, Yuheng Zhang, Ross Maciejewski, and Hanghang Tong. Group fairness via group consensus. InThe 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 1788–1808, 2024
2024
-
[178]
Fair anomaly detection for imbalanced groups, 2024
Ziwei Wu, Lecheng Zheng, Yuancheng Yu, Ruizhong Qiu, John Birge, and Jingrui He. Fair anomaly detection for imbalanced groups, 2024
2024
-
[179]
On the sensitivity of individual fairness: Measures and robust algorithms
Xinyu He, Jian Kang, Ruizhong Qiu, Fei Wang, Jose Sepulveda, and Hanghang Tong. On the sensitivity of individual fairness: Measures and robust algorithms. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management, pages 829–838, 2024
2024
-
[180]
Networked time series imputation via position-aware graph enhanced variational autoencoders
Dingsu Wang, Yuchen Yan, Ruizhong Qiu, Yada Zhu, Kaiyu Guan, Andrew Margenot, and Hanghang Tong. Networked time series imputation via position-aware graph enhanced variational autoencoders. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,...
2023
-
[181]
Telograf: Temporal logic planning via graph-encoded flow matching
Yue Meng and Chuchu Fan. Telograf: Temporal logic planning via graph-encoded flow matching. arXiv preprint arXiv:2505.00562, 2025
2025
-
[182]
Ruizhe Zhong, Xingbo Du, Shixiong Kai, Zhentao Tang, Siyuan Xu, Jianye Hao, Mingxuan Yuan, and Junchi Yan. Flexplanner: Flexible 3d floorplanning via deep reinforcement learning in hybrid action space with multi-modality representation.Advances in Neural Information Processing...
2024
-
[183]
Benchmarking multimodal retrieval augmented generation with dynamic vqa dataset and self-adaptive planning agent.arXiv preprint arXiv:2411.02937, 2024
Yangning Li, Yinghui Li, Xinyu Wang, Yong Jiang, Zhen Zhang, Xinran Zheng, Hui Wang, Hai-Tao Zheng, Philip S Yu, Fei Huang, et al. Benchmarking multimodal retrieval augmented generation with dynamic vqa dataset and self-adaptive planning agent.arXiv preprint arXiv:2411.02937, 2024
2024
-
[184]
Rag over tables: Hierarchical memory index, multi-stage retrieval, and benchmarking, 2025
Jiaru Zou, Dongqi Fu, Sirui Chen, Xinrui He, Zihao Li, Yada Zhu, Jiawei Han, and Jingrui He. Rag over tables: Hierarchical memory index, multi-stage retrieval, and benchmarking, 2025. URL https://arxiv.org/abs/2504.01346. 87 Agentic Reasoning for Large Language Models
2025
-
[185]
Agent planning with world knowledge model.Advances in Neural Information Processing Systems, 37:114843–114871, 2024
Shuofei Qiao, Runnan Fang, Ningyu Zhang, Yuqi Zhu, Xiang Chen, Shumin Deng, Yong Jiang, Pengjun Xie, Fei Huang, and Huajun Chen. Agent planning with world knowledge model.Advances in Neural Information Processing Systems, 37:114843–114871, 2024
2024
-
[186]
Continual reinforcement learning by planning with online world models.arXiv preprint arXiv:2507.09177, 2025
Zichen Liu, Guoji Fu, Chao Du, Wee Sun Lee, and Min Lin. Continual reinforcement learning by planning with online world models.arXiv preprint arXiv:2507.09177, 2025
2025
-
[187]
Adawm: Adaptive world model based planning for autonomous driving.arXiv preprint arXiv:2501.13072, 2025
Hang Wang, Xin Ye, Feng Tao, Chenbin Pan, Abhirup Mallik, Burhaneddin Yaman, Liu Ren, and Junshan Zhang. Adawm: Adaptive world model based planning for autonomous driving.arXiv preprint arXiv:2501.13072, 2025
2025
-
[188]
Rational decision-making agent with internalized utility judgment.arXiv preprint arXiv:2308.12519, 2023
Yining Ye, Xin Cong, Shizuo Tian, Yujia Qin, Chong Liu, Yankai Lin, Zhiyuan Liu, and Maosong Sun. Rational decision-making agent with internalized utility judgment.arXiv preprint arXiv:2308.12519, 2023
2023
-
[189]
Scaling autonomous agents via automatic reward modeling and planning.arXiv preprint arXiv:2502.12130, 2025
Zhenfang Chen, Delin Chen, Rui Sun, Wenjun Liu, and Chuang Gan. Scaling autonomous agents via automatic reward modeling and planning.arXiv preprint arXiv:2502.12130, 2025
2025
-
[190]
Strategic planning: A top-down approach to option generation
Max Ruiz Luyten, Antonin Berthon, and Mihaela van der Schaar. Strategic planning: A top-down approach to option generation. InForty-second International Conference on Machine Learning, 2025
2025
-
[191]
Non-myopic generation of language models for reasoning and planning.arXiv preprint arXiv:2410.17195, 2024
Chang Ma, Haiteng Zhao, Junlei Zhang, Junxian He, and Lingpeng Kong. Non-myopic generation of language models for reasoning and planning.arXiv preprint arXiv:2410.17195, 2024
2024
-
[192]
Physics-informed temporal difference metric learning for robot motion planning.arXiv preprint arXiv:2505.05691, 2025
Ruiqi Ni, Zherong Pan, and Ahmed H Qureshi. Physics-informed temporal difference metric learning for robot motion planning.arXiv preprint arXiv:2505.05691, 2025
2025
-
[193]
Generalizable motion planning via operator learning.arXiv preprint arXiv:2410.17547, 2024
Sharath Matada, Luke Bhan, Yuanyuan Shi, and Nikolay Atanasov. Generalizable motion planning via operator learning.arXiv preprint arXiv:2410.17547, 2024
2024
-
[194]
Toolorchestra: Elevating intelligence via efficient model and tool orchestration, 2025
Hongjin Su, Shizhe Diao, Ximing Lu, Mingjie Liu, Jiacheng Xu, Xin Dong, Yonggan Fu, Peter Belcak, Hanrong Ye, Hongxu Yin, Yi Dong, Evelina Bakhturina, Tao Yu, Yejin Choi, Jan Kautz, and Pavlo Molchanov. Toolorchestra: Elevating intelligence via efficient model and tool orchest...
2025
-
[195]
Latent diffusion planning for imitation learning.arXiv preprint arXiv:2504.16925, 2025
Amber Xie, Oleh Rybkin, Dorsa Sadigh, and Chelsea Finn. Latent diffusion planning for imitation learning.arXiv preprint arXiv:2504.16925, 2025
2025
-
[196]
Safedif- fuser: Safe planning with diffusion probabilistic models
Wei Xiao, Tsun-Hsuan Wang, Chuang Gan, Ramin Hasani, Mathias Lechner, and Daniela Rus. Safedif- fuser: Safe planning with diffusion probabilistic models. InThe Thirteenth International Conference on Learning Representations, 2023
2023
-
[197]
Contradiff: Planningtowardshighreturnstatesviacontrastivelearning
Yixiang Shan, Zhengbang Zhu, Ting Long, Liang Qifan, Yi Chang, Weinan Zhang, and Liang Yin. Contradiff: Planningtowardshighreturnstatesviacontrastivelearning. InTheThirteenthInternational Conference on Learning Representations, 2025
2025
-
[198]
Amortized planning with large- scale transformers: A case study on chess.Advances in Neural Information Processing Systems, 37: 65765–65790, 2024
Anian Ruoss, Grégoire Delétang, Sourabh Medapati, Jordi Grau-Moya, Li K Wenliang, Elliot Catt, John Reid, Cannada A Lewis, Joel Veness, and Tim Genewein. Amortized planning with large- scale transformers: A case study on chess.Advances in Neural Information Processing Systems,...
2024
-
[199]
Art: Automatic multi-step reasoning and tool-use for large language models,
Bhargavi Paranjape, Scott Lundberg, Sameer Singh, Hannaneh Hajishirzi, Luke Zettlemoyer, and Marco Tulio Ribeiro. Art: Automatic multi-step reasoning and tool-use for large language models,
-
[200]
URLhttps://arxiv.org/abs/2303.09014
-
[201]
ChatCoT: Tool-augmented chain-of-thought reasoning on chat-based large language models
Zhipeng Chen, Kun Zhou, Beichen Zhang, Zheng Gong, Xin Zhao, and Ji-Rong Wen. ChatCoT: Tool-augmented chain-of-thought reasoning on chat-based large language models. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Findings of the Association for Computational Linguistics...
2023 doi
-
[202]
GEAR: Augmenting language models with generalizable and efficient tool resolution
Yining Lu, Haoping Yu, and Daniel Khashabi. GEAR: Augmenting language models with generalizable and efficient tool resolution. In Yvette Graham and Matthew Purver, editors,Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguisti...
2024 doi
-
[203]
Ioannidis, Karthik Subbian, Jure Leskovec, and James Zou
Shirley Wu, Shiyu Zhao, Qian Huang, Kexin Huang, Michihiro Yasunaga, Kaidi Cao, Vassilis N. Ioannidis, Karthik Subbian, Jure Leskovec, and James Zou. Avatar: Optimizing llm agents for tool usage via contrastive reasoning. InAdvances in Neural Information Processing Systems, vo...
2024
-
[204]
Toolllm: Facilitating large language models to master 16000+ real-world apis
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. Toolllm: Facilitating large language models to m...
2024
-
[205]
Toolalpaca: Generalized tool learning for language models with 3000 simulated cases.CoRR, abs/2306.05301, 2023
QiaoyuTang,ZiliangDeng,HongyuLin,XianpeiHan,QiaoLiang,andLeSun. Toolalpaca: Generalized tool learning for language models with 3000 simulated cases.CoRR, abs/2306.05301, 2023. doi: 10.48550/ARXIV.2306.05301. URLhttps://doi.org/10.48550/arXiv.2306.05301
-
[206]
Learning to reason with search for llms via reinforcement learning.arXiv preprint arXiv:2503.19470, 2025
Mingyang Chen, Tianpeng Li, Haoze Sun, Yijie Zhou, Chenzheng Zhu, Haofen Wang, Jeff Z Pan, Wen Zhang, Huajun Chen, Fan Yang, et al. Learning to reason with search for llms via reinforcement learning.arXiv preprint arXiv:2503.19470, 2025
2025 arXiv
-
[207]
Reinforcement pre-training.arXiv preprint arXiv:2506.08007, 2025
Qingxiu Dong, Li Dong, Yao Tang, Tianzhu Ye, Yutao Sun, Zhifang Sui, and Furu Wei. Reinforcement pre-training.arXiv preprint arXiv:2506.08007, 2025
2025
-
[208]
Toolrl: Reward is all tool learning needs.arXiv preprint arXiv:2504.13958, 2025
Cheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang, Xiusi Chen, Dilek Hakkani-Tür, Gokhan Tur, and Heng Ji. Toolrl: Reward is all tool learning needs.arXiv preprint arXiv:2504.13958, 2025
2025 arXiv
-
[209]
Taskmatrix.ai: Completing tasks by connecting foundation models with millions of apis.CoRR, abs/2303.16434, 2023
Yaobo Liang, Chenfei Wu, Ting Song, Wenshan Wu, Yan Xia, Yu Liu, Yang Ou, Shuai Lu, Lei Ji, Shaoguang Mao, Yun Wang, Linjun Shou, Ming Gong, and Nan Duan. Taskmatrix.ai: Completing tasks by connecting foundation models with millions of apis.CoRR, abs/2303.16434, 2023. doi: 10....
2023 doi
-
[210]
Octotools: An agentic framework with extensible tools for complex reasoning, 2025
Pan Lu, Bowen Chen, Sheng Liu, Rahul Thapa, Joseph Boen, and James Zou. Octotools: An agentic framework with extensible tools for complex reasoning, 2025. URLhttps://arxiv.org/abs/ 2502.11271. 89 Agentic Reasoning for Large Language Models
2025 arXiv
-
[211]
Toolexpnet: Optimizing multi-tool selection in llms with similarity and dependency-aware experience networks
ZijingZhang, ZhanpengChen, HeZhu,ZiyangChen, NanDu, andXiaolongLi. Toolexpnet: Optimizing multi-tool selection in llms with similarity and dependency-aware experience networks. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,Findings of ...
2025
-
[212]
Bursztyn, Ryan A
Yuchen Zhuang, Xiang Chen, Tong Yu, Saayan Mitra, Victor S. Bursztyn, Ryan A. Rossi, Somdeb Sarkhel, and Chao Zhang. Toolchain*: Efficient action space navigation in large language models with a* search. InThe Twelfth International Conference on Learning Representations, ICLR ...
2024
-
[213]
MultiTool-CoT: GPT-3 can use multiple external tools with chain of thought prompting
Tatsuro Inaba, Hirokazu Kiyomaru, Fei Cheng, and Sadao Kurohashi. MultiTool-CoT: GPT-3 can use multiple external tools with chain of thought prompting. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors,Proceedings of the 61st Annual Meeting of the Association for...
2023 doi
-
[214]
Interleaving re- trieval with chain-of-thought reasoning for knowledge-intensive multi-step questions.arXiv preprint arXiv:2212.10509, 2022
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. Interleaving re- trieval with chain-of-thought reasoning for knowledge-intensive multi-step questions.arXiv preprint arXiv:2212.10509, 2022
2022
-
[215]
Tool documentation enables zero-shot tool-usage with large language models, 2023
Cheng-Yu Hsieh, Si-An Chen, Chun-Liang Li, Yasuhisa Fujii, Alexander Ratner, Chen-Yu Lee, Ranjay Krishna, and Tomas Pfister. Tool documentation enables zero-shot tool-usage with large language models, 2023. URLhttps://arxiv.org/abs/2308.00675
2023
-
[216]
Easytool: Enhancing llm-based agents with concise tool instruction.arXiv preprint arXiv:2401.06201, 2024
Siyu Yuan, Kaitao Song, Jiangjie Chen, Xu Tan, Yongliang Shen, Ren Kan, Dongsheng Li, and Deqing Yang. Easytool: Enhancing llm-based agents with concise tool instruction.arXiv preprint arXiv:2401.06201, 2024
2024
-
[217]
Tool learning with large language models: A survey.Frontiers of Computer Science, 19(8): 198343, 2025
Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, Jun Xu, and Ji-Rong Wen. Tool learning with large language models: A survey.Frontiers of Computer Science, 19(8): 198343, 2025
2025
-
[218]
Tool learning in the wild: Empowering language models as automatic tool agents
Zhengliang Shi, Shen Gao, Lingyong Yan, Yue Feng, Xiuyi Chen, Zhumin Chen, Dawei Yin, Suzan Verberne, and Zhaochun Ren. Tool learning in the wild: Empowering language models as automatic tool agents. InProceedings of the ACM on Web Conference 2025, pages 2222–2237, 2025
2025
-
[219]
Empowering large language models: Tool learning for real-world interaction
Hongru Wang, Yujia Qin, Yankai Lin, Jeff Z Pan, and Kam-Fai Wong. Empowering large language models: Tool learning for real-world interaction. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2983–2986, 2024
2024
-
[220]
Buffer of thoughts: Thought-augmented reasoning with large language models.Advances in Neural Information Processing Systems, 37:113519–113544, 2024
Ling Yang, Zhaochen Yu, Tianjun Zhang, Shiyi Cao, Minkai Xu, Wentao Zhang, Joseph E Gonzalez, and Bin Cui. Buffer of thoughts: Thought-augmented reasoning with large language models.Advances in Neural Information Processing Systems, 37:113519–113544, 2024
2024
-
[221]
Seeclick: Harnessing gui grounding for advanced visual gui agents.arXiv preprint arXiv:2401.10935, 2024
Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang, and Zhiyong Wu. Seeclick: Harnessing gui grounding for advanced visual gui agents.arXiv preprint arXiv:2401.10935, 2024. 90 Agentic Reasoning for Large Language Models
2024 arXiv
-
[222]
Reasonflux- prm: Trajectory-aware prms for long chain-of-thought reasoning in llms, 2025
Jiaru Zou, Ling Yang, Jingwen Gu, Jiahao Qiu, Ke Shen, Jingrui He, and Mengdi Wang. Reasonflux- prm: Trajectory-aware prms for long chain-of-thought reasoning in llms, 2025. URLhttps://arxiv. org/abs/2506.18896
2025
-
[223]
Using an llm to help with code understanding
Daye Nam, Andrew Macvean, Vincent Hellendoorn, Bogdan Vasilescu, and Brad Myers. Using an llm to help with code understanding. InProceedings of the IEEE/ACM 46th International Conference on Software Engineering, pages 1–13, 2024
2024
-
[224]
Agentic reasoning: A streamlined framework for enhancing llm reasoning with agentic tools
Junde Wu, Jiayuan Zhu, Yuyuan Liu, Min Xu, and Yueming Jin. Agentic reasoning: A streamlined framework for enhancing llm reasoning with agentic tools. 2025. URLhttps://arxiv.org/abs/ 2502.04644
2025
-
[225]
Chameleon: Plug-and-play compositional reasoning with large language models
Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. Chameleon: Plug-and-play compositional reasoning with large language models. Advances in Neural Information Processing Systems, 36:43447–43478, 2023
2023
-
[226]
Restgpt: Connecting large language models with real-world restful apis
Yifan Song, Weimin Xiong, Dawei Zhu, Wenhao Wu, Han Qian, Mingbo Song, Hailiang Huang, Cheng Li, Ke Wang, Rong Yao, et al. Restgpt: Connecting large language models with real-world restful apis. arXiv preprint arXiv:2306.06624, 2023
2023
-
[227]
Adapt: As-needed decomposition and planning with language models.arXiv preprint arXiv:2311.05772, 2023
Archiki Prasad, Alexander Koller, Mareike Hartmann, Peter Clark, Ashish Sabharwal, Mohit Bansal, and Tushar Khot. Adapt: As-needed decomposition and planning with language models.arXiv preprint arXiv:2311.05772, 2023
2023
-
[228]
Agent lumos: Unified and modular training for open-source language agents.arXiv preprint arXiv:2311.05657, 2023
Da Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Chandu, Kai-Wei Chang, Yejin Choi, and Bill Yuchen Lin. Agent lumos: Unified and modular training for open-source language agents.arXiv preprint arXiv:2311.05657, 2023
2023
-
[229]
Learning to use tools via cooperative and interactive agents
Zhengliang Shi, Shen Gao, Xiuyi Chen, Yue Feng, Lingyong Yan, Haibo Shi, Dawei Yin, Pengjie Ren, Suzan Verberne, and Zhaochun Ren. Learning to use tools via cooperative and interactive agents. arXiv preprint arXiv:2403.03031, 2024
2024
-
[230]
Understanding the effects of rlhf on llm generalisation and diversity.arXiv preprint arXiv:2310.06452, 2023
Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. Understanding the effects of rlhf on llm generalisation and diversity.arXiv preprint arXiv:2310.06452, 2023
2023
-
[231]
Preserving diversity in supervised fine-tuning of large language models.arXiv preprint arXiv:2408.16673, 2024
ZiniuLi, CongliangChen, TianXu, ZeyuQin, JiancongXiao, Zhi-QuanLuo, andRuoyuSun. Preserving diversity in supervised fine-tuning of large language models.arXiv preprint arXiv:2408.16673, 2024
2024
-
[232]
Attributing mode collapse in the fine-tuning of large language models
Laura O’Mahony, Leo Grinsztajn, Hailey Schoelkopf, and Stella Biderman. Attributing mode collapse in the fine-tuning of large language models. InICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models, volume 2, 2024
2024
-
[233]
itool: Reinforced fine-tuning with dynamic deficiency calibration for advanced tool use
Yirong Zeng, Xiao Ding, Yuxian Wang, Weiwen Liu, Yutai Hou, Wu Ning, Xu Huang, Duyu Tang, Dandan Tu, Bing Qin, et al. itool: Reinforced fine-tuning with dynamic deficiency calibration for advanced tool use. InProceedings of the 2025 Conference on Empirical Methods in Natural L...
2025
-
[234]
Demystifying reinforcement learning in agentic reasoning.arXiv preprint arXiv:2510.11701, 2025
Zhaochen Yu, Ling Yang, Jiaru Zou, Shuicheng Yan, and Mengdi Wang. Demystifying reinforcement learning in agentic reasoning.arXiv preprint arXiv:2510.11701, 2025. 91 Agentic Reasoning for Large Language Models
2025
-
[235]
Sweet-rl: Training multi-turn llm agents on collaborative reasoning tasks.arXiv preprint arXiv:2503.15478, 2025
Yifei Zhou, Song Jiang, Yuandong Tian, Jason Weston, Sergey Levine, Sainbayar Sukhbaatar, and Xian Li. Sweet-rl: Training multi-turn llm agents on collaborative reasoning tasks.arXiv preprint arXiv:2503.15478, 2025
2025
-
[236]
Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution.arXiv preprint arXiv:2502.18449, 2025
Yuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux, Lingming Zhang, Daniel Fried, Gabriel Synnaeve, Rishabh Singh, and Sida I Wang. Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution.arXiv preprint arXiv:2502.18449, 2025
2025 arXiv
-
[237]
Rlvmr: Reinforcement learning with verifiable meta-reasoning rewards for robust long-horizon agents.arXiv preprint arXiv:2507.22844, 2025
Zijing Zhang, Ziyang Chen, Mingxiao Li, Zhaopeng Tu, and Xiaolong Li. Rlvmr: Reinforcement learning with verifiable meta-reasoning rewards for robust long-horizon agents.arXiv preprint arXiv:2507.22844, 2025
2025
-
[238]
Autotool: Dynamic tool selection and integration for agentic reasoning, 2025
Jiaru Zou, Ling Yang, Yunzhe Qi, Sirui Chen, Mengting Ai, Ke Shen, Jingrui He, and Mengdi Wang. Autotool: Dynamic tool selection and integration for agentic reasoning, 2025. URLhttps://arxiv. org/abs/2512.13278
2025
-
[239]
Retool: Reinforcement learning for strategic tool use in llms.arXiv preprint arXiv:2504.11536, 2025
Jiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang, Yujia Qin, Baoquan Zhong, Chengquan Jiang, Jinxin Chi, and Wanjun Zhong. Retool: Reinforcement learning for strategic tool use in llms.arXiv preprint arXiv:2504.11536, 2025
2025 arXiv
-
[240]
Zerosearch: Incentivize the search capability of llms without searching
Hao Sun, Zile Qiao, Jiayan Guo, Xuanbo Fan, Yingyan Hou, Yong Jiang, Pengjun Xie, Yan Zhang, Fei Huang, and Jingren Zhou. Zerosearch: Incentivize the search capability of llms without searching. arXiv preprint arXiv:2505.04588, 2025
2025
-
[241]
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al. Kimi k1. 5: Scaling reinforcement learning with llms.arXiv preprint arXiv:2501.12599, 2025
2025 arXiv
-
[242]
Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507.06261, 2025
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabil...
2025 arXiv
-
[243]
Kimi k2: Open agentic intelligence.arXiv preprint arXiv:2507.20534, 2025
Kimi Team, Yifan Bai, Yiping Bao, Guanduo Chen, Jiahao Chen, Ningxin Chen, Ruijue Chen, Yanru Chen, Yuankun Chen, Yutian Chen, et al. Kimi k2: Open agentic intelligence.arXiv preprint arXiv:2507.20534, 2025
2025 arXiv
-
[244]
Glm-4.5: Agentic, reasoning, and coding (arc) foundation models.arXiv preprint arXiv:2508.06471, 2025
Aohan Zeng, Xin Lv, Qinkai Zheng, Zhenyu Hou, Bin Chen, Chengxing Xie, Cunxiang Wang, Da Yin, Hao Zeng, Jiajie Zhang, et al. Glm-4.5: Agentic, reasoning, and coding (arc) foundation models.arXiv preprint arXiv:2508.06471, 2025
2025 arXiv
-
[245]
Tattoo: Tool-grounded thinking prm for test-time scaling in tabular reasoning.arXiv preprint arXiv:2510.06217, 2025
Jiaru Zou, Soumya Roy, Vinay Kumar Verma, Ziyi Wang, David Wipf, Pan Lu, Sumit Negi, James Zou, and Jingrui He. Tattoo: Tool-grounded thinking prm for test-time scaling in tabular reasoning.arXiv preprint arXiv:2510.06217, 2025
2025
-
[246]
Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings
Shibo Hao, Tianyang Liu, Zhen Wang, and Zhiting Hu. Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors,Advances in Neural Information Processi...
2023
-
[247]
Advanc- ing tool-augmented large language models via meta-verification and reflection learning.CoRR, abs/2506.04625, 2025
Zhiyuan Ma, Jiayu Liu, Xianzhen Luo, Zhenya Huang, Qingfu Zhu, and Wanxiang Che. Advanc- ing tool-augmented large language models via meta-verification and reflection learning.CoRR, abs/2506.04625, 2025. doi: 10.48550/ARXIV.2506.04625. URLhttps://doi.org/10.48550/ arXiv.2506.04625
2025 doi
-
[248]
Chain- of-tools: Utilizing massive unseen tools in the cot reasoning of frozen language models.CoRR, abs/2503.16779, 2025
Mengsong Wu, Tong Zhu, Han Han, Xiang Zhang, Wenbiao Shao, and Wenliang Chen. Chain- of-tools: Utilizing massive unseen tools in the cot reasoning of frozen language models.CoRR, abs/2503.16779, 2025. doi: 10.48550/ARXIV.2503.16779. URLhttps://doi.org/10.48550/ arXiv.2503.16779
2025 doi
-
[249]
Pyvision: Agentic vision with dynamic tooling.CoRR, abs/2507.07998, 2025
Shitian Zhao, Haoquan Zhang, Shaoheng Lin, Ming Li, Qilong Wu, Kaipeng Zhang, and Chen Wei. Pyvision: Agentic vision with dynamic tooling.CoRR, abs/2507.07998, 2025. doi: 10.48550/ARXIV. 2507.07998. URLhttps://doi.org/10.48550/arXiv.2507.07998
2025 doi
-
[250]
Cheng, Abdulrahman Aldossary, Jiaru Bai, Shi Xuan Leong, Jorge A
Yunheng Zou, Austin H. Cheng, Abdulrahman Aldossary, Jiaru Bai, Shi Xuan Leong, Jorge A. Campos Gonzalez Angulo, Changhyeok Choi, Cher Tian Ser, Gary Tom, Andrew Wang, Zijian Zhang, Ilya Yakavets, Han Hao, Chris Crebolder, Varinia Bernales, and Alán Aspuru-Guzik. El agente: An...
2025 doi
-
[251]
Tˆ2agent A tool-augmented multimodal misinformation detection agent with monte carlo tree search.CoRR, abs/2505.19768, 2025
Xing Cui, Yueying Zou, Zekun Li, Pei-Pei Li, Xinyuan Xu, Xuannan Liu, Huaibo Huang, and Ran He. Tˆ2agent A tool-augmented multimodal misinformation detection agent with monte carlo tree search.CoRR, abs/2505.19768, 2025. doi: 10.48550/ARXIV.2505.19768. URLhttps://doi.org/ 10.4...
2025 doi
-
[252]
Toolrerank: Adaptive and hierarchy-aware reranking for tool retrieval
Yuanhang Zheng, Peng Li, Wei Liu, Yang Liu, Jian Luan, and Bin Wang. Toolrerank: Adaptive and hierarchy-aware reranking for tool retrieval. In Nicoletta Calzolari, Min-Yen Kan, Véronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue, editors,Proceedings of the 2024 ...
2024
-
[253]
Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems,...
2020
-
[254]
Crag-comprehensive rag benchmark.Advances in Neural Information Processing Systems, 37:10470–10490, 2024
Xiao Yang, Kai Sun, Hao Xin, Yushi Sun, Nikita Bhalla, Xiangsen Chen, Sajal Choudhary, Rongze Gui, Ziran Jiang, Ziyu Jiang, et al. Crag-comprehensive rag benchmark.Advances in Neural Information Processing Systems, 37:10470–10490, 2024
2024
-
[255]
Measuring and narrowing the compositionality gap in language models.arXiv preprint arXiv:2210.03350, 2022
Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis. Measuring and narrowing the compositionality gap in language models.arXiv preprint arXiv:2210.03350, 2022
2022
-
[256]
Self-rag: Self-reflective retrieval augmented generation
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. Self-rag: Self-reflective retrieval augmented generation. InNeurIPS 2023 workshop on instruction tuning and instruction following, 2023
2023
-
[257]
Deeprag: Thinking to retrieve step by step for large language models.arXiv preprint arXiv:2502.01142, 2025
Xinyan Guan, Jiali Zeng, Fandong Meng, Chunlei Xin, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun, and Jie Zhou. Deeprag: Thinking to retrieve step by step for large language models.arXiv preprint arXiv:2502.01142, 2025. 93 Agentic Reasoning for Large Language Models
2025
-
[258]
Inters: Unlocking the power of large language models in search with instruction tuning.arXiv preprint arXiv:2401.06532, 2024
Yutao Zhu, Peitian Zhang, Chenghao Zhang, Yifei Chen, Binyu Xie, Zheng Liu, Ji-Rong Wen, and Zhicheng Dou. Inters: Unlocking the power of large language models in search with instruction tuning.arXiv preprint arXiv:2401.06532, 2024
2024
-
[259]
Webgpt: Browser-assistedquestion-answering with human feedback.arXiv preprint arXiv:2112.09332, 2021
ReiichiroNakano, JacobHilton, SuchirBalaji, JeffWu, LongOuyang, ChristinaKim, ChristopherHesse, ShantanuJain, VineetKosaraju, WilliamSaunders, etal. Webgpt: Browser-assistedquestion-answering with human feedback.arXiv preprint arXiv:2112.09332, 2021
2021 arXiv
-
[260]
Rag-rl: Advancing retrieval-augmented generation via rl and curriculum learning.arXiv preprint arXiv:2503.12759, 2025
Jerry Huang, Siddarth Madala, Risham Sidhu, Cheng Niu, Hao Peng, Julia Hockenmaier, and Tong Zhang. Rag-rl: Advancing retrieval-augmented generation via rl and curriculum learning.arXiv preprint arXiv:2503.12759, 2025
2025 arXiv
-
[261]
Deepresearcher: Scaling deep research via reinforcement learning in real-world environments.arXiv preprint arXiv:2504.03160, 2025
Yuxiang Zheng, Dayuan Fu, Xiangkun Hu, Xiaojie Cai, Lyumanshan Ye, Pengrui Lu, and Pengfei Liu. Deepresearcher: Scaling deep research via reinforcement learning in real-world environments.arXiv preprint arXiv:2504.03160, 2025
2025 arXiv
-
[262]
Rearter: Retrieval-augmented reasoning with trustworthy process rewarding
Zhongxiang Sun, Qipeng Wang, Weijie Yu, Xiaoxue Zang, Kai Zheng, Jun Xu, Xiao Zhang, Yang Song, and Han Li. Rearter: Retrieval-augmented reasoning with trustworthy process rewarding. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in I...
2025
-
[263]
Agent-g: An agentic framework for graph retrieval augmented generation
Meng-Chieh Lee, Qi Zhu, Costas Mavromatis, Zhen Han, Soji Adeshina, Vassilis N Ioannidis, Huzefa Rangwala, and Christos Faloutsos. Agent-g: An agentic framework for graph retrieval augmented generation
-
[264]
Mc-search: Benchmarking multimodal agentic rag with structured reasoning chains
Xuying Ning, Dongqi Fu, Tianxin Wei, Mengting Ai, Jiaru Zou, Ting-Wei Li, and Jingrui He. Mc-search: Benchmarking multimodal agentic rag with structured reasoning chains. InNeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scal...
2025
-
[265]
Gear: Graph-enhanced agent for retrieval- augmented generation
Zhili Shen, Chenxin Diao, Pavlos Vougiouklis, Pascual Merita, Shriram Piramanayagam, Enting Chen, Damien Graux, Andre Melo, Ruofei Lai, Zeren Jiang, et al. Gear: Graph-enhanced agent for retrieval- augmented generation. InFindings of the Association for Computational Linguisti...
2025
-
[266]
Learning to retrieve and reason on knowledge graph through active self-reflection.arXiv preprint arXiv:2502.14932, 2025
Han Zhang, Langshi Zhou, and Hanfang Yang. Learning to retrieve and reason on knowledge graph through active self-reflection.arXiv preprint arXiv:2502.14932, 2025
2025
-
[267]
Rag-studio: Towards in-domain adaptation of retrieval augmented generation through self-alignment
Kelong Mao, Zheng Liu, Hongjin Qian, Fengran Mo, Chenlong Deng, and Zhicheng Dou. Rag-studio: Towards in-domain adaptation of retrieval augmented generation through self-alignment. InFindings of the Association for Computational Linguistics: EMNLP 2024, pages 725–735, 2024
2024
-
[268]
Raft: Adapting language model to domain specific rag.arXiv preprint arXiv:2403.10131, 2024
Tianjun Zhang, Shishir G Patil, Naman Jain, Sheng Shen, Matei Zaharia, Ion Stoica, and Joseph E Gonzalez. Raft: Adapting language model to domain specific rag.arXiv preprint arXiv:2403.10131, 2024
2024
-
[269]
Ra-dit: Retrieval-augmented dual instruction tuning
Xi Victoria Lin, Xilun Chen, Mingda Chen, Weijia Shi, Maria Lomeli, Richard James, Pedro Rodriguez, Jacob Kahn, Gergely Szilvasy, Mike Lewis, et al. Ra-dit: Retrieval-augmented dual instruction tuning. InThe Twelfth International Conference on Learning Representations, 2023
2023
-
[270]
Sfr-rag: Towards contextually faithful llms.arXiv preprint arXiv:2409.09916, 2024
Xuan-Phi Nguyen, Shrey Pandit, Senthil Purushwalkam, Austin Xu, Hailin Chen, Yifei Ming, Zixuan Ke, Silvio Savarese, Caiming Xong, and Shafiq Joty. Sfr-rag: Towards contextually faithful llms.arXiv preprint arXiv:2409.09916, 2024. 94 Agentic Reasoning for Large Language Models
2024
-
[271]
Self-refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594, 2023
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. Self-refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594, 2023
2023
-
[272]
Enable language models to implicitly learn self-improvement from data
Ziqi Wang, Le Hou, Tianjian Lu, Yuexin Wu, Yunxuan Li, Hongkun Yu, and Heng Ji. Enable language models to implicitly learn self-improvement from data. InProc. The Twelfth International Conference on Learning Representations (ICLR2024), 2024
2024
-
[273]
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. InThe Eleventh International Conference on Learning Representations, 2023. URLhttps: //ope...
2023
-
[274]
Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks.Transactions on Machine Learning Research, 2023
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks.Transactions on Machine Learning Research, 2023. URLhttps://openreview.net/forum?id=YfZ4ZPt8zd
2023
-
[275]
Agenttuning: Enablinggeneralizedagentabilitiesforllms
Aohan Zeng, Mingdao Liu, Rui Lu, Bowen Wang, Xiao Liu, Yuxiao Dong, and Jie Tang. Agenttuning: Enablinggeneralizedagentabilitiesforllms. InFindingsoftheAssociationforComputationalLinguistics: ACL 2024, pages 3053–3077, 2024
2024
-
[276]
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes
Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes. InFindings of the Associati...
2023
-
[277]
Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017
2017
-
[278]
Manning, and Chelsea Finn
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. InAdvances in Neural Information Processing Systems (NeurIPS), 2023
2023
-
[279]
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. InAdvances in Neural Information Processing Systems (NeurIPS), 2022
2022
-
[280]
Reflectevo: Improving meta introspection of small llms by learning self- reflection
Jiaqi Li, Xinyi Dong, Yang Liu, Zhizhuo Yang, Quansen Wang, Xiaobo Wang, Song-Chun Zhu, Zixia Jia, and Zilong Zheng. Reflectevo: Improving meta introspection of small llms by learning self- reflection. InFindings of the Association for Computational Linguistics (ACL), 2025. UR...
2025
-
[281]
Reasoning-cv: Fine-tuning powerful reasoning llms for knowledge- assisted claim verification.arXiv preprint arXiv:2505.12348, 2025
Zhi Zheng and Wee Sun Lee. Reasoning-cv: Fine-tuning powerful reasoning llms for knowledge- assisted claim verification.arXiv preprint arXiv:2505.12348, 2025
2025
-
[282]
Rezero: Enhancing llm search ability by trying one-more-time.arXiv preprint arXiv:2504.11001, 2025
Alan Dao and Thinh Le. Rezero: Enhancing llm search ability by trying one-more-time.arXiv preprint arXiv:2504.11001, 2025
2025
-
[283]
Are retrials all you need? enhancing large language model reasoning without verbalized feedback.arXiv preprint arXiv:2504.12951, 2025
Nearchos Potamitis and Akhil Arora. Are retrials all you need? enhancing large language model reasoning without verbalized feedback.arXiv preprint arXiv:2504.12951, 2025. 95 Agentic Reasoning for Large Language Models
2025
-
[284]
Hung Le, Yue Wang, Akhilesh Deepak Yu, Thanh-Tung Nguyen, Zhiwei Sun, Nan Jiang, Quoc Viet Le, and Steven C. H. Hoi. Coderl: Mastering code generation through pretrained models and deep reinforcement learning. InAdvances in Neural Information Processing Systems (NeurIPS), 2022
2022
-
[285]
Lever: Learning to verify language-to-code generation with execution
Ansong Ni, Srini Iyer, Dragomir Radev, Veselin Stoyanov, Wen-tau Yih, Sida Wang, and Xi Victoria Lin. Lever: Learning to verify language-to-code generation with execution. InInternational Conference on Machine Learning, pages 26106–26128. PMLR, 2023
2023
-
[286]
Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R. Narasimhan. Swe-bench: Can language models resolve real-world github issues? InInternational Conference on Learning Representations (ICLR), 2024
2024
-
[287]
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. InInternational Conference on Machine Learning, pages 8469–8488. PMLR, 2023
2023
-
[288]
Reflect, retry, reward: Self-improving llms via reinforcement learning.arXiv preprint arXiv:2505.24726, 2025
Shelly Bensal, Umar Jamil, Christopher Bryant, Melisa Russak, Kiran Kamble, Dmytro Mozolevskyi, Muayad Ali, and Waseem AlShikh. Reflect, retry, reward: Self-improving llms via reinforcement learning.arXiv preprint arXiv:2505.24726, 2025
2025
-
[289]
Rlaif vs
Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, and Sushant Prakash. Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback. 2024
2024
-
[290]
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Potsawee Manakul, Adian Liusie, and Mark Gales. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. InProceedings of the 2023 conference on empirical methods in natural language processing, pages 9004–9017, 2023
2023
-
[291]
Zero-shot verification-guided chain of thoughts.arXiv preprint arXiv:2501.13122, 2025
Jishnu Ray Chowdhury and Cornelia Caragea. Zero-shot verification-guided chain of thoughts.arXiv preprint arXiv:2501.13122, 2025
2025
-
[292]
Ascot: An adaptive self-correctionchain-of-thoughtmethodforlate-stagefragilityinllms.arXivpreprintarXiv:2508.05282, 2025
Dongxu Zhang, Ning Yang, Jihua Zhu, Jinnan Yang, Miao Xin, and Baoliang Tian. Ascot: An adaptive self-correctionchain-of-thoughtmethodforlate-stagefragilityinllms.arXivpreprintarXiv:2508.05282, 2025
2025
-
[293]
Mm-verify: Enhancing multimodal reasoning with chain-of-thought verification
Linzhuang Sun, Hao Liang, Jingxuan Wei, Bihui Yu, Tianpeng Li, Fan Yang, Zenan Zhou, and Wentao Zhang. Mm-verify: Enhancing multimodal reasoning with chain-of-thought verification. InACL, 2025. URLhttps://aclanthology.org/2025.acl-long.689/
2025
-
[294]
Patil, Kevin Lin, Sarah Wooders, and Joseph Gonza- lez
Charles Packer, Vivian Fang, Shishir G. Patil, Kevin Lin, Sarah Wooders, and Joseph Gonza- lez. Memgpt: Towards llms as operating systems.ArXiv, abs/2310.08560, 2023. URL https: //api.semanticscholar.org/CorpusID:263909014
2023 arXiv
-
[295]
Re-rest: Reflection- reinforced self-training for language agents
Zi-Yi Dou, Cheng-Fu Yang, Xueqing Wu, Kai-Wei Chang, and Nanyun Peng. Re-rest: Reflection- reinforced self-training for language agents. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 15394–15411, 2024
2024
-
[296]
Langchain library
LangChain AI. Langchain library. 2023. URLhttps://www.langchain.com/
2023
-
[297]
LlamaIndex, 11 2022
Jerry Liu. LlamaIndex, 11 2022. URLhttps://github.com/jerryjliu/llama_index. 96 Agentic Reasoning for Large Language Models
2022
-
[298]
Memorybank: Enhancing large language models with long-term memory.ArXiv, abs/2305.10250, 2023
Wanjun Zhong, Lianghong Guo, Qi-Fei Gao, He Ye, and Yanlin Wang. Memorybank: Enhancing large language models with long-term memory.ArXiv, abs/2305.10250, 2023. URLhttps://api. semanticscholar.org/CorpusID:258741194
2023 arXiv
-
[299]
Agent workflow memory.ArXiv, abs/2409.07429, 2024
Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried, and Graham Neubig. Agent workflow memory.ArXiv, abs/2409.07429, 2024. URLhttps://api.semanticscholar.org/CorpusID:272592995
2024 arXiv
-
[300]
Lightmem: Lightweight and efficient memory-augmented generation.arXiv preprint arXiv:2510.18866, 2025
Jizhan Fang, Xinle Deng, Haoming Xu, Ziyan Jiang, Yuqi Tang, Ziwen Xu, Shumin Deng, Yunzhi Yao, Mengru Wang, Shuofei Qiao, et al. Lightmem: Lightweight and efficient memory-augmented generation.arXiv preprint arXiv:2510.18866, 2025
2025
Reviewed May 17, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.