REVIEW 4 major objections 5 minor 47 references
One STEP at a time: Language Agents are Stepwise Planners
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Language agents can be stepwise planners: the STEP framework reaches 67.4 on ScienceWorld by combining a Planner, Executor, Evaluator, and Memory.
desk verdict STEP's Planner is a genuinely useful addition to CLIN, but the 'outperforms SOTA' claim is built on a backbone confound and the evaluation protocol needs tightening. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Planner component, which at each step receives the main task, a suggested strategy from the most recent trial, and the current history, then either refines the previous subtask or derives a new one and retrieves relevant insights from the last attempt's learning summary. Around it sit the Executor, which generates action candidates from the current subtask and insight; the Evaluator, which checks proposed actions against causal rules from memory and sends feedback or approval; and Memory, which stores causal-abstraction insights of the form 'X is necessary for Y' or 'X may not contribute to Y' together with a suggested strategy for the next trial. The argument's force comes from the ablation: without the Planner, the architecture reduces to an Executor-Evaluator-Memory loop that behaves like CLIN, while with the Planner it follows the intended subtask order and completes far more tasks.
What would settle it
Run SayCan, ReAct, and Reflexion on ScienceWorld with the same gpt-4o-mini backbone and the same five-episode protocol used for STEP; if STEP does not still exceed them, the claimed state-of-the-art advantage is an artifact of model choice rather than the framework.
Extended reading notes
Core claim
STEP's result on ScienceWorld is an overall score of 67.4, with 79.9 on short tasks and 54.9 on long tasks, compared with CLIN's 57.2, 71.7, and 42.7. The Planner is the main source of the gain: removing it lowers the average by 10.9 points overall and the task success rate by 16.7 percentage points, bringing the model back to roughly CLIN's level. The authors interpret this as showing that LLM agents are stepwise planners rather than autonomous planners: by isolating a prerequisite subtask, such as finding a hidden thermometer before measuring substance B, the Planner prevents the Executor from prematurely interacting with a known target and incurring a -100 penalty. The framework is designed to keep task order while using distilled in-context insights from memory, and it is positioned as agreeing with the view that LLMs cannot plan end-to-end on their own.
Load-bearing premise
The head-to-head advantage over previous agents rests on the assumption that published baseline scores, some produced with different underlying models, are directly comparable to scores produced with gpt-4o-mini.
Editorial extensions
If this is right
- STEP ranks first on 11 of the 18 ScienceWorld tasks and completes 12, so the gain is not confined to one easy subset of the benchmark.
- Removing the Planner lowers the average score by 10.9 points overall and success by 16.7 percentage points, which directly implicates task decomposition and memory distillation as the engine of the improvement.
- The long-task score improves by 28.6% over CLIN while the short-task score improves by 11.4%, suggesting the stepwise planning advantage grows with task length.
- The Evaluator alone cannot recover the lost performance when the Planner is removed, meaning action filtering without task-order guidance is not enough to match STEP.
Reading between the lines
- A natural extension is to transplant the Planner into other partially observable text environments where acting on a known goal location too early causes reset; the expected outcome is a similar reduction in premature -100/penalty actions.
- Because the published baselines were not run on the same backbone, rerunning all baselines with the same model is the cleanest test of whether the 67.4 score is a framework effect or partly a model effect.
- The failure mode of locking onto an irrelevant subtask suggests that a third check on subtask relevance before execution, rather than only action alignment, could be the next increment of gain.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces STEP, a four-component framework (Planner, Executor, Evaluator, Memory) for language agents, and evaluates it on the ScienceWorld benchmark using gpt-4o-mini as the backbone. The Planner decomposes the main task into subtasks and retrieves relevant insights from memory; the Executor generates actions; the Evaluator checks action alignment with learned rules; and Memory stores insights and strategies across episodes. The authors report an overall score of 67.4, claim that STEP consistently outperforms prior state-of-the-art agents, and present an ablation showing that removing the Planner degrades performance by about 10.9 points on average. The paper includes a GitHub repository and a qualitative analysis of planner-driven behavior on a temperature-measurement task.
Significance. If the controlled results hold, the paper makes a useful empirical contribution: it shows that adding an explicit stepwise Planner to a CLIN-style continual-learning agent can improve task completion in a dynamic text environment, and the ablation in Table 2 isolates the Planner as the main source of gain. The controlled comparison of STEP versus CLIN under the same gpt-4o-mini backbone is a fair and informative design, and the paper is honest in its limitation section about failure modes. However, the headline claim of outperforming all prior SOTA is not currently supported, because the SOTA baselines were not run under the same backbone or evaluation protocol, and the evaluation choices (best-of-five scoring, post hoc exclusion of false positives) create a risk of inflated point estimates. The central architectural insight may be sound, but the paper overstates its empirical scope.
major comments (4)
- [Section 4.1 'Configurations' and Table 1] The SOTA comparison is confounded by different backbone LLMs. Only CLIN and STEP are explicitly marked as using gpt-4o-mini; the SayCan, ReAct, and Reflexion scores are imported from prior work that used substantially weaker models. Because ScienceWorld performance is highly sensitive to model capability, the aggregate gap (e.g., 67.4 vs 39.4 for Reflexion) cannot be attributed to the STEP architecture. The abstract's claim that STEP 'consistently outperforms state-of-the-art models' and the corresponding statements in Section 4.2.1 are therefore not established. To support this claim, the authors must rerun SayCan, ReAct, and Reflexion under the same gpt-4o-mini backbone, the same best-of-5 episode protocol, and the same task split, or explicitly restrict the comparison to CLIN.
- [Section 4.1 'Evaluation Protocol' and Section 4.2.2] The evaluation protocol selects the highest score across 5 episodes, and Section 4.2.2 further excludes 'false-positive cases' where the agent accidentally completes the task. These choices are applied to STEP and CLIN but are not defined for the imported baselines, and the exclusion criterion is not specified a priori. This can materially inflate point estimates and make cross-model comparisons non-comparable. The authors should report the full distribution of episode scores, state the exact rule for identifying false positives, report how many traces were excluded for each condition, and confirm that the same rule was applied uniformly to all baselines.
- [Table 1 and Section 4.2.1] No measure of variance or statistical significance is reported. With only 18 tasks, many individual task differences are small (e.g., Chemistry1: STEP 61.0 vs Reflexion 70.4; Biology2: STEP 46.5 vs CLIN 59.3), and the aggregate gaps (79.9 vs 71.7 for short tasks, 54.9 vs 42.7 for long tasks) have no confidence intervals or significance tests. The claim that STEP 'consistently outperforms' needs error bars across repeated runs or a paired test across tasks; otherwise, the reported improvements may reflect noise in the evaluation procedure rather than a reliable advantage.
- [Section 5 'Limitation'] The authors acknowledge that the Planner can generate poor subtasks (e.g., the Biology wolf-painting failure) that lead to repetitive loops, and that strategy generation often fails to eliminate redundant traces. These failure modes are central to the proposed architecture, yet the paper does not quantify how often they occur or how much they lower the headline scores. Without such a frequency estimate or a sensitivity analysis, the reader cannot assess whether the Planner's average benefit is robust or dominated by a few successful tasks. The paper should report the incidence of these failure modes and their impact on the aggregate results.
minor comments (5)
- [Abstract and Section 4.2.1] The abstract states that STEP 'successfully completes 12 out of 18 tasks,' while Section 4.2.1 says STEP 'ranks first in 11 out of 18 tasks.' Please clarify the distinction between task completion (score 100) and ranking first, and ensure both numbers are defined consistently.
- [Section 4.2.2] The sentence 'we report the best traces while excluding false-positive cases' introduces a data-exclusion step without defining what constitutes a false positive. Provide a concrete definition and a count of excluded cases per condition.
- [Algorithm 1, line 3] The notation 's′ = rule(Sk−1)' is confusing because 's′' is later used for the Evaluator's rules, while 's' is used for the Planner's insights. Please rename one of these to avoid a notation collision.
- [Table 1 and Section 4.1] The abbreviations 'S' and 'L' are used both for task types (short/long) and for memory insights/strategies in the algorithm description; this makes some passages hard to follow. Please use distinct symbols or spell out the terms in each context.
- [Appendix B] The example trajectory says 'pick up cup containing lead' although the action space in Table 3 lists 'pick up OBJ'; please align the wording with the defined action syntax. Also, the sequence 'focus on thermometer' followed by 'focus on lead' followed again by 'focus on thermometer' appears redundant; verify that the transcript is representative.
Circularity Check
No circularity: the STEP results are an external-benchmark evaluation; the baseline-backbone mismatch is a validity concern, not circularity.
full rationale
The paper's central claim is an empirical comparison in ScienceWorld, an external benchmark, and no component of the derivation reduces to its own inputs by construction. STEP is presented as an agent architecture with Planner, Executor, Evaluator, and Memory components; no equations are derived, no parameters are fitted to a subset of the benchmark and then relabeled as predictions, and no 'uniqueness theorem' or load-bearing self-citation is used to force the architecture. The memory mechanism is self-referential in operation (the agent learns from its own prior episodes), but that is a mechanism description rather than circular validation, because the reported 67.4 score is measured against an independent environment with defined rewards and task goals. The ablation study (removing the Planner, Table 2) is an internal comparison whose outcome is not guaranteed by construction; it is an empirical measurement on the same external benchmark. The one substantive weakness, that SayCan, ReAct, and Reflexion baselines in Table 1 were not rerun with the same gpt-4o-mini backbone, is a comparability and external-validity concern about whether the 'outperforms SOTA' claim is established, not a circularity defect: the numbers still come from an independent benchmark rather than from the model's own outputs or from the authors' prior claims. CLIN, the closest predecessor, is authored by other researchers and is rerun under the same backbone as STEP, so the only controlled comparison is also the most directly relevant one. No step in the claimed derivation reduces by definition to its inputs, so the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (2)
- episode count and score selection =
5 episodes, best score taken
- maximum step limits =
37 steps short, 70 steps long
assumptions (4)
- domain assumption A frozen LLM (gpt-4o-mini) can generate useful subtasks, action candidates, and verbal reflections without parameter updates.
- domain assumption The ScienceWorld scoring protocol (including the -100 penalty for wrong focus and best-of-5 reporting) measures planning ability.
- domain assumption Causal-abstraction insights from CLIN, such as 'X is necessary for Y', are a useful memory format.
- ad hoc to paper Baseline scores from Lin et al. (2023) are comparable to gpt-4o-mini runs without being rerun under the same backbone.
invented entities (2)
-
Planner module
-
Evaluator module
Cite this review
Pith. "Pith review of One STEP at a time: Language Agents are Stepwise Planners." pith.science (2026). https://pith.science/paper/K2TYLP2W
@misc{pith2026241108432,
author = {Pith},
title = {Pith review of: One STEP at a time: Language Agents are Stepwise Planners},
year = {2026},
howpublished = {\url{https://pith.science/paper/K2TYLP2W}},
note = {Machine review of arXiv:2411.08432}
}
read the original abstract
Language agents have shown promising adaptability in dynamic environments to perform complex tasks. However, despite the versatile knowledge embedded in large language models, these agents still fall short when it comes to tasks that require planning. We introduce STEP, a novel framework designed to efficiently learn from previous experiences to enhance the planning capabilities of language agents in future steps. Concretely, STEP functions through four interconnected components. First, the Planner takes on the task, breaks it down into subtasks and provides relevant insights. Then the Executor generates action candidates, while the Evaluator ensures the actions align with learned rules from previous experiences. Lastly, Memory stores experiences to inform future decisions. In the ScienceWorld benchmark, our results show that STEP consistently outperforms state-of-the-art models, achieving an overall score of 67.4 and successfully completing 12 out of 18 tasks. These findings highlight STEP's potential as a framework for enhancing planning capabilities in language agents, paving the way for more sophisticated task-solving in dynamic environments.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, K...
arXiv 2022
-
[4]
Prithviraj Ammanabrolu and Matthew Hausknecht. 2020. https://arxiv.org/abs/2001.08837 Graph constrained reinforcement learning for natural language action spaces . Preprint, arXiv:2001.08837
arXiv 2020
-
[5]
Gabor Angeli, Melvin Jose Johnson Premkumar, and Christopher D. Manning. 2015. https://doi.org/10.3115/v1/P15-1034 Leveraging linguistic structure for open domain information extraction . In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Vol...
-
[6]
Brown, Miljan Martic, Shane Legg, and Dario Amodei
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2023. https://arxiv.org/abs/1706.03741 Deep reinforcement learning from human preferences . Preprint, arXiv:1706.03741
arXiv 2023
-
[7]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. https://arxiv.org/abs/2110.14168 Training verifiers to solve math word problems . Preprint, arXiv:2110.14168
arXiv 2021
-
[8]
Gautier Dagan, Frank Keller, and Alex Lascarides. 2023. https://arxiv.org/abs/2308.06391 Dynamic planning with a llm . Preprint, arXiv:2308.06391
arXiv 2023
Show all 47 references
-
[9]
Jiazhan Feng, Ruochen Xu, Junheng Hao, Hiteshi Sharma, Yelong Shen, Dongyan Zhao, and Weizhu Chen. 2023. https://arxiv.org/abs/2311.06158 Language models can be logical solvers . Preprint, arXiv:2311.06158
2023 arXiv
-
[10]
Yao Fu, Dong-Ki Kim, Jaekyeom Kim, Sungryull Sohn, Lajanugen Logeswaran, Kyunghoon Bae, and Honglak Lee. 2024. https://arxiv.org/abs/2403.08978 Autoguide: Automated generation and selection of state-aware guidelines for large language model agents . Preprint, arXiv:2403.08978
2024 arXiv
-
[11]
Majid Ghasemi, Amir Hossein Moosavi, Ibrahim Sorkhoh, Anjali Agrawal, Fadi Alzhouri, and Dariush Ebrahimi. 2024. https://arxiv.org/abs/2408.07712 An introduction to reinforcement learning: Fundamental concepts and practical applications . Preprint, arXiv:2408.07712
2024 arXiv
-
[12]
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2024 a . https://arxiv.org/abs/2305.11738 Critic: Large language models can self-correct with tool-interactive critiquing . Preprint, arXiv:2305.11738
2024 arXiv
-
[13]
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Minlie Huang, Nan Duan, and Weizhu Chen. 2024 b . https://arxiv.org/abs/2309.17452 Tora: A tool-integrated reasoning agent for mathematical problem solving . Preprint, arXiv:2309.17452
2024 arXiv
-
[14]
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. 2023. https://arxiv.org/abs/2305.14992 Reasoning with language model is planning with world model . Preprint, arXiv:2305.14992
2023 arXiv
-
[15]
Ji He, Jianshu Chen, Xiaodong He, Jianfeng Gao, Lihong Li, Li Deng, and Mari Ostendorf. 2016. https://arxiv.org/abs/1511.04636 Deep reinforcement learning with a natural language action space . Preprint, arXiv:1511.04636
2016 arXiv
-
[16]
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. https://arxiv.org/abs/2103.03874 Measuring mathematical problem solving with the math dataset . Preprint, arXiv:2103.03874
2021 arXiv
-
[17]
Xu Huang, Weiwen Liu, Xiaolong Chen, Xingmei Wang, Hao Wang, Defu Lian, Yasheng Wang, Ruiming Tang, and Enhong Chen. 2024. https://arxiv.org/abs/2402.02716 Understanding the planning of llm agents: A survey . Preprint, arXiv:2402.02716
2024 arXiv
-
[18]
Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. https://arxiv.org/abs/2310.06770 Swe-bench: Can language models resolve real-world github issues? Preprint, arXiv:2310.06770
2024 arXiv
-
[19]
Subbarao Kambhampati, Karthik Valmeekam, Lin Guan, Mudit Verma, Kaya Stechly, Siddhant Bhambri, Lucas Saldyt, and Anil Murthy. 2024. https://arxiv.org/abs/2402.01817 Llms can't plan, but can help planning in llm-modulo frameworks . Preprint, arXiv:2402.01817
2024 arXiv
-
[20]
Jikun Kang, Romain Laroche, Xingdi Yuan, Adam Trischler, Xue Liu, and Jie Fu. 2024. https://arxiv.org/abs/2305.16338 Think before you act: Decision transformers with working memory . Preprint, arXiv:2305.16338
2024 arXiv
-
[21]
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2023. https://arxiv.org/abs/2205.11916 Large language models are zero-shot reasoners . Preprint, arXiv:2205.11916
2023 arXiv
-
[22]
Bill Yuchen Lin, Yicheng Fu, Karina Yang, Faeze Brahman, Shiyu Huang, Chandra Bhagavatula, Prithviraj Ammanabrolu, Yejin Choi, and Xiang Ren. 2023. https://arxiv.org/abs/2305.17390 Swiftsage: A generative agent with fast and slow thinking for complex interactive tasks . Prepri...
2023 arXiv
-
[23]
Bo Liu, Yuqian Jiang, Xiaohan Zhang, Qiang Liu, Shiqi Zhang, Joydeep Biswas, and Peter Stone. 2023. https://arxiv.org/abs/2304.11477 Llm+p: Empowering large language models with optimal planning proficiency . Preprint, arXiv:2304.11477
2023 arXiv
-
[24]
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023. https://arxiv.o...
2023 arXiv
-
[25]
Bodhisattwa Prasad Majumder, Bhavana Dalvi Mishra, Peter Jansen, Oyvind Tafjord, Niket Tandon, Li Zhang, Chris Callison-Burch, and Peter Clark. 2023. Clin: A continually learning language agent for rapid task adaptation and generalization. arXiv preprint arXiv:2310.10134
2023 arXiv
-
[26]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
2022 arXiv
-
[27]
Abulhair Saparov and He He. 2023. https://arxiv.org/abs/2210.01240 Language models are greedy reasoners: A systematic formal analysis of chain-of-thought . Preprint, arXiv:2210.01240
2023 arXiv
-
[28]
Tran, Yi Tay, and Donald Metzler
Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Q. Tran, Yi Tay, and Donald Metzler. 2022. https://arxiv.org/abs/2207.07061 Confident adaptive language modeling . Preprint, arXiv:2207.07061
2022 arXiv
-
[29]
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023. https://arxiv.org/abs/2303.17580 Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face . Preprint, arXiv:2303.17580
2023 arXiv
-
[30]
Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. https://arxiv.org/abs/2303.11366 Reflexion: Language agents with verbal reinforcement learning . Preprint, arXiv:2303.11366
2023 arXiv
-
[31]
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. 2021. https://arxiv.org/abs/2010.03768 Alfworld: Aligning text and embodied environments for interactive learning . Preprint, arXiv:2010.03768
2021 arXiv
-
[32]
Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg. 2022. https://arxiv.org/abs/2209.11302 Progprompt: Generating situated robot task plans using large language models . Preprint, arXiv:2209.11302
2022 arXiv
-
[33]
Theodore R Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas L Griffiths. 2023. Cognitive architectures for language agents. arXiv preprint arXiv:2309.02427
2023 arXiv
-
[34]
Oyvind Tafjord, Bhavana Dalvi Mishra, and Peter Clark. 2021. https://arxiv.org/abs/2012.13048 Proofwriter: Generating implications, proofs, and abductive statements over natural language . Preprint, arXiv:2012.13048
2021 arXiv
-
[35]
Karthik Valmeekam, Matthew Marquez, and Subbarao Kambhampati. 2023. https://arxiv.org/abs/2310.08118 Can large language models really improve by self-critiquing their own plans? Preprint, arXiv:2310.08118
2023 arXiv
-
[36]
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023 a . https://arxiv.org/abs/2305.16291 Voyager: An open-ended embodied agent with large language models . Preprint, arXiv:2305.16291
2023 arXiv
-
[37]
Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. 2023 b . https://arxiv.org/abs/2305.04091 Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models . Preprint, arXiv:2305.04091
2023 arXiv
-
[38]
Ruoyao Wang, Peter Jansen, Marc-Alexandre Côté, and Prithviraj Ammanabrolu. 2022. https://arxiv.org/abs/2203.07540 Scienceworld: Is your agent smarter than a 5th grader? Preprint, arXiv:2203.07540
2022 arXiv
-
[39]
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2023 a . https://arxiv.org/abs/2207.01206 Webshop: Towards scalable real-world web interaction with grounded language agents . Preprint, arXiv:2207.01206
2023 arXiv
-
[40]
Griffiths, Yuan Cao, and Karthik Narasimhan
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023 b . https://arxiv.org/abs/2305.10601 Tree of thoughts: Deliberate problem solving with large language models . Preprint, arXiv:2305.10601
2023 arXiv
-
[41]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023 c . https://arxiv.org/abs/2210.03629 React: Synergizing reasoning and acting in language models . Preprint, arXiv:2210.03629
2023 arXiv
-
[42]
Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. 2023 a . https://arxiv.org/abs/2308.10144 Expel: Llm agents are experiential learners . Preprint, arXiv:2308.10144
2023 arXiv
-
[43]
Zirui Zhao, Wee Sun Lee, and David Hsu. 2023 b . https://arxiv.org/abs/2305.14078 Large language models as commonsense knowledge for large-scale task planning . Preprint, arXiv:2305.14078
2023 arXiv
-
[44]
Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, and Yu-Xiong Wang. 2024 a . https://arxiv.org/abs/2310.04406 Language agent tree search unifies reasoning acting and planning in language models . Preprint, arXiv:2310.04406
2024 arXiv
-
[45]
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, and Ed Chi. 2023. https://arxiv.org/abs/2205.10625 Least-to-most prompting enables complex reasoning in large language models . Preprint, arXiv...
2023 arXiv
-
[46]
Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig
Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. 2024 b . https://arxiv.org/abs/2307.13854 Webarena: A realistic web environment for building autonomous agents . Prepri...
2024 arXiv
-
[47]
Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu, Wenhao Yu, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, Simon Brunner, Chen Gong, Thong Hoang, Armel Randy Zebaze, Xiaoheng Hong, Wen-Ding Li, Jean Kaddour, Ming Xu, Zhihan Zhang, Prateek Ya...
2024 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.