REVIEW 3 major objections 6 minor 52 references
Conflicting memory traps stronger AI agents harder: they comply at similar rates, then crash to a shared low success floor.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 10:29 UTC pith:SREDRLJA
load-bearing objection Clear consumption-side diagnostic with a real multi-model regularity; the trap is well measured under constructed traps, less proven for ordinary retrieval. the 3 major comments →
The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Across models, conflicting memory produces similar compliance rates (roughly 63–72 percent), but after compliance success rates fall to a shared low floor (about 17–31 percent). Because stronger models start from higher baselines yet land on the same floor, they suffer larger absolute damage per compliance event. The dominant failure begins at entry, is amplified by repeated exposure, and is hard to reverse.
What carries the argument
Entry–Propagation–Recovery (E-P-R): a trajectory-level diagnostic that measures (1) whether retrieved memory first changes an action at the earliest exposed decision, (2) whether that change persists under continued exposure, and (3) whether the agent can realign after memory-induced divergence. Paired schedule interventions and the Recommendation Compliance Rate / Damage Per Compliance metrics turn the three phases into measurable quantities.
Load-bearing premise
The paper’s traps rely on hand-authored or model-drafted conflicting memories written in a matched template and on a memory-sensitive WebArena subset; the claim that the same entry-propagation-recovery dynamics hold for real retrieved experience rests on how representative those constructed traps are.
What would settle it
If agents of different strengths, given the same conflicting memories on long-horizon browser tasks, either show substantially different compliance rates or, after complying, retain success rates near their no-memory baselines instead of collapsing to a common low floor, the compliance-trap claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper shifts agent-memory research from supply (write/store/retrieve) to consumption, introducing Entry–Propagation–Recovery (E-P-R) as a trajectory-level diagnostic. On WebArena (schedule interventions on a 77-task memory-sensitive subset plus full-distribution pre-checks) and MemTrapBench (231 a-priori trap tasks with DECOY/UPTAKE/GROUNDING/OVERRIDE families), it reports a compliance trap: under persistent conflicting memory, Recommendation Compliance Rate (RCR) is high and roughly scale-independent (~63–72%), post-compliance success collapses to a low floor (~17–31%), and Damage Per Compliance (DPC) therefore grows with baseline capability (up to ~25.5 pp on WebArena and ~28.6 pp on MemTrapBench). Helpful memory, cross-task controls, placement ablations, horizon controls (MiniWoB++), format sweeps, and a 2×2 Entry×Grounding factorial support the pattern and the gate-chain reading.
Significance. If the compliance-trap regularity holds under realistic memory use, it is a high-value diagnostic result for long-horizon agents: it explains why stronger models can lose more absolute performance from the same wrong experience, and it motivates evaluating consumption (entry, propagation, recovery) rather than retrieval quality or final SR alone. Strengths include paired no-memory baselines, multi-model coverage, sign-flip permutation tests with bootstrap CIs, content-matched helpful/conflicting/cross-task memories, placement and horizon controls, full-distribution format sweeps, and an a-priori MemTrapBench that avoids pure outcome selection. The E-P-R framing and the RCR/DPC separation are clean, falsifiable contributions that future memory systems can use as evaluation targets even if they disagree on remedies.
major comments (3)
- [§4.1, Tables 1–3; Appendix C Table 17] §4.1 Table 1 and Appendix C Table 17 show that online-retrieved memory on the full 684-task WebArena distribution yields near-zero or mixed unconditional SR deltas, while the headline trap (Tables 2–3, §4.3–4.4) is measured under always-available, length-matched, high-salience DO/DON’T conflicting passages that reference executable UI elements. The central claim about consumption in memory-augmented agents therefore needs a sharper scope statement: either (i) additional experiments with intermittent, partial, or noisy retrieved experience (e.g., episodic summaries as in Reflexion/AWM, or relevance-gated intermittent injection), or (ii) explicit qualification that the trap characterizes the unsafe regime when conflicting content is persistently exposed, not ordinary retrieval on average. Without one of these, the transfer from constructed traps to the supply pipelines named in the introdu
- [§3.2; Table 2; Appendix B] RCR is defined via an LLM judge (Gemini-3-Flash, Appendix B) that is also an evaluated model in the main tables. The paper reports a looser {YES, UNCLEAR} robustness check and an EDR action-divergence subset (Table 8), but does not report inter-judge agreement, human agreement on a sample, or a fully non-Gemini judge for the headline CRCR subset used in Table 2. Because RCR is load-bearing for the claim that compliance rates are scale-independent while DPC tracks baseline, a short human-audit or second-judge agreement table on the WebArena CRCR labels is needed before the ~63–72% constancy claim can be treated as model-independent rather than judge-dependent.
- [§3.3; §4.3; Figure 4; Appendix A] The 77-task WebArena subset is filtered by outcome change under early helpful/conflicting memory on Qwen3.5-9B/27B (§3.3, Appendix A). MemTrapBench and the full-distribution sweeps mitigate selection bias for the existence of the trap, but several WebArena trajectory figures (Figure 4) and recovery tables (Tables 9–11) still rest on this filter. Absolute effect sizes and recovery rates on WebArena should be labeled as mechanism-diagnosis quantities, not population estimates, and any claim that compares WebArena DPC magnitudes to MemTrapBench DPC should state the selection difference explicitly in the main text (not only in the appendix).
minor comments (6)
- [Figure 3; Table 2] Figure 3 caption says n=74 for Gemini-3-Flash while Table 2 uses N=75/67/77 depending on model; reconcile denominators and state exclusion criteria in one place.
- [§3.3; Appendix E Table 21] The multiplicative gate chain ∆helpful = P(novel)·P(adopt)·P(ground|adopt)·P(lift|ground) is asserted in §3.3 and Appendix E; a short numerical check that the product matches observed helpful deltas (as claimed for WebArena models) would make the decomposition more than schematic.
- [Appendix A; References [8]] Gemma-4 technical report is cited as a 2026 model release without a stable paper; for reproducibility, pin exact checkpoint IDs and serving stack versions in Appendix A (vLLM version is given; model hashes or HF IDs would help).
- [Abstract; §1–2] Typo/grammar: abstract and intro use “asupply” / “Thisconsumptionprocess” style missing spaces in a few places in the provided text; also “V oyager” and “WebV oyager” have stray spaces—clean for camera-ready.
- [§5 Table 6] Table 6 MiniWoB++ median steps is listed as 3 in the prose of §5 but “30.0” appears in the table header row for Med. steps—clarify which is correct (prose says short-horizon ~2–3 steps).
- [§4.4; Appendix C] Cross-task control is sometimes called “irrelevant” (Table 13) and sometimes “cross-task”; use one term consistently in main text and tables.
Circularity Check
No significant circularity: the compliance trap is an empirical regularity from paired interventions, not a quantity forced by definition or fit.
full rationale
This paper’s load-bearing claims are measured outcomes under controlled memory injection, not first-principles derivations that collapse into their inputs. Entry–Propagation–Recovery is a diagnostic decomposition of trajectories, not a uniqueness theorem or fitted law. Recommendation Compliance Rate is an LLM-judge classification of whether step-1 actions execute the memory’s first primitive; Damage Per Compliance is the paired success difference on that compliant subset—a conditional statistic that avoids comparing an all-task baseline to a behavior-selected subset, not a parameter fit renamed as prediction. The similar RCR (~63–72%) and low post-compliance floor across models, and the larger absolute DPC for stronger agents, are contingent empirical findings: nothing in the metric definitions forces RCR to be scale-independent or post-compliance success to a common floor. MemTrapBench traps are authored a priori to isolate gates (DECOY/UPTAKE/GROUNDING/OVERRIDE), but whether agents adopt them at similar rates and fail to recover is still measured against no-memory baselines, not true by construction. The 77-task WebArena subset is outcome-selected (acknowledged by the authors) and can inflate magnitudes; that is selection bias / external-validity risk, not circular derivation, and is triangulated with full-distribution format sweeps and MemTrapBench. The four-gate product for helpful lift is explanatory accounting of successive necessary conditions, not a load-bearing prediction of the compliance trap. No self-citation uniqueness chain, no fitted input called prediction, no ansatz smuggled via citation. The work is self-contained empirical diagnosis.
Axiom & Free-Parameter Ledger
free parameters (3)
- Late-injection start step (step 3)
- RCR LLM-judge decision boundary (YES/NO/UNCLEAR)
- Memory-sensitive 77-task WebArena filter
axioms (3)
- domain assumption BrowserGym accessibility-tree observations plus greedy ReAct decoding are sufficient to study memory consumption dynamics of web agents.
- domain assumption Short DO/DON’T textual passages with matched length/structure are a fair proxy for external textual memories used by Reflexion/ExpeL/AWM-style systems.
- ad hoc to paper Trajectory divergence and recovery can be operationalized via action-string disagreement and re-alignment to the paired no-memory path.
invented entities (3)
-
Entry–Propagation–Recovery (E-P-R) framework
no independent evidence
-
MemTrapBench
no independent evidence
-
Compliance trap (RCR + DPC)
no independent evidence
read the original abstract
Memory is becoming a core component of long-horizon AI agents, allowing agents to reuse past experience when operating web browsers, software tools, and other interactive environments. Existing work mostly treats memory as a supply problem, asking what experience to write, how to store it, and which entry to retrieve for the next task. Yet we still lack a clear account of how models consume retrieved memory across a multi-step action trajectory. This consumption process matters because it determines not only what memories should be retrieved, but also what models and control policies are needed to use them safely. To diagnose this process, we propose Entry--Propagation--Recovery (E-P-R), a trajectory-level framework that asks where memory first changes an action, whether that change carries forward, and whether the agent can recover after leaving a correct path. We instantiate E-P-R on WebArena and on MemTrapBench, a controlled benchmark we build to isolate these phases. We find that the main failure often begins at entry: agents adopt conflicting memory at the first exposed decision point even when it is task-wrong. Repeated exposure then amplifies this early error, while recovery after divergence is weak. Together, these effects create a compliance trap: across models, conflicting memory induces similar compliance rates, but once agents comply, their success rates collapse to a low floor. Stronger agents therefore suffer larger absolute damage because each compliance event erases more baseline capability. These results suggest that memory-augmented agents should be evaluated not only by retrieval quality or final success rate, but by how they consume memory throughout the trajectory.
Figures
Reference graph
Works this paper leans on
-
[1]
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, K...
Pith/arXiv arXiv 2022
-
[2]
SeeClick: Harnessing GUI grounding for advanced visual GUI agents
Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang, and Zhiyong Wu. SeeClick: Harnessing GUI grounding for advanced visual GUI agents. InProceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2024
2024
-
[3]
Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. Mem0: Building production-ready AI agents with scalable long-term memory.arXiv preprint arXiv:2504.19413, 2025
Pith/arXiv arXiv 2025
-
[4]
Mind2Web: Towards a generalist agent for the web
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2Web: Towards a generalist agent for the web. InAdvances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2023
2023
-
[5]
Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste. The BrowserGym ecosystem for web agent research.arXiv preprint arXiv:2412.05467, 2024
Pith/arXiv arXiv 2024
-
[6]
Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste. WorkArena: How capable are web agents at solving common knowledge work tasks?arXiv preprint arXiv:2403.07718, 2024
Pith/arXiv arXiv 2024
-
[7]
Gemma 3 technical report.arXiv preprint arXiv:2503.19786, 2025
Gemma Team, Google DeepMind. Gemma 3 technical report.arXiv preprint arXiv:2503.19786, 2025
Pith/arXiv arXiv 2025
-
[8]
Gemma Team, Google DeepMind. Gemma 4. Google DeepMind model release, 2026. https: //deepmind.google/models/gemma/gemma-4/; technical report not yet available
2026
-
[9]
CUB: Benchmarking context utilisation techniques for language models
Lovisa Hagström, Youna Kim, Haeun Yu, Sang goo Lee, Richard Johansson, Hyunsoo Cho, and Isabelle Augenstein. CUB: Benchmarking context utilisation techniques for language models. arXiv preprint arXiv:2505.16518, 2025
Pith/arXiv arXiv 2025
-
[10]
WebV oyager: Building an end-to-end web agent with large multimodal models
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. WebV oyager: Building an end-to-end web agent with large multimodal models. InProceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2024
2024
-
[11]
CogAgent: A visual language model for GUI agents
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, and Jie Tang. CogAgent: A visual language model for GUI agents. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[12]
Large language models cannot self-correct reasoning yet
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xiny- ing Song, and Denny Zhou. Large language models cannot self-correct reasoning yet. In Proceedings of the International Conference on Learning Representations (ICLR), 2024
2024
-
[13]
Tomoyuki Kagaya, Thong Jing Yuan, Yuxuan Lou, Jayashree Karlekar, Sugiri Pranata, Akira Kinose, Koki Oguri, Felix Wick, and Yang You. RAP: Retrieval-augmented planning with contextual memory for multimodal LLM agents.arXiv preprint arXiv:2402.03610, 2024. 10
Pith/arXiv arXiv 2024
-
[14]
When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs.Transactions of the Association for Computational Linguistics, 12:1417–1440, 2024
Ryo Kamoi, Yusen Zhang, Nan Zhang, Jiawei Han, and Rui Zhang. When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs.Transactions of the Association for Computational Linguistics, 12:1417–1440, 2024
2024
-
[15]
VisualWebArena: Evaluating multimodal agents on realistic visual web tasks
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried. VisualWebArena: Evaluating multimodal agents on realistic visual web tasks. InProceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2024
2024
-
[16]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with PagedAttention. InProceedings of the ACM SIGOPS Symposium on Operating Systems Principles (SOSP), 2023
2023
-
[17]
Retrieval-augmented generation for knowledge-intensive NLP tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrieval-augmented generation for knowledge-intensive NLP tasks. InAdvances in Neural Information Processing Systems (NeurIPS), 2020
2020
-
[18]
WebSuite: Systematically evaluating why web agents fail
Eric Li and Jim Waldo. WebSuite: Systematically evaluating why web agents fail. 2024
2024
-
[19]
Genglin Liu, Shijie Geng, et al. WebCoach: Self-evolving web agents with cross-session memory guidance.arXiv preprint arXiv:2511.12997, 2025. Author list to be confirmed at camera-ready
arXiv 2025
-
[20]
Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the middle: How language models use long contexts.Transactions of the Association for Computational Linguistics (TACL), 12:157–173, 2024
2024
-
[21]
AgentBench: Evaluating LLMs as agents
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang. AgentBench: Evaluating LLMs as agents. InProceedings of the International Conference on Lea...
2024
-
[22]
AgentBoard: An analytical evaluation board of multi-turn LLM agents
Chang Ma, Junlei Zhang, Zhihao Zhu, Cheng Yang, Yujiu Yang, Yaohui Jin, Zhenzhong Lan, Lingpeng Kong, and Junxian He. AgentBoard: An analytical evaluation board of multi-turn LLM agents. InAdvances in Neural Information Processing Systems (NeurIPS), 2024
2024
-
[23]
Self- refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. Self- refine: Iterative refinement with self-feedback. InAdvances in Neural Information Processing Sy...
2023
-
[24]
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christo- pher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schul- man. WebGPT: Browser-assisted question-answering with human feedback.arXiv preprint a...
Pith/arXiv arXiv 2021
-
[25]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human fee...
2022
-
[26]
Patil, Ion Stoica, and Joseph E
Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. MemGPT: Towards LLMs as operating systems. 2024
2024
-
[27]
O’Brien, Carrie J
Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the ACM Symposium on User Interface Software and Technology (UIST), 2023. 11
2023
-
[28]
Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan
Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tyre, Ethan Durmus, Guy Gur-Ari, Jackson Kern...
2023
-
[29]
Bowman, Esin Durmus, Zac Hatfield-Dodds, Scott R
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. Towards understanding sycophancy in language models. Proceedings of the Intern...
2024
-
[30]
World of bits: An open-domain platform for web-based agents
Tianlin Shi, Andrej Karpathy, Linxi Fan, Jonathan Hernandez, and Percy Liang. World of bits: An open-domain platform for web-based agents. InProceedings of the International Conference on Machine Learning (ICML), 2017
2017
-
[31]
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning. InAdvances in Neural Information Processing Systems (NeurIPS), 2023
2023
-
[32]
V oyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. V oyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023
Pith/arXiv arXiv 2023
-
[33]
A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6), 2024
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Ji-Rong Wen. A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6), 2024
2024
-
[34]
Le, Ed H
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V . Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. InProceedings of the International Conference on Learning Representations (ICLR), 2023
2023
-
[35]
Agent workflow memory
Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried, and Graham Neubig. Agent workflow memory. InProceedings of the International Conference on Machine Learning (ICML), 2025
2025
-
[36]
Chi, Quoc V
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V . Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. InAdvances in Neural Information Processing Systems (NeurIPS), 2022
2022
-
[37]
Jerry Wei, Da Huang, Yifeng Lu, Denny Zhou, and Quoc V . Le. Simple synthetic data reduces sycophancy in large language models.arXiv preprint arXiv:2308.03958, 2024
Pith/arXiv arXiv 2024
-
[38]
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, Rui Zheng, Xiaoran Fan, Xiao Wang, Limao Xiong, Yuhao Zhou, Weiran Wang, Changhao Jiang, Yicheng Zou, Xiangyang Liu, Zhangyue Yin, Shihan Dou, Rongxiang Weng, Wensen Cheng, Qi Zhang, Wenjuan Qin, Yongyan Zheng, Xipeng Qiu, Xuanjing Huang, a...
Pith/arXiv arXiv 2023
-
[39]
OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. InAdvances in Neural Information P...
2024
-
[40]
Zidi Xiong, Yuping Lin, Wenya Xie, Pengfei He, Zirui Liu, Jiliang Tang, Himabindu Lakkaraju, and Zhen Xiang. How memory management impacts LLM agents: An empirical study of experience-following behavior.arXiv preprint arXiv:2505.16067, 2025
arXiv 2025
-
[41]
Frank F. Xu, Yufan Song, Boxuan Li, Yuxuan Tang, Kritanjali Jain, Mengxue Bao, Zora Zhiruo Wang, Xuhui Zhou, Zhe Zhang, Mingyang Yang, Haotian Zhu, Yueqi Song, Bowen Li, Xinyuan Pan, Daniel Fried, and Graham Neubig. TheAgentCompany: Benchmarking LLM agents on consequential real world tasks.arXiv preprint arXiv:2412.14161, 2024
Pith/arXiv arXiv 2024
-
[42]
Knowledge conflicts for LLMs: A survey.Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024
Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang, Yue Zhang, and Wei Xu. Knowledge conflicts for LLMs: A survey.Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024
2024
-
[43]
A-MEM: Agentic memory for LLM agents.arXiv preprint arXiv:2502.12110, 2025
Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang. A-MEM: Agentic memory for LLM agents.arXiv preprint arXiv:2502.12110, 2025
Pith/arXiv arXiv 2025
-
[44]
Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, ...
Pith/arXiv arXiv 2025
-
[45]
Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, and Jianfeng Gao. Set- of-mark prompting unleashes extraordinary visual grounding in GPT-4V.arXiv preprint arXiv:2310.11441, 2023
Pith/arXiv arXiv 2023
-
[46]
WebShop: Towards scalable real-world web interaction with grounded language agents
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. WebShop: Towards scalable real-world web interaction with grounded language agents. InAdvances in Neural Information Processing Systems (NeurIPS), 2022
2022
-
[47]
Griffiths, Yuan Cao, and Karthik Narasimhan
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. In Advances in Neural Information Processing Systems (NeurIPS), 2023
2023
-
[48]
ReAct: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. InProceedings of the International Conference on Learning Representations (ICLR), 2023
2023
-
[49]
ExpeL: LLM agents are experiential learners
Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. ExpeL: LLM agents are experiential learners. InProceedings of the AAAI Conference on Artificial Intelligence, 2024
2024
-
[50]
GPT-4V(ision) is a generalist web agent, if grounded
Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. GPT-4V(ision) is a generalist web agent, if grounded. InProceedings of the International Conference on Machine Learning (ICML), 2024
2024
-
[51]
MemoryBank: Enhancing large language models with long-term memory.Proceedings of the AAAI Conference on Artificial Intelligence, 2024
Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang. MemoryBank: Enhancing large language models with long-term memory.Proceedings of the AAAI Conference on Artificial Intelligence, 2024
2024
-
[52]
Write a Review
Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. WebArena: A realistic web environment for building autonomous agents. InProceedings of the International Conference on Learning Representations (ICLR), 2024. 13 A Experimental setup and reproducibility This appe...
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.