Pith. sign in

REVIEW 3 major objections 6 minor 52 references

Conflicting memory traps stronger AI agents harder: they comply at similar rates, then crash to a shared low success floor.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 10:29 UTC pith:SREDRLJA

load-bearing objection Clear consumption-side diagnostic with a real multi-model regularity; the trap is well measured under constructed traps, less proven for ordinary retrieval. the 3 major comments →

arxiv 2607.10608 v1 pith:SREDRLJA submitted 2026-07-12 cs.AI

The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory

classification cs.AI
keywords AI agentsmemory consumptioncompliance trapEntry-Propagation-RecoveryWebArenaMemTrapBenchconflicting memorylong-horizon agents
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Long-horizon AI agents that browse the web or use tools are increasingly given retrieved memories of past experience, but most research only asks what to store and retrieve. This paper asks a different question: once a memory is already in the agent’s context, how does the model actually use it across many steps? The authors introduce a three-phase diagnostic called Entry–Propagation–Recovery. Entry asks whether the memory changes the first action where it is visible; Propagation asks whether that change keeps shaping later steps under repeated exposure; Recovery asks whether the agent can get back on a correct path after it has already diverged. On WebArena and on a new controlled benchmark, MemTrapBench, they find that the main failure usually starts at entry: agents adopt task-wrong but plausible memory at the first decision point. Repeated exposure amplifies the error, and recovery after divergence is weak. The result is a compliance trap: different models comply with conflicting memory at roughly the same rate, but once they comply their success collapses to a similar low floor. Stronger agents therefore lose more absolute performance because each compliance event erases more of their baseline capability. The practical message is that memory systems should be judged by how agents consume memory throughout a trajectory, not only by retrieval quality or final success rate.

Core claim

Across models, conflicting memory produces similar compliance rates (roughly 63–72 percent), but after compliance success rates fall to a shared low floor (about 17–31 percent). Because stronger models start from higher baselines yet land on the same floor, they suffer larger absolute damage per compliance event. The dominant failure begins at entry, is amplified by repeated exposure, and is hard to reverse.

What carries the argument

Entry–Propagation–Recovery (E-P-R): a trajectory-level diagnostic that measures (1) whether retrieved memory first changes an action at the earliest exposed decision, (2) whether that change persists under continued exposure, and (3) whether the agent can realign after memory-induced divergence. Paired schedule interventions and the Recommendation Compliance Rate / Damage Per Compliance metrics turn the three phases into measurable quantities.

Load-bearing premise

The paper’s traps rely on hand-authored or model-drafted conflicting memories written in a matched template and on a memory-sensitive WebArena subset; the claim that the same entry-propagation-recovery dynamics hold for real retrieved experience rests on how representative those constructed traps are.

What would settle it

If agents of different strengths, given the same conflicting memories on long-horizon browser tasks, either show substantially different compliance rates or, after complying, retain success rates near their no-memory baselines instead of collapsing to a common low floor, the compliance-trap claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper shifts agent-memory research from supply (write/store/retrieve) to consumption, introducing Entry–Propagation–Recovery (E-P-R) as a trajectory-level diagnostic. On WebArena (schedule interventions on a 77-task memory-sensitive subset plus full-distribution pre-checks) and MemTrapBench (231 a-priori trap tasks with DECOY/UPTAKE/GROUNDING/OVERRIDE families), it reports a compliance trap: under persistent conflicting memory, Recommendation Compliance Rate (RCR) is high and roughly scale-independent (~63–72%), post-compliance success collapses to a low floor (~17–31%), and Damage Per Compliance (DPC) therefore grows with baseline capability (up to ~25.5 pp on WebArena and ~28.6 pp on MemTrapBench). Helpful memory, cross-task controls, placement ablations, horizon controls (MiniWoB++), format sweeps, and a 2×2 Entry×Grounding factorial support the pattern and the gate-chain reading.

Significance. If the compliance-trap regularity holds under realistic memory use, it is a high-value diagnostic result for long-horizon agents: it explains why stronger models can lose more absolute performance from the same wrong experience, and it motivates evaluating consumption (entry, propagation, recovery) rather than retrieval quality or final SR alone. Strengths include paired no-memory baselines, multi-model coverage, sign-flip permutation tests with bootstrap CIs, content-matched helpful/conflicting/cross-task memories, placement and horizon controls, full-distribution format sweeps, and an a-priori MemTrapBench that avoids pure outcome selection. The E-P-R framing and the RCR/DPC separation are clean, falsifiable contributions that future memory systems can use as evaluation targets even if they disagree on remedies.

major comments (3)
  1. [§4.1, Tables 1–3; Appendix C Table 17] §4.1 Table 1 and Appendix C Table 17 show that online-retrieved memory on the full 684-task WebArena distribution yields near-zero or mixed unconditional SR deltas, while the headline trap (Tables 2–3, §4.3–4.4) is measured under always-available, length-matched, high-salience DO/DON’T conflicting passages that reference executable UI elements. The central claim about consumption in memory-augmented agents therefore needs a sharper scope statement: either (i) additional experiments with intermittent, partial, or noisy retrieved experience (e.g., episodic summaries as in Reflexion/AWM, or relevance-gated intermittent injection), or (ii) explicit qualification that the trap characterizes the unsafe regime when conflicting content is persistently exposed, not ordinary retrieval on average. Without one of these, the transfer from constructed traps to the supply pipelines named in the introdu
  2. [§3.2; Table 2; Appendix B] RCR is defined via an LLM judge (Gemini-3-Flash, Appendix B) that is also an evaluated model in the main tables. The paper reports a looser {YES, UNCLEAR} robustness check and an EDR action-divergence subset (Table 8), but does not report inter-judge agreement, human agreement on a sample, or a fully non-Gemini judge for the headline CRCR subset used in Table 2. Because RCR is load-bearing for the claim that compliance rates are scale-independent while DPC tracks baseline, a short human-audit or second-judge agreement table on the WebArena CRCR labels is needed before the ~63–72% constancy claim can be treated as model-independent rather than judge-dependent.
  3. [§3.3; §4.3; Figure 4; Appendix A] The 77-task WebArena subset is filtered by outcome change under early helpful/conflicting memory on Qwen3.5-9B/27B (§3.3, Appendix A). MemTrapBench and the full-distribution sweeps mitigate selection bias for the existence of the trap, but several WebArena trajectory figures (Figure 4) and recovery tables (Tables 9–11) still rest on this filter. Absolute effect sizes and recovery rates on WebArena should be labeled as mechanism-diagnosis quantities, not population estimates, and any claim that compares WebArena DPC magnitudes to MemTrapBench DPC should state the selection difference explicitly in the main text (not only in the appendix).
minor comments (6)
  1. [Figure 3; Table 2] Figure 3 caption says n=74 for Gemini-3-Flash while Table 2 uses N=75/67/77 depending on model; reconcile denominators and state exclusion criteria in one place.
  2. [§3.3; Appendix E Table 21] The multiplicative gate chain ∆helpful = P(novel)·P(adopt)·P(ground|adopt)·P(lift|ground) is asserted in §3.3 and Appendix E; a short numerical check that the product matches observed helpful deltas (as claimed for WebArena models) would make the decomposition more than schematic.
  3. [Appendix A; References [8]] Gemma-4 technical report is cited as a 2026 model release without a stable paper; for reproducibility, pin exact checkpoint IDs and serving stack versions in Appendix A (vLLM version is given; model hashes or HF IDs would help).
  4. [Abstract; §1–2] Typo/grammar: abstract and intro use “asupply” / “Thisconsumptionprocess” style missing spaces in a few places in the provided text; also “V oyager” and “WebV oyager” have stray spaces—clean for camera-ready.
  5. [§5 Table 6] Table 6 MiniWoB++ median steps is listed as 3 in the prose of §5 but “30.0” appears in the table header row for Med. steps—clarify which is correct (prose says short-horizon ~2–3 steps).
  6. [§4.4; Appendix C] Cross-task control is sometimes called “irrelevant” (Table 13) and sometimes “cross-task”; use one term consistently in main text and tables.

Circularity Check

0 steps flagged

No significant circularity: the compliance trap is an empirical regularity from paired interventions, not a quantity forced by definition or fit.

full rationale

This paper’s load-bearing claims are measured outcomes under controlled memory injection, not first-principles derivations that collapse into their inputs. Entry–Propagation–Recovery is a diagnostic decomposition of trajectories, not a uniqueness theorem or fitted law. Recommendation Compliance Rate is an LLM-judge classification of whether step-1 actions execute the memory’s first primitive; Damage Per Compliance is the paired success difference on that compliant subset—a conditional statistic that avoids comparing an all-task baseline to a behavior-selected subset, not a parameter fit renamed as prediction. The similar RCR (~63–72%) and low post-compliance floor across models, and the larger absolute DPC for stronger agents, are contingent empirical findings: nothing in the metric definitions forces RCR to be scale-independent or post-compliance success to a common floor. MemTrapBench traps are authored a priori to isolate gates (DECOY/UPTAKE/GROUNDING/OVERRIDE), but whether agents adopt them at similar rates and fail to recover is still measured against no-memory baselines, not true by construction. The 77-task WebArena subset is outcome-selected (acknowledged by the authors) and can inflate magnitudes; that is selection bias / external-validity risk, not circular derivation, and is triangulated with full-distribution format sweeps and MemTrapBench. The four-gate product for helpful lift is explanatory accounting of successive necessary conditions, not a load-bearing prediction of the compliance trap. No self-citation uniqueness chain, no fitted input called prediction, no ansatz smuggled via citation. The work is self-contained empirical diagnosis.

Axiom & Free-Parameter Ledger

3 free parameters · 3 axioms · 3 invented entities

Empirical agent paper. Load-bearing content is experimental design choices and operational definitions rather than free physical constants. The main invented constructs are the diagnostic framework, metrics, and benchmark; domain assumptions concern observation interfaces and memory formats common in the literature.

free parameters (3)
  • Late-injection start step (step 3)
    Chosen by hand relative to short WebArena horizons; used to probe history resistance / recovery complement.
  • RCR LLM-judge decision boundary (YES/NO/UNCLEAR)
    Compliance subset depends on a temperature-0 Gemini judge rubric; affects which trajectories enter DPC.
  • Memory-sensitive 77-task WebArena filter
    Tasks retained if outcome flips under early helpful/conflicting on Qwen3.5-9B/27B; inflates effect sizes relative to full distribution (authors triangulate with MemTrapBench and full 684-task sweeps).
axioms (3)
  • domain assumption BrowserGym accessibility-tree observations plus greedy ReAct decoding are sufficient to study memory consumption dynamics of web agents.
    All primary experiments share this interface; visual grounding and stochastic decoding are largely out of scope.
  • domain assumption Short DO/DON’T textual passages with matched length/structure are a fair proxy for external textual memories used by Reflexion/ExpeL/AWM-style systems.
    Stated as following common external-memory design; format sweeps partially test robustness.
  • ad hoc to paper Trajectory divergence and recovery can be operationalized via action-string disagreement and re-alignment to the paired no-memory path.
    Core E-P-R measurement choice; type-text differences alone do not count as divergence.
invented entities (3)
  • Entry–Propagation–Recovery (E-P-R) framework no independent evidence
    purpose: Decompose memory consumption into first action change, persistence under exposure, and post-divergence correction.
    Central diagnostic contribution; not a physical entity but a new measurement ontology for agent trajectories.
  • MemTrapBench no independent evidence
    purpose: Controlled long-horizon browser tasks isolating DECOY/UPTAKE/GROUNDING/OVERRIDE gates without outcome-based selection.
    New benchmark constructed for mechanism diagnosis; independent of WebArena subset selection.
  • Compliance trap (RCR + DPC) no independent evidence
    purpose: Name the regularity that compliance rates are similar while conditional damage scales with baseline capability.
    Primary empirical claim package; defined via LLM-judged recommendation compliance and paired success drop.

pith-pipeline@v1.1.0-grok45 · 31562 in / 2769 out tokens · 38470 ms · 2026-07-14T10:29:44.754732+00:00 · methodology

0 comments
read the original abstract

Memory is becoming a core component of long-horizon AI agents, allowing agents to reuse past experience when operating web browsers, software tools, and other interactive environments. Existing work mostly treats memory as a supply problem, asking what experience to write, how to store it, and which entry to retrieve for the next task. Yet we still lack a clear account of how models consume retrieved memory across a multi-step action trajectory. This consumption process matters because it determines not only what memories should be retrieved, but also what models and control policies are needed to use them safely. To diagnose this process, we propose Entry--Propagation--Recovery (E-P-R), a trajectory-level framework that asks where memory first changes an action, whether that change carries forward, and whether the agent can recover after leaving a correct path. We instantiate E-P-R on WebArena and on MemTrapBench, a controlled benchmark we build to isolate these phases. We find that the main failure often begins at entry: agents adopt conflicting memory at the first exposed decision point even when it is task-wrong. Repeated exposure then amplifies this early error, while recovery after divergence is weak. Together, these effects create a compliance trap: across models, conflicting memory induces similar compliance rates, but once agents comply, their success rates collapse to a low floor. Stronger agents therefore suffer larger absolute damage because each compliance event erases more baseline capability. These results suggest that memory-augmented agents should be evaluated not only by retrieval quality or final success rate, but by how they consume memory throughout the trajectory.

Figures

Figures reproduced from arXiv: 2607.10608 by Alan Yuille, Xinyi Bai, Yixiong Chen.

Figure 1
Figure 1. Figure 1: Motivation and Overview. (A) Prior work studies the memory supply side. (B) We analyze memory consumption: how trajectories change with/without memory injection. it later. Prior work has observed related failures in experience following [40], but we still lack a trajectory-level account of where memory first changes behavior, how far that change carries forward, and whether the agent can undo it [PITH_FUL… view at source ↗
Figure 2
Figure 2. Figure 2: Entry–Propagation–Recovery diagnostic framework. Entry: does retrieved memory change the agent’s action? Propagation: does the change persist under later observations? Recovery: can the agent correct a harmful deviation? Helpful memory improves trajectories when grounded; conflicting memory harms when adopted and not recovered. By fixing the memory passage and changing only its availability over time, we t… view at source ↗
Figure 3
Figure 3. Figure 3: WebArena schedule deltas on the 77-task subset (5 models, including Gemini-3-Flash, n=74). (a) Helpful memory: modest gains, mostly persistent. (b) Persistent conflicting hurts every model; peaks on Qwen3.5-27B (−20.8 pp) and reaches −14.9 pp on Gemini-3-Flash. Late conflicting is weak on WebArena’s ≈ 7-step median; MemTrapBench (Section 4.4) restores late helpful. Cross￾task control within ±3 pp (Appendix… view at source ↗
Figure 4
Figure 4. Figure 4: Entry and recovery on WebArena under persistent injection. (a) Memory enters early. P(tentry ≤ t) reaches 0.86–0.94 by step 7 on all models. (b) The recovery rates are low. Conditional on divergence, helpful trajectories re-align (27–42%) more often than conflicting (7–15%). the compliance trap. (i) RCR is high and approximately scale-independent (63–72% across the five models). (ii) Conditional success af… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

52 extracted references · 15 linked inside Pith

  1. [1]

    Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, K...

  2. [2]

    SeeClick: Harnessing GUI grounding for advanced visual GUI agents

    Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang, and Zhiyong Wu. SeeClick: Harnessing GUI grounding for advanced visual GUI agents. InProceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2024

  3. [3]

    Mem0: Building production-ready AI agents with scalable long-term memory.arXiv preprint arXiv:2504.19413, 2025

    Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. Mem0: Building production-ready AI agents with scalable long-term memory.arXiv preprint arXiv:2504.19413, 2025

  4. [4]

    Mind2Web: Towards a generalist agent for the web

    Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2Web: Towards a generalist agent for the web. InAdvances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2023

  5. [5]

    Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste

    Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste. The BrowserGym ecosystem for web agent research.arXiv preprint arXiv:2412.05467, 2024

  6. [6]

    Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste

    Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste. WorkArena: How capable are web agents at solving common knowledge work tasks?arXiv preprint arXiv:2403.07718, 2024

  7. [7]

    Gemma 3 technical report.arXiv preprint arXiv:2503.19786, 2025

    Gemma Team, Google DeepMind. Gemma 3 technical report.arXiv preprint arXiv:2503.19786, 2025

  8. [8]

    Gemma Team, Google DeepMind. Gemma 4. Google DeepMind model release, 2026. https: //deepmind.google/models/gemma/gemma-4/; technical report not yet available

  9. [9]

    CUB: Benchmarking context utilisation techniques for language models

    Lovisa Hagström, Youna Kim, Haeun Yu, Sang goo Lee, Richard Johansson, Hyunsoo Cho, and Isabelle Augenstein. CUB: Benchmarking context utilisation techniques for language models. arXiv preprint arXiv:2505.16518, 2025

  10. [10]

    WebV oyager: Building an end-to-end web agent with large multimodal models

    Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. WebV oyager: Building an end-to-end web agent with large multimodal models. InProceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2024

  11. [11]

    CogAgent: A visual language model for GUI agents

    Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, and Jie Tang. CogAgent: A visual language model for GUI agents. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024

  12. [12]

    Large language models cannot self-correct reasoning yet

    Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xiny- ing Song, and Denny Zhou. Large language models cannot self-correct reasoning yet. In Proceedings of the International Conference on Learning Representations (ICLR), 2024

  13. [13]

    RAP: Retrieval-augmented planning with contextual memory for multimodal LLM agents.arXiv preprint arXiv:2402.03610, 2024

    Tomoyuki Kagaya, Thong Jing Yuan, Yuxuan Lou, Jayashree Karlekar, Sugiri Pranata, Akira Kinose, Koki Oguri, Felix Wick, and Yang You. RAP: Retrieval-augmented planning with contextual memory for multimodal LLM agents.arXiv preprint arXiv:2402.03610, 2024. 10

  14. [14]

    When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs.Transactions of the Association for Computational Linguistics, 12:1417–1440, 2024

    Ryo Kamoi, Yusen Zhang, Nan Zhang, Jiawei Han, and Rui Zhang. When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs.Transactions of the Association for Computational Linguistics, 12:1417–1440, 2024

  15. [15]

    VisualWebArena: Evaluating multimodal agents on realistic visual web tasks

    Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried. VisualWebArena: Evaluating multimodal agents on realistic visual web tasks. InProceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2024

  16. [16]

    Gonzalez, Hao Zhang, and Ion Stoica

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with PagedAttention. InProceedings of the ACM SIGOPS Symposium on Operating Systems Principles (SOSP), 2023

  17. [17]

    Retrieval-augmented generation for knowledge-intensive NLP tasks

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrieval-augmented generation for knowledge-intensive NLP tasks. InAdvances in Neural Information Processing Systems (NeurIPS), 2020

  18. [18]

    WebSuite: Systematically evaluating why web agents fail

    Eric Li and Jim Waldo. WebSuite: Systematically evaluating why web agents fail. 2024

  19. [19]

    WebCoach: Self-evolving web agents with cross-session memory guidance.arXiv preprint arXiv:2511.12997, 2025

    Genglin Liu, Shijie Geng, et al. WebCoach: Self-evolving web agents with cross-session memory guidance.arXiv preprint arXiv:2511.12997, 2025. Author list to be confirmed at camera-ready

  20. [20]

    Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

    Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the middle: How language models use long contexts.Transactions of the Association for Computational Linguistics (TACL), 12:157–173, 2024

  21. [21]

    AgentBench: Evaluating LLMs as agents

    Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang. AgentBench: Evaluating LLMs as agents. InProceedings of the International Conference on Lea...

  22. [22]

    AgentBoard: An analytical evaluation board of multi-turn LLM agents

    Chang Ma, Junlei Zhang, Zhihao Zhu, Cheng Yang, Yujiu Yang, Yaohui Jin, Zhenzhong Lan, Lingpeng Kong, and Junxian He. AgentBoard: An analytical evaluation board of multi-turn LLM agents. InAdvances in Neural Information Processing Systems (NeurIPS), 2024

  23. [23]

    Self- refine: Iterative refinement with self-feedback

    Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. Self- refine: Iterative refinement with self-feedback. InAdvances in Neural Information Processing Sy...

  24. [24]

    WebGPT: Browser-assisted question-answering with human feedback.arXiv preprint arXiv:2112.09332, 2021

    Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christo- pher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schul- man. WebGPT: Browser-assisted question-answering with human feedback.arXiv preprint a...

  25. [25]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human fee...

  26. [26]

    Patil, Ion Stoica, and Joseph E

    Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. MemGPT: Towards LLMs as operating systems. 2024

  27. [27]

    O’Brien, Carrie J

    Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the ACM Symposium on User Interface Software and Technology (UIST), 2023. 11

  28. [28]

    Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan

    Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tyre, Ethan Durmus, Guy Gur-Ari, Jackson Kern...

  29. [29]

    Bowman, Esin Durmus, Zac Hatfield-Dodds, Scott R

    Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. Towards understanding sycophancy in language models. Proceedings of the Intern...

  30. [30]

    World of bits: An open-domain platform for web-based agents

    Tianlin Shi, Andrej Karpathy, Linxi Fan, Jonathan Hernandez, and Percy Liang. World of bits: An open-domain platform for web-based agents. InProceedings of the International Conference on Machine Learning (ICML), 2017

  31. [31]

    Reflexion: Language agents with verbal reinforcement learning

    Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning. InAdvances in Neural Information Processing Systems (NeurIPS), 2023

  32. [32]

    V oyager: An open-ended embodied agent with large language models

    Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. V oyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023

  33. [33]

    A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6), 2024

    Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Ji-Rong Wen. A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6), 2024

  34. [34]

    Le, Ed H

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V . Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. InProceedings of the International Conference on Learning Representations (ICLR), 2023

  35. [35]

    Agent workflow memory

    Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried, and Graham Neubig. Agent workflow memory. InProceedings of the International Conference on Machine Learning (ICML), 2025

  36. [36]

    Chi, Quoc V

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V . Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. InAdvances in Neural Information Processing Systems (NeurIPS), 2022

  37. [37]

    Jerry Wei, Da Huang, Yifeng Lu, Denny Zhou, and Quoc V . Le. Simple synthetic data reduces sycophancy in large language models.arXiv preprint arXiv:2308.03958, 2024

  38. [38]

    The rise and potential of large language model based agents: A survey.arXiv preprint arXiv:2309.07864, 2023

    Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, Rui Zheng, Xiaoran Fan, Xiao Wang, Limao Xiong, Yuhao Zhou, Weiran Wang, Changhao Jiang, Yicheng Zou, Xiangyang Liu, Zhangyue Yin, Shihan Dou, Rongxiang Weng, Wensen Cheng, Qi Zhang, Wenjuan Qin, Yongyan Zheng, Xipeng Qiu, Xuanjing Huang, a...

  39. [39]

    OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments

    Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. InAdvances in Neural Information P...

  40. [40]

    How memory management impacts LLM agents: An empirical study of experience-following behavior.arXiv preprint arXiv:2505.16067, 2025

    Zidi Xiong, Yuping Lin, Wenya Xie, Pengfei He, Zirui Liu, Jiliang Tang, Himabindu Lakkaraju, and Zhen Xiang. How memory management impacts LLM agents: An empirical study of experience-following behavior.arXiv preprint arXiv:2505.16067, 2025

  41. [41]

    Frank F. Xu, Yufan Song, Boxuan Li, Yuxuan Tang, Kritanjali Jain, Mengxue Bao, Zora Zhiruo Wang, Xuhui Zhou, Zhe Zhang, Mingyang Yang, Haotian Zhu, Yueqi Song, Bowen Li, Xinyuan Pan, Daniel Fried, and Graham Neubig. TheAgentCompany: Benchmarking LLM agents on consequential real world tasks.arXiv preprint arXiv:2412.14161, 2024

  42. [42]

    Knowledge conflicts for LLMs: A survey.Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024

    Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang, Yue Zhang, and Wei Xu. Knowledge conflicts for LLMs: A survey.Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024

  43. [43]

    A-MEM: Agentic memory for LLM agents.arXiv preprint arXiv:2502.12110, 2025

    Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang. A-MEM: Agentic memory for LLM agents.arXiv preprint arXiv:2502.12110, 2025

  44. [44]

    Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, ...

  45. [45]

    Set- of-mark prompting unleashes extraordinary visual grounding in GPT-4V.arXiv preprint arXiv:2310.11441, 2023

    Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, and Jianfeng Gao. Set- of-mark prompting unleashes extraordinary visual grounding in GPT-4V.arXiv preprint arXiv:2310.11441, 2023

  46. [46]

    WebShop: Towards scalable real-world web interaction with grounded language agents

    Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. WebShop: Towards scalable real-world web interaction with grounded language agents. InAdvances in Neural Information Processing Systems (NeurIPS), 2022

  47. [47]

    Griffiths, Yuan Cao, and Karthik Narasimhan

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. In Advances in Neural Information Processing Systems (NeurIPS), 2023

  48. [48]

    ReAct: Synergizing reasoning and acting in language models

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. InProceedings of the International Conference on Learning Representations (ICLR), 2023

  49. [49]

    ExpeL: LLM agents are experiential learners

    Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. ExpeL: LLM agents are experiential learners. InProceedings of the AAAI Conference on Artificial Intelligence, 2024

  50. [50]

    GPT-4V(ision) is a generalist web agent, if grounded

    Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. GPT-4V(ision) is a generalist web agent, if grounded. InProceedings of the International Conference on Machine Learning (ICML), 2024

  51. [51]

    MemoryBank: Enhancing large language models with long-term memory.Proceedings of the AAAI Conference on Artificial Intelligence, 2024

    Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang. MemoryBank: Enhancing large language models with long-term memory.Proceedings of the AAAI Conference on Artificial Intelligence, 2024

  52. [52]

    Write a Review

    Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. WebArena: A realistic web environment for building autonomous agents. InProceedings of the International Conference on Learning Representations (ICLR), 2024. 13 A Experimental setup and reproducibility This appe...