Pith. sign in

REVIEW 3 major objections 6 minor 84 references

Coding agent teams built on Claude Code's Agent Teams can be made to survive session restarts by keeping each agent's working state in files—so a closed terminal or a compaction no longer erases the team's knowledge.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 04:08 UTC pith:XUTNDPHR

load-bearing objection A genuinely useful design manual for a real problem—team-level agent persistence—but the central recovery claim is unmeasured and rests on an untested prompt-following assumption. the 3 major comments →

arxiv 2607.22917 v2 pith:XUTNDPHR submitted 2026-07-24 cs.AI cs.LGcs.MA

Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams

classification cs.AI cs.LGcs.MA
keywords multi-agent LLM systemsagent persistencecheckpointingClaude Codefilesystem-based operationsagent teamssoftware engineering agentsstate reconstruction
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Agent Team Work Zone (ATWZ) is a filesystem-based operations layer that wraps Claude Code's Agent Teams to give multi-agent coding teams a property they otherwise lack: continuity across sessions. Its central rule is 'files over context': everything a future agent might need—role, decisions, current task, commitments—is written to a per-agent directory called a workstation, instead of living only in the volatile chat context, which ends with the session and is thinned by compaction (the summarization that condenses a long conversation). With automatic checkpoints triggered by an idle hook, and a team registry that maps who is on the team, the paper claims a team can be reactivated with one command after a process ends, with each fresh teammate reconstructing its predecessor's state by reading the workstation. The paper also argues the file-based message flow turns agent-to-agent handoffs into inspectable, auditable documents and cuts the need for heavy prompt rewriting. It presents this as a design manual built by using the layer on itself, and explicitly reports no benchmarks, leaving empirical evaluation to future work.

Core claim

The paper's central claim is that the loss of state when an agent team's session ends—or when compaction thins a long conversation—is not an unavoidable platform limitation but a problem that can be engineered around without changing how the agents run. The design keeps agents ephemeral and reconstructs them from durable artifacts: each teammate owns a workstation holding a role file, working notes, a checkpoint snapshot, and commitments; the lead owns a registry of the team; and an idle hook forces periodic checkpoint writes. On reactivation, a fresh teammate's opening instruction is to read its own workstation and resume from that state. The paper is careful to call this reconstruction, no

What carries the argument

The workstation: a private directory per agent holding role definition, working notes, checkpoint, and commitment files, together with the team registry (TEAMMATE_INFO.json) and the TeammateIdle hook that ages timestamps and forces checkpoints. The workstation is the unit of durability; the registry is the input to reactivation; and the hook makes saving automatic. The whole design reduces to one rule: anything a future agent will need is written to a file outside the session's context window.

Load-bearing premise

The recovery mechanism assumes a freshly spawned agent will faithfully read and adopt the files in its workstation—role, checkpoint, commitments—and that the idle hook fires as documented; if a new teammate ignores its files, reactivation returns a hollow team.

What would settle it

A simple experiment would settle it: run a three-agent team on a small coding task, close the session mid-task, reactivate the team, and check whether the revived teammate's next message reflects the checkpointed task state. If agents frequently ignore their workstation files and restart from scratch, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Long-lived coding projects can sustain a standing team across multiple sessions; a human is no longer required to be the team's memory and scheduler.
  • Compaction stops eroding working detail, since the checkpoint is a separately written, authored snapshot plus an append-only journal of recent turns.
  • A team interrupted by a closed terminal, SSH drop, or restart can be rebuilt with a single command, provided the revived teammates read their workstations.
  • Agent-to-agent reports become durable, inspectable documents, giving the project an audit trail of decisions and changes that survives closed chats.
  • Upgrades are made safe by a strict ownership split: framework files are replaced, user workstations are never touched.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Outside the paper's own scope: the same file-first pattern could plausibly be adapted to other coding agents whose subagents are ephemeral, though the paper only notes the similarity without exploring it.
  • The load-bearing untested step is whether fresh agents reliably adopt their workstation files; a small empirical study measuring reactivation success under different interruption depths would settle it.
  • If the approach holds, the unit of persistence in multi-agent systems shifts from conversation history to something like 'agency'—role plus commitments—which other agent frameworks could adopt as a checkpoint design.
  • The paper flags as open work a consistency test between checkpoint writer and reader; automating that contract check would be a concrete next step a maintainer could add.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper presents ATWZ (Agent Team Work Zone), a filesystem-based persistence layer for Claude Code's Agent Teams. The design treats each agent as an ephemeral process whose durable state lives in per-agent 'workstation' directories: role files, checkpoints with journals, commitments, a team registry, and file-based meeting rooms and roundtables. The stated contributions are (1) automatic checkpointing so that an agent's knowledge can be recovered after context compaction, (2) one-command team reactivation after a session ends, (3) mitigation of 'agentic technical debt' via durable, inspectable decision records, and (4) reduced prompt-writing overhead through file-based handoffs and reports. The manuscript is explicitly framed as a developer-oriented design manual: it describes the architecture, covers both pre- and post-2.1.178 Claude Code APIs, catalogs failure modes encountered in dogfooding, and explicitly disclaims any benchmarks or user studies (§12). The central recovery mechanism is reconstruction from files: a freshly spawned agent is instructed to read its own workstation and resume from the checkpoint (§6.2).

Significance. If the central claims hold, ATWZ addresses a real and currently unserved gap: the irrecoverability of in-process teammates when a Claude Code session ends and the loss of working detail under compaction. The design is coherent and unusually candid: the file-over-context principle is applied consistently, the failure catalog in §10 is specific and plausible, the version-aware treatment of the Claude Code API change is careful, and the paper ships open-source code. The paper is also honest about its status as a design manual with no empirical evaluation. However, the significance is bounded by the fact that the headline claims—'knowledge can be recovered after compaction' and 'the team can be restored with a single command'—are capability assertions that rest entirely on an untested prompt-following assumption. The design's value is therefore prospective: it is a well-reasoned engineering proposal whose central mechanism has not been demonstrated to work. What would make the contribution credible is a minimal recovery-reliability measurement or a sharpened claim that explicitly labels these as untested design goals.

major comments (3)
  1. [Abstract and §6.2] The central claims—'knowledge can be recovered after compaction' and 'the team can be restored with a single command'—are capability assertions, yet the only success criterion presented is that the revived teammate 'sends a message back in the current session.' A message receipt does not establish that the agent adopted the checkpoint: a spawned agent that failed to read its workstation could still send a generic acknowledgment, and the manuscript provides no measurement of how often fresh agents actually follow the opening instruction to read their role file, checkpoint, and commitments. Section 12 disclaims evaluation, but the abstract does not carry that hedge. This is load-bearing: the entire reactivation flow reduces to this untested instruction-following reliability. Either run a minimal recovery study (spawn N revived agents from checkpoints and score whether they act on checkpoin
  2. [§6.4] Automatic checkpointing depends on the TeammateIdle hook firing as documented in the lead session. If the hook silently fails to fire on the target Claude Code versions, the 'periodically back up its working state' feature silently stops, and nothing in the design detects that absence: the release gate and liveness rule cover other failure modes but not hook non-firing. The manuscript should either provide evidence that the hook fires as documented on the versions targeted, or add a watchdog—for example, a checkpoint-freshness check at reactivation or a session-end checkpoint—so that a silent hook failure is observable and recoverable.
  3. [§4.5 and §10] The hardening narrative is central to the paper's method ('use it on itself, then fix what breaks'), and §10 presents a catalog of failure modes 'learned in use,' including two releases shipped without migration scripts and a Name-2 drift incident. However, no artifact evidence ties this catalog to actual runs: there are no logs, no commit history excerpts, and no incident records in the manuscript, so the claims that these failures occurred and were fixed durably are unfalsifiable from the paper alone. Since the repository is open-source, a short appendix pointing to specific commits or a maintained incident log would substantiate the dogfooding claim and let a reader verify the hardening history.
minor comments (6)
  1. [§7.1] Message filenames encode a 'minute-precise timestamp.' Two messages from the same author within the same minute would collide; consider second precision or a sequence suffix.
  2. [§5.2] The upgrade-safety guarantee relies on comment markers delimiting framework-owned blocks in README.md. It would help to state explicitly what happens if a user edits inside a marked block, since the boundary is then no longer mechanically enforceable.
  3. [§6.2] The sentence 'Disk files cannot confirm a particular reactivation, since a spawn that dies before messaging leaves every file unchanged; the receipt, not the artifact, is the evidence' is honest but highlights the weakness noted in the major comment: the symmetric statement—that a receipt cannot confirm reconstruction—should also appear here or in §12.
  4. [§11 and references] Several references are the author's own preprints and appear tangential to the persistence question (e.g., [27]–[30] on reasoning modes, [31] on protein language models, [40] on long-term memory probing, [50] on benchmark configuration). Consider trimming to those that directly frame continuity or memory, or adding a sentence explaining each citation's relevance.
  5. [Throughout] The marginal-footnote style (e.g., '1:In-context memory = the session's live context window...') is unusual and sometimes visually interrupts the narrative; conventional footnotes would improve readability.
  6. [§8.1] The model-assignment rationale ('mechanical polling on the cheapest capable model... adversarial work on the strongest') is plausible but unsubstantiated; a citation or a short explanation of the cost-quality tradeoff would strengthen it.

Circularity Check

0 steps flagged

No circularity found: the persistence-layer design is a self-contained construction; the untested prompt-following reliability assumption is an evaluation gap, not a circular reduction.

full rationale

The paper is a developer-oriented design manual rather than an empirical study, and I find no step in its derivation chain that reduces to its own inputs. The central mechanism—state written to per-agent workstation files, automatic checkpointing via the TeammateIdle hook, and reactivation by spawning fresh teammates that read their workstations—is a concrete file/hook/skill construction built on Claude Code's documented Agent Teams limitations. There are no fitted parameters, equations, or renamed empirical patterns. The 'files over context' principle is explicitly acknowledged as the same move Claude Code's subagent memory directory already makes (Section 2.1, note 2: 'the memory directory is, in fact, the same move this manual generalises'), which is an honest statement of lineage rather than a disguised equivalence. The main concern one could raise—that a revived teammate's knowledge is unverified because Section 6.2 defines success as the teammate sending a message back and Section 12 admits 'we have not measured how much [the mechanisms] help'—is an empirical gap and an untested prompt-following assumption, not circularity: nothing about the design defines recovery as message receipt by construction, and the paper explicitly frames recovery as 'reconstruction from durable files' (Section 3). The dogfooding note and the failure catalog in Section 10 are self-referential but are presented as anecdotal hardening history ('not a proof that the design is complete'), not as the evidence for the central claim. Self-citations in Section 11 (e.g., [27]-[30], [40], [50]) and Section A are contextual related-work references and are not load-bearing for the persistence layer. Accordingly the circularity score is 0.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 3 invented entities

No free parameters: the design involves no numerical fitting. The axioms are platform-behavior assumptions and the untested reconstruction assumption, which together form the load-bearing basis for recovery. Invented entities are software design artifacts, not physical postulates; they lack independent falsifiable evidence.

axioms (4)
  • domain assumption A fresh teammate that reads its workstation files will recover the prior teammate's working state.
    Central to reactivation (Section 6.2); unmeasured dependency on LLM instruction-following.
  • domain assumption Claude Code's hook system fires the TeammateIdle hook as documented and can block an idle to force a checkpoint.
    The checkpoint gate (Section 6.4) depends on platform hook semantics; docs may change.
  • domain assumption The documented limitations of Agent Teams (teammates not restored on resume) hold for the versions targeted.
    Cited from Anthropic docs [1]; if the platform later persists teammates natively, the layer's rationale weakens, though the design remains useful.
  • ad hoc to paper Recovery via reconstruction from files is an acceptable substitute for live-process preservation.
    The paper explicitly chooses reconstruction (Section 3) but does not evaluate whether reconstructed agents perform equivalently.
invented entities (3)
  • workstation no independent evidence
    purpose: Per-agent persistent file directory holding role, notes, TODO, and commitments.
    The paper introduces it as the core primitive; no external falsifiable signal.
  • checkpoint file (with journal) no independent evidence
    purpose: Periodic snapshot of agent state to survive compaction and restart.
    Introduced by the paper as the basis for recovery; no external benchmark verifies its efficacy.
  • meeting room / roundtable no independent evidence
    purpose: File-based asynchronous messaging between agents.
    Software artifact introduced by the paper; no independent handle outside the design.

pith-pipeline@v1.3.0-alltime-deepseek · 20668 in / 8893 out tokens · 83047 ms · 2026-08-01T04:08:00.327901+00:00 · methodology

0 comments
read the original abstract

Large Language Model (LLM) agents have significantly improved coding and programming workflows. Claude Code, in particular, is one of the most powerful LLM coding agents and is capable of conducting complex coding tasks. However, several drawbacks can undermine long-term agentic workflows. (1) Irrecoverable agent teams: The Agent Teams feature is powerful, but the working state accumulated by each teammate is lost and cannot be resumed once the process stops, for example, when a terminal is closed. (2) Compaction erodes working detail: Compaction condenses the conversation into a summary, causing an agent's working details to become vague. (3) Agentic "technical debt": Over time, a user's decisions and the agents' operations become trapped in compacted old chats, making the project increasingly difficult to maintain and review. (4) Heavy prompt writing: Assigning or handing off tasks requires users to repeatedly write long prompts to achieve the expected agentic performance. We propose ATWZ (Agent Team Work Zone), a filesystem-based operations layer built around Claude Code's native Agent Teams that addresses these problems. Its central design principle is to treat each agent and teammate as a human employee and preserve their important working state in files stored in a dedicated directory called a "workstation," together with the skills, hooks, and scripts that use and maintain these files. With ATWZ, an agent team can periodically back up its working state, allowing an agent's knowledge to be recovered after compaction. After a process ends, the team can be restored with a single command. These features also substantially mitigate the agentic "technical debt" described above. Moreover, within ATWZ, agent "employees" can send documents to one another, greatly reducing the effort required to write prompts.

Figures

Figures reproduced from arXiv: 2607.22917 by Shouren Wang.

Figure 1
Figure 1. Figure 1: Agent Team Work Zone: a file￾based operations layer over Claude Code. Agents remain ephemeral (middle); the durable workspace (bottom) lets an in￾terrupted team be rebuilt. be resumed (/resume), which restores the lead, but—as Claude Code’s documentation makes explicit—not its in-process teammates, which must be rebuilt [1]. Compaction—the summarization that keeps a long conversation within the context win… view at source ↗
Figure 2
Figure 2. Figure 2: What a session end destroys (drawn for pre-2.1.178 releases): the team￾mate processes, the team’s in-memory co￾ordination state, and any working con￾text held only in a context window. The lead’s session can be /resumed; on pre￾2.1.178 releases the on-disk team direc￾tory lingers as a stale ghost, whereas on CC ≥ 2.1.178 the session-scoped team config is cleaned up on exit. Either way the live team does no… view at source ↗
Figure 3
Figure 3. Figure 3: Keeping state in files rather than in the volatile, in-context memory; and how a per-agent workstation note differs from Claude Code’s relevance￾recalled memory/. volatile—it ends with the session and is thinned by compaction—and it cannot be inspected from outside. For teammates this loss is permanent (they cannot be resumed); a lead may recover its own context with /resume. The design treats it as workin… view at source ↗
Figure 4
Figure 4. Figure 4: Reactivating a team on CC ≥ 2.1.178: read the team registry, spawn each teammate by <slug>-<role> name into the automatic session team, and let each teammate recover from its own checkpoint. No registration-clearing step is needed. Prior versions (pre-2.1.178). On earlier releases ( [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Reactivating a team on prior versions: clear the stale registration, read the registry, and respawn each teammate so it recovers from its own checkpoint. rebuild; and a receipt expires—a teammate that answered several turns ago may be gone now, so every judgment needs a fresh check. A lighter third state, benched, lets a teammate be taken offline to free a slot with its workstation preserved, then woken by… view at source ↗
Figure 6
Figure 6. Figure 6: The TeammateIdle checkpoint gate: an age-based hook blocks a team￾mate’s idle and forces a checkpoint, then self-brakes once the timestamp refreshes. bugs. Below a threshold (15 minutes) idle proceeds; at or above it, the hook returns a non-zero exit that blocks the idle and feeds the teammate a reminder to checkpoint now. Writing the checkpoint refreshes the file’s timestamp, so the next idle passes and t… view at source ↗
Figure 7
Figure 7. Figure 7: File-based messaging: an agent writes a message addressed to another; the recipient reads it with /check-inbox, does the task and marks it resolved; the issuer later archives it. 7.1. Two channels, one protocol Communication runs over the two file-based spaces from Section 5: a top-level meeting room that any flat agent or team lead can use, and a per￾team roundtable for a lead and its teammates. Both run … view at source ↗
Figure 8
Figure 8. Figure 8: A team’s lifecycle (baseline, shared across versions): spawn, an active loop of work and automatic checkpoints, bench or offboard, and reactivation after a session ends; plus the role roster of a lead and specialists (role names in the figure are illustrative). See [PITH_FULL_IMAGE:figures/full_fig_p016_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Lifecycle on CC ≥ 2.1.178: <slug>-<role> is the naming conven￾tion the skills apply, letting the idle hook derive each teammate’s workstation path from the slug prefix without ambiguity; on prior versions the hook read it from the idle event’s team name instead [PITH_FULL_IMAGE:figures/full_fig_p017_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

84 extracted references · 1 canonical work pages

  1. [1]

    Claude Code documentation

    Agent Teams. Claude Code documentation. Anthropic.url:https: //code.claude.com/docs/en/agent-teams(visitedon07/21/2026)

  2. [2]

    Subagents.Codexdocumentation;wordingasofJune2026archived at https://web.archive.org/web/20260613202301/https:// developers.openai.com/codex/subagents.OpenAI.url: https: //learn.chatgpt.com/docs/agent-configuration/subagents (visited on 07/22/2026)

  3. [3]

    Claude Code documentation

    Subagents. Claude Code documentation. Anthropic.url:https:// code.claude.com/docs/en/sub-agents (visitedon07/22/2026)

  4. [4]

    Claude Code documentation

    Run agents in parallel. Claude Code documentation. Anthropic. url: https : / / code . claude . com / docs / en / agents(visited on 07/22/2026)

  5. [5]

    com/anthropics/claude-code/issues/7317(visitedon07/21/2026)

    Resumable subagent sessions (feature request).url:https://github. com/anthropics/claude-code/issues/7317(visitedon07/21/2026)

  6. [6]

    Codex documentation

    Codex Cloud. Codex documentation. OpenAI.url:https://learn. chatgpt.com/docs/cloud(visited on 07/22/2026)

  7. [7]

    Claude Code documentation

    Settings. Claude Code documentation. Anthropic.url: https : //code.claude.com/docs/en/settings(visited on 07/22/2026)

  8. [8]

    Claude Code documentation

    Memory. Claude Code documentation. Anthropic.url: https : //code.claude.com/docs/en/memory(visited on 07/22/2026)

  9. [9]

    Anthropic

    Introducing Claude Opus 4.8. Anthropic. 2026.url:https://www. anthropic.com/news/claude-opus-4-8

  10. [10]

    OpenAI.OpenAI GPT-5 System Card. 2025. arXiv: 2601 . 03267 [cs.CL]

  11. [11]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey,AbhishekKadian,AhmadAl-Dahle,AieshaLetman,Akhil Mathur, Alan Schelten, Alex Vaughan, et al.The Llama 3 Herd of Models. 2024. arXiv:2407.21783 [cs.AI]

  12. [12]

    AutoGen: Enabling Next-Gen LLM Applications via Multi- Agent Conversation

    Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, et al. “AutoGen: Enabling Next-Gen LLM Applications via Multi- Agent Conversation”. In:Conference on Language Modeling (COLM)

  13. [13]

    MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework

    Sirui Hong, Mingchen Zhuge, Jiaqi Chen, Xiawu Zheng, Yuheng Cheng,CeyaoZhang,JinlinWang,ZiliWang,StevenKaShingYau, Zijuan Lin, et al. “MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework”. In:International Conference on Learning Representations (ICLR). 2024. arXiv:2308.00352 [cs.AI]

  14. [14]

    CAMEL: Communicative Agents for “Mind

    Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. “CAMEL: Communicative Agents for “Mind” Exploration of Large Language Model Society”. In:Advances in Neural Information Processing Systems (NeurIPS)

  15. [15]

    ChatDev: Communicative Agents for Software Development

    Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, et al. “ChatDev: Communicative Agents for Software Development”. In: Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL). 2024.doi: 10.18653/v1/2024.acl- long.810 . arXiv:2307.07924 [cs.SE]

  16. [16]

    AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors

    WeizeChen,Yusheng Su,JingweiZuo, ChengYang,ChenfeiYuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Chen Qian, et al. “AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors”. In:International Conference on Learning Representations (ICLR). 2024. arXiv:2308.10848 [cs.CL]

  17. [17]

    Improving Factuality and Reasoning in Language Models through Multiagent Debate

    Yilun Du, ShuangLi, Antonio Torralba,Joshua B.Tenenbaum, and Igor Mordatch. “Improving Factuality and Reasoning in Language Models through Multiagent Debate”. In:International Conference on Machine Learning (ICML). 2024. arXiv:2305.14325 [cs.CL]

  18. [18]

    LangGraph documentation

    Persistence. LangGraph documentation. LangChain.url:https: //docs.langchain.com/oss/python/langgraph/persistence (visited on 07/21/2026)

  19. [19]

    AutoGen AgentChat documentation, stable chan- nel

    Managing State. AutoGen AgentChat documentation, stable chan- nel. Microsoft.url: https://microsoft.github.io/autogen/ stable/user-guide/agentchat-user-guide/tutorial/state. html(visited on 07/22/2026)

  20. [20]

    Towards Reasoning in Large Language Models: A Survey

    Jie Huang and Kevin Chen-Chuan Chang. “Towards Reasoning in Large Language Models: A Survey”. In:Findings of the Association for Computational Linguistics: ACL 2023. 2023.doi:10.18653/v1/ 2023.findings-acl.67. arXiv:2212.10403 [cs.CL]

  21. [21]

    Chain-of- Thought Prompting Elicits Reasoning in Large Language Models

    JasonWei,XuezhiWang,DaleSchuurmans,MaartenBosma,Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. “Chain-of- Thought Prompting Elicits Reasoning in Large Language Models”. In:Advances in Neural Information Processing Systems (NeurIPS)

  22. [22]

    Self- Consistency Improves Chain of Thought Reasoning in Language Models

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. “Self- Consistency Improves Chain of Thought Reasoning in Language Models”. In:International Conference on Learning Representations (ICLR). 2023. arXiv:2203.11171 [cs.CL]

  23. [23]

    Tree of Thoughts: Deliberate Problem Solving with Large Language Models

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. “Tree of Thoughts: Deliberate Problem Solving with Large Language Models”. In: Advances in Neural Information Processing Systems (NeurIPS). 2023. arXiv:2305.10601 [cs.CL]

  24. [24]

    ReAct: Synergizing Reasoning and ActinginLanguageModels

    ShunyuYao,JeffreyZhao,DianYu,NanDu,IzhakShafran,Karthik Narasimhan, and Yuan Cao. “ReAct: Synergizing Reasoning and ActinginLanguageModels”.In:InternationalConferenceonLearning Representations (ICLR). 2023. arXiv:2210.03629 [cs.CL]

  25. [25]

    Reflexion: Language AgentswithVerbalReinforcementLearning

    NoahShinn,FedericoCassano,EdwardBerman,AshwinGopinath, Karthik Narasimhan, and Shunyu Yao. “Reflexion: Language AgentswithVerbalReinforcementLearning”.In:AdvancesinNeural Information Processing Systems (NeurIPS). 2023. arXiv:2303.11366 [cs.AI]

  26. [26]

    Self-Refine: Iterative Refinement with Self-Feedback

    Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prab- humoye, Yiming Yang, et al. “Self-Refine: Iterative Refinement with Self-Feedback”. In:Advances in Neural Information Processing Systems (NeurIPS). 2023. arXiv:2303.17651 [cs.CL]

  27. [27]

    arXiv:2510

    Shouren Wang, Wang Yang, Xianxuan Long, Qifan Wang, Vipin Chaudhary, and Xiaotian Han.Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?2025. arXiv:2510. 12680 [cs.CL]. 26

  28. [28]

    Shouren Wang, Wang Yang, Chuang Ma, Debargha Ganguly, Vikash Singh, Chaoda Song, Xinpeng Li, Xianxuan Long, Vipin Chaudhary, and Xiaotian Han.Path-Lock Expert: Separating Reason- ing Mode in Hybrid Thinking via Architecture-Level Separation. 2026. arXiv:2604.27201 [cs.CL]

  29. [29]

    Mid- Think: Training-Free Intermediate-Budget Reasoning via Token- Level Triggers

    VanYang,ShourenWang,DebarghaGanguly,XinpengLi,Chaoda Song, Vikash Singh, Vipin Chaudhary, and Xiaotian Han. “Mid- Think: Training-Free Intermediate-Budget Reasoning via Token- Level Triggers”. In:Findings of the Association for Computational Linguistics: ACL 2026. Ed. by Maria Liakata, Viviane P. Moreira, Jia- junZhang,andDavidJurgens.SanDiego,California,...

  30. [30]

    Wang Yang, Shouren Wang, Chaoda Song, Chuang Ma, Xinpeng Li, Nengbo Wang, Kaixiong Zhou, Vipin Chaudhary, and Xiaotian Han.When Domains Interact: Asymmetric and Order-Sensitive Cross- Domain Effects in Reinforcement Learning for Reasoning. 2026. arXiv: 2602.01365 [cs.LG]

  31. [31]

    A Survey on Efficient Protein Language Models

    Shouren Wang, Debargha Ganguly, Vinooth Rao Kulkarni, Van Yang, Zhuoran Qiao, Daniel Blankenberg, Vipin Chaudhary, and Xiaotian Han. “A Survey on Efficient Protein Language Models”. In:Transactions on Machine Learning Research(2026). Survey Certi- fication.issn: 2835-8856.url: https://openreview.net/forum? id=PTReuOwsXz

  32. [32]

    Toolformer: Language Models Can Teach Themselves to Use Tools

    Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. “Toolformer: Language Models Can Teach Themselves to Use Tools”. In:Advances in Neural Information Processing Systems (NeurIPS). 2023. arXiv:2302.04761 [cs.CL]

  33. [33]

    Gorilla:LargeLanguageModelConnectedwithMassiveAPIs

    ShishirG.Patil,TianjunZhang,XinWang,andJosephE.Gonzalez. “Gorilla:LargeLanguageModelConnectedwithMassiveAPIs”.In: Advances in Neural Information Processing Systems (NeurIPS). 2024. arXiv:2305.15334 [cs.CL]

  34. [34]

    HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

    Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. “HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face”. In:Advances in Neural Information Processing Systems (NeurIPS). 2023. arXiv:2303.17580 [cs.CL]

  35. [35]

    ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

    Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et al. “ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs”.In:InternationalConferenceonLearningRepresentations(ICLR)

  36. [36]

    Lost in the Middle: How Language Models Use Long Contexts

    Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. “Lost in the Middle: How Language Models Use Long Contexts”. In:Transactions of the Association for Computational Linguistics (TACL)(2024).doi: 10.1162/tacl_a_00638. arXiv:2307.03172 [cs.CL]

  37. [37]

    MemGPT: Towards LLMs as Operating Systems

    Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. “MemGPT: Towards LLMs as Operating Systems”. In:Conference on Language Modeling (COLM). 2024. arXiv:2310.08560 [cs.AI]. 27

  38. [38]

    GenerativeAgents: Interactive Simulacra of Human Behavior

    Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris,PercyLiang,andMichaelS.Bernstein.“GenerativeAgents: Interactive Simulacra of Human Behavior”. In:ACM Symposium on User Interface Software and Technology (UIST). 2023. arXiv:2304. 03442 [cs.HC]

  39. [39]

    arXiv:2307.16789 [cs.AI]

  40. [40]

    Xianxuan Long, Zhikai Chen, Shenglai Zeng, Shouren Wang, Kai Guo, and Jiliang Tang.MemTrace: Probing What Final Accuracy Misses in Long-Term Memory. 2026. arXiv:2606.17328 [cs.CL]

  41. [41]

    Voyager: An Open-Ended Embodied Agent with Large Language Models

    Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. “Voyager: An Open-Ended Embodied Agent with Large Language Models”. In:Transactions on Machine Learning Research (TMLR)(2024). arXiv: 2305.16291 [cs.AI]

  42. [42]

    SSRN preprint 6746498

    Van Yang, Chaoda Song, Xinpeng Li, Shouren Wang, Nengbo Wang, Yanyan Zhang, Chuang Ma, Debargha Ganguly, Vikash Singh, Shuai Xu, et al.A Survey on Agent Skills for LLMs: A Lifecycle Perspective from Construction to Ecosystems. SSRN preprint 6746498. 2026.url: https : / / papers . ssrn . com / sol3 / papers . cfm ? abstract_id=6746498

  43. [43]

    CognitiveArchitecturesforLanguageAgents

    Theodore R. Sumers, Shunyu Yao, Karthik Narasimhan, and ThomasL.Griffiths.“CognitiveArchitecturesforLanguageAgents”. In:Transactions on Machine Learning Research (TMLR)(2024). arXiv: 2309.02427 [cs.AI]

  44. [44]

    SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineer- ing

    John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. “SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineer- ing”.In:AdvancesinNeuralInformationProcessingSystems(NeurIPS)

  45. [45]

    OpenHands: An Open Platform for AI Software Developers as Generalist Agents

    Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, et al. “OpenHands: An Open Platform for AI Software Developers as Generalist Agents”. In:International Conference on Learning Representations (ICLR). 2025. arXiv:2407. 16741 [cs.SE]

  46. [46]

    ExecutableCodeActionsElicitBetterLLM Agents

    XingyaoWang,YangyiChen,LifanYuan,YizheZhang,YunzhuLi, HaoPeng,andHengJi.“ExecutableCodeActionsElicitBetterLLM Agents”. In:International Conference on Machine Learning (ICML)

  47. [47]

    SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

    Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. “SWE-bench: Can Language Models Resolve Real-World GitHub Issues?” In: International Conference on Learning Representations (ICLR). 2024. arXiv:2310.06770 [cs.SE]

  48. [48]

    Agentless: Demystifying LLM-based Software Engineer- ing Agents

    Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang. “Agentless: Demystifying LLM-based Software Engineer- ing Agents”. In:Proceedings of the ACM on Software Engineering (FSE)(2025). arXiv:2407.01489 [cs.SE]. 28

  49. [49]

    arXiv:2405.15793 [cs.SE]

  50. [50]

    WangYang,ChaodaSong,XinpengLi,DebarghaGanguly,Chuang Ma, Shouren Wang, Zhihao Dou, Yuli Zhou, Vipin Chaudhary, and Xiaotian Han.AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments. 2026. arXiv:2604.06111 [cs.AI]

  51. [51]

    AgentBench: Evaluating LLMs as Agents

    Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al. “AgentBench: Evaluating LLMs as Agents”. In:International Confer- ence on Learning Representations (ICLR). 2024. arXiv:2308.03688 [cs.AI]

  52. [52]

    arXiv:2402.01030 [cs.CL]

  53. [53]

    AutoCodeRover: Autonomous Program Improvement

    Yuntong Zhang, Haifeng Ruan, Zhiyu Fan, and Abhik Roychoud- hury. “AutoCodeRover: Autonomous Program Improvement”. In: ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA). 2024.doi: 10 . 1145 / 3650212 . 3680384. arXiv: 2404.05427 [cs.SE]

  54. [54]

    OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

    Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Jing Hua Toh, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. “OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments”. In:Advances in Neural Information Processing Systems (NeurIPS)

  55. [55]

    Large Language Model-Based Agents for Software Engineering: A Survey

    JunweiLiu,KaixinWang,YixuanChen,XinPeng,ZhenpengChen, Lingming Zhang, and Yiling Lou. “Large Language Model-Based Agents for Software Engineering: A Survey”. In:ACM Transactions on Software Engineering and Methodology (TOSEM)(2026). arXiv: 2409.02977 [cs.SE]

  56. [56]

    Large Language Model based Multi-Agents: A Survey of Progress and Challenges

    Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V. Chawla, Olaf Wiest, and Xiangliang Zhang. “Large Language Model based Multi-Agents: A Survey of Progress and Challenges”.In:InternationalJointConferenceonArtificialIntelligence (IJCAI). 2024. arXiv:2402.01680 [cs.CL]

  57. [57]

    A Survey on Evaluation of Large Language Models

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, KaijieZhu,HaoChen,XiaoyuanYi,CunxiangWang,YidongWang, et al. “A Survey on Evaluation of Large Language Models”. In: ACM Transactions on Intelligent Systems and Technology15.3 (2024). doi:10.1145/3641289. arXiv:2307.03109 [cs.CL]

  58. [58]

    WebArena: A Realistic Web Environment for Build- ing Autonomous Agents

    Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al. “WebArena: A Realistic Web Environment for Build- ing Autonomous Agents”. In:International Conference on Learning Representations (ICLR). 2024. arXiv:2307.13854 [cs.AI]

  59. [59]

    GAIA: a benchmark for Gen- eral AI Assistants

    Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. “GAIA: a benchmark for Gen- eral AI Assistants”. In:International Conference on Learning Repre- sentations (ICLR). 2024. arXiv:2311.12983 [cs.CL]

  60. [60]

    Emergent Abilities of Large Language Models

    Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, SebastianBorgeaud,DaniYogatama,MaartenBosma,DennyZhou, Donald Metzler, et al. “Emergent Abilities of Large Language Models”. In:Transactions on Machine Learning Research (TMLR) (2022). arXiv:2206.07682 [cs.CL]

  61. [61]

    arXiv:2404.07972 [cs.AI]

  62. [62]

    A Survey on Large Language Model based Autonomous Agents

    Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, et al. “A Survey on Large Language Model based Autonomous Agents”. In:FrontiersofComputerScience(2024).arXiv: 2308.11432 [cs.AI]

  63. [63]

    Finetuned Language Models Are Zero-Shot Learners

    Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. “Finetuned Language Models Are Zero-Shot Learners”. In: International Conference on Learning Representations (ICLR). 2022. arXiv:2109.01652 [cs.CL]

  64. [64]

    Deep Reinforcement Learning from Human Preferences

    Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. “Deep Reinforcement Learning from Human Preferences”. In:Advances in Neural Information Processing Systems (NeurIPS). 2017. arXiv:1706.03741 [stat.ML]

  65. [65]

    Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei.Scaling Laws for Neural Language Models

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei.Scaling Laws for Neural Language Models

  66. [66]

    Direct Preference Optimization: Your Language Model is Secretly a Reward Model

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. “Direct Preference Optimization: Your Language Model is Secretly a Reward Model”. In:Advances in Neural Information Processing Systems (NeurIPS)

  67. [67]

    Training Compute-OptimalLargeLanguageModels

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, LisaAnneHendricks,JohannesWelbl,AidanClark,etal.“Training Compute-OptimalLargeLanguageModels”.In:AdvancesinNeural Information Processing Systems (NeurIPS). 2022. arXiv:2203.15556 [cs.CL]. 29

  68. [68]

    FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

    Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness”. In:Advances in Neural Information Processing Systems (NeurIPS). 2022. arXiv:2205.14135 [cs.LG]

  69. [69]

    LanguageModelsare Few-Shot Learners

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam,GirishSastry,AmandaAskell,etal.“LanguageModelsare Few-Shot Learners”. In:Advances in Neural Information Processing Systems (NeurIPS). 2020. arXiv:2005.14165 [cs.CL]

  70. [70]

    ASurvey on In-context Learning

    Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, HemingXia,JingjingXu,ZhiyongWu,TianyuLiu,etal.“ASurvey on In-context Learning”. In:Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2024. arXiv:2301.00234 [cs.CL]

  71. [71]

    Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S

    Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg,AntoineBosselut,EmmaBrunskill,etal.OntheOpportunities and Risks of Foundation Models. 2021. arXiv:2108.07258 [cs.LG]

  72. [72]

    A Survey of Large Language Models

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. “A Survey of Large Language Models”. In:Frontiers of Computer Science(2026). arXiv:2303.18223 [cs.CL]

  73. [73]

    Training Language Models to Follow Instructions with Human Feedback

    Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wain- wright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Kata- rina Slama, Alex Ray, et al. “Training Language Models to Follow Instructions with Human Feedback”. In:Advances in Neural In- formation Processing Systems (NeurIPS). 2022. arXiv:2203.02155 [cs.CL]

  74. [74]

    ToolLearningwithFoundationModels

    Yujia Qin, Shengding Hu, Yankai Lin, Weize Chen, Ning Ding, Ganqu Cui, Zheni Zeng, Yufei Huang, Chaojun Xiao, Chi Han, etal.“ToolLearningwithFoundationModels”.In:ACMComputing Surveys(2025). arXiv:2304.08354 [cs.CL]

  75. [75]

    arXiv:2305.18290 [cs.LG]

  76. [76]

    EfficientMemoryManagementforLargeLanguageModel Serving with PagedAttention

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica.“EfficientMemoryManagementforLargeLanguageModel Serving with PagedAttention”. In:ACM Symposium on Operating Systems Principles (SOSP). 2023. arXiv:2309.06180 [cs.LG]

  77. [78]

    GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

    Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alis- tarh. “GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers”. In:International Conference on Learning Representations (ICLR). 2023. arXiv:2210.17323 [cs.LG]

  78. [79]

    Switch Trans- formers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

    William Fedus, Barret Zoph, and Noam Shazeer. “Switch Trans- formers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity”. In:Journal of Machine Learning Research (JMLR) (2022). arXiv:2101.03961 [cs.LG]. 30

  79. [82]

    Tula Masterman, Sandi Besen, Mason Sawtell, and Alex Chao.The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey. 2024. arXiv:2404.11584 [cs.AI]

  80. [84]

    A Survey on the Memory Mechanism of Large Language Model based Agents

    Zeyu Zhang, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Quanyu Dai, Jieming Zhu, Zhenhua Dong, and Ji-Rong Wen. “A Survey on the Memory Mechanism of Large Language Model based Agents”. In:ACM Transactions on Information Systems (TOIS)(2025).doi: 10.1145/3748302. arXiv:2404.13501 [cs.AI]. 31

Showing first 80 references.