Pith. sign in

REVIEW 3 major objections 3 minor 76 references

This review argues that modern language agents have already migrated most individual control mechanisms from classical cognitive architectures, and that the open research target is five residual couplings — runtime invariants connecting sta

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-07-31 23:26 UTC pith:73CBU76W

load-bearing objection A genuinely reusable synthesis framework whose headline map of five residual bundles is a provisional research agenda, not an established finding — and the paper mostly knows this. the 3 major comments →

arxiv 2607.23942 v1 pith:73CBU76W submitted 2026-07-27 cs.AI

From Cognitive Architectures to Language Agents: A Mechanism-Level Review of Lineage, Convergence, and Migration Gaps

classification cs.AI
keywords mechanism-level reviewcognitive architectureslanguage agentsmigration depthruntime invariantsresidual couplingscontrol semanticsevidence coding
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper's bounded synthesis claim: modern language agents have operationalized most of the cognitive substrate — adaptive memory policies, typed failure recovery, dynamic team selection, workflow search, skill induction, resource scheduling, and uncertainty-conditioned action — but often by independent convergence rather than documented inheritance. What remains are not missing mechanisms but missing couplings: the paper derives, via a seven-field mechanism representation, an evidence-depth coding, and explicit merge/split rules, exactly five residual bundles worth testing as composable runtime invariants, plus one candidate it closes to an existing system. A sympathetic reader should care because this reframes cognitive-architecture research as a source of precise, falsifiable control hypotheses rather than feature lists.

Core claim

On the paper's own terms, the discovery is uneven migration depth: for each of ten historical cognitive architectures, the strongest reviewed modern system achieves D3 or D4 on the individual control edge (autonomous trigger plus outcome-dependent learning), yet no reviewed system composes the historical couplings — and the one composition candidate that did emerge, skill governance, is already implemented by a system combining calibrated routing, typed verification, bounded repair, and replanning/fallback. The five residual bundles are therefore not feature requests; they are specific state-trigger-authority-transition-learning packets, each with insertion points, closest baselines, and fal

What carries the argument

The central object is the mechanism tuple K = ⟨S, C, T, P, F, L, R⟩ — state substrate, control locus, transition trigger, persistence boundary, failure semantics, learning operator, and resource/governance policy — used to reconstruct every historical and modern system, plus two independent coding axes: E1–E4 for whether a correspondence is documented lineage, structural migration, functional approximation, or convergence, and D0–D4 for how deeply the mapped mechanism is implemented, from a conceptual label to an experience-adapted control law. Residual extraction merges atomic missing invariants only when they share an insertion boundary and one supplies state or authority for the other; cl

Load-bearing premise

The partition into five residual bundles rests on a single coder's provisional classifications of every D3/D4 claim; the paper states that independent inter-rater reliability has not been established, and any reclassification could open or close a bundle.

What would settle it

Have a second coder, blind to the provisional codes, re-classify the D3/D4 evidence packets the paper provides: if any reviewed system already composes one of the five missing couplings, that bundle closes, just as the skill-governance candidate closed. Alternatively, run the paper's own preregistered protocol for the memory bundle against a recency-based baseline; if the simpler selector is non-inferior on recall and overhead, the bundle dissolves empirically.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Future agent evaluations should measure internal transition integrity — state integrity, interruption behavior, repeated failure, drift, transfer, and cost — alongside task success, not just final outputs.
  • Adaptive memory research must now be compared against systems that already learn memory operations and admission under value-per-byte governance, so the incremental contribution is the activation–latency–utility coupling, not adaptive admission itself.
  • Skill governance as a composed chain is already implemented by a reviewed system combining calibrated routing, typed verification, bounded repair, and replanning/fallback; any proposed extension must use that system as the mandatory baseline.
  • Each residual bundle yields a preregistered intervention with a runtime insertion point, an added control law, a closest baseline, and a primary falsifier — not an open-ended architecture proposal.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the governor semantics is right, it gives a cheap diagnostic: any control decision left to a prompt rather than an explicit state-and-trigger law is likely still at D2–D3, so migration depth can be estimated from interface design alone.
  • The pattern of independent convergence suggests a prediction: as these bundles are implemented, agent runtimes will increasingly resemble the classical architectures' couplings — a convergence that would validate the historical link without any documented lineage.
  • The closed skill-governance case implies a methodological extension: every residual bundle should be periodically re-screened against closest baselines as new systems appear, so some of the five bundles may already be closed or narrowed by systems published after this review's cutoff.
  • A testable extension the paper does not pursue is applying the same E/D coding to commercial agent platforms with persistent memory, interruption, and approval hooks, whose released interfaces may compose parts of the commitment and resource bundles faster than research prototypes.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper presents a mechanism-level review connecting ten historical cognitive architectures, eight language-agent runtime families, and forty-two mechanism-focused modern systems. Each mechanism is reconstructed via a seven-field tuple (state, control, transition, persistence, failure, learning, resource), and correspondences are coded on two independent axes: evidence relation (E1–E4) and migration depth (D0–D4). The central finding is that modern agents have operationalized many individual control mechanisms, while five residual control bundles remain open: activation–latency–utility (B1), typed impasse–substate–compilation (B2), content competition–workspace–broadcast learning (B3), intention–reconsideration–method authority (B4), and uncertainty–resources–interruption–stopping (B6). A sixth candidate, B5 (skill governance), is reported as closed by GraSP. The paper also contributes a catalog, an auditable coding framework, and preregistration-ready intervention protocols. It explicitly disclaims independent inter-rater reliability, systematic recall estimates, and completed experiments, and repeatedly states that the five-bundle partition is conditional on the current provisional codes.

Significance. If the residual-bundle claim holds, this is a substantial contribution: it moves the comparison of cognitive architectures and language agents from feature labels to control semantics and provides a falsifiable, corpus-relative research agenda. The paper's methodological strengths are real and should be credited: E/D codes are separated, coding rules are operationalized in Table II, runtime evidence is version-pinned in Appendix E, query templates are documented in Appendix F, and the intervention protocols in Appendix D specify ablations and rejection conditions. The closed B5 case is a useful demonstration that the framework can eliminate, not merely generate, research gaps. The main risk is not internal inconsistency but the gap between the headline 'five residual bundles remain' and the evidence base: the codes are provisional and single-coder, and the absence claims are bounded by an unmeasured search recall.

major comments (3)
  1. [§II-E, Appendix C, Abstract] The central result 'Five residual bundles remain' is generated from the E/D codes, but Section II-E states the manuscript 'does not claim independent inter-rater reliability' and Appendix C says scores 'must remain provisional until disagreements are recorded and adjudicated.' Every D3/D4 row in Table IV and Fig. 6 is single-coder. A shift in one D3/D4 grading can close or open a bundle, as the paper itself notes for GraSP/B5. Because the abstract and Section VI present the five-bundle partition as the headline outcome rather than as a provisional output of one coder, this is not a peripheral validity caveat. Please either complete the promised second-author recoding and report disagreements, or rephrase the headline and Section VI throughout as conditional on the current provisional codes.
  2. [§VI-B, §VI-C(e), Appendix C R7e] The closure of B5 rests entirely on one source, GraSP [63], described in Appendix C/O as combining calibrated multi-skill routing, typed DAG verification, bounded repair, and replanning/ReAct fallback. The review's own protocol (§II-C) requires official implementation evidence when a mechanism depends on executable behavior, and Appendix E provides version-pinned repository evidence for runtime families but not for GraSP. Without inspecting the GraSP release or otherwise verifying the four claimed edges against the paper's artifacts, the statement that 'B5 is therefore an implemented E3/D4 convergence case' is a single-source inference. Since B5's closure changes the residual count from six to five, this point is load-bearing. To fix: inspect and pin the GraSP implementation, or mark B5 as 'reported but not independently verified.'
  3. [§II-B, Appendix F] The paper states that result counts were not frozen and that no search-recall estimate is made. The five residual bundles are absence claims: they assert that no reviewed system composes the edges. Without a recall estimate for the six query blocks in Table XIV, the abstract's 'Five residual bundles remain' overstates what the search can certify relative to the broader literature. The paper already says it makes no PRISMA completeness claim, so the fix is largely presentation-level if the claims are consistently limited to 'within the frozen corpus and the documented search.' But the abstract and Section VI-C state the bundles unconditionally. Please align the wording throughout with the actual evidence boundary.
minor comments (3)
  1. [Throughout (e.g., §IV-B, Fig. 7)] The name 'Voyager' appears as 'V oyager' in several places, apparently a rendering artifact. Please fix.
  2. [Table I] The row 'This review' marks 'Primary' in all five columns. Consider a heading or note clarifying that these labels are self-assessments of analytical focus, not independent evaluations.
  3. [Appendix F] The paper says a dated research log supplies replayable query templates and record-level dispositions, but Appendix F only summarizes the blocks. If space permits, include the template list or a stable link, since the auditability claim depends on access to those templates.

Circularity Check

0 steps flagged

No significant circularity: the residual bundles are generated by explicit, externally grounded coding rules over a frozen corpus; disclosed provisionality is a validity limit, not a circular step.

full rationale

This is a qualitative mapping review, not a derivation with fitted parameters. The five residual bundles are produced by explicit rules applied to externally published systems: atomic residuals are merged only when they share a runtime insertion boundary and one atom supplies state or authority for another, and split when triggers, baselines, or falsifiers are independent (Sec. II-D, Table V). The paper does not fit a parameter and then predict the same quantity; the E1-E4/D0-D4 codes are relation and depth classifications of external evidence, and the residual is defined as the coupling absent after the strongest reviewed D3/D4 precedent is considered. GraSP's closure of B5 is an external-corpus finding from a different author group, not a self-citation. The limitations the paper itself states — "The present manuscript does not claim independent inter-rater reliability" (Sec. II-E), "scores must remain provisional until disagreements are recorded and adjudicated" (Appendix C), and the absence of a search-recall estimate (Appendix F) — affect confidence and completeness of an absence claim, but an unverified or provisional evidence base is not circularity. The paper explicitly says "This partition is auditable conditional on the current provisional codes, not a claim that five is the unique partition of the literature," which undercuts any reading that the five bundles are forced by the framework's own definitions. No equation reduces to its own input and no fitted quantity is renamed as a prediction. Score 0.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 5 invented entities

The paper introduces no numeric free parameters; its load-bearing assumptions are the adequacy of the seven-field representation, the representativeness of the frozen corpus, the correctness of the provisional single-coder E/D codes, and the external characterization of GraSP. The five residual bundles are new postulated constructs with explicit falsifiers.

axioms (4)
  • domain assumption The seven-field mechanism representation K=⟨S,C,T,P,F,L,R⟩ is sufficient to capture the control semantics that determine how an agent runs.
    The entire comparison and all residual bundles are defined over this tuple; if control semantics outside these fields (e.g., interface conventions, prompt formatting, social context) affect agent behavior, the gaps may be misidentified. Invoked in Section II-C.
  • domain assumption The frozen corpus of ten historical architectures, eight runtime families, and 42 mechanism-focused systems is representative of the field at the 26 July 2026 cutoff.
    The authors state the corpus was frozen and make no PRISMA completeness claim; missing systems could close a residual bundle or reveal additional ones. Sections II-B and VII-C.
  • domain assumption The E1–E4 and D0–D4 codes assigned by the single coder are correct.
    The paper explicitly does not claim inter-rater reliability and labels scores provisional; if a different coder assigns different levels, the migration ledger and residual bundles would change. Section II-E and Appendix C.
  • domain assumption The characterization of GraSP in reference [63] is accurate: it composes calibrated routing, typed verification, bounded repair, and replanning/ReAct fallback.
    B5's closure as a negative case hinges on this external characterization; if GraSP does not actually implement the full chain, B5 would become a residual bundle again. Section V-A-d and Appendix C-P.
invented entities (5)
  • B1: activation–latency–utility coupling independent evidence
    purpose: Proposed runtime invariant for memory control: runtime memory items carry activation, predicted retrieval latency, downstream value, and maintenance cost; context assembly triggers competition.
    Postulated in §VI-C-a; falsifier: lower relevant recall or overhead without downstream use (Table VI).
  • B2: typed impasse–substate–compilation independent evidence
    purpose: An execution ledger classifies failure, freezes parent state, opens a bounded recovery subgraph, returns a typed resolution, and compiles a reusable recovery rule.
    Postulated in §VI-C-b; falsifier: misdiagnosis, state corruption, or no reuse (Table VI).
  • B3: content competition–workspace–broadcast learning independent evidence
    purpose: Typed candidates from memory, agents, tools, monitors, and plans compete for a bounded workspace; the winner becomes globally available and downstream utility updates later admission.
    Postulated in §VI-C-c; falsifier: no gain over GWA selection or top-k retrieval (Table VI).
  • B4: intention–reconsideration–method authority independent evidence
    purpose: A persistent intention record stores goal, rationale, continuation conditions, and termination status; a governor decides whether to continue, delegate, suspend, abandon, or change method.
    Postulated in §VI-C-d; falsifier: stale commitment, oscillation, premature stopping, or method changes without diagnostic value (Table VI).
  • B6: uncertainty–resources–interruption–stopping independent evidence
    purpose: One calibrated uncertainty object propagates through memory, planning, tool authorization, execution monitoring, and stopping, allocating budgets and authorizing interruption.
    Postulated in §VI-C-f; falsifier: miscalibration or cost without risk reduction (Table VI).

pith-pipeline@v1.3.0-alltime-deepseek · 29901 in / 12407 out tokens · 118029 ms · 2026-07-31T23:26:52.012087+00:00 · methodology

0 comments
read the original abstract

Memory, planning, reflection, and tool use are often compared as feature labels, obscuring the control semantics that determine how an agent actually runs. This review connects ten historical cognitive architectures, eight language-agent runtime families, and forty-two mechanism-focused modern systems. We reconstruct each mechanism through state, control, transition, persistence, failure, learning, and resource governance, then code evidence relation (E1-E4) separately from migration depth (D0-D4). The resulting landscape is uneven. Modern agents have operationalized substantial parts of adaptive memory, failure recovery, dynamic team selection, workflow search, skill induction, resource scheduling, and uncertainty-conditioned action, although often through independent convergence rather than documented inheritance. The strongest remaining opportunities lie in couplings among mechanisms. Closest-baseline screening closes one proposed gap: GraSP already combines calibrated multi-skill selection, typed compilation, verification, bounded repair, and replanning or ReAct fallback. Five residual bundles remain: activation with latency and action utility; typed impasse with isolated substates and resolution compilation; bounded content competition with broadcast and admission learning; persistent intention with reconsideration and live method authority; and uncertainty with resource allocation, interruption, and stopping. We contribute a distinctive-mechanism catalog, an auditable evidence-depth framework, and a falsifiable agenda for testing these bundles as composable runtime invariants.

Figures

Figures reproduced from arXiv: 2607.23942 by Haodi Fan, Zucong Lan.

Figure 1
Figure 1. Figure 1: Review logic. A similarity claim is admitted only after architecture reconstruction and a node–edge–invariant test; migration gaps [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Executable control cycles for ACT-R and Soar. ACT-R exposes typed buffers to production matching and updates retrieval and [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Four distinct execution organizations. CLARION coordinates explicit and implicit control; LIDA requires competition before selective [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Two additional control organizations. EPIC makes processor overlap, duration, and contention explicit; Sigma propagates evidence [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: The modern unit of analysis. A policy such as ReAct or reflection runs inside a framework that owns authoritative state, transition [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Provisional single-coder migration-depth landscape for all ten historical families in the frozen corpus. Most reach operational control [PITH_FULL_IMAGE:figures/full_fig_p011_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Four deep mappings: three residual couplings and one closed candidate. Dashed arrows lead to surviving residuals; the solid green [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Dependencies among five residual bundles and the migrated B5 skill governor. B4 scopes recovery; B6 governs risk and resources; [PITH_FULL_IMAGE:figures/full_fig_p017_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Synthesis architecture for testing the five residual integration hypotheses while reusing a migrated GraSP-like skill governor. Arrowheads [PITH_FULL_IMAGE:figures/full_fig_p018_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

76 extracted references · 38 linked inside Pith

  1. [1]

    A review of 40 years of cognitive architecture research: Core cognitive abilities and practical applications,

    I. Kotseruba and J. K. Tsotsos, “A review of 40 years of cognitive architecture research: Core cognitive abilities and practical applications,”Artificial Intelligence Review, vol. 53, pp. 17–94, 2020

  2. [2]

    Cognitive architectures for language agents,

    T. R. Sumers, S. Yao, K. Narasimhan, and T. L. Griffiths, “Cognitive architectures for language agents,”arXiv preprint arXiv:2309.02427, 2023

  3. [3]

    The rise and potential of large language model based agents: A survey,

    Z. Xi, W. Chen, X. Guo, W. He, Y . Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou,et al., “The rise and potential of large language model based agents: A survey,” arXiv preprint arXiv:2309.07864, 2023

  4. [4]

    Agent harness for large language model agents: A survey,

    Q. Meng, Y . Wang, L. Chen, Y . Li, W. Wu, W. Jiang, Q. Wang, C. Lu, Y . Gao, Y . Wu, and Y . Hu, “Agent harness for large language model agents: A survey,”Preprints,

  5. [5]

    From prompt-response to goal-directed systems: The evolution of agentic AI software architecture,

    M. Alenezi, “From prompt-response to goal-directed systems: The evolution of agentic AI software architecture,”arXiv preprint arXiv:2602.10479, 2026

  6. [6]

    An integrated theory of the mind,

    J. R. Anderson, “An integrated theory of the mind,”Psychological Review, vol. 111, no. 4, pp. 1036–1060, 2004

  7. [7]

    ACT-R 7.30+ reference manual,

    ACT-R Research Group, “ACT-R 7.30+ reference manual,” 2024

  8. [8]

    Soar: An architecture for general intelligence,

    J. E. Laird, A. Newell, and P. S. Rosenbloom, “Soar: An architecture for general intelligence,”Artificial Intelligence, vol. 33, no. 1, pp. 1–64, 1987

  9. [9]

    The soar architecture,

    Soar Research Group, “The soar architecture,” 2025

  10. [10]

    Cognition and multi-agent interaction: The CLARION cognitive architecture: Extending cognitive modeling to social simulation,

    R. Sun, “Cognition and multi-agent interaction: The CLARION cognitive architecture: Extending cognitive modeling to social simulation,” 2005

  11. [11]

    Global workspace theory, its LIDA model and the underlying neuroscience,

    S. Franklin, S. Strain, J. Snaider, R. McCall, and U. Faghihi, “Global workspace theory, its LIDA model and the underlying neuroscience,”Biologically Inspired Cognitive Architectures, 2012

  12. [12]

    The Hearsay-II speech-understanding system: A tutorial,

    L. D. Erman, F. Hayes-Roth, V . R. Lesser, and D. R. Reddy, “The Hearsay-II speech-understanding system: A tutorial,”ACM Computing Surveys, vol. 12, no. 2, pp. 213–253, 1980

  13. [13]

    Modeling rational agents within a BDI-architecture,

    A. S. Rao and M. P. Georgeff, “Modeling rational agents within a BDI-architecture,” inProceedings of the Second International Conference on Principles of Knowledge Representation and Reasoning, pp. 473–484, Morgan Kaufmann, 1991

  14. [14]

    Commitment and effectiveness of situated agents,

    D. N. Kinny and M. P. Georgeff, “Commitment and effectiveness of situated agents,” inProceedings of the Twelfth International Joint Conference on Artificial Intelligence, pp. 82–88, 1991

  15. [15]

    BDI agents: From theory to practice,

    A. S. Rao and M. P. Georgeff, “BDI agents: From theory to practice,” inProceedings of the First International Conference on Multi-Agent Systems, pp. 312–319, 1995

  16. [16]

    MIDCA: A metacognitive, integrated dual-cycle architecture for self-regulated autonomy,

    M. T. Cox, Z. Alavi, D. Dannenhauer, V . Eyorokon, H. Munoz-Avila, and D. Perlis, “MIDCA: A metacognitive, integrated dual-cycle architecture for self-regulated autonomy,” 2016

  17. [17]

    Evolution of the Icarus cognitive architecture,

    D. Choi and P. Langley, “Evolution of the Icarus cognitive architecture,” 2018

  18. [18]

    An overview of the EPIC architecture for cognition and performance with application to human-computer interaction,

    D. E. Kieras and D. E. Meyer, “An overview of the EPIC architecture for cognition and performance with application to human-computer interaction,”Human-Computer Interaction, vol. 12, no. 4, pp. 391–438, 1997

  19. [19]

    The Sigma cognitive architecture and system: Toward functionally elegant grand unification,

    P. S. Rosenbloom, A. Demski, and V . Ustun, “The Sigma cognitive architecture and system: Toward functionally elegant grand unification,” 2016

  20. [20]

    Memgpt: Towards llms as operating systems,

    C. Packer, S. Wooders, K. Lin, V . Fang, S. G. Patil, I. Stoica, and J. E. Gonzalez, “Memgpt: Towards llms as operating systems,”arXiv preprint arXiv:2310.08560, 2023. 28

  21. [21]

    Letta: Stateful agents with advanced memory

    Letta, Inc., “Letta: Stateful agents with advanced memory.” GitHub repository, 2026. Commit b76da9092518cbaa2d09042e52fdcbde69243e18; accessed 2026-07-25

  22. [22]

    Langgraph: Build resilient language agents as graphs

    LangChain AI, “Langgraph: Build resilient language agents as graphs.” GitHub repository, 2026. Commit 30c4d58db86455128e42ddec96b1ba53c553ba22; accessed 2026-07-25

  23. [23]

    Autogen: Enabling next-gen llm applications via multi-agent conversation,

    Q. Wu, G. Bansal, J. Zhang, Y . Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, A. H. Awadallah, R. W. White, D. Burger, and C. Wang, “Autogen: Enabling next-gen llm applications via multi-agent conversation,”arXiv preprint arXiv:2308.08155, 2023

  24. [24]

    Autogen: A programming framework for agentic ai

    Microsoft, “Autogen: A programming framework for agentic ai.” GitHub repository, 2026. Commit 027ecf0a379bcc1d09956d46d12d44a3ad9cee14; accessed 2026-07-25

  25. [25]

    Agentscope 1.0: A developer-centric framework for building agentic applications,

    D. Gao, Z. Li, Y . Xie, W. Kuang, L. Yao, B. Qian, Z. Ma, Y . Cui, H. Luo, S. Li, L. Yi, Y . Yu, S. He, Z. Luo, W. Zhou, Z. Zhang, X. He, Z. Chen, W. Liao, F. I. Kushnazarov, Y . Li, B. Ding, and J. Zhou, “Agentscope 1.0: A developer-centric framework for building agentic applications,”arXiv preprint arXiv:2508.16279, 2025

  26. [26]

    Openhands: An open platform for ai software developers as generalist agents,

    X. Wang, B. Li, Y . Song, F. F. Xu, X. Tang, M. Zhuge, J. Pan, Y . Song, B. Li, J. Singh, H. H. Tran, F. Li, R. Ma, M. Zheng, B. Qian, Y . Shao, N. Muennighoff, Y . Zhang, B. Hui, J. Lin, R. Brennan, H. Peng, H. Ji, and G. Neubig, “Openhands: An open platform for ai software developers as generalist agents,”arXiv preprint arXiv:2407.16741, 2024

  27. [27]

    The openhands software agent sdk: A composable and extensible foundation for production agents,

    X. Wang, S. Rosenberg, J. Michelini, C. Smith, H. Tran, E. Nyst, R. Malhotra, X. Zhou, V . Chen, R. Brennan, and G. Neubig, “The openhands software agent sdk: A composable and extensible foundation for production agents,”arXiv preprint arXiv:2511.03690, 2025

  28. [28]

    Microsoft agent framework

    Microsoft, “Microsoft agent framework.” GitHub repository, 2026. Commit c6442de52882a47fa6796fb380c213cd65f2fc8e; accessed 2026-07-25

  29. [29]

    Openai agents sdk for python

    OpenAI, “Openai agents sdk for python.” GitHub repository, 2026. Commit c1b423749e2bf8ca5f89cad13e2a144c9683a6ee; accessed 2026-07-25

  30. [30]

    AIOS: LLM agent operating system,

    K. Mei, X. Zhu, W. Xu, W. Hua, M. Jin, Z. Li, S. Xu, R. Ye, Y . Ge, and Y . Zhang, “AIOS: LLM agent operating system,”arXiv preprint arXiv:2403.16971, 2024

  31. [31]

    Magentic-one: A generalist multi-agent system for solving complex tasks,

    A. Fourney, G. Bansal, H. Mozannar, C. Tan, E. Salinas, Erkang, Zhu, F. Niedtner, G. Proebsting, G. Bassman, J. Gerrits, J. Alber, P. Chang, R. Loynd, R. West, V . Dibia, A. Awadallah, E. Kamar, R. Hosn, and S. Amershi, “Magentic-one: A generalist multi-agent system for solving complex tasks,”arXiv preprint arXiv:2411.04468, 2024

  32. [32]

    Open agent specification (agent spec): A unified representation for ai agents,

    S. Amini, Y . Benajiba, C. Bernardis, P. Cayet, H. Chafi, A. Fathan, L. Faucon, D. Hilloulin, S. Hong, I. Kossyk, T. M. S. Le, R. Patra, S. Ravi, J. Schweizer, J. Singh, S. Singh, W. Sun, K. Talamadupula, and J. Xu, “Open agent specification (agent spec): A unified representation for ai agents,”arXiv preprint arXiv:2510.04173, 2025

  33. [33]

    Agentic memory: Learning unified long-term and short-term memory management for large language model agents,

    Y . Yu, L. Yao, Y . Xie, Q. Tan, J. Feng, Y . Li, and L. Wu, “Agentic memory: Learning unified long-term and short-term memory management for large language model agents,”arXiv preprint arXiv:2601.01885, 2026

  34. [34]

    Memory-R1: Enhancing large language model agents to manage and utilize memories via reinforcement learning,

    S. Yan, X. Yang, Z. Huang, E. Nie, Z. Ding, Z. Li, X. Ma, J. Bi, K. Kersting, J. Z. Pan, H. Schuetze, V . Tresp, and Y . Ma, “Memory-R1: Enhancing large language model agents to manage and utilize memories via reinforcement learning,”arXiv preprint arXiv:2508.19828, 2025

  35. [35]

    Memory as a controlled process: Learned adaptive memory management for LLM agents,

    E. H. Jiang, Z. Zhang, Y . Wu, L. Li, D. Liu, X. Liang, R. Sun, Y . Li, E. Sun, H. Luo, Z. Kang, A. Caliskan, K.-W. Chang, and Y . N. Wu, “Memory as a controlled process: Learned adaptive memory management for LLM agents,”arXiv preprint arXiv:2607.13591, 2026

  36. [36]

    Deltamem: Towards agentic memory management via reinforcement learning,

    Q. Zhang, S. Huang, C. Liu, S. Yang, J. Zhao, H. Wang, and P. Xie, “Deltamem: Towards agentic memory management via reinforcement learning,”arXiv preprint arXiv:2604.01560, 2026

  37. [37]

    Beyond heuristics: A decision-theoretic framework for agent memory management,

    C. Sun, X. Chen, J. Luo, D. Zhang, and X. Li, “Beyond heuristics: A decision-theoretic framework for agent memory management,”arXiv preprint arXiv:2512.21567, 2025

  38. [38]

    Adaptive memory admission control for LLM agents,

    G. Zhang, W. Jiang, X. Wang, A. Behr, K. Zhao, J. Friedman, X. Chu, and A. Anoun, “Adaptive memory admission control for LLM agents,”arXiv preprint arXiv:2603.04549, 2026

  39. [39]

    Forget to improve: On-device LLM-agent continual learning via budget-curated memory,

    B. Wu, Z. Ding, J. Huang, and Y . Zhao, “Forget to improve: On-device LLM-agent continual learning via budget-curated memory,”arXiv preprint arXiv:2606.25115, 2026

  40. [40]

    PALADIN: Self-correcting language model agents to cure tool-failure cases,

    S. V . Vuddanti, A. Shah, S. K. Chittiprolu, T. Song, S. Dev, K. Zhu, and M. Chaudhary, “PALADIN: Self-correcting language model agents to cure tool-failure cases,” arXiv preprint arXiv:2509.25238, 2025

  41. [41]

    AgentHER: Hindsight experience replay for LLM agent trajectory relabeling,

    L. Ding, “AgentHER: Hindsight experience replay for LLM agent trajectory relabeling,”arXiv preprint arXiv:2603.21357, 2026

  42. [42]

    Reflexion: Language agents with verbal reinforcement learning,

    N. Shinn, F. Cassano, A. Berman, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” inAdvances in Neural Information Processing Systems, 2023

  43. [43]

    AgentDebugX: An open-source toolkit for failure observability, attribution, and recovery in LLM agents,

    K. Zhu, X. Ye, Z. Han, Y . Zhao, B. Li, W. Zhang, M. Tian, X. Tang, P. Lu, J. Zou, J. You, and H. Ji, “AgentDebugX: An open-source toolkit for failure observability, attribution, and recovery in LLM agents,”arXiv preprint arXiv:2607.18754, 2026

  44. [44]

    Shepherd: Enabling programmable meta-agents via reversible agentic execution traces,

    S. Yu, D. Chong, A. Nandi, D. Soylu, J. Sun, C. D. Manning, and W. Shi, “Shepherd: Enabling programmable meta-agents via reversible agentic execution traces,”arXiv preprint arXiv:2605.10913, 2026

  45. [45]

    A dynamic LLM-powered agent network for task-oriented agent collaboration,

    Z. Liu, Y . Zhang, P. Li, Y . Liu, and D. Yang, “A dynamic LLM-powered agent network for task-oriented agent collaboration,”arXiv preprint arXiv:2310.02170, 2023

  46. [46]

    Adaptive graph pruning for multi-agent communication,

    B. Li, Z. Zhao, D.-H. Lee, and G. Wang, “Adaptive graph pruning for multi-agent communication,”arXiv preprint arXiv:2506.02951, 2025

  47. [47]

    Theater of Mind for LLMs: A cognitive architecture based on global workspace theory,

    W. Shang, “Theater of Mind for LLMs: A cognitive architecture based on global workspace theory,”arXiv preprint arXiv:2604.08206, 2026

  48. [48]

    Controlled yet natural: A hybrid BDI–LLM conversational agent for child helpline training,

    M. A. Owayyed, A. Denga, and W.-P. Brinkman, “Controlled yet natural: A hybrid BDI–LLM conversational agent for child helpline training,” inProceedings of the 25th ACM International Conference on Intelligent Virtual Agents, 2025

  49. [49]

    A self-aware BDI agent for LLM-driven reasoning: Architecture, design, and preliminary evaluation,

    C. Lawton, D. Sacharny, and T. C. Henderson, “A self-aware BDI agent for LLM-driven reasoning: Architecture, design, and preliminary evaluation,” Tech. Rep. UUCS-26-001, University of Utah, 2026

  50. [50]

    Devil’s advocate: Anticipatory reflection for LLM agents,

    H. Wang, T. Li, Z. Deng, D. Roth, and Y . Li, “Devil’s advocate: Anticipatory reflection for LLM agents,”arXiv preprint arXiv:2405.16334, 2024

  51. [51]

    Cognitive control architecture (CCA): A lifecycle supervision framework for robustly aligned AI agents,

    Z. Liang, T. Hu, Z. Chen, and M. Tang, “Cognitive control architecture (CCA): A lifecycle supervision framework for robustly aligned AI agents,”arXiv preprint arXiv:2512.06716, 2025

  52. [52]

    Evaluating goal drift in language model agents,

    R. Arike, E. Donoway, H. Bartsch, and M. Hobbhahn, “Evaluating goal drift in language model agents,”arXiv preprint arXiv:2505.02709, 2025

  53. [53]

    When agents commit too soon: Diagnosing premature commitment in llm agents,

    A. Mehta, “When agents commit too soon: Diagnosing premature commitment in llm agents,”arXiv preprint arXiv:2606.22936, 2026

  54. [54]

    Automated design of agentic systems,

    S. Hu, C. Lu, and J. Clune, “Automated design of agentic systems,”arXiv preprint arXiv:2408.08435, 2024

  55. [55]

    AFlow: Automating agentic workflow generation,

    J. Zhang, J. Xiang, Z. Yu, F. Teng, X. Chen, J. Chen, M. Zhuge, X. Cheng, S. Hong, J. Wang, B. Zheng, and B. Liu, “AFlow: Automating agentic workflow generation,” arXiv preprint arXiv:2410.10762, 2024

  56. [56]

    Metareflection: Learning instructions for language agents using past reflections,

    P. Gupta, S. Kirtania, A. Singha, S. Gulwani, A. Radhakrishna, S. Shi, and G. Soares, “Metareflection: Learning instructions for language agents using past reflections,” arXiv preprint arXiv:2405.13009, 2024

  57. [57]

    Online skill learning for web agents via state-grounded dynamic retrieval,

    J. Li, K. Deng, Y . Wang, J. Huang, Y . Shi, Q. Tan, J. Lu, and N. Liu, “Online skill learning for web agents via state-grounded dynamic retrieval,”arXiv preprint arXiv:2606.04391, 2026

  58. [58]

    SCALAR: Learning and composing skills through LLM-guided symbolic planning and deep RL grounding,

    R. Zabounidis, Y . Wu, S. Stepputtis, W. Kim, Y . Li, T. Mitchell, and K. Sycara, “SCALAR: Learning and composing skills through LLM-guided symbolic planning and deep RL grounding,”arXiv preprint arXiv:2603.09036, 2026

  59. [59]

    V oyager: An open-ended embodied agent with large language models,

    G. Wang, Y . Xie, Z. Jiang, A. Mandlekar, Y . Xiao, C. Zhu, Y . Fan, and A. Anandkumar, “V oyager: An open-ended embodied agent with large language models,”arXiv preprint arXiv:2305.16291, 2023

  60. [60]

    Agent workflow memory,

    Z. Z. Wang, J. Mao, D. Fried, and G. Neubig, “Agent workflow memory,”arXiv preprint arXiv:2409.07429, 2024

  61. [61]

    Generative skill composition for LLM agents,

    X. Zhao, Z. Tan, V . Tadiparthi, N. Agarwal, K. Lee, E. M. Pari, H. N. Mahjoub, and T. Chen, “Generative skill composition for LLM agents,”arXiv preprint arXiv:2606.32025, 2026

  62. [62]

    SkillOps: Managing LLM agent skill libraries as self-maintaining software ecosystems,

    X. Song, H. Pu, and L. Zhao, “SkillOps: Managing LLM agent skill libraries as self-maintaining software ecosystems,”arXiv preprint arXiv:2605.13716, 2026

  63. [63]

    GraSP: Graph-structured skill compositions for LLM agents,

    T. Xia, L. Hu, Y . Sun, M. Xu, L. Xu, S. Wang, W. Xu, and J. Jiang, “GraSP: Graph-structured skill compositions for LLM agents,”arXiv preprint arXiv:2604.17870, 2026

  64. [64]

    Agentic compilation: Mitigating the LLM rerun crisis for minimized-inference-cost web automation,

    J. Chundru, “Agentic compilation: Mitigating the LLM rerun crisis for minimized-inference-cost web automation,”arXiv preprint arXiv:2604.09718, 2026

  65. [65]

    SkVM: Revisiting language VM for skills across heterogenous LLMs and harnesses,

    L. Chen, E. Feng, Y . Xia, and H. Chen, “SkVM: Revisiting language VM for skills across heterogenous LLMs and harnesses,”arXiv preprint arXiv:2604.03088, 2026

  66. [66]

    SkillSmith: Compiling agent skills into boundary-guided runtime interfaces,

    D. Xu, Z. Chen, Z. Pan, J. Guan, D. Dong, J. Li, and B. Pu, “SkillSmith: Compiling agent skills into boundary-guided runtime interfaces,”arXiv preprint arXiv:2605.15215, 2026

  67. [67]

    SkCC: Portable and secure skill compilation for cross-framework LLM agents,

    Y . Ouyang, Y . Xiao, Y . Gu, and X. Zhang, “SkCC: Portable and secure skill compilation for cross-framework LLM agents,”arXiv preprint arXiv:2605.03353, 2026

  68. [68]

    AgentRM: An OS-inspired resource manager for LLM agent systems,

    J. She, “AgentRM: An OS-inspired resource manager for LLM agent systems,”arXiv preprint arXiv:2603.13110, 2026

  69. [69]

    Spend less, reason better: Budget-aware value tree search for LLM agents,

    Y . Li, W. Deng, J. Li, and X. Li, “Spend less, reason better: Budget-aware value tree search for LLM agents,”arXiv preprint arXiv:2603.12634, 2026

  70. [70]

    Towards uncertainty-aware language agent,

    J. Han, W. Buntine, and E. Shareghi, “Towards uncertainty-aware language agent,”arXiv preprint arXiv:2401.14016, 2024

  71. [71]

    Robots that ask for help: Uncertainty alignment for large language model planners,

    A. Z. Ren, A. Dixit, A. Bodrova, S. Singh, S. Tu, N. Brown, P. Xu, L. Takayama, F. Xia, J. Varley, Z. Xu, and D. Sadigh, “Robots that ask for help: Uncertainty alignment for large language model planners,”arXiv preprint arXiv:2307.01928, 2023

  72. [72]

    Calibrate-then-act: Cost-aware exploration in LLM agents,

    W. Ding, N. Tomlin, and G. Durrett, “Calibrate-then-act: Cost-aware exploration in LLM agents,”arXiv preprint arXiv:2602.16699, 2026

  73. [73]

    Utility-guided agent orchestration for efficient LLM tool use,

    B. Liu, G. Zhao, and H. Xu, “Utility-guided agent orchestration for efficient LLM tool use,”arXiv preprint arXiv:2603.19896, 2026

  74. [74]

    Cognitive LLMs: Toward human-like artificial intelligence by integrating cognitive architectures and large language models for manufacturing decision-making,

    S. Wu, A. Oltramari, J. Francis, C. L. Giles, and F. E. Ritter, “Cognitive LLMs: Toward human-like artificial intelligence by integrating cognitive architectures and large language models for manufacturing decision-making,”Neurosymbolic Artificial Intelligence, 2025

  75. [75]

    Bootstrapping cognitive agents with a large language model,

    F. Zhu and R. Simmons, “Bootstrapping cognitive agents with a large language model,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 1, 2024

  76. [76]

    Procedural knowledge learning in soar,

    Soar Research Group, “Procedural knowledge learning in soar,” 2025