Pith. sign in

REVIEW 3 major objections 5 minor 300 references

Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence

T0 review · 3 major / 5 minor · reviewed 2026-08-27 · deepseek-v4-flash

Pith's one-line read Graph Engineering asserts that system intelligence is an organizational achievement: explicit, dynamic graphs of tasks, agents, and runtime states, not more capable individual agents, are what coordinate LLM agent systems.

desk verdict A useful survey whose taxonomy is the real contribution, but whose central claim that graph structures are the necessary foundation for System Intelligence is asserted rather than demonstrated. read the letter →

arxiv 2608.21156 v2 pith:6INSLI3A submitted 2026-08-21 cs.IR cs.AIcs.ET

classification cs.IRcs.AIcs.ET
keywords GraphEngineeringSystemIntelligenceLLMagentsmulti-agentcoordinationtaskorganizationruntimestatemanagementagentevolutionsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that LLM-based agent systems are approaching a structural limit: real-world tasks routinely require heterogeneous expertise, interdependent subtasks, parallel execution, independent verification, and persistent state, and these demands exceed what any single agent can organize. The authors introduce Graph Engineering as the response, building explicit, dynamic, evolving graph structures that represent tasks, agents, and system states, so relationships among them can be scheduled, monitored, diagnosed, and recovered. Their central claim is that System Intelligence—the ability of an agent system to coordinate multiple intelligent components into a coherent adaptive whole—emerges from systematic governance of those relationships, not from adding more agents or more context. If the paper is right, next-generation agent systems should be designed around graph-native substrates and auditable runtime state, with individual models treated as components inside a larger organizational structure.

What carries the argument

The central object is the operational graph: nodes stand for subtasks, agents, capabilities, or state snapshots, and typed edges encode dependencies, delegation, information flow, provenance, and validity constraints. The paper's argument is carried by treating three such graphs as coupled, executable structures rather than static pictures: a Work Organization Graph for what to do, an Agent Team Graph for who does it and how information moves, and a State Evolution Graph for how the system tracks and recovers from what has happened. The load-bearing mechanism is externalization—moving the relationships out of a single agent's context and into an auditable substrate where parallel work, independent verification, fault localization, and rollback become well-defined operations.

What would settle it

A controlled comparison would settle it: run the same long-horizon, parallel, fault-prone tasks under matched execution budgets in two conditions—a single state-of-the-art agent with long context and broad tool access, versus a graph-engineered multi-agent system with explicit task, team, and state graphs. If the single agent matches the system on task success, failure recovery, and wall-clock cost, the paper's central claim is wrong; if the system wins and the gap grows with task complexity, the claim is supported.

Watch

Extended reading notes

Core claim

The paper asks what makes a multi-agent LLM system intelligent and answers: its organization, not its headcount. System Intelligence is defined as the ability to decompose complex objectives, allocate responsibilities across heterogeneous agents, coordinate interdependent execution, and maintain system-level state; and the paper argues this capability is not reducible to stronger models, longer contexts, or more agents. Graph Engineering is the claim that this capability is achieved by externalizing relationships among tasks, components, and runtime states into explicit, dynamic graph structures that can be optimized and revised. The survey organizes the evidence into three coupled graph views: Task Organization graphs make goals decomposable and schedulable; Agent Coordination graphs bind capabilities to roles and shape communication; and Runtime State Management graphs record what actually happened, localize the first invalid state, and enable recovery from validated boundaries. The paper concludes that persistent System Evolution—turning accumulated execution evidence into validated structural change—is the necessary next step, and that shared ontology semantics are required for graphs to be interpreted consistently across components.

Load-bearing premise

The load-bearing premise is that complex tasks fundamentally exceed the organizational capacity of any single agent, no matter how much context or how many tools it is given; if one strong agent with a long enough context could handle the same tasks, the case for distributing intelligence across a graph-engineered system collapses.

Editorial extensions

If this is right

  • Agent systems should be organized around explicit task graphs whose edges encode dependency, concurrency, and verification constraints, making parallel scheduling a graph operation rather than an implicit property of one loop.
  • Agent teams should be assembled from capability and communication graphs that can be rewired in response to runtime feedback, so role assignments and information flow adapt to changing conditions.
  • Failures should be handled through recorded state: the system localizes the first invalid transition, marks a recovery boundary, and re-executes only the affected region instead of restarting from scratch.
  • Benchmarking should separate model-level gains from system-organization gains, measuring structural fidelity, operational correctness, and evolution in addition to end-task success.
  • The engineering frontier becomes persistent system evolution—converting runtime evidence into validated, versioned, and rollback-safe changes to the graphs that govern future runs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper leaves implicit: if the organizational story is right, a graph-native runtime that provides scheduling, state, and provenance as first-class services should outperform ad-hoc multi-agent frameworks on long-horizon, fault-prone tasks under matched execution budgets; head-to-head comparisons of that kind have not yet been run.
  • The same logic predicts a shift in scaling curves: for tasks that exceed a single agent's organizational capacity, success should scale more smoothly with the structure of the graph (number of nodes, edges, and state transitions) than with context length or model size.
  • The paper's move toward ontology engineering suggests a concrete design constraint: in open, long-running systems, agents that disagree about what counts as task completion, sufficient evidence, or a valid state will need a shared machine-interpretable semantic model, which makes ontology consistency a testable requirement rather than a philosophical one.
  • If system intelligence is genuinely about relationships, then evaluation should include structural perturbation tests—removing or rewiring edges in the task or team graph and measuring how much performance degrades—since that isolates the contribution of organization from the contribution of individual components.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Request a human review

A listed scientist reviews the paper for a fee and the review publishes here regardless of verdict. See the reviewers or get listed.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This survey proposes a new framing for LLM-based agent systems. It defines a progression from Model Intelligence to Individual Intelligence and then to System Intelligence, and introduces Graph Engineering as the paradigm that uses explicit, dynamic graph structures for task organization, agent coordination, and runtime state management. The paper reviews the literature along these axes, catalogs benchmarks, open-source libraries, and applications, and identifies ontology engineering as a future direction. Its central claim, stated in Section 4.1, is that system intelligence requires systematic governance of relationships among tasks, components, and runtime states, and that graph structures provide the natural and unifying substrate for that governance.

Significance. If the central claim is accepted, the paper offers a useful synthesis and design agenda for the field: it provides a coherent taxonomy, a comprehensive resource collection, and a clear set of research challenges. The authors deserve credit for including candid limitation statements: Section 5.1 acknowledges that end-task success cannot distinguish organizational gains from gains due to model, context, or compute, and Section 9.7 notes that persistent system evolution remains rare. They also ship a curated GitHub resource list. However, the load-bearing premise that complex tasks inherently exceed any single agent's organizational capacity is asserted rather than demonstrated, and the definitions of Individual and System Intelligence are partly circular because the individual-agent abstraction in Eq. (1) already permits graph-aware harnesses. The value of the survey is therefore real but the central claim, as stated, is not yet established.

major comments (3)
  1. [Sections 1, 3.5, and 4.1] The paper's central premise is that complex tasks 'exceed the organizational capacity of any single agent' and that 'simply augmenting an individual agent's capabilities or context cannot resolve this problem.' This claim is load-bearing: if it is false, Graph Engineering is one useful technique rather than the 'structural foundation' of System Intelligence. The manuscript provides no controlled evidence for this premise. It cites no comparison of a graph-engineered multi-agent system against a single agent with a graph-aware harness under matched compute, context, and tool budgets. Section 5.1 explicitly concedes that end-task success cannot separate organizational gains from other factors, and Section 7.4 calls for matched budgets and structural ablations as future work. To support the central claim, the paper should either synthesize existing controlled results (e.g., BenchAgent [96], MASEval [79], MAS-PromptBench [17]) or soften the claim to describe Graph Engineering as a promising and increasingly prevalent approach. As written, the necessity claim is not supported.
  2. [Section 2.1, Eq. (1) and Section 3.3] The definition of an Individual Agent, A_i = Loop(F_i, H_i; s_i), includes an Agent Harness H_i that, per Section 3.3, can contain runtime orchestration, memory, tools, external state, and graph-based state machines (as exemplified by Burr, LangGraph, and Pydantic AI in Table 2). Under this definition, a single agent whose harness maintains an explicit task graph, spawns parallel tool calls, isolates execution states, and invokes independent verifiers already satisfies the three Graph Engineering dimensions of task organization, agent coordination, and runtime state management described in Section 4. The three limitations listed in Section 3.5 (serial execution, role confusion, fragile state) are properties of a naive single-context loop, not of the general abstraction A_i. Consequently, the claimed dichotomy between Individual Intelligence and System Intelligence is partly definitional: it rests on counting agents rather than on the organizational substrate. The authors should give an operational criterion that distinguishes system-level intelligence from harness-engineering-level intelligence, such as distributed control or the absence of a shared central context, and apply it consistently.
  3. [Section 4.1 and Section 9.7] Graph Engineering is defined so broadly that the taxonomy risks becoming vacuous. Any explicit representation of task dependencies, such as ReWOO variable references, LLMCompiler dataflow DAGs, state machines, or event streams, is classified as a graph structure, and systems that merely use such structures are included in the 'System Intelligence' tables. The paper itself introduces a useful distinction between being 'graph-structured' and being 'graph-engineered' only in Section 9.7, after the earlier sections have already used the broad definition. This conflation makes the claim that graphs are a 'unified foundation' true by construction, because almost any structured system can be described as a graph. The paper should provide an operational definition of Graph Engineering early in Section 4 that requires the graph to be dynamic, executable, and subject to structural optimization or evolution, and then apply that definition consistently in Tables 2 and 3. Otherwise the proposed paradigm cannot be falsified.
minor comments (5)
  1. [Section 3.5] The sentence beginning 'Despite the advances in Harness and Advances in Harness and Loop Engineering...' contains a duplicated phrase and needs copyediting.
  2. [Section 5.2.1] The citation for the LAMP system, listed as reference [7], does not match the cited source; reference [7] is an Anthropic document on orchestrating Claude Code sessions. The citation appears to be misaligned and should be corrected or replaced.
  3. [References] Reference [1] contains the Editorial note 'verify authors and final publication metadata before submission.' This internal comment should be removed and the reference completed.
  4. [Table 1] Several rows show malformed link placeholders such as '/da◎abase' (e.g., AppWorld, GateMem). These should be replaced with the intended URLs or a uniform placeholder.
  5. [Section 2.1, Eq. (1)] The notation s t i in Eq. (1) would be clearer as s_i(t), or with an explicit sentence explaining that the superscript is a time index rather than an exponent.

Circularity Check

2 steps flagged · score 4.0 of 10

The central framing is partly definitional: System Intelligence and Graph Engineering name the same organizing/coordinating/state-management capacities, and the 'single agent cannot do it' limitation is asserted for a narrower loop than the paper's own harness definition.

  1. self definitional [Section 4.1, Overview of Graph Engineering]
    "At its core, system intelligence requires the systematic governance of relationships among tasks, components, and runtime states. ... To this end, we introduce Graph Engineering as a structure-centered engineering foundation for system intelligence: it uses graph structures as the core substrate for externalizing relationships among tasks, components, and runtime states, thereby supporting system-level organization, coordination, monitoring, recovery, and optimization."

    System Intelligence is defined as the ability to organize objectives, coordinate agents, and maintain system state, while Graph Engineering is defined as the graph-based realization of exactly those three functions. The conclusion that Graph Engineering provides 'the structural foundation for System Intelligence' is therefore an analytic consequence of the two definitions: any system that organizes tasks, coordinates agents, and manages state through graph structures is, by definition, an instance of Graph Engineering. No experiment, theorem, or falsifiable prediction is offered to show that graph structures are necessary or uniquely sufficient; the central claim is built into the definitions rather than independently derived.

  2. self definitional [Section 2.1 Eq. (1); Section 3.5 Limitations of Individual Intelligence]
    "An Individual Agent can therefore be abstracted as Ai = Loop(Fi, Hi; s_i) ... The Agent Harness extends these intrinsic capabilities through interfaces for perception and context construction, memory and knowledge access, tool invocation, reusable skills, and runtime governance. ... Although a single agent can use tools or call specialist models, these capabilities remain coordinated within the same control loop rather than organized into stable and independent roles. ... An individual's context is not an organized or persistent state."

    Equation (1) defines an individual agent to include a harness containing memory, tools, skills, and runtime governance. Under that definition, a single agent whose harness maintains an explicit task graph, spawns parallel tool calls, isolates sub-execution states, and invokes external verifiers is still a single agent. Yet Section 3.5 attributes the limitations of serial execution, role confusion, and fragile state to 'a single-agent loop' and 'an individual's context', which is a narrower reading than the formal definition.

full rationale

There is no equation-level fitted parameter renamed as a prediction, and no load-bearing uniqueness theorem is imported from the authors' prior work; the extensive reference list is survey coverage rather than an argument. The paper is self-contained as a taxonomy, and the surveyed systems exist independently of the proposed framework. The circularity that does exist is at the level of the central definitions: System Intelligence is defined as organizing work, coordinating agents, and maintaining state, while Graph Engineering is defined as the graph-based realization of exactly those three functions, so the headline claim that Graph Engineering is the structural foundation of System Intelligence is analytic rather than empirically established. Similarly, the limitation argument for individual intelligence relies on a narrower reading of Eq. (1) than the paper's own harness definition, since a harness may contain task graphs, parallel executors, isolated state, and verifiers while still being a single agent. The paper's own Sections 5.1 and 7.4 concede that end-task success cannot separate organizational gains from model, context, or compute gains, which further weakens the empirical force of the framework but is an honest limitation rather than concealment. Overall, the central claim has substantial definitional content, but the survey still provides independent descriptive value in organizing existing systems; a score of 4 reflects partial, not total, circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 2 invented entities

No free parameters appear in this survey. The conceptual framework rests on the four assumptions listed above; the first two are domain assumptions about agent architecture, and the third is the paper's own postulate that graphs are the right substrate. The paper's 'System Intelligence' and 'Graph Engineering' are introduced as labels rather than independently evidenced entities.

assumptions (4)
  • domain assumption An LLM-based agent can be abstracted as Agent = Loop(LLM + Harness) with runtime state.
    Section 1 and equation (1) in Section 2.1 present this as the definition of an individual agent; the whole survey builds on this abstraction, but it is a modeling choice rather than an empirical finding.
  • domain assumption Complex real-world tasks exceed the organizational capacity of any single agent, and more context or more capabilities cannot fix this.
    Section 1 and Section 3.5 state this as a fundamental limitation and use it to motivate System Intelligence; no experiment is provided that compares a sufficiently augmented single agent against a multi-agent system.
  • ad hoc to paper Graph structures are a natural and sufficient substrate for organizing tasks, agents, and runtime state.
    Section 4.1 asserts that graphs 'provide a natural structure' and that Graph Engineering is a structure-centered foundation; the paper does not prove representational completeness or compare graphs with non-graph formalisms.
  • domain assumption System Intelligence can be separated from component model quality and measured at the system level.
    Sections 5.1 and 7.4 call for evaluation that separates system organization from model capability, but the paper does not provide a validated metric that achieves this separation.
invented entities (2)
  • System Intelligence
    purpose: A label for the hypothesized system-level capability of coordinating multiple agents into a coherent whole.
    Introduced as a conceptual construct to motivate the survey; it is not independently measured and the paper notes that end-task success cannot distinguish it from stronger models or extra compute.
  • Graph Engineering
    purpose: A proposed engineering paradigm giving a name to the use of explicit graph structures for task organization, agent coordination, and runtime state management.
    A framing device that reorganizes existing graph-based agent systems under one term; it makes no falsifiable prediction that could confirm or refute the paradigm itself.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence." pith.science (2026). https://pith.science/paper/6INSLI3A

@misc{pith2026260821156,
  author       = {Pith},
  title        = {Pith review of: Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6INSLI3A}},
  note         = {Machine review of arXiv:2608.21156}
}
read the original abstract

LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage information access, Harness Engineering to organize external tools and resources, and Loop Engineering to support continual reflection and self-improvement. Yet as tasks grow more complex, individual intelligence faces a fundamental limit: many tasks require heterogeneous expertise, interdependent subtasks, parallel execution, independent verification, and persistent state, exceeding any single agent's organizational capacity. Augmenting one agent's capabilities or context cannot resolve this architectural mismatch; intelligence must instead be distributed across specialized agents and organized at the system level. We call this System Intelligence: an agent system's ability to organize and coordinate multiple intelligent components into a coherent, adaptive whole pursuing a shared objective. Achieving it requires more than adding agents; it demands explicit structures to organize work, coordinate heterogeneous agents, and maintain evolving execution states. We introduce Graph Engineering, an emerging paradigm for next-generation agent systems. Unlike prior paradigms that mainly optimize individual interactions or agent-level behavior, Graph Engineering constructs explicit, dynamic, evolving graph structures representing tasks, agents, and system states. These abstractions provide a unified foundation for organizing complex objectives, orchestrating heterogeneous agents, modeling system dynamics, and enabling scalable agent evolution. We systematically review the principles, methodologies, and applications of Graph Engineering for LLM agents. Related papers, open-source data, and projects are collected at https://github.com/DEEP-JLU/Awesome-Graph-Engineering.

Figures

Figures reproduced from arXiv: 2608.21156 by the authors.

Figure 1
Figure 1. Overview From Model Intelligence to System Intelligence. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. A Taxonomy of Evolving Techniques in the Era of LLM Agents. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. From model intelligence to individual intelligence. [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: An Illustrative Conceptualization of System Intelligence and Its Related Technologies. [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: Overview of Graph Engineering. Task Organization structures the objective into explicit subtasks and executable workflows; Agent Coordination matches capabilities to work, defines team topology, and routes communication among agents; Runtime State Management records ex…
Figure 6
Figure 6. Figure 6: Overview of Task Organization. Goal Decomposition translates a high-level objective into explicit subtasks, exposes their dependencies for scheduling, and refines the remaining task graph using intermediate execution results. Workflow Optimization compiles semantic sub…
Figure 7
Figure 7. Figure 7: Overview of Agent Coordination. The Agent Capability Graph maps agents to their capabilities and accessible resources; the Agent Team Graph organizes agents into task-dependent collaboration structures; and the Communication Graph specifies and adapts information flow …
Figure 8
Figure 8. Figure 8: Overview of Runtime State Management. Runtime State Management structures how to manage state by recording consistent and traceable runtime views, localizing failures from execution evidence, and recovering from validated states. These capabilities turn distributed exe…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

300 extracted references · 167 canonical work pages

  1. [7]

    Orchestrate teams of claude code sessions

    Anthropic. Orchestrate teams of claude code sessions. URLhttps://code.claude.com/docs/en/ agent-teams

  2. [10]

    Claude 3.7 sonnet and claude code

    Anthropic. Claude 3.7 sonnet and claude code. Anthropic, Feb. 2025

  3. [1]

    ESWC Special Track on LLMs for Knowledge Engi- neering; verify authors and final publication metadata before submission

    LLMs4OM: Matching ontologies with large language models, 2024. ESWC Special Track on LLMs for Knowledge Engi- neering; verify authors and final publication metadata before submission

  4. [96]

    Y . Fu, R. Fang, J. Shao, H. Zheng, Z. Zhu, B. Luo, and T. Lin. Do more agents help? controlled and protocol-aligned evaluation of LLM agent workflows.arXiv preprint arXiv:2606.05670, 2026. URLhttps://arxiv.org/abs/2606. 05670

  5. [79]

    C. Emde, A. Rubinstein, A. Goel, A. Heakl, S. Yun, S. J. Oh, and M. Gubri. MASEval: Extending multi-agent evaluation from models to systems. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pages 345–356. Association for Computational Linguistics, 2026. doi: 10.18653/v1/2026.acl-demo. 34

  6. [17]

    MAS-PromptBench: When Does Prompt Optimization Improve Multi-Agent LLM Systems?

    J. Bai and L. Shi. MAS-PromptBench: When does prompt optimization improve multi-agent LLM systems?arXiv preprint arXiv:2606.23664, 2026

  7. [2]

    Agashe, Y

    S. Agashe, Y . Fan, A. Reyna, and X. E. Wang. Llm-coordination: Evaluating and analyzing multi-agent coordination abilities in large language models. InFindings of the Association for Computational Linguistics: NAACL 2025, pages 8053–8072,

  8. [3]

    L. A. Agrawal, S. Tan, D. Soylu, N. Ziems, R. Khare, K. Opsahl-Ong, A. Singhvi, H. Shandilya, M. J. Ryan, M. Jiang, C. Potts, K. Sen, A. G. Dimakis, I. Stoica, D. Klein, M. Zaharia, and O. Khattab. GEPA: Reflective prompt evolution can outperform reinforcement learning.arXiv preprint arXiv:2507.19457, 2025. 35

Show all 300 references
  1. [4]

    Ahn and M

    J. Ahn and M. Kim. From prompts to contracts: Harness engineering for auditable enterprise LLM agents.arXiv preprint arXiv:2607.08028, 2026

  2. [5]

    Alonso, S

    P. Alonso, S. Yovine, and V . A. Braberman. Tdad: Test-driven agentic development - reducing code regressions in ai coding agents via graph-based impact analysis.arXiv preprint arXiv:2603.17973, 2026

  3. [6]

    OpenCode: Agents and subagents

    Anomaly. OpenCode: Agents and subagents. OpenCode documentation, 2026. URLhttps://opencode.ai/docs/ agents/

  4. [8]

    Introducing the model context protocol

    Anthropic. Introducing the model context protocol. Anthropic, Nov. 2024

  5. [9]

    Equipping agents for the real world with agent skills

    Anthropic. Equipping agents for the real world with agent skills. Anthropic Engineering, Oct. 2025

  6. [11]

    Effective harnesses for long-running agents

    Anthropic. Effective harnesses for long-running agents. Anthropic Engineering, Nov. 2025

  7. [12]

    Claude Code: Anthropic’s agentic coding system

    Anthropic. Claude Code: Anthropic’s agentic coding system. Anthropic product documentation, 2026. URLhttps: //www.anthropic.com/product/claude-code

  8. [13]

    Claude Agent SDK for Python

    Anthropic. Claude Agent SDK for Python. GitHub repository and documentation, 2026. URLhttps://github.com/ anthropics/claude-agent-sdk-python

  9. [14]

    Apache Burr: Stateful application and agent framework

    Apache Software Foundation. Apache Burr: Stateful application and agent framework. GitHub repository and documentation,

  10. [15]

    A. Asai, Z. Wu, Y . Wang, A. Sil, and H. Hajishirzi. Self-RAG: Learning to retrieve, generate, and critique through self- reflection. InInternational Conference on Learning Representations, 2024

  11. [16]

    Babaei Giglou, J

    H. Babaei Giglou, J. D’Souza, and S. Auer. LLMs4OL: Large language models for ontology learning. InThe Semantic Web – ISWC 2023, Lecture Notes in Computer Science, 2023. doi: 10.1007/978-3-031-47240-4_22

  12. [18]

    T. Bai, Z. Wan, P. Zhou, X. Yu, Y . You, and I. W. Tsang. Skilldag: Self-evolving typed skill graphs for llm skill selection at scale.arXiv preprint arXiv:2606.03056, 2026

  13. [19]

    X. Bai, H. Lin, C. Liu, Y . Zhang, X. Jin, X. Cao, and Y . Li. SkillZip: Evaluation-free skill compression for self-evolving agents by discovering reusable structure.arXiv preprint arXiv:2608.11079, 2026

  14. [20]

    Y . Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al. Constitutional AI: Harmlessness from AI feedback.arXiv preprint arXiv:2212.08073, 2022

  15. [21]

    Balaji, P

    S. Balaji, P. Mishra, A. Sachdeva, and S. Agrawal. Beyond ivr: Benchmarking customer support llm agents for business- adherence. InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Industry Track), pages 193–208, 2026....

  16. [22]

    Barke, A

    S. Barke, A. Goyal, A. Khare, A. Singh, S. Nath, and C. Bansal. AgentRx: Diagnosing AI agent failures from execution trajectories.arXiv preprint arXiv:2602.02475, 2026. doi: 10.48550/arXiv.2602.02475. URLhttps://arxiv.org/ abs/2602.02475

  17. [23]

    Barres, H

    V . Barres, H. Dong, S. Ray, X. Si, and K. Narasimhan.τ 2-bench: Evaluating conversational agents in a dual-control environ- ment.arXiv preprint arXiv:2506.07982, 2025

  18. [24]

    Barres, H

    V . Barres, H. Dong, S. Ray, X. Si, and K. Narasimhan.τ 2-Bench: Evaluating conversational agents in a dual-control environment, 2025

  19. [25]

    Y . Bei, W. Zhang, S. Wang, W. Chen, S. Zhou, H. Chen, Y . Li, J. Bu, S. Pan, Y . Yu, I. King, F. Karray, and P. S. Yu. Graphs meet AI agents: Taxonomy, progress, and future opportunities.arXiv preprint arXiv:2506.18019, 2025. URL https://arxiv.org/abs/2506.18019. 36

  20. [26]

    Y . Bei, T. Wei, X. Ning, Y . Zhao, Z. Liu, X. Lin, Y . Zhu, H. Hamann, J. He, and H. Tong. Mem-gallery: Benchmarking multimodal long-term conversational memory for mllm agents. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1...

  21. [27]

    Belov, A

    V . Belov, A. Sosedka, A. Sakhovskiy, E. Kovtun, A. Boyarskikh, and S. Budennyy. Llm agents factory: Retrieval of domain- specific llm agents. InProceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 4474–4479, 2026

  22. [28]

    Besta, N

    M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, M. Podstawski, L. Gianinazzi, J. Gajda, T. Lehmann, H. Niewiadomski, P. Nyczyk, and T. Hoefler. Graph of thoughts: Solving elaborate problems with large language models. InProceedings of the AAAI Conference on Artificial Intell...

  23. [29]

    Bonagiri, D

    A. Bonagiri, D. Borkar, G. J. Anderias, S. Rafatirad, and H. Homayoun. Causalflow: Causal attribution and counterfactual repair for llm agent failures.arXiv preprint arXiv:2605.25338, 2026

  24. [30]

    L. C. Borro, L. A. B. Macarini, G. Tindall, M. Montero, and A. B. Struck. Memori: A persistent memory layer for efficient, context-aware llm agents.arXiv preprint arXiv:2603.19935, 2026

  25. [31]

    T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners. InAdvances in Neural Information Processing Systems, volume 33, pages 1877–1901, 2020

  26. [32]

    S. Cao, J. He, and F. Tan. Higmem: A hierarchical and llm-guided memory system for long-term conversational agents. In Findings of the Association for Computational Linguistics: ACL 2026, pages 33853–33862, 2026

  27. [33]

    J. H. Caufield, H. Hegde, V . Emonet, N. L. Harris, M. P. Joachimiak, N. Matentzoglu, H. Kim, S. Moxon, J. Reese, M. A. Haendel, P. N. Robinson, and C. J. Mungall. Structured prompt interrogation and recursive extraction of semantics (SPIRES): A method for populating knowledge...

  28. [34]

    Cemri, M

    M. Cemri, M. Z. Pan, S. Yang, L. A. Agrawal, B. Chopra, R. Tiwari, K. Keutzer, A. Parameswaran, D. Klein, K. Ramchandran, M. A. Zaharia, J. E. Gonzalez, and I. Stoica. Why do multi-agent LLM systems fail? InAdvances in Neural Information Processing Systems, volume 38, 2025. do...

  29. [35]

    E. Y . Chang and L. Geng. Sagallm: Context management, validation, and transaction guarantees for multi-agent LLM planning.Proceedings of the VLDB Endowment, 18(12):4874–4886, 2025. doi: 10.14778/3750601.3750611

  30. [36]

    G. Chen, S. Dong, Y . Shu, G. Zhang, J. Sesay, B. F. Karlsson, J. Fu, and Y . Shi. Autoagents: A framework for automatic agent generation.arXiv preprint arXiv:2309.17288, 2023

  31. [37]

    H. Chen, X. Song, J. Jin, P. Ren, and L.-J. Zhang. Toward an organizational science of multi-agent llm systems: Decoupling who, how, and which algorithm.arXiv preprint arXiv:2607.25446, 2026

  32. [38]

    J. Chen, A. Prasad, S. Saha, E. Stengel-Eskin, and M. Bansal. Magicore: Multi-agent, iterative, coarse-to-fine refinement for reasoning. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 32651– 32674, 2025

  33. [39]

    M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y . Burda, N. Joseph, G. Brockman, et al. Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374, 2021. URLhttps://arxiv.org/ abs/2107.03374

  34. [40]

    M. Chen, J. Wang, Z. Liu, Y . Wang, and Q. Wang. From failed trajectories to reliable LLM agents: Diagnosing and repairing harness flaws.arXiv preprint arXiv:2606.06324, 2026

  35. [41]

    M. Chen, J. Wang, F. Mu, Y . Wang, Z. Liu, H. Feng, and Q. Wang. Seeing the whole elephant: A benchmark for failure attribution in llm-based multi-agent systems. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics, pages 19888–19905, 2026....

  36. [42]

    S. Chen, C. Zhou, Z. Yuan, Q. Zhang, Z. Cui, H. Chen, Y . Xiao, J. Cao, and X. Huang. You don’t need pre-built graphs for rag: Retrieval augmented generation with adaptive reasoning structures. InProceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 3...

  37. [43]

    W. Chen, Y . Su, J. Zuo, C. Yang, C. Yuan, C.-M. Chan, H. Yu, Y . Lu, Y .-H. Hung, C. Qian, et al. Agentverse: Facilitating multi- agent collaboration and exploring emergent behaviors. InInternational Conference on Learning Representations, volume 2024, pages 20094–20136, 2024. 37

  38. [44]

    W. Chen, Z. You, R. Li, C. Qian, C. Zhao, C. Yang, R. Xie, Z. Liu, M. Sun, et al. Internet of agents: Weaving a web of heterogeneous agents for collaborative intelligence. InInternational Conference on Learning Representations, volume 2025, pages 36374–36411, 2025

  39. [45]

    W. Chen, D. Yao, W. Li, X. Meng, C. Gong, and J. Bi. Gtool: Graph enhanced tool planning with large language model. In International Conference on Learning Representations, volume 2026, pages 69247–69269, 2026

  40. [46]

    X. Chen, H. Yi, M. You, W. Liu, L. Wang, H. Li, X. Zhang, Y . Guo, L. Fan, G. Chen, Q. Lao, W. Fu, K. Li, and J. Li. Enhancing diagnostic capability with multi-agents conversational large language models.npj Digital Medicine, 8:159, 2025. doi: 10.1038/s41746-025-01550-0. URLht...

  41. [47]

    Y . Chen, H. Lai, Y . Feng, C. Han, Q. Zhang, B. Lu, M. Li, X. Wang, Z. Wang, S. Xu, Z. Li, Z. Jin, H. Wu, C. Li, and Q. Chen. Beyond semantic organization: Memory as execution state management for long-horizon agents.arXiv preprint arXiv:2606.06090, 2026

  42. [48]

    Z. Chen, H. Liu, D. Xu, D. Dong, J. Li, B. Pu, and J. Zhai. Cordon: Semantic transactions for tool-using llm agents.arXiv preprint arXiv:2606.17573, 2026

  43. [49]

    Z. Chen, Z. Peng, X. Liang, C. Wang, P. Liang, L. Zeng, M. Ju, and Y . Yuan. MAP: Evaluation and multi-agent enhancement of large language models for inpatient pathways.npj Health Systems, 3:37, 2026. doi: 10.1038/s44401-026-00085-0. URL https://doi.org/10.1038/s44401-026-00085-0

  44. [50]

    Z. Chen, Q. Zhang, Z. Xiang, Z. Wei, L. Gao, X. Huang, Z. Zhang, and J. Su. Legalgraphrag: Multi-agent graph retrieval- augmented generation for reliable legal reasoning. InProceedings of the 64th Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Lon...

  45. [51]

    Cheng, J

    M. Cheng, J. Ouyang, S. Yu, R. Yan, Y . Luo, Z. Liu, D. Wang, Q. Liu, and E. Chen. Agent-r1: Training powerful llm agents with end-to-end reinforcement learning.arXiv preprint arXiv:2511.14460, 2025

  46. [52]

    Chhikara, D

    P. Chhikara, D. Khant, S. Aryan, T. Singh, and D. Yadav. Mem0: Building production-ready AI agents with scalable long- term memory.arXiv preprint arXiv:2504.19413, 2025

  47. [53]

    Chirkova, T

    N. Chirkova, T. Formal, V . Nikoulina, and S. Clinchant. Provence: Efficient and robust context pruning for retrieval- augmented generation. InInternational Conference on Learning Representations, 2025

  48. [54]

    Chowdhery, S

    A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al. PaLM: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1–113, 2023

  49. [55]

    H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y . Tay, W. Fedus, Y . Li, X. Wang, M. Dehghani, S. Brahma, A. Webson, S. S. Gu, Z. Dai, M. Suzgun, X. Chen, A. Chowdhery, et al. Scaling instruction-finetuned language models.Journal of Machine Learning Research, 25(70):1–53, 2024

  50. [56]

    Cline: Multi-agent teams

    Cline. Cline: Multi-agent teams. Cline documentation, 2026. URLhttps://docs.cline.bot/sdk/guides/ multi-agent-teams

  51. [57]

    CrewAI: Multi-agent automation framework

    CrewAI, Inc. CrewAI: Multi-agent automation framework. GitHub repository and documentation, 2026. URLhttps: //github.com/crewAIInc/crewAI

  52. [58]

    D. Dai, C. Deng, C. Zhao, R. Xu, H. Gao, D. Chen, J. Li, W. Zeng, X. Yu, Y . Wu, Z. Xie, Y . K. Li, P. Huang, F. Luo, C. Ruan, Z. Sui, and W. Liang. DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models. In Proceedings of the 62nd Annual Mee...

  53. [59]

    Y . Dang, C. Qian, X. Luo, J. Fan, Z. Xie, R. Shi, W. Chen, C. Yang, X. Che, Y . Tian, et al. Multi-agent collaboration via evolving orchestration.Advances in neural information processing systems, 38:165025–165059, 2026

  54. [60]

    S. O. de Macedo. What makes a harness a harness: Necessary and sufficient conditions for an agent harness.arXiv preprint arXiv:2606.10106, 2026

  55. [61]

    Debenedetti, J

    E. Debenedetti, J. Zhang, M. Balunovi ´c, L. Beurer-Kellner, M. Fischer, and F. Tramèr. AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents.arXiv preprint arXiv:2406.13352, 2024

  56. [62]

    Debenedetti, I

    E. Debenedetti, I. Shumailov, T. Fan, J. Hayes, N. Carlini, D. Fabian, C. Kern, C. Shi, A. Terzis, and F. Tramèr. Defeating prompt injections by design.arXiv preprint arXiv:2503.18813, 2025

  57. [63]

    DeepSeek-V3 technical report.arXiv preprint arXiv:2412.19437, 2024

    DeepSeek-AI. DeepSeek-V3 technical report.arXiv preprint arXiv:2412.19437, 2024. 38

  58. [64]

    Haystack: Open-source ai orchestration framework

    deepset. Haystack: Open-source ai orchestration framework. GitHub repository and documentation, 2026. URLhttps: //github.com/deepset-ai/haystack

  59. [65]

    X. Deng, J. Da, E. Pan, Y . Y . He, C. Ide, K. Garg, N. Lauffer, A. Park, N. Pasari, C. Rane, K. Sampath, M. Krishnan, S. Kundurthy, S. Hendryx, Z. Wang, C. B. C. Zhang, N. Jacobson, B. Liu, and B. Kenstler. SWE-Bench Pro: Can ai agents solve long-horizon software engineering ...

  60. [66]

    H. Ding, P. Liu, J. Wang, Z. Ji, M. Cao, R. Zhang, L. Ai, E. Yang, T. Shi, and L. Yu. DynaWeb: Model-based reinforcement learning of web agents.arXiv preprint arXiv:2601.22149, 2026

  61. [67]

    Dochkina

    V . Dochkina. Drop the hierarchy and roles: How self-organizing llm agents outperform designed structures.arXiv preprint arXiv:2603.28990, 2026

  62. [68]

    G. Dong, H. Yuan, K. Lu, C. Li, M. Xue, D. Liu, W. Wang, Z. Yuan, C. Zhou, and J. Zhou. How abilities in large language models are affected by supervised fine-tuning data composition. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vo...

  63. [69]

    G. Dong, L. Bao, Z. Wang, K. Zhao, X. Li, J. Jin, J. Yang, H. Mao, F. Zhang, K. Gai, et al. Agentic entropy-balanced policy optimization.arXiv preprint arXiv:2510.14545, 2025

  64. [70]

    G. Dong, J. Lu, J. Huang, W. Zhong, L. Liu, S. Huang, Z. Li, Y . Zhao, X. Song, X. Li, et al. Agent-world: Scaling real-world environment synthesis for evolving general agent intelligence.arXiv preprint arXiv:2604.18292, 2026

  65. [71]

    G. Dong, X. Song, Y . Hu, J. Jin, C. Zhang, Y . Chen, X. Li, H. Yuan, X. Yang, T. Wen, et al. Towards long-horizon agents: A survey. 2026

  66. [72]

    Dong, I.-W

    J.-K. Dong, I.-W. Huang, C.-T. Wu, and Y .-t. Tsai. Etom: A five-level benchmark for evaluating tool orchestration within the mcp ecosystem. InFindings of the Association for Computational Linguistics: EACL 2026, pages 1453–1488, 2026. doi: 10.18653/v1/2026.findings-eacl.75. U...

  67. [73]

    Y . Dong, X. Zhu, Z. Pan, L. Zhu, and Y . Yang. Villageragent: A graph-based multi-agent framework for coordinating complex task dependencies in minecraft. InFindings of the Association for Computational Linguistics: ACL 2024, 2024. doi: 10.18653/v1/2024.findings-acl.964. URLh...

  68. [74]

    Y . Dong, J. He, Y . Hou, D. Du, Z. Xu, S. Yu, Y . Xia, and H. Chen. DeltaBox: Scaling stateful AI agents with millisecond- level sandbox checkpoint/rollback.arXiv preprint arXiv:2605.22781, 2026. doi: 10.48550/arXiv.2605.22781. URLhttps: //arxiv.org/abs/2605.22781

  69. [75]

    Y . Du, Y . Wang, H. Xu, J. Xu, S. Tan, B. Zhao, B. Yang, Z. Xu, M. Kong, H. Wei, J. Liu, and Q. Zhu. Living-harness is an interactive-agent evolver.arXiv preprint arXiv:2607.26598, 2026

  70. [76]

    W. Duan, J. Lu, and J. Xuan. Bayesian ego-graph inference for networked multi-agent reinforcement learning.Advances in Neural Information Processing Systems, 38:73072–73105, 2026

  71. [77]

    D. Edge, H. Trinh, N. Cheng, J. Bradley, A. Chao, A. Mody, S. Truitt, and J. Larson. From local to global: A graph RAG approach to query-focused summarization.arXiv preprint arXiv:2404.16130, 2024

  72. [78]

    Ekelhart, K

    A. Ekelhart, K. Kurniawan, F. J. Ekaputra, and E. Kiesling. Agento: An ontology for modeling agentic ai systems. In European Semantic Web Conference, pages 298–320. Springer, 2026

  73. [80]

    prompt engineering

    EvoMap. Evomap: From “prompt engineering” to “epigenetic engineering”. URLhttps://evomap.ai/blog/ epigenetic-engineering

  74. [81]

    R. Fang, Y . Liang, X. Wang, J. Wu, S. Qiao, P. Xie, F. Huang, H. Chen, and N. Zhang. Memp: Exploring agent procedural memory. InFindings of the Association for Computational Linguistics: ACL 2026, pages 17490–17502, 2026

  75. [82]

    Fathallah, M

    A. Fathallah, M. Das, S. Giorgis, A. Poltronieri, P. Haase, and N. Kovriguina. NeOn-GPT: A large language model-powered pipeline for ontology learning. InThe Semantic Web – ESWC 2025, Lecture Notes in Computer Science, 2025. doi: 10.1007/ 978-3-031-78952-6_4

  76. [83]

    Fedus, B

    W. Fedus, B. Zoph, and N. Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022. 39

  77. [84]

    J. Feng, S. Huang, X. Qu, G. Zhang, Y . Qin, B. Zhong, C. Jiang, J. Chi, and W. Zhong. Retool: Reinforcement learning for strategic tool use in llms.arXiv preprint arXiv:2504.11536, 2025

  78. [85]

    J. Feng, Y . Yu, A. Dong, X. Hu, S. Niu, P. Li, Y . Dang, A. Abdelhameed, J. Bian, X. Jiang, et al. Ontocodex: a multi-agent biomedical ontology enrichment framework.npj Health Systems, 3(1):73, 2026

  79. [86]

    L. Feng, Z. Xue, T. Liu, and B. An. ToRL: Scaling tool-integrated reinforcement learning.arXiv preprint arXiv:2503.23383, 2025

  80. [87]

    S. Feng, Z. Wang, P. Goyal, Y . Wang, W. Shi, H. Xia, H. Palangi, L. Zettlemoyer, Y . Tsvetkov, C.-Y . Lee, et al. Heterogeneous swarms: Jointly optimizing model roles and weights for multi-llm systems.Advances in Neural Information Processing Systems, 38:114319–114351, 2026

  81. [88]

    T. Feng, H. Zhang, Z. Lei, P. Han, and J. You. Graphplanner: Graph memory-augmented agentic routing for multi-agent llms. arXiv preprint arXiv:2604.23626, 2026

  82. [89]

    X. Feng, X. Song, L. Li, G. Liu, and J. Shao. Searl: Joint optimization of policy and tool graph memory for self-evolving agents. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 24518–24535, 2026

  83. [90]

    Y . Feng, J. Sun, Z. Yang, J. Ai, C. Li, Z. Li, F. Zhang, K. He, R. Ma, J. Lin, J. Sun, Y . Xiao, S. Zhou, W. Wu, Y . Liu, P. Liu, S. Zhang, and K. Zhang. LongCLI-Bench: A preliminary benchmark and study for long-horizon agentic programming in command-line interfaces. InFindin...

  84. [91]

    Fernando, D

    C. Fernando, D. Banarse, H. Michalewski, S. Osindero, and T. Rocktäschel. Promptbreeder: Self-referential self- improvement via prompt evolution. InInternational Conference on Learning Representations, 2024

  85. [92]

    Fourney, G

    A. Fourney, G. Bansal, H. Mozannar, C. Tan, E. Salinas, F. Niedtner, G. Proebsting, G. Bassman, J. Gerrits, J. Alber, et al. Magentic-one: A generalist multi-agent system for solving complex tasks.arXiv preprint arXiv:2411.04468, 2024

  86. [93]

    D. Fu, J. Mei, R. Wu, X. Yang, J. Xu, D. Wang, P. Cai, Y . Liu, L. Wen, and B. Shi. The agent’s first day: Benchmarking learn- ing, exploration, and scheduling in the workplace scenarios. InFindings of the Association for Computational Linguistics: ACL 2026, pages 30094–30109, 2026

  87. [94]

    D. Fu, J. Mei, R. Wu, X. Yang, J. Xu, D. Wang, P. Cai, Y . Liu, L. Wen, and B. Shi. The agent’s first day: Benchmarking learn- ing, exploration, and scheduling in the workplace scenarios. InFindings of the Association for Computational Linguistics: ACL 2026, pages 30094–30109....

  88. [95]

    K. Fu, L. Lyu, S. Li, S. Huang, S. Xu, J. Zheng, X. Liu, S. Liu, G. Barbone, Y . Liu, et al. Agentic laboratories of the future: Towards world models for scientific discovery. 2026

  89. [97]

    Galster, S

    M. Galster, S. Mohsenimofidi, J. L. Lulla, M. A. Abubakar, C. Treude, and S. Baltes. Configuring agentic AI coding tools: An exploratory study.arXiv preprint arXiv:2602.14690, 2026

  90. [98]

    H.-a. Gao, J. Geng, W. Hua, M. Hu, X. Juan, H. Liu, S. Liu, J. Qiu, X. Qi, Y . Wu, H. Wang, H. Xiao, Y . Zhou, S. Zhang, J. Zhang, J. Xiang, Y . Fang, Q. Zhao, D. Liu, Q. Ren, C. Qian, Z. Wang, M. Hu, H. Wang, Q. Wu, H. Ji, and M. Wang. A survey of self-evolving agents: What, ...

  91. [99]

    J. Gao, X. Zou, Y . Ai, D. Li, Y . Niu, B. Qi, and J. Liu. Graph counselor: Adaptive graph exploration via multi-agent synergy to enhance llm reasoning. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 246...

  92. [100]

    L. Gao, X. Ma, J. Lin, and J. Callan. Precise zero-shot dense retrieval without relevance labels. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1762–1777. Association for Computational Linguistics, 2023

  93. [101]

    Geng and E

    L. Geng and E. Y . Chang. Alas: Transactional and dynamic multi-agent llm planning.arXiv preprint arXiv:2511.03094, 2025. 40

  94. [102]

    Ghafarollahi and M

    A. Ghafarollahi and M. J. Buehler. SciAgents: Automating scientific discovery through bioinspired multi-agent intelligent graph reasoning.Advanced Materials, 37(22):2413523, 2025. doi: 10.1002/adma.202413523. URLhttps://doi.org/ 10.1002/adma.202413523

  95. [103]

    A. E. Ghareeb, B. Chang, L. Mitchener, A. Yiu, C. J. Szostkiewicz, D. Shved, G. J. Gyimesi, J. M. Laurent, S. M. Wright, M. T. Razzak, A. D. White, S. C. Finnemann, M. M. Hinks, and S. G. Rodriques. A multi-agent system for automating scientific discovery.Nature, 655:497–505, ...

  96. [104]

    Github copilot: Meet the new coding agent

    GitHub. Github copilot: Meet the new coding agent. GitHub Blog, May 2025

  97. [105]

    Gemini cli: Your open-source ai agent

    Google. Gemini cli: Your open-source ai agent. Google, June 2025

  98. [106]

    Agent Development Kit: An open-source framework for ai agents

    Google. Agent Development Kit: An open-source framework for ai agents. GitHub repository and documentation, 2026. URLhttps://github.com/google/adk-python

  99. [107]

    Gottweis, W.-H

    J. Gottweis, W.-H. Weng, A. Daryin, T. Tu, P. Sirkovic, A. Myaskovsky, G. Glowaty, F. Weissenberger, A. Orlandi, et al. Accelerating scientific discovery with Co-Scientist.Nature, 655:487–496, 2026. doi: 10.1038/s41586-026-10644-y. URL https://doi.org/10.1038/s41586-026-10644-y

  100. [108]

    Gou et al

    Z. Gou et al. CRITIC: Large language models can self-correct with tool-interactive critiquing.arXiv preprint arXiv:2305.11738, 2023. doi: 10.48550/arXiv.2305.11738. URLhttps://arxiv.org/abs/2305.11738

  101. [109]

    Grattafiori, A

    A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024

  102. [110]

    Grötschla, L

    F. Grötschla, L. Müller, J. Tönshoff, M. Galkin, and B. Perozzi. Agentsnet: Coordination and collaborative reasoning in multi-agent llms.arXiv preprint arXiv:2507.08616, 2025. URLhttps://arxiv.org/abs/2507.08616

  103. [111]

    T. R. Gruber. A translation approach to portable ontology specifications.Knowledge Acquisition, 5(2):199–220, 1993. doi: 10.1006/knac.1993.1008

  104. [112]

    J.-C. Gu, J. Zhang, D. Wu, Y . Li, K.-W. Chang, and N. Peng. BRIEF-Pro: Universal context compression with short-to-long synthesis for fast and accurate multi-hop reasoning. InFindings of the Association for Computational Linguistics: ACL 2026, pages 14221–14241. Association f...

  105. [113]

    D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning.Nature, 645(8081):633–638, 2025

  106. [114]

    D. Guo, J. Wu, and S. M. Yiu. When to retrieve during reasoning: Adaptive retrieval for large reasoning models.arXiv preprint arXiv:2604.26649, 2026

  107. [115]

    J. Guo, Z. Hao, C. Wang, C. Fan, T. Luo, H. Li, Y . Gao, H. Mei, J. Peng, R. Xu, M. Dong, H. Wu, M. Zheng, K. Han, S. Wang, C. Xu, and Y . Wang. From question answering to task completion: A survey on agent system and harness design, 2026

  108. [116]

    Q. Guo, R. Wang, J. Guo, B. Li, K. Song, X. Tan, G. Liu, J. Bian, and Y . Yang. Connecting large language models with evolutionary algorithms yields powerful prompt optimizers. InInternational Conference on Learning Representations, 2024

  109. [117]

    S. Guo, Y . Wang, Z. Su, Y . Pan, Q. Hu, and T. H. Luan. Agent discovery in internet of agents: Challenges and solutions. IEEE Network, 2026

  110. [118]

    X. Guo, X. Wang, Y . Chen, S. Li, C. Han, M. Li, and H. Ji. Syncmind: Measuring agent out-of-sync recovery in collaborative software engineering. InProceedings of the 42nd International Conference on Machine Learning, volume 267 ofProceed- ings of Machine Learning Research, pa...

  111. [119]

    J. Han, Y . Xu, Y . Liao, X. Wang, Z. Jiang, Z. Di, F. Lu, Z. Hu, and Y . Xiao. Skill-use: Can LLMs actually use skills in agentic harnesses?arXiv preprint arXiv:2608.04828, 2026

  112. [120]

    Z. Hao, H. Wang, J. Luo, J. Zhang, Y . Zhou, Q. Lin, C. Wang, H. Dong, and J. Chen. Recreate: Reasoning and creating domain agents driven by experience. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 310...

  113. [121]

    Z. Hao, T. Wang, H. Dong, Z. Liu, H. Wang, X. Lin, Q. Lin, C. Wang, H. Dong, and J. Chen. Evolve as a team: Collaborative self-evolution for llm-based multi-agent systems.arXiv preprint arXiv:2605.29790, 2026. 41

  114. [122]

    Hasan and P

    N. Hasan and P. BusiReddyGari. Dpbench: Large language models struggle with simultaneous coordination.arXiv preprint arXiv:2602.13255, 2026. URLhttps://arxiv.org/abs/2602.13255

  115. [123]

    He and D

    J. He and D. Yu. Sovereign agentic loops: Decoupling AI reasoning from execution in real-world systems.arXiv preprint arXiv:2604.22136, 2026. doi: 10.48550/arXiv.2604.22136. URLhttps://arxiv.org/abs/2604.22136

  116. [124]

    Y . He, J. Chen, D. Antonyrajah, and I. Horrocks. BERTMap: A BERT-based ontology alignment system. InProceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 5684–5691, 2022. doi: 10.1609/aaai.v36i5.20510

  117. [125]

    Y . He, J. Chen, H. Dong, I. Horrocks, C. Allocca, T. Kim, and B. Sapkota. DeepOnto: A python package for ontology engineering with deep learning.Semantic Web, 15(5):1991–2004, 2024. doi: 10.3233/SW-243568

  118. [126]

    Z. He, Y . Wang, C. Zhi, Y . Hu, T.-P. Chen, L. Yin, Z. Chen, T. A. Wu, S. Ouyang, Z. Wang, J. Pei, J. McAuley, Y . Choi, and A. Pentland. Memoryarena: Benchmarking agent memory in interdependent multi-session agentic tasks.arXiv preprint arXiv:2602.16313, 2026

  119. [127]

    Hendrycks, C

    D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding. InInternational Conference on Learning Representations, 2021. URLhttps://arxiv.org/abs/ 2009.03300

  120. [128]

    Hoffmann, S

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifr...

  121. [129]

    S. Hong, M. Zhuge, J. Chen, X. Zheng, Y . Cheng, J. Wang, C. Zhang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, C. Ran, L. Xiao, C. Wu, and J. Schmidhuber. MetaGPT: Meta programming for a multi-agent collaborative framework. InThe Twelfth International Conference on Learning Repre...

  122. [130]

    Z. Hong, Z. Yuan, Q. Zhang, H. Chen, J. Dong, F. Huang, and X. Huang. Next-generation database interfaces: A survey of llm-based text-to-sql.IEEE Transactions on Knowledge and Data Engineering, 2025

  123. [131]

    X. Hou, S. Wang, Y . Zhao, and H. Wang. When agents do not stop: Uncovering infinite agentic loops in llm agents.arXiv preprint arXiv:2607.01641, 2026

  124. [132]

    X. Hou, S. Wang, Y . Zhao, and H. Wang. When agents do not stop: Uncovering infinite agentic loops in LLM agents.arXiv preprint arXiv:2607.01641, 2026. doi: 10.48550/arXiv.2607.01641. URLhttps://arxiv.org/abs/2607.01641

  125. [133]

    M. Hu, T. Chen, Q. Chen, Y . Mu, W. Shao, and P. Luo. HiAgent: Hierarchical working memory management for solving long-horizon agent tasks with large language model.arXiv preprint arXiv:2408.09559, 2024

  126. [134]

    M. Hu, Y . Zhou, W. Fan, Y . Nie, B. Xia, T. Sun, Z. Ye, Z. Jin, Y . Li, Q. Chen, et al. Owl: Optimized workforce learning for general multi-agent assistance in real-world task automation.arXiv preprint arXiv:2505.23885, 2025

  127. [135]

    S. Hu, C. Lu, and J. Clune. Automated design of agentic systems. InInternational Conference on Learning Representations,

  128. [136]

    W. Hu. From agent loops to structured graphs: A scheduler-theoretic framework for LLM agent execution.arXiv preprint arXiv:2604.11378, 2026. URLhttps://arxiv.org/abs/2604.11378

  129. [137]

    Y . Hu, Q. Zhou, Q. Chen, X. Li, L. Liu, D. Zhang, A. Kachroo, T. Oz, and O. Tripp. Qualityflow: An agentic workflow for program synthesis controlled by LLM quality checks.arXiv preprint arXiv:2501.17167, 2025. URLhttps://arxiv. org/abs/2501.17167

  130. [138]

    URLhttps://arxiv.org/abs/2408.08435

  131. [139]

    Huang, Y

    D. Huang, Y . Ding, B. Liu, Q. Liu, X. Chen, J. Bian, H. Sun, Z. Tu, D. Chu, X. Yu, and D. Sui. SkillWiki: A living knowledge infrastructure for agent skills.arXiv preprint arXiv:2606.16523, 2026

  132. [140]

    Huang, F

    J. Huang, F. Cheng, J. Jiang, Z. Yu, and A. Aizawa. BenchTrace: A benchmark for testing reflection ability and controlled evolution in LLM agents.arXiv preprint arXiv:2605.29225, 2026. URLhttps://arxiv.org/abs/2605.29225

  133. [141]

    Y . Hu, Y . Wang, and J. McAuley. Evaluating memory in llm agents via incremental multi-turn interactions. InThe Four- teenth International Conference on Learning Representations, 2026. URLhttps://openreview.net/forum?id= DT7JyQC3MR

  134. [142]

    Huang, Z

    J. Huang, Z. Zhang, K. Shi, Y . Ye, and C. Zhang. Evolverouter: Co-evolving routing and prompt for multi-agent question answering.arXiv preprint arXiv:2604.05149, 2026

  135. [143]

    Huang, C

    L. Huang, C. Yang, H. Zhou, H. Song, Z. Chen, R. Le, Y . Song, W. X. Zhao, and T. Zhang. Evo-Bench: Can language models improve agent harness?arXiv preprint arXiv:2608.09096, 2026

  136. [144]

    Huang, J

    J. Huang, J. Hsia, J. Sun, F. Shi, W. Huang, and I. H. White. Proof-or-stop: Don’t trust the agent, trust the evidence – loop engineering for verifiable evidence-gated lifecycle control, 2026. URLhttps://arxiv.org/abs/2607.14890. 42

  137. [145]

    Y . In, M. Tanjim, J. Subramanian, S. Kim, U. Bhattacharya, W. Kim, S. Park, S. Sarkhel, and C. Park. Rethinking failure attribution in multi-agent systems: A multi-perspective benchmark and evaluation.arXiv preprint arXiv:2603.25001, 2026. URLhttps://arxiv.org/abs/2603.25001

  138. [146]

    Izacard and E

    G. Izacard and E. Grave. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 874–880. Association for Computat...

  139. [147]

    Huang, H

    Z. Huang, H. Que, H. Zeng, G. Zhang, Z. Wang, J. Chen, H. Wang, Z. Hou, C. Pu, S. Yan, and W. Huang. Harness-IF: Evaluating instruction following across instruction surfaces in coding agents.arXiv preprint arXiv:2608.11727, 2026

  140. [148]

    S. Ji, Y . Li, and B. Hooi. Memory is reconstructed, not retrieved: Graph memory for llm agents.arXiv preprint arXiv:2606.06036, 2026

  141. [149]

    Q. Jia, Y . Shen, X. Song, K. Zhang, S. Wang, D. Pei, X. Zhu, and G. Zhai. One battle after another: Probing LLMs’ limits on multi-turn instruction following with a benchmark evolving framework. InProceedings of the 64th Annual Meeting of the Association for Computational Ling...

  142. [150]

    J. Ji, Y . Li, H. Liu, Z. Du, Z. Wei, Q. Qi, W. Shen, and Y . Lin. SRAP-Agent: Simulating and optimizing scarce resource allocation policy with LLM-based agent. InFindings of the Association for Computational Linguistics: EMNLP 2024, pages 267–293. Association for Computationa...

  143. [151]

    Jiang, Q

    H. Jiang, Q. Wu, C.-Y . Lin, Y . Yang, and L. Qiu. LLMLingua: Compressing prompts for accelerated inference of large language models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13358–13376. Association for Computational Ling...

  144. [152]

    Jiang, L

    S. Jiang, L. Ma, Z. Hong, K. Wang, Z. Lu, T. Wang, S. Chen, J. Zhang, T. Pan, W. Li, J. Liang, and Y . Xiao. SEA-Eval: A benchmark for evaluating self-evolving agents beyond episodic assessment.arXiv preprint arXiv:2604.08988, 2026

  145. [153]

    A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. Bou Hanna, F. Bressand, et al. Mixtral of experts.arXiv preprint arXiv:2401.04088, 2024

  146. [154]

    C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan. SWE-bench: Can language models resolve real-world GitHub issues? InInternational Conference on Learning Representations, 2024. URLhttps://arxiv.org/ abs/2310.06770

  147. [155]

    B. Jin, H. Zeng, Z. Yue, D. Wang, H. Zamani, and J. Han. Search-r1: Training llms to reason and leverage search engines with reinforcement learning.arXiv preprint arXiv:2503.09516, 2025

  148. [156]

    Jiang, F

    Z. Jiang, F. Huang, H. Xing, X. Wu, Y . Gao, R. Cao, M. Wang, S. Liu, and Y . Li. Demystifying agent skills: Why they work—until they don’t.arXiv preprint arXiv:2608.14036, 2026

  149. [157]

    Kadu and A

    A. Kadu and A. Krishnan. Reflexgrad: Within-episode failure recovery in llm agents via progress-gated dual-process routing. arXiv preprint arXiv:2511.14584, 2025

  150. [158]

    J. Kang, M. Ji, Z. Zhao, and T. Bai. Memory OS of AI agent.arXiv preprint arXiv:2506.06326, 2025

  151. [159]

    Y . Jin, K. Sharma, V . Rakesh, Y . Dou, M. Pan, M. Das, and S. Kumar. SARA: Selective and adaptive retrieval-augmented generation with context compression. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages...

  152. [160]

    Kaplan, S

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361, 2020. 43

  153. [161]

    Karpas, O

    E. Karpas, O. Abend, Y . Belinkov, B. Lenz, O. Lieber, N. Ratner, Y . Shoham, H. Bata, Y . Levine, K. Leyton-Brown, D. Muhl- gay, N. Rozen, E. Schwartz, G. Shachaf, S. Shalev-Shwartz, A. Shashua, and M. Tenenholtz. MRKL systems: A modular, neuro-symbolic architecture that comb...

  154. [162]

    Kang, W.-N

    M. Kang, W.-N. Chen, D. Han, H. A. Inan, L. Wutschitz, Y . Chen, R. Sim, and S. Rajmohan. ACON: Optimizing context compression for long-horizon LLM agents. InInternational Conference on Machine Learning, 2026. arXiv:2510.00615

  155. [163]

    Kavathekar, H

    I. Kavathekar, H. Jain, A. Rathod, P. Kumaraguru, and T. Ganu. Tamas: Benchmarking adversarial risks in multi-agent llm systems. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 31238–31268, 2026. doi: 10....

  156. [164]

    Z. Ke, Y . Ming, A. Xu, R. Chin, X.-P. Nguyen, P. Jwalapuram, J. Wang, S. Yavuz, C. Xiong, and S. Joty. Mas-orchestra: Understanding and improving multi-agent reasoning through holistic orchestration and controlled benchmarks.arXiv preprint arXiv:2601.14652, 2026

  157. [165]

    Karpukhin, B

    V . Karpukhin, B. O˘guz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W.-t. Yih. Dense passage retrieval for open-domain question answering. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, pages 6769–6781. Association for Computati...

  158. [166]

    S. Kim, S. Moon, R. Tabrizi, N. Lee, M. W. Mahoney, K. Keutzer, and A. Gholami. An LLM compiler for parallel function calling. InInternational Conference on Machine Learning, 2024. URLhttps://arxiv.org/abs/2312.04511

  159. [167]

    S. Kim, I. Bang, S. Jang, C. Kim, S. Bae, J. Choi, R. Xuan, and T. Kim. OMHBench: Benchmarking balanced and grounded omni-modal multi-hop reasoning. InFindings of the Association for Computational Linguistics: ACL 2026, pages 18311– 18334. Association for Computational Linguis...

  160. [168]

    S. Khan. Verified detection and prevention of concurrency anomalies in multi-agent large language model systems.arXiv preprint arXiv:2606.17182, 2026

  161. [169]

    Y . Kim, K. Gu, C. Park, C. Park, S. Schmidgall, A. A. Heydari, Y . Yan, Z. Zhang, Y . Zhuang, Y . Liu, et al. Towards a science of scaling agent systems.arXiv preprint arXiv:2512.08296, 2025

  162. [170]

    Kimi K2: Open agentic intelligence.arXiv preprint arXiv:2507.20534, 2025

    Kimi Team. Kimi K2: Open agentic intelligence.arXiv preprint arXiv:2507.20534, 2025

  163. [171]

    Y . Kim, A. Abdelaziz, T. C. Ferreira, M. Al-Badrashiny, and H. Sawaf. Bel esprit: Multi-agent framework for building ai model pipelines. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pages 329–339, 2025

  164. [172]

    Kotliarskyi, V

    A. Kotliarskyi, V . Zhu, and Z. Brock. An open-source spec for codex orchestration: Symphony. OpenAI Engineering, Apr. 2026

  165. [173]

    Krause, L

    B. Krause, L. Chen, and E. Kahembwe. Autograms: Autonomous graphical agent modeling software.arXiv preprint arXiv:2407.10049, 2024

  166. [174]

    Kimi k2.5: Visual agentic intelligence.arXiv preprint arXiv:2602.02276, 2026

    Kimi Team. Kimi k2.5: Visual agentic intelligence.arXiv preprint arXiv:2602.02276, 2026. URLhttps://arxiv.org/ abs/2602.02276

  167. [175]

    Lambert, J

    N. Lambert, J. Morrison, V . Pyatkin, S. Huang, H. Ivison, F. Brahman, L. J. V . Miranda, A. Liu, N. Dziri, S. Lyu, Y . Gu, S. Malik, V . Graf, J. D. Hwang, J. Yang, R. Le Bras, O. Tafjord, C. Wilhelm, L. Soldaini, N. A. Smith, Y . Wang, P. Dasigi, and H. Hajishirzi. Tulu 3: P...

  168. [176]

    LangChain: Agent and application framework

    LangChain, Inc. LangChain: Agent and application framework. GitHub repository and documentation, 2026. URLhttps: //github.com/langchain-ai/langchain

  169. [177]

    W. Kwon, Z. Li, S. Zhuang, Y . Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica. Efficient memory management for large language model serving with PagedAttention. InProceedings of the 29th Symposium on Operating Systems Principles, 2023. URLhttps://arxiv.org/...

  170. [178]

    Langflow: Visual framework for ai agents and workflows

    Langflow. Langflow: Visual framework for ai agents and workflows. GitHub repository and documentation, 2026. URL https://github.com/langflow-ai/langflow

  171. [179]

    Dify: Agentic workflow and llm application platform

    LangGenius, Inc. Dify: Agentic workflow and llm application platform. GitHub repository and documentation, 2026. URL https://github.com/langgenius/dify. 44

  172. [180]

    LangGraph: Low-level orchestration for stateful agents

    LangChain, Inc. LangGraph: Low-level orchestration for stateful agents. GitHub repository and documentation, 2026. URL https://github.com/langchain-ai/langgraph

  173. [181]

    J. Lee. Llm agents: A survey, 2026. Preprints.org preprint, posted August 5, 2026

  174. [182]

    K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini. Deduplicating training data makes language models better. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8424–8445. A...

  175. [183]

    H. Lee, S. Phatale, H. Mansoor, T. Mesnard, J. Ferret, K. Lu, C. Bishop, E. Hall, V . Carbune, A. Rastogi, and S. Prakash. RLAIF vs. RLHF: Scaling reinforcement learning from human feedback with AI feedback.arXiv preprint arXiv:2309.00267, 2023

  176. [184]

    Y . Lee, H. Yen, X. Ye, and D. Chen. Agentic aggregation for parallel scaling of long-horizon agentic tasks.arXiv preprint arXiv:2604.11753, 2026

  177. [185]

    H. Y . Leong, Y . Li, Y . Wu, W. Ouyang, W. Zhu, J. Gao, and W. Han. Amas: Adaptively determining communication topology for llm-based multi-agent system. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 2061–2070, 2025

  178. [186]

    Y . Lee, R. Nair, Q. Zhang, K. Lee, O. Khattab, and C. Finn. Meta-harness: End-to-end optimization of model harnesses. arXiv preprint arXiv:2603.28052, 2026

  179. [187]

    Lewis, E

    P. Lewis, E. Perez, A. Piktus, F. Petroni, V . Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, S. Riedel, and D. Kiela. Retrieval-augmented generation for knowledge-intensive NLP tasks. InAdvances in Neural Information Processing Systems, volume 33, pages...

  180. [188]

    A. Li, Y . Xie, S. Li, F. Tsung, B. Ding, and Y . Li. Agent-oriented planning in multi-agent systems. InInternational Conference on Learning Representations, volume 2025, pages 19495–19517, 2025

  181. [189]

    Letta Agent SDK: Stateful agents with persistent memory

    Letta. Letta Agent SDK: Stateful agents with persistent memory. GitHub repository and documentation, 2026. URL https://github.com/letta-ai/letta-agent-sdk

  182. [190]

    D. Li, Z. Li, H. Du, X. Wu, S. Gui, Y . Kuang, and L. Sun. Graph of skills: Dependency-aware structural retrieval for massive agent skills.arXiv preprint arXiv:2604.05333, 2026

  183. [191]

    D. Li, Z. Liu, J. Wang, J. Huang, F. Li, B. Jia, B. Hu, and M. Zhang. Lycheememory v2: Efficient long-term memory for llm agents via semantic segment-level consolidation.arXiv preprint arXiv:2608.12990, 2026

  184. [192]

    B. Li, Z. Zhao, D.-H. Lee, and G. Wang. Adaptive graph pruning for multi-agent communication.arXiv preprint arXiv:2506.02951, 2025

  185. [193]

    G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem. CAMEL: Communicative agents for “mind” ex- ploration of large language model society.arXiv preprint arXiv:2303.17760, 2023. URLhttps://arxiv.org/abs/ 2303.17760

  186. [194]

    H. Li, Z. Fang, R. Feng, Y . Zhao, J. Liu, P. Gao, H. Ye, D. Lin, Q. Lin, S. Rajmohan, et al. Loopsbench: From harness engineering to loop engineering in benchmarking coding agent.arXiv preprint arXiv:2608.00267, 2026

  187. [195]

    F. Li, J. Wu, T. Fu, N. Jaques, W. Zhou, and M.-Y . Kan. Flowsteer: Prompt-only workflow steering exposes planning-time vulnerabilities in multi-agent LLM systems.arXiv preprint arXiv:2605.11514, 2026. URLhttps://arxiv.org/abs/ 2605.11514

  188. [196]

    J. Li, D. Garijo, and M. Poveda-Villalón. Large language models for ontology engineering: a systematic literature review. Semantic Web, 17(4):22104968261465514, 2026

  189. [197]

    J. Li, Z. Jin, T. Men, Y . Hao, K. Zhu, L. Wang, D. Huang, L. Wang, S. Hua, L. Wang, et al. Agentic environment engineer- ing for large language models: A survey of environment modeling, synthesis, evaluation, and application.arXiv preprint arXiv:2606.12191, 2026

  190. [198]

    J. Li, A. Fang, G. Smyrnis, M. Ivgi, M. Jordan, S. Gadre, H. Bansal, E. Guha, S. Keh, K. Arora, S. Garg, R. Xin, N. Muen- nighoff, et al. DataComp-LM: In search of the next generation of training sets for language models. InAdvances in Neural Information Processing Systems, vo...

  191. [199]

    K. Li, X. Yu, Z. Ni, Y . Zeng, Y . Xu, Z. Zhang, X. Li, J. Sang, X. Duan, X. Wang, et al. Timem: Temporal-hierarchical memory consolidation for long-horizon conversational agents. InFindings of the Association for Computational Linguistics: ACL 2026, pages 21700–21720, 2026

  192. [200]

    M. Li, Y . Zhao, B. Yu, F. Song, H. Li, H. Yu, Z. Li, F. Huang, and Y . Li. API-Bank: A comprehensive benchmark for tool- augmented LLMs. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 3102–3116. Association for Computational Li...

  193. [201]

    K. Li, J. Shi, Y . Xiao, M. Jiang, J. Sun, Y . Wu, D. Fu, S. Xia, X. Cai, T. Xu, W. Si, W. Li, D. Wang, and P. Liu. AgencyBench: Benchmarking the frontiers of autonomous agents in 1m-token real-world contexts. InProceedings of the 64th Annual Meeting of the Association for Com...

  194. [202]

    P. Li, S. Zhang, Y . Zhang, S. He, D. van Dijk, and R. Ying. Morse: Task-oriented multi-agent system with mixture of role-subtask experts.arXiv preprint arXiv:2608.09251, 2026

  195. [203]

    R. Li, X. Wang, D. Berlowitz, J. Mez, H. Lin, and H. Yu. CARE-AD: A multi-agent large language model frame- work for alzheimer’s disease prediction using longitudinal clinical notes.npj Digital Medicine, 8:541, 2025. doi: 10.1038/s41746-025-01940-4. URLhttps://doi.org/10.1038/...

  196. [204]

    N. Li, C. Gao, M. Li, Y . Li, and Q. Liao. EconAgent: Large language model-empowered agents for simulating macroeco- nomic activities. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15523–15536. Associat...

  197. [205]

    X. Li, M. Liu, and C. Yuen. LLM agent communication protocol (LACP) requires urgent standardization: A telecom-inspired protocol is necessary. InProceedings of the NeurIPS 2025 Workshop on AI for Next Generation Communication Networks,

  198. [206]

    X. Li, W. Jiao, J. Jin, G. Dong, J. Jin, Y . Wang, H. Wang, Y . Zhu, J.-R. Wen, Y . Lu, et al. Deepagent: A general reasoning agent with scalable toolsets. InProceedings of the ACM Web Conference 2026, pages 2219–2230, 2026

  199. [207]

    S. Li, Y . Liu, Q. Wen, C. Zhang, and S. Pan. Assemble your crew: Automatic multi-agent communication topology design via autoregressive graph generation. InProceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 23142–23150, 2026

  200. [208]

    X. Li, Y . Wang, H. Lu, Z. Chen, M. Li, P. Song, M. Zheng, and T. Cai. Memtx: Transactional belief commit for stateful agent memory.arXiv preprint arXiv:2607.23929, 2026

  201. [209]

    URLhttps://arxiv.org/abs/2510.13821

    doi: 10.48550/arXiv.2510.13821. URLhttps://arxiv.org/abs/2510.13821

  202. [210]

    Y . Li, J. Yang, Z. Zheng, Z. Hu, Y . Sui, S. Wang, Y . He, and B. Hooi. Apex: Autonomous policy exploration for self-evolving llm agents.arXiv preprint arXiv:2605.21240, 2026

  203. [211]

    X. Li, T. Lyu, Y . Yang, L. Shan, S. Yang, L. Zhang, Z. Huang, Q. Liu, and Y . Li. Escaping the context bottleneck: Active context curation for LLM agents via reinforcement learning.arXiv preprint arXiv:2604.11462, 2026

  204. [212]

    Z. Li, Y . Mi, Z. Zhou, H. Jiang, G. Zhang, K. Wang, and J. Fang. Goal-aware identification and rectification of misinformation in multi-agent systems. InInternational Conference on Learning Representations, volume 2026, pages 24661–24687, 2026

  205. [213]

    Y . Li, S. Ping, X. Chen, X. Qi, Z. Wang, Y . Luo, and X. Zhang. Agentgit: A version control framework for reliable and scalable llm-powered multi-agent systems.arXiv preprint arXiv:2511.00628, 2025

  206. [214]

    Liévin, A

    V . Liévin, A. Palepu, W.-H. Weng, K. Saab, D. Stutz, Y . Cheng, K. Kulkarni, S. S. Mahdavi, J. Barral, D. R. Webster, et al. Towards conversational artificial intelligence for disease management.Nature, 655:1292–1299, 2026. doi: 10.1038/ s41586-026-10764-5. URLhttps://doi.org...

  207. [215]

    Z. Li, S. Xu, K. Mei, W. Hua, B. Rama, O. Raheja, H. Wang, H. Zhu, and Y . Zhang. Autoflow: Automated workflow generation for large language model agents.arXiv preprint arXiv:2407.12821, 2024. URLhttps://arxiv.org/abs/ 2407.12821

  208. [216]

    M. Lin, J. Wu, Z. Wang, Z. Shi, Y . Sang, B. He, Z. Liu, T. Wei, Z. Wu, Z. Zhang, D. Wang, X. Zhang, B. Dumoulin, C. Xie, Y . Zhou, S. Wang, and H. Lu. Harness updating is not harness benefit: Disentangling evolution capabilities in self-evolving LLM agents.arXiv preprint arXi...

  209. [217]

    Liang, H

    Q. Liang, H. Wang, Z. Liang, and Y . Liu. From skill text to skill structure: The scheduling-structural-logical representation for agent skills.arXiv preprint arXiv:2604.24026, 2026

  210. [218]

    OntoExtend: A framework for requirement-driven scalable ontology construction, 2026

    Lippolis et al. OntoExtend: A framework for requirement-driven scalable ontology construction, 2026. Emerging work; verify authors, venue, DOI, and publication status before submission

  211. [219]

    J. Lin, S. Liu, C. Pan, L. Lin, S. Dou, X. Huang, H. Yan, Z. Han, and T. Gui. Agentic harness engineering: Observability- driven automatic evolution of coding-agent harnesses.arXiv preprint arXiv:2604.25850, 2026

  212. [220]

    A. Liu, J. Wang, S. Kaski, J. Wang, and M. Yang. A principle of targeted intervention for multi-agent reinforcement learning. arXiv preprint arXiv:2510.17697, 2025

  213. [221]

    W. Lin. Towards self-driving codebases. URLhttps://cursor.com/blog/self-driving-codebases. 46

  214. [222]

    G. Liu, H. Lin, H. Zeng, H. Wang, and Q. Yao. Mas-on-the-fly: Dynamic adaptation of llm-based multi-agent systems at test time.arXiv preprint arXiv:2602.13671, 2026

  215. [223]

    A. S. Lippolis, M. J. Saeedizade, S. Schmid, S. Blattner, R. Keskisärkkä, A. Gangemi, E. Blomqvist, and A. G. Nuzzolese. Ontoextend: A framework for requirement-driven and scalable ontology extension with llms, 2026. URLhttps://arxiv. org/abs/2607.17963

  216. [224]

    H. Liu, Y . Ming, S. Joty, and C. Zhao. Harnessing LLM agents with skill programs.arXiv preprint arXiv:2605.17734, 2026

  217. [225]

    C. Liu, C. Zhang, Y . Wu, W. Lu, N. Wu, et al. Agentpo: Enhancing multi-agent collaboration via reinforcement learning. In International Conference on Learning Representations, volume 2026, pages 143134–143152, 2026

  218. [226]

    J. Liu, C. S. Xia, Y . Wang, and L. Zhang. Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation. InAdvances in Neural Information Processing Systems, 2023. URLhttps: //arxiv.org/abs/2305.01210

  219. [227]

    H. Liu, R. Li, W. Xiong, Z. Zhou, and W. Peng. WorkTeam: Constructing workflows from natural language with multi-agents. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguis- tics: Human Language Technologies ...

  220. [228]

    URLhttps://aclanthology.org/2025.naacl-industry.3/

    doi: 10.18653/v1/2025.naacl-industry.3. URLhttps://aclanthology.org/2025.naacl-industry.3/

  221. [229]

    S. Liu, J. Yang, B. Jiang, Y . Li, J. Guo, X. Liu, and B. Dai. Context as a tool: Context management for long-horizon SWE-agents.arXiv preprint arXiv:2512.22087, 2025

  222. [230]

    J. Liu, D. Shen, Y . Zhang, B. Dolan, L. Carin, and W. Chen. What makes good in-context examples for GPT-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 100–114. Associ...

  223. [231]

    X. Liu, R. Song, X. Wang, and X. Chen. Select, read, and write: A multi-agent framework of full-text-based related work generation. InFindings of the Association for Computational Linguistics: ACL 2025, pages 7009–7028, 2025

  224. [232]

    J. Liu, H. Xi, S. Zhang, Y . Zeng, T. Yue, C. Wang, J. Kang, Q. Wu, and H. Wang. Who&when pro: Can llms really attribute failures in ai agents?arXiv preprint arXiv:2607.09996, 2026

  225. [233]

    N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang. Lost in the middle: How language models use long contexts.Transactions of the Association for Computational Linguistics, 12:157–173, 2024

  226. [234]

    Z. Liu, Y . Zhang, P. Li, Y . Liu, and D. Yang. A dynamic llm-powered agent network for task-oriented agent collaboration. arXiv preprint arXiv:2310.02170, 2023

  227. [235]

    X. Liu, H. Yu, H. Zhang, Y . Xu, X. Lei, H. Lai, Y . Gu, H. Ding, K. Men, K. Yang, S. Zhang, X. Deng, A. Zeng, Z. Du, C. Zhang, S. Shen, T. Zhang, Y . Su, H. Sun, M. Huang, Y . Dong, and J. Tang. AgentBench: Evaluating LLMs as agents. In International Conference on Learning Re...

  228. [236]

    Z. Liu, Z. Shi, Y . Sang, B. He, M. Lin, T. Wei, D. Wang, B. Dumoulin, W. Jin, and H. Lu. Adaptive auto-harness: Sustained self-improvement for agentic system deployment on open-ended task streams.arXiv preprint arXiv:2606.01770, 2026

  229. [237]

    Y . Liu, G. Zhang, K. Wang, S. Li, and S. Pan. Graph-augmented large language model agents: Current progress and future prospects.arXiv preprint arXiv:2507.21407, 2025. URLhttps://arxiv.org/abs/2507.21407

  230. [238]

    Y . Liu, Y . Liu, X. Yin, B. Wang, C. Zhang, H. Yin, and Z. Han. Openclawbench: Benchmarking process-side anomalies in real-world agent execution trajectories.arXiv preprint arXiv:2605.29253, 2026. URLhttps://arxiv.org/abs/ 2605.29253

  231. [239]

    Lopopolo

    R. Lopopolo. Harness engineering: Leveraging codex in an agent-first world. OpenAI Engineering, Feb. 2026

  232. [240]

    Z. Liu, C. Chen, W. Li, P. Qi, T. Pang, C. Du, W. S. Lee, and M. Lin. Understanding r1-zero-like training: A critical perspective.arXiv preprint arXiv:2503.20783, 2025

  233. [241]

    J. Lu, T. Holleis, Y . Zhang, B. Aumayer, F. Nan, F. Bai, S. Ma, S. Ma, M. Li, G. Yin, Z. Wang, and R. Pang. ToolSandbox: A stateful, conversational, interactive evaluation benchmark for LLM tool use capabilities.arXiv preprint arXiv:2408.04682, 2024

  234. [242]

    LlamaIndex Workflows: Event-driven agent workflows

    LlamaIndex. LlamaIndex Workflows: Event-driven agent workflows. GitHub repository and documentation, 2026. URL https://github.com/run-llama/workflows-py. 47

  235. [243]

    Longpre, L

    S. Longpre, L. Hou, T. Vu, A. Webson, H. W. Chung, Y . Tay, D. Zhou, Q. V . Le, B. Zoph, J. Wei, and A. Roberts. The flan collection: Designing data and methods for effective instruction tuning. InProceedings of the 40th International Conference on Machine Learning, volume 202...

  236. [244]

    Y . Luo, R. Gao, L. Teng, X. Wen, J. Jiang, Q. Zhang, Y . Sun, S. Zhang, J. Feng, T. Liu, W. Zhang, and D. Pei. Graph of states: Solving abductive tasks with large language models. InInternational Conference on Machine Learning, 2026. URL https://arxiv.org/abs/2603.21250

  237. [245]

    C. Lu, C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha. The ai scientist: Towards fully automated open-ended scientific discovery.arXiv preprint arXiv:2408.06292, 2024. doi: 10.48550/arXiv.2408.06292. URLhttps://arxiv.org/abs/ 2408.06292

  238. [246]

    Y . Ma, Z. Wang, Y . Li, Z. Li, X. Guo, W. Sun, C. Zhang, and Y . Ye. Proplay: Procedural world models for self-evolving llm agents.arXiv preprint arXiv:2606.12780, 2026

  239. [247]

    Y . Lu, Y . Hu, X. Zhao, and J. Cao. Dytopo: Dynamic topology routing for multi-agent reasoning via semantic matching. arXiv preprint arXiv:2602.06039, 2026

  240. [248]

    X. Luo, Y . Zhang, Z. He, Z. Wang, S. Zhao, D. Li, L. K. Qiu, and Y . Yang. Agent lightning: Train ANY AI agents with reinforcement learning.arXiv preprint arXiv:2508.03680, 2025

  241. [249]

    Madaan, N

    A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y . Yang, S. Gupta, B. P. Majumder, K. Hermann, S. Welleck, A. Yazdanbakhsh, and P. Clark. Self-refine: Iterative refinement with self-feedback. InAdvances in Neural Informatio...

  242. [250]

    Z. Luo, Z. Shen, W. Yang, Z. Zhao, P. Jwalapuram, A. Saha, D. Sahoo, S. Savarese, C. Xiong, and J. Li. Mcp-universe: Benchmarking large language models with real-world model context protocol servers.arXiv preprint arXiv:2508.14704, 2025

  243. [251]

    Mastra: Typescript framework for ai agents and workflows

    Mastra. Mastra: Typescript framework for ai agents and workflows. GitHub repository and documentation, 2026. URL https://github.com/mastra-ai/mastra

  244. [252]

    Z. Ma, H. Huang, S. Zou, Y . Wang, S. Yang, Y . Hu, F. Wei, and X. Chu. Longhorizon-harness: Advancing long-horizon agents for real-world tasks.arXiv preprint arXiv:2608.01964, 2026

  245. [253]

    S. Macedo. Stop hand-holding your coding agent: Engineering the loops that replace step-by-step prompting.arXiv preprint arXiv:2607.00038, 2026. doi: 10.48550/arXiv.2607.00038. URLhttps://arxiv.org/abs/2607.00038

  246. [254]

    Q. Meng, Y . Wang, L. Chen, Y . Li, W. Wu, W. Jiang, Q. Wang, C. Lu, Y . Gao, Y . Wu, and Y . Hu. Agent harness for large language model agents: A survey, 2026

  247. [255]

    Mahmud, E

    S. Mahmud, E. Bagdasarian, and S. Zilberstein. Collab: A framework for designing scalable benchmarks for agentic llms. In NeurIPS 2025 Workshop on Scaling Environments for Agents, 2025. URLhttps://openreview.net/forum?id= 372FjQy1cF

  248. [256]

    M. A. Merrill, A. G. Shaw, N. Carlini, B. Li, H. Raj, I. Bercovich, L. Shi, J. Y . Shin, T. Walshe, E. K. Buchanan, J. Shen, G. Ye, H. Lin, J. Poulos, M. Wang, M. Nezhurina, J. Jitsev, D. Lu, O. M. Mastromichalakis, Z. Xu, Z. Chen, Y . Liu, R. Zhang, L. L. Chen, A. Kashyap, J....

  249. [257]

    K. Mei, X. Zhu, W. Xu, W. Hua, M. Jin, Z. Li, S. Xu, R. Ye, Y . Ge, and Y . Zhang. AIOS: Llm agent operating system. In Conference on Language Modeling, 2025

  250. [258]

    T. Men, P. Cao, Z. Jin, Y . Chen, K. Liu, and J. Zhao. A troublemaker with contagious jailbreak makes chaos in honest towns. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 17561–17587, 2025

  251. [259]

    Microsoft Agent Framework

    Microsoft. Microsoft Agent Framework. GitHub repository and documentation, 2026. URLhttps://github.com/ microsoft/agent-framework

  252. [260]

    Merrill, A

    M. Merrill, A. Shaw, N. Carlini, B. Li, H. Raj, I. Bercovich, L. Shi, J. Shin, T. Walshe, E. K. Buchanan, et al. Terminal- bench: Benchmarking agents on hard, realistic tasks in command line interfaces. InInternational Conference on Learning Representations, volume 2026, pages...

  253. [261]

    H. Ming, F. Li, X. Wu, and W. Que. Retrieval as reasoning: Self-evolving agent-native retrieval via LLM-Wiki.arXiv preprint arXiv:2605.25480, 2026

  254. [262]

    Mialon, C

    G. Mialon, C. Fourrier, C. Swift, T. Wolf, Y . LeCun, and T. Scialom. GAIA: A benchmark for general AI assistants. In International Conference on Learning Representations, 2024. URLhttps://arxiv.org/abs/2311.12983

  255. [263]

    AutoGen: A programming framework for agentic ai

    Microsoft. AutoGen: A programming framework for agentic ai. GitHub repository, 2026. URLhttps://github.com/ microsoft/autogen. Maintenance mode

  256. [264]

    Mohammadi, N

    B. Mohammadi, N. Potamitis, L. Klein, A. Arora, and L. Bindschaedler. Atomix: Timely, transactional tool use for reliable agentic workflows.arXiv preprint arXiv:2602.14849, 2026

  257. [265]

    S. Min, X. Lyu, A. Holtzman, M. Artetxe, M. Lewis, H. Hajishirzi, and L. Zettlemoyer. Rethinking the role of demonstrations: What makes in-context learning work? InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11048–11064. Associ...

  258. [266]

    M. Murphy. Teams of agents can take the headaches and potential costs out of finding it bugs. IBM Research, Dec. 2025. URL https://research.ibm.com/blog/project-alice-software-bugs-agents. Project ALICE: Agentic Logic for Incident and Codebug Elimination

  259. [267]

    Minimax agent team: Built for long-running tasks and continuous evolution

    MiniMax. Minimax agent team: Built for long-running tasks and continuous evolution. URLhttps://www.minimax. io/blog/minimax-agent-team-long-running-1779893953

  260. [268]

    Model Context Protocol Python SDK

    Model Context Protocol Contributors. Model Context Protocol Python SDK. Official GitHub repository and documentation,

  261. [269]

    URLhttps://github.com/modelcontextprotocol/python-sdk

  262. [270]

    B. Niu, Y . Song, K. Lian, Y . Shen, Y . Yao, K. Zhang, and T. Liu. Flow: Modularized agentic workflow automation. InInter- national Conference on Learning Representations, 2025. URLhttps://proceedings.iclr.cc/paper_files/ paper/2025/hash/ba84da6921f3040b74ee163aa7451f53-Abstr...

  263. [271]

    C. Mu, Y . Zeng, Q. Zhang, K. Shao, C. Chu, H. Guo, D. Jia, Z. Wang, and S. Hu. Adaptive theory of mind for llm-based multi-agent coordination. InProceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 29608–29616, 2026

  264. [272]

    N. F. Noy and D. L. McGuinness. Ontology development 101: A guide to creating your first ontology. Technical re- port, Stanford Knowledge Systems Laboratory, 2001. URLhttps://protege.stanford.edu/publications/ ontology_development/ontology101-noy-mcguinness.html

  265. [273]

    Nakajima

    Y . Nakajima. The log is the agent: Event-sourced reactive graphs for auditable, forkable agentic systems.arXiv preprint arXiv:2605.21997, 2026

  266. [274]

    Z. Nie, R. Shen, X. Yu, B. Yin, J. Zhang, and X. Hu. Skillgraph: Self-evolving multi-agent collaboration with multimodal graph topology.arXiv preprint arXiv:2604.17503, 2026

  267. [275]

    X. Ning, K. Tieu, D. Fu, et al. Code as agent harness.arXiv preprint arXiv:2605.18747, 2026. URLhttps://arxiv. org/abs/2605.18747

  268. [276]

    Introducing the Codex app

    OpenAI. Introducing the Codex app. OpenAI, Feb. 2026. URLhttps://openai.com/index/ introducing-the-codex-app/

  269. [277]

    Hermes Agent: The agent that grows with you

    Nous Research. Hermes Agent: The agent that grows with you. GitHub repository and documentation, 2026. URLhttps: //github.com/NousResearch/hermes-agent

  270. [278]

    OpenClaw: Persistent personal agents and multi-agent routing

    OpenClaw Contributors. OpenClaw: Persistent personal agents and multi-agent routing. GitHub repository and documenta- tion, 2026. URLhttps://github.com/openclaw/openclaw

  271. [279]

    Megatron-LM and Megatron Core: Gpu-optimized training of transformer models at scale

    NVIDIA. Megatron-LM and Megatron Core: Gpu-optimized training of transformer models at scale. GitHub repository and documentation, 2026. URLhttps://github.com/NVIDIA/Megatron-LM

  272. [280]

    GPT-4 technical report.arXiv preprint arXiv:2303.08774, 2023

    OpenAI. GPT-4 technical report.arXiv preprint arXiv:2303.08774, 2023

  273. [281]

    Introducing codex

    OpenAI. Introducing codex. OpenAI, May 2025. 49

  274. [282]

    Ouyang, J

    S. Ouyang, J. Yan, I. Hsu, Y . Chen, K. Jiang, Z. Wang, R. Han, L. Le, S. Daruki, X. Tang, et al. Reasoningbank: Scaling agent self-evolving with reasoning memory. InInternational Conference on Learning Representations, volume 2026, pages 94327–94354, 2026

  275. [283]

    OpenAI Agents SDK

    OpenAI. OpenAI Agents SDK. GitHub repository and documentation, 2026. URLhttps://github.com/openai/ openai-agents-python

  276. [284]

    The ontology system, 2026

    Palantir Technologies. The ontology system, 2026. URLhttps://www.palantir.com/docs/foundry/ architecture-center/ontology-system. Palantir Foundry Architecture Center

  277. [285]

    Opsahl-Ong, M

    K. Opsahl-Ong, M. J. Ryan, J. Purtell, D. Broman, C. Potts, M. Zaharia, and O. Khattab. Optimizing instructions and demonstrations for multi-stage language model programs.arXiv preprint arXiv:2406.11695, 2024

  278. [286]

    Ortaç, E

    U. Ortaç, E. Tosun, A. K. Özbek, F. B. Terzio ˘glu, and R. Bayraktar. Agentology: Ontology-driven operational environments for multi-agent systems.Available at SSRN 6919461, 2026

  279. [287]

    Ouyang, J

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe. Training language models to follow instructio...

  280. [288]

    Pappu, B

    A. Pappu, B. El, H. Cao, C. di Nolfo, Y . Sun, M. Cao, and J. Zou. Multi-agent teams hold experts back.arXiv preprint arXiv:2602.01011, 2026

  281. [289]

    Packer, V

    C. Packer, V . Fang, S. G. Patil, K. Lin, S. Wooders, and J. E. Gonzalez. MemGPT: Towards LLMs as operating systems. arXiv preprint arXiv:2310.08560, 2023

  282. [290]

    J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein. Generative agents: Interactive simulacra of human behavior. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, 2023

  283. [291]

    Q. Pan, Y . Yang, J. Li, J. Zhou, K. Chen, X. Li, Q. Chen, and L. He. Anything2Skill: Compiling external knowledge into reusable skills for agents.arXiv preprint arXiv:2606.09316, 2026

  284. [292]

    W. Pan, S. Liu, C.-Y . Lin, J. Zeng, X. Tang, X. Zhou, Y . Lu, and X. Jia. Evolving agents in the dark: Retrospective harness optimization via self-preference.arXiv preprint arXiv:2606.05922, 2026

  285. [293]

    Papadakis, A

    C. Papadakis, A. Dimitriou, G. Filandrianos, M. Lymperaiou, K. Thomas, and G. Stamou. Atlas: Adaptive trading with llm agents through dynamic prompt optimization and multi-agent coordination.arXiv preprint arXiv:2510.15949, 2025

  286. [294]

    Polat, M

    C. Polat, M. Tuncel, M. Kurban, E. Serpedin, and H. Kurban. xchemagents: Agentic ai for explainable quantum chemistry. arXiv preprint arXiv:2505.20574, 2025

  287. [295]

    Parisi, Y

    A. Parisi, Y . Zhao, and N. Fiedel. TALM: Tool augmented language models.arXiv preprint arXiv:2205.12255, 2022

  288. [297]

    S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez. Gorilla: Large language model connected with massive APIs. InAdvances in Neural Information Processing Systems, volume 37, 2024

  289. [298]

    Penedo, H

    G. Penedo, H. Kydlí ˇcek, L. Ben Allal, A. Lozhkov, M. Mitchell, C. Raffel, L. von Werra, and T. Wolf. The FineWeb datasets: Decanting the web for the finest text data at scale. InAdvances in Neural Information Processing Systems, volume 37, pages 30811–30849, 2024

  290. [299]

    Perera, K

    S. Perera, K. Hapuarachchi, F. Leymann, and R. Khalaf. Robust agent compensation (rac): Teaching ai agents to compensate. InProceedings of the ACM Conference on AI and Agentic Systems, pages 253–262, 2026

  291. [2025]

    URLhttps://aclanthology.org/2025.findings-naacl

    doi: 10.18653/v1/2025.findings-naacl.448. URLhttps://aclanthology.org/2025.findings-naacl. 448/

  292. [2026]

    URLhttps://github.com/apache/burr

Pith tools

Reviewed August 27, 2026 · model on record in the stance chip above.