Pith. sign in

REVIEW 2 major objections 5 minor 62 references

Agent skill libraries are learning systems: evolving, verified stores of reusable procedures—not static toolbags—and their design is best read as an eight-stage lifecycle of evidence, admission, maintenance, and governance.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 14:12 UTC pith:XC5EYVDL

load-bearing objection Useful lifecycle survey that gives the field a shared language for evolving skill libraries; soft spots are the usual survey ones and already owned. the 2 major comments →

arxiv 2607.10113 v1 pith:XC5EYVDL submitted 2026-07-11 cs.AI

Dynamic Agent Skills: A Lifecycle Survey and Taxonomy of Evolving Skill Libraries

classification cs.AI
keywords agent skillsskill librarieslifecycle architectureverification and admissionretrieval scalingskill-aware RLskill safetybenchmarking
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This survey argues that when language-model agents keep reusable procedures outside the model, the skill library itself becomes part of learning. Skills may be code, natural-language lessons, SKILL.md packages, graphs, or adapters, but the shared problem is how the library changes over time: evidence from interaction drives proposals; verification decides what enters; storage and retrieval decide what gets used; maintenance repairs or prunes; and provenance supports rollback and sharing. From a 124-paper 2023–2026 audit set, the authors offer three comparison tools—a six-sense taxonomy of what “skill” means, an eight-stage lifecycle, and a light skill-record schema plus ten update operators—and use them to grade recurring patterns. The practical claim is that write-time discipline (admission, verifier quality, repair, abstraction) repeatedly matters, flat retrieval often worsens as libraries grow, and benchmarks still under-report library trajectories, usage-without-utility, and safety surfaces. A sympathetic reader cares because this reframes agent improvement from better prompts or bigger models toward managing an external procedural store that can grow, go stale, inject risk, or consolidate into weights.

Core claim

Dynamic skill systems should be synthesized as lifecycle-managed, verified, evolving artifact stores: agents collect interaction evidence, propose skill updates, verify and admit candidates, organize them for retrieval and composition, repair or prune stale entries, and govern sharing through provenance and rollback. Across the audited literature, admission and repair are repeatedly important, verifier quality materially affects skill-aware reinforcement learning, flat retrieval can degrade as libraries grow, and current benchmarks still under-report library trajectories, usage–utility gaps, and safety surfaces.

What carries the argument

The eight-stage lifecycle architecture (evidence acquisition → proposal → verification/admission → organization/storage → retrieval/composition → maintenance/repair → distillation/portability → governance), supported by a six-sense skill taxonomy, a seven-field editable skill-record schema, and a ten-operator library-update vocabulary used as a comparison scaffold rather than a new learning method.

Load-bearing premise

That a non-exhaustive, arXiv-heavy iterative search cut off at May 31, 2026 yields a primary paper cluster representative enough to treat the seven graded patterns as field regularities rather than artifacts of inclusion frontiers, naming, and mismatched evaluation harnesses.

What would settle it

A shared trajectory-aware benchmark that logs operator velocities, maintenance-off and admission-off ablations, library size over time, and skill use versus task gain, and then finds that systems without strong write-time admission or repair match or beat gated systems once harness, backbone, and task stream are controlled.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This taxonomy-driven survey synthesizes 124 papers (2023–May 31 2026) on dynamic LLM-agent skill libraries. It argues that such systems are best understood as lifecycle-managed, verified, evolving artifact stores rather than static prompt or tool collections. The paper contributes three comparison tools: a six-sense taxonomy of skill artifacts (Table 1), an eight-stage lifecycle architecture (Figure 1, Table 2), and a lightweight skill-record schema plus ten-operator update vocabulary (§3). Using these, it organizes system families (§6), mechanisms (§7), evaluation (§8), seven evidence-graded patterns with explicit caveats (§9, Table 7), infrastructure and safety surfaces (§§10–11), and open problems with concrete first experiments (§12, Table 10). Limitations of the non-PRISMA, arXiv-heavy corpus and heterogeneous evidence are stated in §§2 and 13.

Significance. If the synthesis holds, the paper supplies a shared language for a fast-moving, overloaded literature: what counts as a dynamic skill, which lifecycle stages a system implements, how library updates should be reported, and which failure modes (admission, flat retrieval, usage–utility gaps, skill injection) matter for evaluation and deployment. Strengths include transparent evidence roles and A–D grades, repeated caveats rather than pooled leaderboards, a concrete reporting checklist, and open problems paired with first experiments. For TMLR-style survey work this is a useful organizational contribution even without new algorithms; the lifecycle and operator framing make method, benchmark, infrastructure, and safety papers comparable without forcing a single ranking.

major comments (2)
  1. §9 and Table 7 are the main empirical payload beyond taxonomy. The abstract and §1.3 present seven patterns as field regularities (admission/repair importance, verifier quality in skill-aware RL, flat-retrieval degradation, etc.), while §2.2–2.5 and §13 correctly refuse PRISMA exhaustiveness and cross-paper leaderboards. For load-bearing claims such as R1 and R6, please add a short sensitivity check or explicit primary-cluster subset: e.g., whether the same qualitative conclusions survive when restricted to papers with controlled ablations (grade A/B only) or when 2026-only registry/safety papers are held out. Without that, readers may over-read convergent 2026 naming conventions as causal regularities.
  2. §3.2–3.3 introduce St=⟨Ct,πt,Tt,Rt,φt,νt,≺t⟩ and Lt→Lt+1 with a fixed ten-operator vocabulary listed as a contribution in §1.3. The manuscript already says this is a comparison scaffold, not a method, but φt is alternately a deterministic map and a proposal kernel, and operators such as Refine vs Rewrite and Abstract vs Distill are only informally distinguished. Please either (i) give brief decision criteria for operator assignment used when coding Table 11, or (ii) demote the vocabulary in the contribution list to “working terms for audit,” so the central claim does not appear to rest on an unvalidated closed operator algebra.
minor comments (5)
  1. Figure 4’s blue curve is helpful, but the dashed green “mitigation” path should be labeled more clearly as a qualitative summary of heterogeneous systems, not a single comparable trajectory (as the caption partly notes).
  2. Table 1 is dense; a one-line pointer early in §1 that the first four senses are in-scope and the last two are boundary cases would help readers who land on the table before §3.1.
  3. §8.2 notes under-reported operator velocity and repair quality; consider adding these two metrics explicitly to the five-item reporting checklist in §15 so the evaluation agenda is fully mirrored in the conclusion.
  4. Occasional near-duplicate system names (SkillFlow-2025 vs SkillFlow-Bench; SAGE vs SAGER) are disambiguated in §2.4; a short glossary or consistent hyphenation in tables would reduce residual confusion.
  5. Typos/style: “recurringtasks” spacing in §1; ensure consistent capitalization of method names across Table 11 and the main text.

Circularity Check

0 steps flagged

No significant circularity: lifecycle taxonomy and evidence-graded patterns are explicit comparison scaffolds grounded in cited within-paper ablations, not predictions forced by definition or self-citation.

full rationale

This is a taxonomy-driven survey, not a first-principles derivation. The six-sense taxonomy, eight-stage lifecycle, skill-record schema ⟨C,π,T,R,φ,ν,≺⟩, and ten-operator vocabulary are introduced as comparison tools for organizing heterogeneous papers (§1.3, §3, §5–6); the paper states the notation is “a comparison scaffold, not a separate model of skill learning” and that the lifecycle is architecture for classification rather than a uniqueness theorem. The seven patterns in §9 are tied to named within-paper ablations and benchmarks (e.g., CoEvoSkills verifier drop, AutoRefine maintenance-off, Single-Agent-Skills size sweep) with explicit A–D grades, caveats, and refusal of cross-paper leaderboards (§2.5, §8–9, §13). There is no fitted parameter re-labeled as prediction, no load-bearing self-citation chain (sole author Yubo Li does not cite prior own uniqueness results to force the taxonomy), and no equation that reduces a claimed prediction to its inputs by construction. Coding systems into stages and then reporting which stages they implement is ordinary survey framing, not circular derivation. Limitations §13 already own corpus non-exhaustiveness and architectural (not causal) conclusions. Score 0 is the honest finding.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 5 invented entities

As a taxonomy survey, the load-bearing commitments are definitional and corpus-construction choices rather than fitted physical constants. The central synthesis depends on treating external invocable artifacts as the object of study, partitioning ‘skill’ into six senses, decomposing dynamics into eight stages and ten operators, and reading patterns from a non-exhaustive 124-paper audit set with qualitative A–D evidence grades. No numerical model is fit; free parameters are protocol knobs (cutoff, inclusion frontier, operator set size).

free parameters (4)
  • May 31 2026 corpus cutoff
    All main-text claims are bounded by this audit date; later papers are excluded by policy and could change pattern grades.
  • 124-paper modern audit set size / inclusion frontier
    Primary cluster membership is determined by iterative search plus three inclusion criteria; no exhaustive PRISMA count is claimed, so pattern support depends on this constructed set.
  • Ten-operator vocabulary cardinality and labels
    Add/Refine/Merge/Split/Prune/Distill/Abstract/Compose/Rewrite/Rerank is an author-chosen coarse basis used throughout comparisons; alternative operator partitions would re-slice the same systems.
  • Qualitative evidence grades A–D thresholds
    Pattern strength in §9 depends on hand-assigned grades (multiple ablations vs single study vs convergent behavior vs architecture-only).
axioms (5)
  • domain assumption A dynamic skill is an externally invocable persistent artifact whose lifecycle participates in library transitions Lt→Lt+1, distinct from ordinary tools and from classical HRL options.
    Stated in §3.1–3.3 and used to include/exclude papers and to separate retrieval-at-inference from library learning.
  • ad hoc to paper The six artifact senses and five dynamic properties (executable, editable, portable, inspectable, verification handle) adequately partition current literature for lifecycle comparison.
    Table 1 and §3.1 introduce this partition as the conceptual entry point; boundary cases (memory, capability labels) are retained only when they illuminate library behavior.
  • ad hoc to paper Eight lifecycle stages plus admission-as-selection capture the recurring design commitments of dynamic skill systems.
    §5 reference architecture and Table 2 organize all later synthesis; stage order is acknowledged as non-waterfall but still treated as the comparison skeleton.
  • domain assumption Within-paper ablations and convergent benchmark behavior can support evidence-graded field patterns without shared harnesses or pooled effect sizes.
    §2.5 and §9 explicitly adopt qualitative grades and refuse cross-paper leaderboards while still asserting seven patterns.
  • standard math Classical options/HRL provide a useful static starting point (initiation, policy, termination) but leave artifact-store machinery implicit.
    §2.6 and §3.2 cite Sutton et al. 1999 and SoK-Skills four-tuple as background, then extend with time index, edit, admission, lineage, and library object.
invented entities (5)
  • Six-sense skill taxonomy with dynamic-property axes no independent evidence
    purpose: Disambiguate overloaded ‘skill’ artifacts before lifecycle comparison
    Introduced in Table 1 as the survey’s conceptual entry point; independent of any single method paper.
  • Eight-stage dynamic-skill lifecycle architecture no independent evidence
    purpose: Make design commitments comparable across heterogeneous systems
    §5 and Table 2 define stages from evidence acquisition through governance; used as the primary organizing object.
  • Editable skill-record schema St=⟨Ct,πt,Tt,Rt,φt,νt,≺t⟩ no independent evidence
    purpose: Compare edit handles, admission, and lineage without claiming a new learning algorithm
    §3.2 extends SoK four-tuple with edit field, verification field, and lineage; scaffold for operator discussion.
  • Ten-operator library-update vocabulary no independent evidence
    purpose: Provide common terms for library transitions without elevating them to a method contribution
    Equation 4 / §3.3; repeatedly used as taxonomic fingerprint of systems.
  • Seven evidence-graded patterns (R1–R7) for dynamic skills no independent evidence
    purpose: Synthesize recurring empirical regularities with explicit caveats and open problems
    §9 Table 7; grades and caveats are paper-authored judgments over the audit set.

pith-pipeline@v1.1.0-grok45 · 52402 in / 3962 out tokens · 46136 ms · 2026-07-14T14:12:00.222855+00:00 · methodology

0 comments
read the original abstract

Large language model agents increasingly store reusable procedures outside the model. These reusable procedures are often called \emph{skills}: they may be code functions, natural-language instructions, SKILL.md packages, workflow graphs, or learned adapters that a future agent can retrieve and invoke. This taxonomy-driven survey asks how such skill libraries change over time. Across a $124$-paper $2023$--$2026$ audit set, we synthesize dynamic skill systems as \emph{lifecycle-managed, verified, evolving artifact stores}: agents collect evidence from interaction, propose skill updates, verify and admit candidates, organize them for retrieval and composition, repair or prune stale entries, and govern sharing through provenance and rollback. We organize the literature around three survey tools. First, a $\text{six}$-sense taxonomy distinguishes the structurally different artifacts called ``skills'' in current papers. Second, an $\text{eight}$-stage lifecycle architecture identifies the recurring design decisions behind evidence acquisition, proposal, verification/admission, storage, retrieval/composition, maintenance, distillation/portability, and governance. Third, a lightweight skill-record schema and $\text{ten}$-operator vocabulary provide common terms for comparing library updates without elevating them into a separate method contribution. Using this structure, we synthesize evidence-graded patterns with explicit caveats: admission and repair are repeatedly important, verifier quality materially affects skill-aware RL, flat retrieval can degrade as libraries grow, and current benchmarks still under-report library trajectories, usage--utility gaps, and safety surfaces. We close with concrete reporting standards and open problems for evaluating dynamic skills as changing libraries rather than static prompt or tool collections.

Figures

Figures reproduced from arXiv: 2607.10113 by Yubo Li.

Figure 1
Figure 1. Figure 1: Dynamic skill systems as lifecycle-managed artifact stores. Interaction evidence drives proposal and verification; admitted artifacts enter an evolving skill store, where retrieval and execution create further evidence. Maintenance repairs, prunes, merges, or reranks the store over time. Governance and provenance wrap the lifecycle through shields, audit logs, lineage records, and rollback handles, while d… view at source ↗
Figure 2
Figure 2. Figure 2: Temporal structure of the dynamic-skills audit set. Counts summarize the 124 modern papers covered by the survey and exclude only older classical-options and HRL background anchors; bound￾ary/context papers are included for scope but are not treated as primary causal evidence in the evidence￾graded patterns. The figure reports annual additions and cumulative coverage through the May 31, 2026 cutoff. Repres… view at source ↗
Figure 3
Figure 3. Figure 3: Lifecycle coverage across dynamic-skill system families. Cell intensity summarizes how centrally each family implements or evaluates a lifecycle stage in the coding sheet. The figure is a synthesis device rather than a ranking: executable libraries concentrate around proposal, verification, admission, storage, and retrieval; graph systems concentrate around storage, retrieval, and maintenance; two-timescal… view at source ↗
Figure 4
Figure 4. Figure 4: Retrieval-scaling evidence behind R3. The blue curve plots the controlled Single-Agent￾Skills sweep reported in the paper: 16–32 skills remain around 96–98%, 64 skills at 92%, 128 skills at 78%, and 256 skills at 64%. The dashed green path summarizes the mitigation family rather than a pooled leaderboard: hierarchy, graph retrieval, and full-text reranking push the failure rightward in Single-Agent￾Skills,… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

62 extracted references · 1 canonical work pages · 1 internal anchor

  1. [1]

    Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data.arXiv:2602.21320,

    Emre Can Acikgoz, Cheng Qian, Jonas Hübotter, Heng Ji, Dilek Hakkani-Tür, and Gokhan Tur. Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data.arXiv:2602.21320,

  2. [2]

    Experiential Reflective Learning for Self-Improving LLM Agents.arXiv:2603.24639,

    Marc-Antoine Allard, Arnaud Teinturier, Victor Xing, and Gautier Viaud. Experiential Reflective Learning for Self-Improving LLM Agents.arXiv:2603.24639,

  3. [3]

    EvoSkill: Automated Skill Discovery for Multi-Agent Systems.arXiv:2603.02766,

    Salaheddin Alzubi, Noah Provenzano, Jaydon Bingham, Weiyuan Chen, and Tu Vu. EvoSkill: Automated Skill Discovery for Multi-Agent Systems.arXiv:2603.02766,

  4. [4]

    Shuzhen Bi, Mengsong Wu, Hao Hao, Keqian Li, Wentao Liu, Siyu Song, Hongbo Zhao, and Aimin Zhou

    doi: 10.1609/aaai.v31i1.10916. Shuzhen Bi, Mengsong Wu, Hao Hao, Keqian Li, Wentao Liu, Siyu Song, Hongbo Zhao, and Aimin Zhou. Automating Skill Acquisition through Large-Scale Mining of Open-Source Agentic Repositories: A Frame- work for Multi-Agent Procedural Knowledge Extraction.arXiv:2603.11808,

  5. [5]

    CODE-SHARP: Continuous Open-ended Discovery and Evolution of Skills as Hierarchical Reward Programs.arXiv:2602.10085,

    Richard Bornemann, Pierluigi Vito Amadori, and Antoine Cully. CODE-SHARP: Continuous Open-ended Discovery and Evolution of Skills as Hierarchical Reward Programs.arXiv:2602.10085,

  6. [6]

    Large Language Models as Tool Makers.arXiv:2305.17126,

    Tianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen, and Denny Zhou. Large Language Models as Tool Makers.arXiv:2305.17126,

  7. [7]

    Building Self-Evolving Agents via Experience-Driven Lifelong Learning: A Framework and Benchmark

    Yuxuan Cai, Yipeng Hao, Jie Zhou, Hang Yan, Zhikai Lei, Rui Zhen, Zhenhua Han, Yutao Yang, Junsong Li, Qianjun Pan, Tianyu Huai, Qin Chen, Xin Li, Kai Chen, Bo Zhang, Xipeng Qiu, and Liang He. Building Self-Evolving Agents via Experience-Driven Lifelong Learning: A Framework and Benchmark. arXiv:2508.19005,

  8. [8]

    SkVM: Revisiting Language VM for Skills across Het- erogenous LLMs and Harnesses.arXiv:2604.03088, 2026a

    Le Chen, Erhu Feng, Yubin Xia, and Haibo Chen. SkVM: Revisiting Language VM for Skills across Het- erogenous LLMs and Harnesses.arXiv:2604.03088, 2026a. Qijia Chen, Andrea Bellucci, Zhida Sun, and Giulio Jacucci. SkillDroid: Compile Once, Reuse Forever. arXiv:2604.14872, 2026b. Shiqi Chen, Jingze Gai, Ruochen Zhou, Jinghan Zhang, Tongyao Zhu, Junlong Li, ...

  9. [9]

    Evolving Medical Imaging Agents via Experience-driven Self-skill Discovery.arXiv:2603.05860,

    39 Published in Transactions on Machine Learning Research (07/2026) Lin Fan, Pengyu Dai, Zhipeng Deng, Haolin Wang, Xun Gong, Yefeng Zheng, and Yafei Ou. Evolving Medical Imaging Agents via Experience-driven Self-skill Discovery.arXiv:2603.05860,

  10. [10]

    A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems.arXiv:2508.07407,

    Jinyuan Fang, Yanwen Peng, Xi Zhang, Yingxu Wang, Xinhao Yi, Guibin Zhang, Yi Xu, Bin Wu, Siwei Liu, Zihao Li, Zhaochun Ren, Nikos Aletras, Xi Wang, Han Zhou, and Zaiqiao Meng. A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems.arXiv:2508.07407,

  11. [11]

    Jingzhi Gong, Ruizhen Gu, Zhiwei Fei, Yazhuo Cao, Lukas Twist, Alina Geiger, Shuo Han, Dominik Sobania, Federica Sarro, and Jie M. Zhang. SkillMOO: Multi-Objective Optimization of Agent Skills for Software Engineering.arXiv:2604.09297,

  12. [12]

    SWE- Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?arXiv:2603.15401,

    Tingxu Han, Yi Zhang, Wei Song, Chunrong Fang, Zhenyu Chen, Youcheng Sun, and Lijie Hu. SWE- Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?arXiv:2603.15401,

  13. [13]

    Ma- licious Or Not: Adding Repository Context to Agent Skill Classification.arXiv:2603.16572,

    Florian Holzbauer, David Schmidt, Gabriel Gegenhuber, Sebastian Schrittwieser, and Johanna Ullrich. Ma- licious Or Not: Adding Repository Context to Agent Skill Classification.arXiv:2603.16572,

  14. [14]

    SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills.arXiv:2604.06550,

    Yinghan Hou and Zongyou Yang. SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills.arXiv:2604.06550,

  15. [15]

    MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills.arXiv:2604.20441,

    Yingyong Hou, Xinyuan Lao, Huimei Wang, Qianyu Yao, Wei Chen, Bocheng Huang, Fei Sun, Yuxian Lv, Weiqi Lei, Xueqian Wen, Shengyang Xie, Pengfei Xia, and Zhujun Tan. MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills.arXiv:2604.20441,

  16. [16]

    Bilevel Optimization of Agent Skills via Monte Carlo Tree Search.arXiv:2604.15709, 2026a

    Chenyi Huang, Haoting Zhang, Jingxu Xu, Zeyu Zheng, and Yunduan Lin. Bilevel Optimization of Agent Skills via Monte Carlo Tree Search.arXiv:2604.15709, 2026a. Ken Huang and Jerry Huang. Audited Skill-Graph Self-Improvement for Agentic LLMs via Verifiable Re- wards, Experience Synthesis, and Continual Memory.arXiv:2512.23760,

  17. [17]

    CASCADE: Cumulative Agentic Skill Creation through Autonomous Development and Evolution.arXiv:2512.23880,

    Xu Huang, Junwu Chen, Yuxing Fei, Zhuohan Li, Philippe Schwaller, and Gerbrand Ceder. CASCADE: Cumulative Agentic Skill Creation through Autonomous Development and Evolution.arXiv:2512.23880,

  18. [18]

    Yi Huang, Bowen Zheng, Yunxi Dong, Hong Tang, Huan Zhao, S. M. Rakibul Hasan Shawon, and Hualiang Zhang. A Self-Evolving Agentic Framework for Metasurface Inverse Design.arXiv:2604.01480, 2026b. Zisu Huang, Jingwen Xu, Yifan Yang, Ziyang Gong, Qihao Yang, Muzhao Tian, Xiaohua Wang, Changze Lv, Xuemei Gao, Qi Dai, Bei Liu, Kai Qiu, Xue Yang, Dongdong Chen,...

  19. [19]

    Guanyu Jiang, Zhaochen Su, Xiaoye Qu, and Yi R. Fung. XSkill: Continual Learning from Experience and Skills in Multimodal Agents.arXiv:2603.12056, 2026a. Yanna Jiang, Delong Li, Haiyu Deng, Baihe Ma, Xu Wang, Qin Wang, and Guangsheng Yu. SoK: Agentic Skills – Beyond Tool Use in LLM Agents.arXiv:2602.20867, 2026b. 40 Published in Transactions on Machine Le...

  20. [20]

    EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents.arXiv:2605.10332,

    Ruofei Ju, Xinrui Wang, Xin Ding, Yifan Yang, Hao Wu, Shiqi Jiang, Qianxi Zhang, Hao Wen, Xiangyu Li, Weijun Wang, Kun Li, Yunxin Liu, Haipeng Dai, Wei Wang, and Ting Cao. EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents.arXiv:2605.10332,

  21. [21]

    Co-Evolving Agents: Learning from Failures as Hard Negatives.arXiv:2511.22254,

    Yeonsung Jung, Trilok Padhi, Sina Shaham, Dipika Khullar, Joonhyun Jeong, Ninareh Mehrabi, and Eunho Yang. Co-Evolving Agents: Learning from Failures as Hard Negatives.arXiv:2511.22254,

  22. [22]

    SkillFlow: Scalable and Efficient Agent Skill Retrieval System.arXiv:2504.06188,

    Fangzhou Li, Pagkratios Tagkopoulos, and Ilias Tagkopoulos. SkillFlow: Scalable and Efficient Agent Skill Retrieval System.arXiv:2504.06188,

  23. [23]

    The World Won’t Stay Still: Programmable Evolution for Agent Benchmarks.arXiv:2603.05910, 2026a

    Guangrui Li, Yaochen Xie, Yi Liu, Ziwei Dong, Xingyuan Pan, Tianqi Zheng, Jason Choi, Michael Morais, Binit Jha, Shaunak Mishra, Bingrou Zhou, Chen Luo, Monica Cheng, and Dawn Song. The World Won’t Stay Still: Programmable Evolution for Agent Benchmarks.arXiv:2603.05910, 2026a. Hao Li, Chunjiang Mu, Jianhao Chen, Siyue Ren, Zhiyao Cui, Yiqun Zhang, Lei Ba...

  24. [24]

    SkillsIn- jector: Dynamic Skill Context Construction for LLM Agents.arXiv:2605.29794, 2026d

    Yanchao Li, Wanhao Liu, Ben Gao, Jiaqing Xie, Zhehong Ai, Na Zou, Yuqiang Li, and Tianfan Fu. SkillsIn- jector: Dynamic Skill Context Construction for LLM Agents.arXiv:2605.29794, 2026d. Yu Li, Rui Miao, Zhengling Qi, and Tian Lan. ARISE: Agent Reasoning with Intrinsic Skill Evolution in Hierarchical Reinforcement Learning.arXiv:2603.16060, 2026e. Zhiyuan...

  25. [25]

    Agent Skills: A Data-Driven Analysis of Claude Skills for Extending Large Language Model Functionality.arXiv:2602.08004,

    George Ling, Shanshan Zhong, and Richard Huang. Agent Skills: A Data-Driven Analysis of Claude Skills for Extending Large Language Model Functionality.arXiv:2602.08004,

  26. [26]

    Graph of Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills.arXiv:2604.05333, 2026a

    41 Published in Transactions on Machine Learning Research (07/2026) Dawei Liu, Zongxia Li, Hongyang Du, Xiyang Wu, Shihang Gui, Yongbei Kuang, and Lichao Sun. Graph of Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills.arXiv:2604.05333, 2026a. Hongjun Liu, Yifei Ming, Shafiq Joty, and Chen Zhao. Harnessing LLM Agents with Skill Program...

  27. [27]

    Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills.arXiv:2603.25158,

    Jingwei Ni, Yihao Liu, Xinpeng Liu, Yutao Sun, Mengyu Zhou, Pengyu Cheng, Dexin Wang, Erchao Zhao, Xiaoxi Jiang, and Guanjun Jiang. Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills.arXiv:2603.25158,

  28. [28]

    SkillGraph: Self-Evolving Multi-Agent Collaboration with Multimodal Graph Topology.arXiv:2604.17503,

    Zheng Nie, Ruolin Shen, Xinlei Yu, Bo Yin, Jiangning Zhang, and Xiaobin Hu. SkillGraph: Self-Evolving Multi-Agent Collaboration with Multimodal Graph Topology.arXiv:2604.17503,

  29. [29]

    AutoRefine: From Trajectories to Reusable Expertise for Continual LLM Agent Refinement

    Libin Qiu, Zhirong Gao, Junfu Chen, Yuhang Ye, Weizhi Huang, Xiaobo Xue, Wenkai Qiu, and Shuo Tang. AutoRefine: From Trajectories to Reusable Expertise for Continual LLM Agent Refinement. arXiv:2601.22758,

  30. [30]

    Supply- Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems.arXiv:2604.03081,

    Yubin Qu, Yi Liu, Tongcheng Geng, Gelei Deng, Yuekang Li, Leo Zhang, Ying Zhang, and Lei Ma. Supply- Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems.arXiv:2604.03081,

  31. [31]

    Skilldex: A Package Manager and Registry for Agent Skill Packages with Hierarchical Scope-Based Distribution.arXiv:2604.16911,

    Sampriti Saha and Pranav Hemanth. Skilldex: A Package Manager and Registry for Agent Skill Packages with Hierarchical Scope-Based Distribution.arXiv:2604.16911,

  32. [32]

    Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry.arXiv:2605.11418,

    Shoumik Saha, Kazem Faghih, and Soheil Feizi. Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry.arXiv:2605.11418,

  33. [33]

    AgentSkillsEnableaNewClassofRealistic and Trivially Simple Prompt Injections.arXiv:2510.26328,

    42 Published in Transactions on Machine Learning Research (07/2026) DavidSchmotz, SaharAbdelnabi, andMaksymAndriushchenko. AgentSkillsEnableaNewClassofRealistic and Trivially Simple Prompt Injections.arXiv:2510.26328,

  34. [34]

    Skill-Inject: Mea- suring Agent Vulnerability to Skill File Attacks.arXiv:2602.20156,

    David Schmotz, Luca Beurer-Kellner, Sahar Abdelnabi, and Maksym Andriushchenko. Skill-Inject: Mea- suring Agent Vulnerability to Skill File Attacks.arXiv:2602.20156,

  35. [35]

    SkillFoundry: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources

    Shuaike Shen, Wenduo Cheng, Mingqian Ma, Alistair Turcan, Martin Jinye Zhang, and Jian Ma. SkillFoundry: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources. arXiv:2604.03964,

  36. [36]

    Evolving Programmatic Skill Networks.arXiv:2601.03509,

    Haochen Shi, Xingdi Yuan, and Bang Liu. Evolving Programmatic Skill Networks.arXiv:2601.03509,

  37. [37]

    ABSTRAL: Automatic Design of Multi-Agent Systems Through Iterative Refinement and Topology Optimization.arXiv:2603.22791, 2026a

    Weijia Song, Jiashu Yue, and Zhe Pang. ABSTRAL: Automatic Design of Multi-Agent Systems Through Iterative Refinement and Topology Optimization.arXiv:2603.22791, 2026a. Xinyuan Song, Hongji Pu, and Liang Zhao. SkillOps: Managing LLM Agent Skill Libraries as Self- Maintaining Software Ecosystems.arXiv:2605.13716, 2026b. Weihang Su, Jianming Long, Qingyao Ai...

  38. [38]

    Yiqun Sun, Pengfei Wei, and Lawrence B. Hsieh. Don’t Retrieve, Navigate: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG.arXiv:2604.14572,

  39. [39]

    SAGER: Self-Evolving User Policy Skills for Recommendation Agent

    doi: 10.1016/S0004-3702(99)00052-1. Zhen Tao, Riwei Lai, Chenyun Yu, Weixin Chen, Li Chen, Beibei Kong, Lei Cheng, Chengxiang Zhuo, Zang Li, and Qingqiang Sun. SAGER: Self-Evolving User Policy Skills for Recommendation Agent. arXiv:2604.14972,

  40. [40]

    BadSkill: Backdoor Attacks on Agent Skills via Model- in-Skill Poisoning.arXiv:2604.09378,

    Guiyao Tie, Jiawen Shi, Pan Zhou, and Lichao Sun. BadSkill: Backdoor Attacks on Agent Skills via Model- in-Skill Poisoning.arXiv:2604.09378,

  41. [41]

    SkillX: Automatically Constructing Skill Knowledge Bases for Agents.arXiv:2604.04804, 2026a

    Chenxi Wang, Zhuoyun Yu, Xin Xie, Wuguannan Yao, Runnan Fang, Shuofei Qiao, Kexin Cao, Guozhou Zheng, Xiang Qi, Peng Zhang, and Shumin Deng. SkillX: Automatically Constructing Skill Knowledge Bases for Agents.arXiv:2604.04804, 2026a. Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: AnO...

  42. [42]

    SkillGrad: Optimizing Agent Skills Like Gradient Descent.arXiv:2605.27760, 2026b

    Hanyu Wang, Yifan Lan, Bochuan Cao, Lu Lin, and Jinghui Chen. SkillGrad: Optimizing Agent Skills Like Gradient Descent.arXiv:2605.27760, 2026b. Jiayu Wang, Yifei Ming, Zixuan Ke, Shafiq Joty, Aws Albarghouthi, and Frederic Sala. SkillOrchestra: Learning to Route Agents via Skill Transfer.arXiv:2602.19672, 2026c. Jingxing Wang, Chenyu Zhou, Zhihui Fu, Jun ...

  43. [43]

    Toward User Comprehension Supports for LLM Agent Skill Specifications

    Zikai Alex Wen. Toward User Comprehension Supports for LLM Agent Skill Specifications. arXiv:2605.19362,

  44. [44]

    Group- Evolving Agents: Open-Ended Self-Improvement via Experience Sharing.arXiv:2602.04837,

    Zhaotian Weng, Antonis Antoniades, Deepak Nathani, Zhen Zhang, Xiao Pu, and Xin Eric Wang. Group- Evolving Agents: Open-Ended Self-Improvement via Experience Sharing.arXiv:2602.04837,

  45. [45]

    EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle.arXiv:2510.16079,

    Rong Wu, Xiaoman Wang, Jianbiao Mei, Pinlong Cai, Daocheng Fu, Cheng Yang, Licheng Wen, Xue- meng Yang, Yufan Shen, Yuxin Wang, and Botian Shi. EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle.arXiv:2510.16079,

  46. [46]

    Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks

    Xiyang Wu, Zongxia Li, Guangyao Shi, Alexander Duffy, Tyler Marques, Matthew Lyle Olson, Tianyi Zhou, and Dinesh Manocha. Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks. arXiv:2604.20987, 2026a. Zhe Wu, Donglin Mo, Hongjin Lu, Junliang Xing, Jianheng Liu, Yuheng Jing, Kai Li, Kun Shao, Jianye Hao, and Yuanchun Shi. K2-Agent: Co-Evol...

  47. [47]

    SkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces.arXiv:2605.15215,

    Duling Xu, Zheng Chen, Zaifeng Pan, Jiawei Guan, Dong Dong, Jialin Li, and Bangzheng Pu. SkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces.arXiv:2605.15215,

  48. [48]

    Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward.arXiv:2602.12430,

    Renjun Xu and Yang Yan. Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward.arXiv:2602.12430,

  49. [49]

    Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks.arXiv:2510.08002, 2025a

    44 Published in Transactions on Machine Learning Research (07/2026) Cheng Yang, Xuemeng Yang, Licheng Wen, Daocheng Fu, Jianbiao Mei, Rong Wu, Pinlong Cai, Yufan Shen, Nianchen Deng, Botian Shi, Yu Qiao, and Haifeng Li. Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks.arXiv:2510.08002, 2025a. Min Yang, Jinghua Piao, Xu ...

  50. [50]

    AgentEvolver: Towards Efficient Self-Evolving Agent System.arXiv:2511.10395,

    Yunpeng Zhai, Shuchang Tao, Cheng Chen, Anni Zou, Ziqian Chen, Qingxu Fu, Shinji Mai, Li Yu, Jiaji Deng, Zouying Cao, Zhaoyang Liu, Bolin Ding, and Jingren Zhou. AgentEvolver: Towards Efficient Self-Evolving Agent System.arXiv:2511.10395,

  51. [51]

    EvoAgent: An Evolvable Agent Framework with Skill Learning and Multi-Agent Delegation.arXiv:2604.20133, 2026a

    Aimin Zhang, Jiajing Guo, Fuwei Jia, Chen Lv, Boyu Wang, and Fangzheng Li. EvoAgent: An Evolvable Agent Framework with Skill Learning and Multi-Agent Delegation.arXiv:2604.20133, 2026a. Barry Zhang, Keith Lazuka, and Mahesh Murag. Equipping Agents for the Real World with Agent Skills. Anthropic Engineering Blog,

  52. [52]

    Published October 16, 2025; updated December 18,

    URLhttps://www.anthropic.com/engineering/ equipping-agents-for-the-real-world-with-agent-skills. Published October 16, 2025; updated December 18,

  53. [53]

    SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards.arXiv:2602.21158, 2026b

    Dengjia Zhang, Xiaoou Liu, Lu Cheng, Yaqing Wang, Kenton Murray, and Hua Wei. SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards.arXiv:2602.21158, 2026b. Di Zhang. AgentDevel: Reframing Self-Evolving LLM Agents as Release Engineering.arXiv:2601.04620,

  54. [54]

    STARS: Skill-Triggered Audit for Request-Conditioned Invocation Safety in Agent Systems.arXiv:2604.10286, 2026c

    Guijia Zhang, Shu Yang, Xilin Gong, and Di Wang. STARS: Skill-Triggered Audit for Request-Conditioned Invocation Safety in Agent Systems.arXiv:2604.10286, 2026c. Hanrong Zhang, Shicheng Fan, Henry Peng Zou, Yankai Chen, Zhenting Wang, Jiayu Zhou, Chengze Li, Wei-Chieh Huang, Yifei Yao, Kening Zheng, Xue Liu, Xiaoxiao Li, and Philip S. Yu. CoEvoSkills: Sel...

  55. [55]

    Shanshan Zhong, Yi Lu, Jingjie Ning, Yibing Wan, Lihan Feng, Yuyi Ao, Leonardo F. R. Ribeiro, Markus Dreyer, Sean Ammirati, and Chenyan Xiong. SkillLearnBench: Benchmarking Continual Learning Meth- ods for Agent Skill Generation on Real-World Tasks.arXiv:2604.20087,

  56. [56]

    Memento-Skills: Let Agents Design Agents.arXiv:2603.18743, 2026a

    Huichi Zhou, Siyuan Guo, Anjie Liu, Zhongwei Yu, Ziqin Gong, Bowen Zhao, Zhixun Chen, Menglong Zhang, Yihang Chen, Jinsong Li, Runyu Yang, Qiangbin Liu, Xinlei Yu, Jianmin Zhou, Na Wang, Chunyang Sun, and Jun Wang. Memento-Skills: Let Agents Design Agents.arXiv:2603.18743, 2026a. Yingli Zhou, Shu Wang, Yaodong Su, Wenchuan Du, Yixiang Fang, and Xuemin Lin...

  57. [57]

    Method comparisons in the main text are based on mechanisms and evidence roles rather than on recency

    46 Published in Transactions on Machine Learning Research (07/2026) A Coding Protocol and Audit Materials This appendix records the audit machinery behind the taxonomy. Method comparisons in the main text are based on mechanisms and evidence roles rather than on recency. A.1 Screening and Coding Fields Each note was coded for: problem framing, artifact ty...

  58. [58]

    Found. Code fast Task A Exec flat Tool-maker + tool-user beats few-shot continued on next page 47 Published in Transactions on Machine Learning Research (07/2026) Table 11 (continued) Method Cluster Artif. Clock Trig. Operators Signal Store Headline Skill-aware RL SAGE(Wang et al., 2025a) RL Code 2TS Task A,R Rew flat Skill-integrated RL on AppWorld Skill...

  59. [59]

    Clock Trig

    ExecLib MD fast Task A,M,D Judge flat Prevalence- weighted consoli- dation continued on next page 48 Published in Transactions on Machine Learning Research (07/2026) Table 11 (continued) Method Cluster Artif. Clock Trig. Operators Signal Store Headline SkillClaw(Ma et al., 2026b) ExecLib MD fast User A,R,M,K XUser flat Cross-user skill evolution AutoSkill...

  60. [60]

    Clock Trig

    Param Mix slow RL D,B Teach flat Co-adapt diffi- culty + env ARISE(Li et al., 2026e) Param LoRA slow RL D,K Rew subsp Swarm PPO + PSO actions continued on next page 49 Published in Transactions on Machine Learning Research (07/2026) Table 11 (continued) Method Cluster Artif. Clock Trig. Operators Signal Store Headline EXIF(Yang et al., 2025b) Param LoRA s...

  61. [61]

    Infra Code fast Per K Judge flat Full-text re- trieval 80K skills SkVM(Chen et al., 2026a) Infra Code fast User A,R,C Exec DAG Skills as compil- able code SkillNet(Liang et al., 2026b) Infra MD fast Per A,K,P Judge graph Ontology + rela- tion graph SkillOrchestra(Wang et al., 2026c) Infra MD fast Per K,C Judge ontol Skill handbooks for routing SkillFlow-2...

  62. [62]

    rewards continued on next page 50 Published in Transactions on Machine Learning Research (07/2026) Table 11 (continued) Method Cluster Artif

    Safety Mix 2TS Task A,R,P,D Rew+Judge graph Audited skill graph + verif. rewards continued on next page 50 Published in Transactions on Machine Learning Research (07/2026) Table 11 (continued) Method Cluster Artif. Clock Trig. Operators Signal Store Headline ClawSafety(Wei et al.,