Pith. sign in

Paper Citation Record · LEDGER

Co-Evolving Skill Generation and Policy Optimization

As of 7 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 1 inbound Pith citation observation for arXiv:2606.08755.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.08755 v1

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T18:37:00.015083Z

measured 97 of 97 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:15:44.387347Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-30T07:24:21.952673Z

Reference resolution

96 of 96 outbound references displayed

  • verified exact73
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87283b5b-078c-4e8d-be7b-1daf5bfc13ce · outbound

This paper cites Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering.

Co-Evolving Skill Generation and Policy Optimization Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.897250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:bfec086f43e30aed146a48c6ea314da487a6040383032a6829f23f1711f6147b

Observation 638b214c-087f-4744-b9ce-c933abdeb97f · outbound

This paper cites Memento-skills: Let agents design agents.

Co-Evolving Skill Generation and Policy Optimization Memento-skills: Let agents design agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:26.451969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:87d39abf8b12f559e6051209ad20211b3ce789450a16ce41acb033d5fa33bd68

Observation a7182a07-895d-418b-8882-1874d3f53b4c · outbound

This paper cites Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents.

Co-Evolving Skill Generation and Policy Optimization Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.841643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:a8f9c9e26bded912a88293d40dc0c4b707904a5d95fbe4d74a0948f09942c865

Observation 0d0cbeb4-d1ea-4e5f-869e-d0ab8c28e771 · outbound

This paper cites AEL: Agent Evolving Learning for Open-Ended Environments.

Co-Evolving Skill Generation and Policy Optimization AEL: Agent Evolving Learning for Open-Ended Environments

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.884218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:6c348989e603600f076ef756fc378c81747ef494ceb6efa9e646a8e05c4c518a

Observation 76d76fec-c167-4b3f-9fb7-5296f2031039 · outbound

This paper cites Arunkumar, G.

Co-Evolving Skill Generation and Policy Optimization Arunkumar, G

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:47:26.424854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:e5b6910a0254208181f0cdadf3cba81dd687fd8ff390031322fffc8af72868c2

Observation f5b1c40d-9b05-4f54-b95d-ebeb54d0a35e · outbound

This paper cites The Landscape of Agentic Reinforcement Learning for LLMs: A Survey.

Co-Evolving Skill Generation and Policy Optimization The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.427426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:b6bf86f884a993516432e7d0d8fe6f73a1242fd2232af4b85f6a7f45ba2f225c

Observation 0497ade2-9dfe-4459-8852-8bb12cc15391 · outbound

This paper cites Agentic large language models, a survey.Journal of Artificial Intelligence Research, 84, 2025.

Co-Evolving Skill Generation and Policy Optimization Agentic large language models, a survey.Journal of Artificial Intelligence Research, 84, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:19a3d39c6e2d69a3065138dda0773a406a042465516a2a311bd0ff6bf3051b30

Observation 8be65db4-28ea-48bf-982b-fca614b9bafc · outbound

This paper cites Agentic Reasoning for Large Language Models.

Co-Evolving Skill Generation and Policy Optimization Agentic Reasoning for Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.432494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:2bb25e2fd6c253333ede69dad83d85d5b703b846793a3b953697a2792fab6aa5

Observation d1a10e44-d33a-427c-8c03-0bbaf2d76db8 · outbound

This paper cites Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens.

Co-Evolving Skill Generation and Policy Optimization Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:12.196403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:ff526e3138e2be3e5e8898133c7f12481679f91fdee5cf884272f23092a1e1db

Observation b1c34d4c-78c4-43c6-878a-8fce6eec9efa · outbound

This paper cites Brain-inspired graph multi- agent systems for llm reasoning,.

Co-Evolving Skill Generation and Policy Optimization Brain-inspired graph multi- agent systems for llm reasoning,

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:26.421959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:497411714dd1c568ad97b1d2f1405c79fea9662fb40a35c9ee819be117503ce0

Observation 90411c73-9e43-4b14-b4dd-32e7726d1c9a · outbound

This paper cites One Interaction Is Worth a Thousand Guesses: Benchmarking the Interactive Capabilities of Deep Research Agents.

Co-Evolving Skill Generation and Policy Optimization One Interaction Is Worth a Thousand Guesses: Benchmarking the Interactive Capabilities of Deep Research Agents

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.435251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:bec3ba637c88130237b5fdb93ae4493244568de654c58302f6ea992e0749ab6b

Observation 89b83765-680a-46f7-99a9-d882e543677b · outbound

This paper cites arXiv:2602.18998 (2026).

Co-Evolving Skill Generation and Policy Optimization arXiv:2602.18998 (2026)

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.892666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:9404f02255ac09350fa01b20a2831f8d42cc1d39aec62a47609147a873c5f9cf

Observation 68da68c9-854a-459a-ba3e-270e619acb1b · outbound

This paper cites Agentic reasoning: A streamlined framework for enhancing llm reasoning with agentic tools.

Co-Evolving Skill Generation and Policy Optimization Agentic reasoning: A streamlined framework for enhancing llm reasoning with agentic tools

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:784742e592a2dd4f70b26244bb2f6aa858acb7dfd414ee710c3e254cd62dffcb

Observation 16ee237f-0240-42ad-9f34-91489e651eb3 · outbound

This paper cites LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios.

Co-Evolving Skill Generation and Policy Optimization LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.899087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:25967b4486893de43160b8b73d84d8a3b01f2ce9844911eba670aaeea8494bc0

Observation 57afa478-200c-4c24-a405-40d2ad93c06e · outbound

This paper cites Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward.

Co-Evolving Skill Generation and Policy Optimization Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.876462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:d74803b7728b102fcfc9bbc8f0e73a17c0573fdf487f145c31ef9f65c97c221a

Observation 85cba283-49fb-438a-a9d3-bc994d2c039e · outbound

This paper cites SoK: Agentic Skills -- Beyond Tool Use in LLM Agents.

Co-Evolving Skill Generation and Policy Optimization SoK: Agentic Skills -- Beyond Tool Use in LLM Agents

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.836483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:82e57d43b16dfab9c31de1ed5f4e052ebedc0ceb1b995dd4bda1eb652fa7f719

Observation 14eee291-7dab-4fbf-9a2e-cdfdb0209697 · outbound

This paper cites SkillX: Automatically Constructing Skill Knowledge Bases for Agents.

Co-Evolving Skill Generation and Policy Optimization SkillX: Automatically Constructing Skill Knowledge Bases for Agents

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.853843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:dcbd40bf9bb8def7967f2d543224bc239488b24e321a7b2b208bbbcd2faa642a

Observation e647e662-9de1-42a9-9dfe-d168bdaa51d2 · outbound

This paper cites SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning.

Co-Evolving Skill Generation and Policy Optimization SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.850746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:e90e4e25cf75f6d9491df2d621a38151f4fc0e4b04ce267d45c36d546587a8c1

Observation 05dcb7c0-96c8-4bcb-888d-f4afca18a68c · outbound

This paper cites Tooltree: Efficient llm agent tool planning via dual-feedback monte carlo tree search and bidirectional pruning.

Co-Evolving Skill Generation and Policy Optimization Tooltree: Efficient llm agent tool planning via dual-feedback monte carlo tree search and bidirectional pruning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.840525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:6d099066a94a4910012e3ff893001d79e2c7ed9e1183402d7edfda298726403d

Observation b5c9c2ff-84b8-48a9-bb4f-ba78743248be · outbound

This paper cites Agentic Tool Use in Large Language Models.

Co-Evolving Skill Generation and Policy Optimization Agentic Tool Use in Large Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.824054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:a21332a4e742f1b1d2db2f8af7f9a433e6afae329ca4575b04e1100678a4b112

Observation ed9f1dee-1e6e-4a63-8d81-9c960b78a778 · outbound

This paper cites arXiv preprint , year =.

Co-Evolving Skill Generation and Policy Optimization arXiv preprint , year =

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:57:25.890044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:60b831816b4c1726c325f7accd7a6e7246ae5a32c9893ed2010456cced42c94d

Observation 54e3a7a8-b2d5-46fa-b0f2-dff021b7812b · outbound

This paper cites an unresolved cited work.

Co-Evolving Skill Generation and Policy Optimization Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:46c6a12f9de118800ccbf0cc3f4723d4dfb55c96e19ea91390aec2026ff04053

Observation dc31ce23-ccac-4721-b49b-127489c56fe4 · outbound

This paper cites CAS- CADE: Cumulative agentic skill creation through autonomous development and evolution.

Co-Evolving Skill Generation and Policy Optimization CAS- CADE: Cumulative agentic skill creation through autonomous development and evolution

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.834911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:e15366705425477e28ef528479b99eabc4e6f9025c59698f4355b68aff886180

Observation 325bec0c-3a15-4f56-91aa-7b14ab9159da · outbound

This paper cites SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills.

Co-Evolving Skill Generation and Policy Optimization SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.851525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:2dd310c62b7f22a8ee4ec36f5133c9f2c30f0fc139e3aafcef20f7070d0ae012

Observation f519e037-47cb-4232-93ec-7bcf3a0b64f0 · outbound

This paper cites Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback.

Co-Evolving Skill Generation and Policy Optimization Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.856666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:c5716123a4edd0ac199e80eaa0a994ae7fe4bbb521138f11b80901ebd9fd01bb

Observation 37f93bd7-0806-4ebd-b779-d35f1b75ac20 · outbound

This paper cites Proposer-agent-evaluator (pae): Autonomous skill discovery for foundation model internet agents.

Co-Evolving Skill Generation and Policy Optimization Proposer-agent-evaluator (pae): Autonomous skill discovery for foundation model internet agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:265e340228e1545ab2c1c89a11b8cdb2151e05c849b7cdab08418d9f8376e287

Observation eeeba68a-3cfd-4114-aa41-032163ffaf83 · outbound

This paper cites Memp: Exploring Agent Procedural Memory.

Co-Evolving Skill Generation and Policy Optimization Memp: Exploring Agent Procedural Memory

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.438012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:8ec19adf380812d62b357fdf01df499bed0b7a37063acb94f296d54d0d3340ef

Observation 0d90c69c-de88-4c3e-b763-96f3c8733781 · outbound

This paper cites Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution.

Co-Evolving Skill Generation and Policy Optimization Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.817912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:42c5a310cb17a67edabf99f08621dec11d91c75e7c1a538e1f9ea8f885505f6f

Observation 2ca10401-ff03-4fdf-9869-f16180eb46d0 · outbound

This paper cites Meta Context Engineer- ing via Agentic Skill Evolution, February 2026.https://arxiv.org/abs/2601.21557.

Co-Evolving Skill Generation and Policy Optimization Meta Context Engineer- ing via Agentic Skill Evolution, February 2026.https://arxiv.org/abs/2601.21557

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.906423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:f4adb5ae9f1b1126712072619c424d28080014c3ad46eefd41657281c7457468

Observation 744d9e5e-7c9f-4e9e-9484-f09d527eebc6 · outbound

This paper cites SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization.

Co-Evolving Skill Generation and Policy Optimization SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.880988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:70df4b31e7556cde7a96cca4c8ef2ed0da3f230016baa2180621a8b8a8dfc37e

Observation 7deaae6e-99eb-4aad-b5ef-4e2392a88cf6 · outbound

This paper cites Dynamic Dual-Granularity Skill Bank for Agentic RL.

Co-Evolving Skill Generation and Policy Optimization Dynamic Dual-Granularity Skill Bank for Agentic RL

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.894872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:1c8bddd8534b40da177c924fc5ec332a95ef1037fae8dce4850d1a29eba9e2fb

Observation 8a421fcc-bf46-49f7-ae95-a1b50906259f · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Co-Evolving Skill Generation and Policy Optimization Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.826927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:af70d529a1557ed047d1299cde405b7cbbbc711623373ad04e00ce2de43fe461

Observation b719c3d5-d8ab-40ae-ac4a-d5cda4723c48 · outbound

This paper cites Skillact: Using skill abstractions improves llm agents.

Co-Evolving Skill Generation and Policy Optimization Skillact: Using skill abstractions improves llm agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:d9065dd31d385cd88fcced7e11cccf32fcad37fcb564b4206df2fd61abfb9ea8

Observation 6c42ef35-cb9c-409d-9caa-89719d3b8d7e · outbound

This paper cites Inducing Programmatic Skills for Agentic Tasks.

Co-Evolving Skill Generation and Policy Optimization Inducing Programmatic Skills for Agentic Tasks

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:26.441022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:d727fe69e8b550456fafe95bd5490b42592823ee58ed00e0c1e12b46ec3935ba

Observation 1ba98e6f-a887-4a81-933d-b8eee3dc8955 · outbound

This paper cites Agent skills from the perspective of procedural memory: A survey.Authorea Preprints, 2026.

Co-Evolving Skill Generation and Policy Optimization Agent skills from the perspective of procedural memory: A survey.Authorea Preprints, 2026

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:b1b80c8c0cb36d10c56324bc653e695dded64c53d80725485645d696461e80f2

Observation e68be63c-88a4-4e21-b367-9d7217c0f6b9 · outbound

This paper cites Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents.

Co-Evolving Skill Generation and Policy Optimization Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.839040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:6e16e779e0607d7b51a8badd25bcabcdc3affb95ca438d38309bcbdc8c084a4b

Observation d062d8a4-04bc-4c49-85a0-2825b2cb9cad · outbound

This paper cites Automating skill acquisition through large-scale mining of open-source agentic repositories: A framework for multi-agent procedural knowledge extraction.

Co-Evolving Skill Generation and Policy Optimization Automating skill acquisition through large-scale mining of open-source agentic repositories: A framework for multi-agent procedural knowledge extraction

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.873879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:6c2b5e6209bd1851c0c9e943e623a2c39c9ec199012928ef085ae062834517bb

Observation dcf9c615-0ad7-4c20-8730-d3d04af93226 · outbound

This paper cites SKILLFOUNDRY: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources.

Co-Evolving Skill Generation and Policy Optimization SKILLFOUNDRY: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.903852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:f34ce9b8f90769fb5855d3b4a5bb0e8e9f227fa860043d6b7139fbcee8a64db6

Observation 3cc258b4-1d36-4cfc-a613-e57d83e97bb6 · outbound

This paper cites Autoskill: Experience-driven lifelong learning via skill self-evolution.

Co-Evolving Skill Generation and Policy Optimization Autoskill: Experience-driven lifelong learning via skill self-evolution

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.859355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:5ed0867d859529a9dc7ea3834a3a35373fa783734d4f98e6fa0e4947f99ccf61

Observation dc25d14d-789c-43ba-8341-3bc55a6f7c60 · outbound

This paper cites Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills.

Co-Evolving Skill Generation and Policy Optimization Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.894014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:49c264d751d5d0101afe48961c18106e5adaf5c1b13ab65d1d66c1ff24482736

Observation 7984dccf-c9dd-4962-8f93-a7bf6651068e · outbound

This paper cites Reinforcement Learning for Self-Improving Agent with Skill Library.

Co-Evolving Skill Generation and Policy Optimization Reinforcement Learning for Self-Improving Agent with Skill Library

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.443531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:bc8bd86d4b6b57cc10cbe9b77a32386437362ec222d1be3f04748c6cfbf7bc29

Observation 3a2fc74d-8077-4fb6-8813-a65f968861c7 · outbound

This paper cites Arise: Agent reasoning with intrinsic skill evolution in hierarchical reinforcement learning.

Co-Evolving Skill Generation and Policy Optimization Arise: Agent reasoning with intrinsic skill evolution in hierarchical reinforcement learning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.862516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:f5c4162cbc7dacd4b143bad40ff2270cef6aba14faa434b2104b3aa27ab067f5

Observation 1f58cbe1-ae32-40d0-8ecf-ab1507d6ccd1 · outbound

This paper cites CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification.

Co-Evolving Skill Generation and Policy Optimization CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.901517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:c1f80348ca544bafa4401051cf21521e4ca6b240894c501297a74f6bfc1a1454

Observation 63b4c900-7527-4b2f-8f13-65129af8004e · outbound

This paper cites SkillClaw: Let Skills Evolve Collectively with Agentic Evolver.

Co-Evolving Skill Generation and Policy Optimization SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.830923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:f80d677616bc470294047aad12a06cac44f35a918ca9e03a85041eefe4e58cd4

Observation 600aaf1c-be28-41ff-8ac2-792c42c3f4b2 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

Co-Evolving Skill Generation and Policy Optimization SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.878550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:27f9304c6c0aaed82804958f8890d8fd6c7fd680e407f674ad0e62aa0a762056

Observation 004a7c61-0a79-441a-9ffb-62bc2fd2846c · outbound

This paper cites SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks.

Co-Evolving Skill Generation and Policy Optimization SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.440552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:38ecda0ac626fbbfdae5c9bcee223ea36c9c2b4cabec79b39861383d3d0cc4d4

Observation 48386583-dbbb-4203-814d-6928691314a9 · outbound

This paper cites How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings.

Co-Evolving Skill Generation and Policy Optimization How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.849136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:726436d6be2ccc8482fac5f22494ec5807890f431cabfb031eb12250f24555eb

Observation a1255679-a3c9-425d-99e2-91210c1e4572 · outbound

This paper cites Skilltester: Benchmarking utility and security of agent skills.

Co-Evolving Skill Generation and Policy Optimization Skilltester: Benchmarking utility and security of agent skills

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:26.457806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:46f4bf02f40235df97ede7c200e8e0a26ce14dd6a5a934fcb6b2420a7c7a5791

Observation 59a2cabd-dc1b-4498-b3ae-41a918821918 · outbound

This paper cites SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents.

Co-Evolving Skill Generation and Policy Optimization SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.460447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:230f91378ab68cdc5b57527c64a72d8e1bb89af676f50ee517d120279d79560f

Observation d08258a8-df2e-4639-ac6a-e3f2b12b5e61 · outbound

This paper cites Welcome to the era of experience.Google AI, 1:11, 2025.

Co-Evolving Skill Generation and Policy Optimization Welcome to the era of experience.Google AI, 1:11, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:054106a527da47d83aa57fdb8856757d7962fb47a10c0e33c4c7cdda3c79bf5a

Observation bd293901-17f6-4896-b9cc-d6087f836954 · outbound

This paper cites Memory for autonomous llm agents: Mechanisms, evaluation, and emerging frontiers.

Co-Evolving Skill Generation and Policy Optimization Memory for autonomous llm agents: Mechanisms, evaluation, and emerging frontiers

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:26.449273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:3a48854c290a303cbcdc5541af9e7ee48f44a98300231e64efdb52459ccb2724

Observation cc670e79-6e3e-4094-a0c4-a1ebd6290b17 · outbound

This paper cites Memento: Fine-tuning LLM Agents without Fine-tuning LLMs.

Co-Evolving Skill Generation and Policy Optimization Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.867397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:39fadb5cfc2dc1d03ed0271de43c3d3a5661bd5fcb050cc94627ce54d8e70062

Observation b44203e5-9781-40bf-b35f-be9e65c24968 · outbound

This paper cites Scaling agent learning via experience synthesis.

Co-Evolving Skill Generation and Policy Optimization Scaling agent learning via experience synthesis

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.904457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:476adc2fb81f2111e09ab096b0640ccce47fdb813dc4a34adf60e3278955c9f9

Observation 3a71f265-0b05-4144-be95-d1611ff72898 · outbound

This paper cites Agent Learning via Early Experience.

Co-Evolving Skill Generation and Policy Optimization Agent Learning via Early Experience

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.878799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:f155d7bb64a9cea9a890a940645b3722512807e3b95e1730a13e4272cc8d5bcb

Observation a5e360c6-d55c-426a-85c8-ea046b3da241 · outbound

This paper cites MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory.

Co-Evolving Skill Generation and Policy Optimization MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.899564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:8f27e025bb3557459adfbdc473164266c75fa7705615158e6e18fdd215d27ade

Observation 4c6afe53-99af-4150-b9d3-2959a043f709 · outbound

This paper cites Trajectory-informed memory generation for self-improving agent systems.

Co-Evolving Skill Generation and Policy Optimization Trajectory-informed memory generation for self-improving agent systems

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.887645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:3ab362ececb0b36b4e8d947339677685ef1b52b0d81759ac24cc98122d20d06c

Observation a5db41fc-e95d-4093-b10e-50799bead76e · outbound

This paper cites ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory.

Co-Evolving Skill Generation and Policy Optimization ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.910403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:ea330077915604761433c64b42d323bb464a484ff206f33994a98954e751ca01

Observation 05e77c32-6688-4ac0-aeff-e9378f78f76f · outbound

This paper cites R2d2: Remembering, replaying and dynamic decision making with a reflective agentic memory.

Co-Evolving Skill Generation and Policy Optimization R2d2: Remembering, replaying and dynamic decision making with a reflective agentic memory

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:8c308ae9afd9bfb7d1df4a85345bbf41131a0dcac5560653755fa596b5c53ab5

Observation 24ce572e-c007-471f-8b30-94f42ad805d2 · outbound

This paper cites Dynamic cheatsheet: Test-time learning with adaptive memory.

Co-Evolving Skill Generation and Policy Optimization Dynamic cheatsheet: Test-time learning with adaptive memory

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:649409343026f887065c7d9f24607795097966ade6a864158ad390ca62d9f555

Observation ed5a25d2-4d6f-49a6-89ec-81ce22cd7023 · outbound

This paper cites Flex: Continuous agent evolution via forward learning from experience.

Co-Evolving Skill Generation and Policy Optimization Flex: Continuous agent evolution via forward learning from experience

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.776640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:bbddf957a308875c2cf52d9388cb565dc697c03c17bcf57c4005e94ccdce2d19

Observation d4072ae0-91a9-4c30-b42b-3f320d3ec772 · outbound

This paper cites Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents.

Co-Evolving Skill Generation and Policy Optimization Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.828262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:15e7e5b0172e50ed4fa9499edc60a299b6bf2156cf3657bb517e1caf73d06492

Observation 9c57175e-7d8f-47db-9a70-d85c37009f51 · outbound

This paper cites Legomem: Modular procedural memory for multi-agent llm systems for workflow automation.

Co-Evolving Skill Generation and Policy Optimization Legomem: Modular procedural memory for multi-agent llm systems for workflow automation

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.837741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:a42f2df3cac318edd2584e9491b33ad8644e223c7b59ba97ca180e2d37e37d09

Observation 5b8faa6b-4820-41d3-bea7-a7fa4983d446 · outbound

This paper cites Agent kb: Leveraging cross-domain experience for agentic problem solving.

Co-Evolving Skill Generation and Policy Optimization Agent kb: Leveraging cross-domain experience for agentic problem solving

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.907637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:ea976eeae504bbabcf6713fdf93489d14affb5555bb13c308fbd1e5bdc8555cf

Observation f0fe70a6-bcd7-432b-95d9-ee64fd3b2e0e · outbound

This paper cites G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems.

Co-Evolving Skill Generation and Policy Optimization G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:26.421753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:b5f5106a33b7e4a318163307c80ee032654b694c6f24dd0069ace8947117578a

Observation 81f77580-c744-4f1a-95c8-82f1823fb822 · outbound

This paper cites EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle.

Co-Evolving Skill Generation and Policy Optimization EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.424338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:90230202545aeb728591bd43d24fa8e5a5c6495deab8cb993bcf376b2ae4c54f

Observation 838d8488-37bb-4416-af6b-c11213ffba73 · outbound

This paper cites Graph-based agent memory: Taxonomy, techniques, and applications.

Co-Evolving Skill Generation and Policy Optimization Graph-based agent memory: Taxonomy, techniques, and applications

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.845493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:0c942ce36b900f24d9c827fb7e6dcf41909d9cc455026e6ce764b4aa0ac56273

Observation 1934cf68-119b-4c96-95c2-1cd1d79aea29 · outbound

This paper cites MemEvolve: Meta-Evolution of Agent Memory Systems.

Co-Evolving Skill Generation and Policy Optimization MemEvolve: Meta-Evolution of Agent Memory Systems

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.890807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:7252e42510be690aa6ef8011838b8f8387949eead6fbe6e91687c8917998d173

Observation 8d55a5ce-5ba9-4013-9431-9584deaa2662 · outbound

This paper cites Agentevolver: Towards efficient self-evolving agent system.

Co-Evolving Skill Generation and Policy Optimization Agentevolver: Towards efficient self-evolving agent system

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:26.454755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:e837fd01a707e9c3222e2eff30f7ec14e4c5733cd5aa7218ed243e6e0e044262

Observation 8b5c4b91-481f-4016-804d-6b844681a852 · outbound

This paper cites Building self-evolving agents via experience-driven lifelong learning: A framework and benchmark.

Co-Evolving Skill Generation and Policy Optimization Building self-evolving agents via experience-driven lifelong learning: A framework and benchmark

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:26.449485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:30c728437f35f173efd9b0488daa466082fe57eed4b0607d8bccf28fdfada062

Observation 987f4751-8af1-4d9d-9111-2a728b5bbf26 · outbound

This paper cites Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning.

Co-Evolving Skill Generation and Policy Optimization Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.901885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:67c21dbbd2f79ba96c6f17ea53eefdde9e9122e41cf5b9f3d9edac804cfeffa6

Observation c858693e-e2d5-4671-b494-3ad29d1ddd0c · outbound

This paper cites Adaptive Memory Crystallization for Autonomous AI Agent Learning in Dynamic Environments.

Co-Evolving Skill Generation and Policy Optimization Adaptive Memory Crystallization for Autonomous AI Agent Learning in Dynamic Environments

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.413524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:c5f1fd5dbbaa5ae02d12f089c026e3426a56d10a80c7b54db7b854cea2d8168b

Observation 8560c48b-187d-43bd-a22a-43e1a5f0de09 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Co-Evolving Skill Generation and Policy Optimization ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.416212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:bcef0297958c9c0ca707ced6ecf2406e8eb40f5ced46e818b3d8f9a0e6172f79

Observation 3cec899c-61a7-4953-a26d-fb19336f1751 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757, 2022.

Co-Evolving Skill Generation and Policy Optimization Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757, 2022

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:f494875ff7337c1c0940b5356e0356a6d8fb3087566814f8bd38c54d855641dd

Observation 194a0efe-d2d0-4622-8392-25a94e9e0617 · outbound

This paper cites Qwen3 Technical Report.

Co-Evolving Skill Generation and Policy Optimization Qwen3 Technical Report

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.429860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:07e13558fc528949813d948180f57ef27c6a025027814ee3d40f778be2aaa456

Observation 07bcbd80-c110-4399-90e7-0132421a4958 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Co-Evolving Skill Generation and Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.443218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:23cd99e561345db1252f9fe7874fecaa775b76f895f93ddd7b9537dd4e726279

Observation 42349223-83ac-4946-a1e6-3342dd2fcbd8 · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347.

Co-Evolving Skill Generation and Policy Optimization The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:26.426988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:003c11031a28fbd9dd26d1dd18e92e7e426ce66fa23b145fc73f4dc58a43ce37

Observation 4ab0146b-dcbc-4e50-9d96-a814996eb7d1 · outbound

This paper cites Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019.

Co-Evolving Skill Generation and Policy Optimization Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:233a8a28b60f1e75b926debd9ef0e41a142a03b32969cf402433e44921293ee3

Observation cd5e89d8-8b27-44f1-bc96-f2302069de34 · outbound

This paper cites Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension.

Co-Evolving Skill Generation and Policy Optimization Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:f89e0b982bf53dc3ea9a86d8d6629c36bdd9a854944f5bc137f771e707c66001

Observation ae7693b3-c8b7-4137-930a-2a8a59b385fe · outbound

This paper cites When not to trust language models: Investigating effectiveness of parametric and non-parametric memories.

Co-Evolving Skill Generation and Policy Optimization When not to trust language models: Investigating effectiveness of parametric and non-parametric memories

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:e672e7254244962ada7df8cbef6b689a808ec1d9a56a1a8b843351d960a1c5d0

Observation 801c0f63-9afa-4f4b-ad94-e8d52133ed60 · outbound

This paper cites Hotpotqa: A dataset for diverse, explainable multi-hop question answering.

Co-Evolving Skill Generation and Policy Optimization Hotpotqa: A dataset for diverse, explainable multi-hop question answering

Reference 80

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:61782d8354eb19d2f6f44436ea015a820801858fc952484d3326283154e187f1

Observation 1bc986e8-6f15-47fd-8abc-00cd3cb2434e · outbound

This paper cites Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps.

Co-Evolving Skill Generation and Policy Optimization Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps

Reference 81

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:ff4ee42649029912989df3a7b8ef06686c9fc5d494dc38b0a4f5d9186dc25cb3

Observation d04772b8-43f2-4702-b346-6f8d16b35b3d · outbound

This paper cites Musique: Multi-hop questions via single-hop question composition.Transactions of the Association for Computational Linguistics, 10:539–554, 2022.

Co-Evolving Skill Generation and Policy Optimization Musique: Multi-hop questions via single-hop question composition.Transactions of the Association for Computational Linguistics, 10:539–554, 2022

Reference 82

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:9cc92bd17315c22030119a4d5d848a489f71ae6f7e2499185fe331cf840b0d1d

Observation ce80d306-3640-4842-81c4-ca6fef390c23 · outbound

This paper cites Measuring and narrowing the compositionality gap in language models.

Co-Evolving Skill Generation and Policy Optimization Measuring and narrowing the compositionality gap in language models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:bcee987d6e075a518ca06ac4429dcf102e8f967563810630aaead67f2d806a2d

Observation 8af126c3-526c-4b8e-817a-37f560c4602d · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Co-Evolving Skill Generation and Policy Optimization Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.429486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:326c0a7bd26cc4adc8036f4cf74f1721de9ef7ecf23a35b25d23999fc6f514bc

Observation 6009b2cf-53a8-4e43-a23f-e0faac7ac119 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Co-Evolving Skill Generation and Policy Optimization ReAct: Synergizing Reasoning and Acting in Language Models

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.432235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:b88791339ecda02a8947394a04b137a8d9df83d6351a65bae96b5626e20c2391

Observation ae2ea331-8cbb-4e54-a599-8aa4334fc6c1 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652, 2023.

Co-Evolving Skill Generation and Policy Optimization Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652, 2023

Reference 86

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:7dbcc7b02b31577a27a41d2c90f81701db37d84efa4918fc76b99451b797498a

Observation f0a149d7-9968-4fa6-8131-d788eec6be4d · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

Co-Evolving Skill Generation and Policy Optimization Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.446146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:6715617a220e0810fa1ed5200b4d4f968a951300edcdcaed87a4107202e7600f

Observation f940e363-7245-4592-b44f-e2ee9638a819 · outbound

This paper cites Expel: Llm agents are experiential learners.

Co-Evolving Skill Generation and Policy Optimization Expel: Llm agents are experiential learners

Reference 88

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:62112ff57610c1162b920d11d72b44ff699022e84999d0d37fd26c7b16b6f3f2

Observation 6cbc82e2-1e9c-4321-b6ca-3c2c303f78aa · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms.

Co-Evolving Skill Generation and Policy Optimization Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms

Reference 89

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:8659f2c81ec5f1d8d242221eaa9d453aed7cb7276f4bf637a5e873c853a3f8f2

Observation 0901fb9c-d856-4d7b-9198-fb19603db4d6 · outbound

This paper cites SimpleMem: Efficient Lifelong Memory for LLM Agents.

Co-Evolving Skill Generation and Policy Optimization SimpleMem: Efficient Lifelong Memory for LLM Agents

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.435624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:6bd329f71ecb2fa9a058a9f9ce8863bcb0b6344a65fee7a417eaa7ad2296264b

Observation 8c140f9f-5923-49fb-992a-73b8cdbd086e · outbound

This paper cites Search-o1: Agentic search-enhanced large reasoning models.

Co-Evolving Skill Generation and Policy Optimization Search-o1: Agentic search-enhanced large reasoning models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-06-27T18:37:00.015083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:05679d995e9929c88821f875a2e98d106b3721ac1b0e55b202c7636203805887

Observation b11de2d0-1b76-4da9-b8f7-3e8f9e911bbb · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Co-Evolving Skill Generation and Policy Optimization Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.842847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:51687d122dac17b14823c599146e507ea0c82b99e2048e630696db9b7522076c

Observation 9dc86ad0-4ac1-4ba5-9ff1-3600745d6288 · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

Co-Evolving Skill Generation and Policy Optimization ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.864373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:54768737e554c2797ea41636cb7d936d1ff4ec57cde8f5973377c68fa932d321

Observation ffaea00e-6581-4880-a5f1-68dd30ab359c · outbound

This paper cites StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization.

Co-Evolving Skill Generation and Policy Optimization StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.881525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:61c4d76bed11246736400091fd3f7fa56bae605ea50c17358acb66ae4d1826eb

Observation ac9f0a0b-3ab0-49aa-a391-3851bf17818b · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Co-Evolving Skill Generation and Policy Optimization Group-in-Group Policy Optimization for LLM Agent Training

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.896537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:ea85c097197c2da8d30224d117fd8aa94b5d576fc22958c6728d84ab1cf105c5

Observation 889cc33b-1ce3-4984-829d-c17a53a6fbaf · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

Co-Evolving Skill Generation and Policy Optimization Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.861682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:3c0792fe1f0bff3a499f10588a659927ac9a9dfa6f5cea416e94e6d7174b27f3

Pith citing papers

Observation 22bfe044-6bb1-4c4c-9bbd-3c7adba51a78 · inbound

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation cites this paper.

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation Co-Evolving Skill Generation and Policy Optimization

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:24:21.953954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-30T07:15:44.387347Z digest=sha256:e568bcc914d503b6ca86f6d14d21fbf8d1f44eca874a9c7294ecc07f49f91353