Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T18:37:00.015083Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 1 inbound Pith citation observation for arXiv:2606.08755.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T18:37:00.015083Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T07:15:44.387347Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-30T07:24:21.952673Z
96 of 96 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 87283b5b-078c-4e8d-be7b-1daf5bfc13ce · outbound
Co-Evolving Skill Generation and Policy Optimization Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 638b214c-087f-4744-b9ce-c933abdeb97f · outbound
Co-Evolving Skill Generation and Policy Optimization Memento-skills: Let agents design agents
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a7182a07-895d-418b-8882-1874d3f53b4c · outbound
Co-Evolving Skill Generation and Policy Optimization Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0d0cbeb4-d1ea-4e5f-869e-d0ab8c28e771 · outbound
Co-Evolving Skill Generation and Policy Optimization AEL: Agent Evolving Learning for Open-Ended Environments
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 76d76fec-c167-4b3f-9fb7-5296f2031039 · outbound
Co-Evolving Skill Generation and Policy Optimization Arunkumar, G
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f5b1c40d-9b05-4f54-b95d-ebeb54d0a35e · outbound
Co-Evolving Skill Generation and Policy Optimization The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0497ade2-9dfe-4459-8852-8bb12cc15391 · outbound
Co-Evolving Skill Generation and Policy Optimization Agentic large language models, a survey.Journal of Artificial Intelligence Research, 84, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8be65db4-28ea-48bf-982b-fca614b9bafc · outbound
Co-Evolving Skill Generation and Policy Optimization Agentic Reasoning for Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d1a10e44-d33a-427c-8c03-0bbaf2d76db8 · outbound
Co-Evolving Skill Generation and Policy Optimization Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b1c34d4c-78c4-43c6-878a-8fce6eec9efa · outbound
Co-Evolving Skill Generation and Policy Optimization Brain-inspired graph multi- agent systems for llm reasoning,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 90411c73-9e43-4b14-b4dd-32e7726d1c9a · outbound
Co-Evolving Skill Generation and Policy Optimization One Interaction Is Worth a Thousand Guesses: Benchmarking the Interactive Capabilities of Deep Research Agents
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 89b83765-680a-46f7-99a9-d882e543677b · outbound
Co-Evolving Skill Generation and Policy Optimization arXiv:2602.18998 (2026)
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 68da68c9-854a-459a-ba3e-270e619acb1b · outbound
Co-Evolving Skill Generation and Policy Optimization Agentic reasoning: A streamlined framework for enhancing llm reasoning with agentic tools
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16ee237f-0240-42ad-9f34-91489e651eb3 · outbound
Co-Evolving Skill Generation and Policy Optimization LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 57afa478-200c-4c24-a405-40d2ad93c06e · outbound
Co-Evolving Skill Generation and Policy Optimization Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 85cba283-49fb-438a-a9d3-bc994d2c039e · outbound
Co-Evolving Skill Generation and Policy Optimization SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 14eee291-7dab-4fbf-9a2e-cdfdb0209697 · outbound
Co-Evolving Skill Generation and Policy Optimization SkillX: Automatically Constructing Skill Knowledge Bases for Agents
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e647e662-9de1-42a9-9dfe-d168bdaa51d2 · outbound
Co-Evolving Skill Generation and Policy Optimization SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 05dcb7c0-96c8-4bcb-888d-f4afca18a68c · outbound
Co-Evolving Skill Generation and Policy Optimization Tooltree: Efficient llm agent tool planning via dual-feedback monte carlo tree search and bidirectional pruning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b5c9c2ff-84b8-48a9-bb4f-ba78743248be · outbound
Co-Evolving Skill Generation and Policy Optimization Agentic Tool Use in Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ed9f1dee-1e6e-4a63-8d81-9c960b78a778 · outbound
Co-Evolving Skill Generation and Policy Optimization arXiv preprint , year =
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 54e3a7a8-b2d5-46fa-b0f2-dff021b7812b · outbound
Co-Evolving Skill Generation and Policy Optimization Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc31ce23-ccac-4721-b49b-127489c56fe4 · outbound
Co-Evolving Skill Generation and Policy Optimization CAS- CADE: Cumulative agentic skill creation through autonomous development and evolution
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 325bec0c-3a15-4f56-91aa-7b14ab9159da · outbound
Co-Evolving Skill Generation and Policy Optimization SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f519e037-47cb-4232-93ec-7bcf3a0b64f0 · outbound
Co-Evolving Skill Generation and Policy Optimization Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 37f93bd7-0806-4ebd-b779-d35f1b75ac20 · outbound
Co-Evolving Skill Generation and Policy Optimization Proposer-agent-evaluator (pae): Autonomous skill discovery for foundation model internet agents
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeeba68a-3cfd-4114-aa41-032163ffaf83 · outbound
Co-Evolving Skill Generation and Policy Optimization Memp: Exploring Agent Procedural Memory
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0d90c69c-de88-4c3e-b763-96f3c8733781 · outbound
Co-Evolving Skill Generation and Policy Optimization Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2ca10401-ff03-4fdf-9869-f16180eb46d0 · outbound
Co-Evolving Skill Generation and Policy Optimization Meta Context Engineer- ing via Agentic Skill Evolution, February 2026.https://arxiv.org/abs/2601.21557
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 744d9e5e-7c9f-4e9e-9484-f09d527eebc6 · outbound
Co-Evolving Skill Generation and Policy Optimization SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7deaae6e-99eb-4aad-b5ef-4e2392a88cf6 · outbound
Co-Evolving Skill Generation and Policy Optimization Dynamic Dual-Granularity Skill Bank for Agentic RL
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a421fcc-bf46-49f7-ae95-a1b50906259f · outbound
Co-Evolving Skill Generation and Policy Optimization Voyager: An Open-Ended Embodied Agent with Large Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b719c3d5-d8ab-40ae-ac4a-d5cda4723c48 · outbound
Co-Evolving Skill Generation and Policy Optimization Skillact: Using skill abstractions improves llm agents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c42ef35-cb9c-409d-9caa-89719d3b8d7e · outbound
Co-Evolving Skill Generation and Policy Optimization Inducing Programmatic Skills for Agentic Tasks
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1ba98e6f-a887-4a81-933d-b8eee3dc8955 · outbound
Co-Evolving Skill Generation and Policy Optimization Agent skills from the perspective of procedural memory: A survey.Authorea Preprints, 2026
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68be63c-88a4-4e21-b367-9d7217c0f6b9 · outbound
Co-Evolving Skill Generation and Policy Optimization Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d062d8a4-04bc-4c49-85a0-2825b2cb9cad · outbound
Co-Evolving Skill Generation and Policy Optimization Automating skill acquisition through large-scale mining of open-source agentic repositories: A framework for multi-agent procedural knowledge extraction
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dcf9c615-0ad7-4c20-8730-d3d04af93226 · outbound
Co-Evolving Skill Generation and Policy Optimization SKILLFOUNDRY: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3cc258b4-1d36-4cfc-a613-e57d83e97bb6 · outbound
Co-Evolving Skill Generation and Policy Optimization Autoskill: Experience-driven lifelong learning via skill self-evolution
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dc25d14d-789c-43ba-8341-3bc55a6f7c60 · outbound
Co-Evolving Skill Generation and Policy Optimization Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7984dccf-c9dd-4962-8f93-a7bf6651068e · outbound
Co-Evolving Skill Generation and Policy Optimization Reinforcement Learning for Self-Improving Agent with Skill Library
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3a2fc74d-8077-4fb6-8813-a65f968861c7 · outbound
Co-Evolving Skill Generation and Policy Optimization Arise: Agent reasoning with intrinsic skill evolution in hierarchical reinforcement learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1f58cbe1-ae32-40d0-8ecf-ab1507d6ccd1 · outbound
Co-Evolving Skill Generation and Policy Optimization CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 63b4c900-7527-4b2f-8f13-65129af8004e · outbound
Co-Evolving Skill Generation and Policy Optimization SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 600aaf1c-be28-41ff-8ac2-792c42c3f4b2 · outbound
Co-Evolving Skill Generation and Policy Optimization SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 004a7c61-0a79-441a-9ffb-62bc2fd2846c · outbound
Co-Evolving Skill Generation and Policy Optimization SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 48386583-dbbb-4203-814d-6928691314a9 · outbound
Co-Evolving Skill Generation and Policy Optimization How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a1255679-a3c9-425d-99e2-91210c1e4572 · outbound
Co-Evolving Skill Generation and Policy Optimization Skilltester: Benchmarking utility and security of agent skills
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 59a2cabd-dc1b-4498-b3ae-41a918821918 · outbound
Co-Evolving Skill Generation and Policy Optimization SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d08258a8-df2e-4639-ac6a-e3f2b12b5e61 · outbound
Co-Evolving Skill Generation and Policy Optimization Welcome to the era of experience.Google AI, 1:11, 2025
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd293901-17f6-4896-b9cc-d6087f836954 · outbound
Co-Evolving Skill Generation and Policy Optimization Memory for autonomous llm agents: Mechanisms, evaluation, and emerging frontiers
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cc670e79-6e3e-4094-a0c4-a1ebd6290b17 · outbound
Co-Evolving Skill Generation and Policy Optimization Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b44203e5-9781-40bf-b35f-be9e65c24968 · outbound
Co-Evolving Skill Generation and Policy Optimization Scaling agent learning via experience synthesis
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3a71f265-0b05-4144-be95-d1611ff72898 · outbound
Co-Evolving Skill Generation and Policy Optimization Agent Learning via Early Experience
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a5e360c6-d55c-426a-85c8-ea046b3da241 · outbound
Co-Evolving Skill Generation and Policy Optimization MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4c6afe53-99af-4150-b9d3-2959a043f709 · outbound
Co-Evolving Skill Generation and Policy Optimization Trajectory-informed memory generation for self-improving agent systems
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a5db41fc-e95d-4093-b10e-50799bead76e · outbound
Co-Evolving Skill Generation and Policy Optimization ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 05e77c32-6688-4ac0-aeff-e9378f78f76f · outbound
Co-Evolving Skill Generation and Policy Optimization R2d2: Remembering, replaying and dynamic decision making with a reflective agentic memory
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24ce572e-c007-471f-8b30-94f42ad805d2 · outbound
Co-Evolving Skill Generation and Policy Optimization Dynamic cheatsheet: Test-time learning with adaptive memory
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed5a25d2-4d6f-49a6-89ec-81ce22cd7023 · outbound
Co-Evolving Skill Generation and Policy Optimization Flex: Continuous agent evolution via forward learning from experience
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d4072ae0-91a9-4c30-b42b-3f320d3ec772 · outbound
Co-Evolving Skill Generation and Policy Optimization Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9c57175e-7d8f-47db-9a70-d85c37009f51 · outbound
Co-Evolving Skill Generation and Policy Optimization Legomem: Modular procedural memory for multi-agent llm systems for workflow automation
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5b8faa6b-4820-41d3-bea7-a7fa4983d446 · outbound
Co-Evolving Skill Generation and Policy Optimization Agent kb: Leveraging cross-domain experience for agentic problem solving
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f0fe70a6-bcd7-432b-95d9-ee64fd3b2e0e · outbound
Co-Evolving Skill Generation and Policy Optimization G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 81f77580-c744-4f1a-95c8-82f1823fb822 · outbound
Co-Evolving Skill Generation and Policy Optimization EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 838d8488-37bb-4416-af6b-c11213ffba73 · outbound
Co-Evolving Skill Generation and Policy Optimization Graph-based agent memory: Taxonomy, techniques, and applications
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1934cf68-119b-4c96-95c2-1cd1d79aea29 · outbound
Co-Evolving Skill Generation and Policy Optimization MemEvolve: Meta-Evolution of Agent Memory Systems
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8d55a5ce-5ba9-4013-9431-9584deaa2662 · outbound
Co-Evolving Skill Generation and Policy Optimization Agentevolver: Towards efficient self-evolving agent system
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8b5c4b91-481f-4016-804d-6b844681a852 · outbound
Co-Evolving Skill Generation and Policy Optimization Building self-evolving agents via experience-driven lifelong learning: A framework and benchmark
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 987f4751-8af1-4d9d-9111-2a728b5bbf26 · outbound
Co-Evolving Skill Generation and Policy Optimization Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c858693e-e2d5-4671-b494-3ad29d1ddd0c · outbound
Co-Evolving Skill Generation and Policy Optimization Adaptive Memory Crystallization for Autonomous AI Agent Learning in Dynamic Environments
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8560c48b-187d-43bd-a22a-43e1a5f0de09 · outbound
Co-Evolving Skill Generation and Policy Optimization ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3cec899c-61a7-4953-a26d-fb19336f1751 · outbound
Co-Evolving Skill Generation and Policy Optimization Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757, 2022
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 194a0efe-d2d0-4622-8392-25a94e9e0617 · outbound
Co-Evolving Skill Generation and Policy Optimization Qwen3 Technical Report
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 07bcbd80-c110-4399-90e7-0132421a4958 · outbound
Co-Evolving Skill Generation and Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 42349223-83ac-4946-a1e6-3342dd2fcbd8 · outbound
Co-Evolving Skill Generation and Policy Optimization The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4ab0146b-dcbc-4e50-9d96-a814996eb7d1 · outbound
Co-Evolving Skill Generation and Policy Optimization Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd5e89d8-8b27-44f1-bc96-f2302069de34 · outbound
Co-Evolving Skill Generation and Policy Optimization Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae7693b3-c8b7-4137-930a-2a8a59b385fe · outbound
Co-Evolving Skill Generation and Policy Optimization When not to trust language models: Investigating effectiveness of parametric and non-parametric memories
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 801c0f63-9afa-4f4b-ad94-e8d52133ed60 · outbound
Co-Evolving Skill Generation and Policy Optimization Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bc986e8-6f15-47fd-8abc-00cd3cb2434e · outbound
Co-Evolving Skill Generation and Policy Optimization Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04772b8-43f2-4702-b346-6f8d16b35b3d · outbound
Co-Evolving Skill Generation and Policy Optimization Musique: Multi-hop questions via single-hop question composition.Transactions of the Association for Computational Linguistics, 10:539–554, 2022
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce80d306-3640-4842-81c4-ca6fef390c23 · outbound
Co-Evolving Skill Generation and Policy Optimization Measuring and narrowing the compositionality gap in language models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8af126c3-526c-4b8e-817a-37f560c4602d · outbound
Co-Evolving Skill Generation and Policy Optimization Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6009b2cf-53a8-4e43-a23f-e0faac7ac119 · outbound
Co-Evolving Skill Generation and Policy Optimization ReAct: Synergizing Reasoning and Acting in Language Models
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ae2ea331-8cbb-4e54-a599-8aa4334fc6c1 · outbound
Co-Evolving Skill Generation and Policy Optimization Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652, 2023
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0a149d7-9968-4fa6-8131-d788eec6be4d · outbound
Co-Evolving Skill Generation and Policy Optimization Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f940e363-7245-4592-b44f-e2ee9638a819 · outbound
Co-Evolving Skill Generation and Policy Optimization Expel: Llm agents are experiential learners
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cbc82e2-1e9c-4321-b6ca-3c2c303f78aa · outbound
Co-Evolving Skill Generation and Policy Optimization Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0901fb9c-d856-4d7b-9198-fb19603db4d6 · outbound
Co-Evolving Skill Generation and Policy Optimization SimpleMem: Efficient Lifelong Memory for LLM Agents
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8c140f9f-5923-49fb-992a-73b8cdbd086e · outbound
Co-Evolving Skill Generation and Policy Optimization Search-o1: Agentic search-enhanced large reasoning models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b11de2d0-1b76-4da9-b8f7-3e8f9e911bbb · outbound
Co-Evolving Skill Generation and Policy Optimization Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9dc86ad0-4ac1-4ba5-9ff1-3600745d6288 · outbound
Co-Evolving Skill Generation and Policy Optimization ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ffaea00e-6581-4880-a5f1-68dd30ab359c · outbound
Co-Evolving Skill Generation and Policy Optimization StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ac9f0a0b-3ab0-49aa-a391-3851bf17818b · outbound
Co-Evolving Skill Generation and Policy Optimization Group-in-Group Policy Optimization for LLM Agent Training
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 889cc33b-1ce3-4984-829d-c17a53a6fbaf · outbound
Co-Evolving Skill Generation and Policy Optimization Text Embeddings by Weakly-Supervised Contrastive Pre-training
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 22bfe044-6bb1-4c4c-9bbd-3c7adba51a78 · inbound
UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation Co-Evolving Skill Generation and Policy Optimization
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.