Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-03T18:47:46.719344Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2607.01120.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-03T18:47:46.719344Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T09:49:31.023088Z
A source-named dated measurement, never combined with another source.
Source: cited_works
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e1c90530-8511-499e-a8b0-d26b7a6805f1 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Openclaw: The ai that actually does things, 2026
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1b3b92dd-efc9-48a8-b07c-4b4fe1ddfee2 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents OpenClaw-RL: Train Any Agent Simply by Talking
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 967125cd-f96c-4e9c-b929-22c36418c295 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Metaclaw: Just talk–an agent that meta-learns and evolves in the wild
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 04c9d290-5979-48eb-9048-d581db4657d6 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2fe8fb89-55af-4473-881c-7459d52f566f · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Reflexion: Language agents with verbal reinforcement learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d6e51a11-f1ef-438b-b3f4-2e2e0894cdd5 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Memento-skills: Let agents design agents
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ae361a64-5409-4815-89ff-015f892e6109 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Agentic context engineering: Evolving contexts for self-improving language models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dca7b9d2-8929-46f6-8200-13b76856f345 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Areal: A large-scale asynchronous reinforcement learning system for language reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b1e742bc-aef7-445d-8132-a304f92dcdbd · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents A Survey of Reinforcement Learning for Large Reasoning Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation efe2340e-a09f-4d24-ac53-05179d597600 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Training language models to follow instructions with human feedback
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8df964e6-5775-4b93-897b-8a6b9c73b844 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Constitutional AI: Harmlessness from AI Feedback
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 96ec482e-174a-4b88-b947-04ac42f8395d · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Proximal Policy Optimization Algorithms
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 06eec497-5edc-48f7-81a0-5663b9d3136a · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b7e9d4cf-9adc-4873-9c94-9cb7193c9153 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 76af6b6d-0be3-4282-b302-959eb43ce550 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4b4ace6b-a685-4d6c-878a-839beb201afc · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents ReAct: Synergizing Reasoning and Acting in Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e6c8c4c1-7b6f-4ead-a109-8ba3a0fb3888 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents TextGrad: Automatic "Differentiation" via Text
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7311449a-4606-4d17-8358-8fec45566b0d · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Unlocking long-horizon agentic search with large-scale end-to-end rl
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 06870b1d-8722-4fea-8bb3-1141e75b3305 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7dcf6429-605c-42ca-925d-cafe84bc77f9 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Optimizing {RLHF} training for large language models with stage fusion
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9487bfb8-2fb1-45b3-a0a8-28adfd9e3a54 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents G-Core: A Simple, Scalable and Balanced RLHF Trainer
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 79e78776-42d1-406d-8f8c-808815f3e29f · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Hybridflow: A flexible and efficient rlhf framework
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 082f5c95-f61d-4dd9-8c06-0462846c29b6 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 17173bbf-93a2-485d-9c69-c7d3c3888201 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7c103216-7fca-42f0-8da6-95001fddfca1 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Introducing the Model Context Protocol
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 07fc8c61-7a8c-402a-a713-6ed78e051e15 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Agent2agent (a2a) protocol.https://a2a-protocol.org/, 2025
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 79ec63d6-1e71-4b4b-b3e7-b1c6f0664c4b · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents A Survey of AI Agent Protocols
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aad236f2-dc72-4606-960e-cccab34bc2f6 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9092a0e7-8054-4066-95f6-fc2fc0a351e6 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3aa93142-cd4f-4de2-8eb6-38292f3c8ac6 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Agent data protocol: Unifying datasets for diverse, effective fine-tuning of llm agents
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c8bf950d-8111-425e-9de1-2afa6e8be718 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents LangChain: The agent engineering platform
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2db5f868-cab5-4c31-865f-3fed96a8be40 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents LangGraph: Build resilient language agents as graphs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 41e9af60-311a-4abb-8ce8-1c387b0e3477 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents CrewAI: Framework for orchestrating role-playing, autonomous AI agents
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eeafd810-b69f-4294-8a62-ab7eb4ff6b8c · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents OpenAI Agents SDK: A lightweight, powerful framework for multi-agent workflows
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 639b0062-83c4-4c49-a225-30d408f222ee · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Claude Agent SDK.https://github.com/anthropics/claude-agent-sdk-python, 2025
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 73883c96-3a86-45eb-b16f-3cc6483ff40b · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Agentprm: Process reward models for llm agents via step-wise promise and progress
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e4742587-7e2f-4d8b-99c3-607576e104d8 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Rlanything: Forge environment, policy, and reward model in completely dynamic rl system
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5e8abccd-d426-4fc2-8f39-021a16556997 · outbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Hermes agent: The self-improving ai agent built by nous research
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6a869fb7-4c00-430a-973c-afe7325b5047 · inbound
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.