Pith. sign in

Paper Citation Record · LEDGER

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents

As of 6 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2607.01120.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.01120 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-03T18:47:46.719344Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T09:49:31.023088Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact20
  • verified fuzzy18
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1c90530-8511-499e-a8b0-d26b7a6805f1 · outbound

This paper cites Openclaw: The ai that actually does things, 2026.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Openclaw: The ai that actually does things, 2026

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.196295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:7bc856a6c9cc4543678fa3fa7df396263b2928f51242d78e478a456dc62480b4

Observation 1b3b92dd-efc9-48a8-b07c-4b4fe1ddfee2 · outbound

This paper cites OpenClaw-RL: Train Any Agent Simply by Talking.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents OpenClaw-RL: Train Any Agent Simply by Talking

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:48:49.543357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:c2eb2936f206cd211c7b65616a3f060ac6b98454d75a4d9875c55fb9d2d790f3

Observation 967125cd-f96c-4e9c-b929-22c36418c295 · outbound

This paper cites Metaclaw: Just talk–an agent that meta-learns and evolves in the wild.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Metaclaw: Just talk–an agent that meta-learns and evolves in the wild

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:48:49.564921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:05a95f33ef193ccc3a5bb747263431aaf521f0780cc89132b2bf826f2e3ea181

Observation 04c9d290-5979-48eb-9048-d581db4657d6 · outbound

This paper cites SkillClaw: Let Skills Evolve Collectively with Agentic Evolver.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:48:49.546339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:56c56a9c03e0cf57939d0b22389f35c2cc67b25d162a5ca90f95b91af937f879

Observation 2fe8fb89-55af-4473-881c-7459d52f566f · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Reflexion: Language agents with verbal reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.206472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:0dd338ec2016a7ca72ad2574e3d110b810a532886961584cf2f2e95fd227552b

Observation d6e51a11-f1ef-438b-b3f4-2e2e0894cdd5 · outbound

This paper cites Memento-skills: Let agents design agents.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Memento-skills: Let agents design agents

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:48:49.535633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:55d77d6a08e8bc2e2ed7bd51560b595a10afacd0a76efac276b09ad8cf68148b

Observation ae361a64-5409-4815-89ff-015f892e6109 · outbound

This paper cites Agentic context engineering: Evolving contexts for self-improving language models.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Agentic context engineering: Evolving contexts for self-improving language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.208651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:b1fbedfcfff710a0c51486868cd64fb3ea967303f702df2d21c91d5996ae55e5

Observation dca7b9d2-8929-46f6-8200-13b76856f345 · outbound

This paper cites Areal: A large-scale asynchronous reinforcement learning system for language reasoning.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Areal: A large-scale asynchronous reinforcement learning system for language reasoning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.208460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:11efbd5a699a41b81c0fffe0eccdb1beff8f657a3486c2efbc6056df62e7f246

Observation b1e742bc-aef7-445d-8132-a304f92dcdbd · outbound

This paper cites A Survey of Reinforcement Learning for Large Reasoning Models.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents A Survey of Reinforcement Learning for Large Reasoning Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:48:49.532549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:0e7069266b66b2d3c1029e37410743ec166fb2c1ea078d264fc914b1c84c9550

Observation efe2340e-a09f-4d24-ac53-05179d597600 · outbound

This paper cites Training language models to follow instructions with human feedback.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Training language models to follow instructions with human feedback

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.218276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:009b5f4bb4febec147c6db8782741e670de7146cbf2b7bddf01c3b90ac21b61a

Observation 8df964e6-5775-4b93-897b-8a6b9c73b844 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Constitutional AI: Harmlessness from AI Feedback

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:48:49.527409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:80e22a55e034624e3f561b6ca53fc0c980dfde5c6609e549dc87eeaecaf18700

Observation 96ec482e-174a-4b88-b947-04ac42f8395d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Proximal Policy Optimization Algorithms

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:48:49.529749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:77b55eb40a4ab56f76fd3bc4796ce9cdfbe7ced3f3099027ad0b7393db428e28

Observation 06eec497-5edc-48f7-81a0-5663b9d3136a · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.210686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:6ed110723362754e828cf3e4da2749d51c102c80b87c5573064f4b9075c02595

Observation b7e9d4cf-9adc-4873-9c94-9cb7193c9153 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:48:49.540942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:9d0496f6e861ca2a85d5012e436ed70cd136b23864435b5b1b1571df569ddd17

Observation 76af6b6d-0be3-4282-b302-959eb43ce550 · outbound

This paper cites The Landscape of Agentic Reinforcement Learning for LLMs: A Survey.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:48:49.516258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:71f8be67db0fef19dad1d89dbd01908384aa44be35bcec3d43e56f8fd22884ef

Observation 4b4ace6b-a685-4d6c-878a-839beb201afc · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:48:49.519138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:607d8bbeba453f1320a86f285c08643d4f379da1b829b12e84da6511cfc890ba

Observation e6c8c4c1-7b6f-4ead-a109-8ba3a0fb3888 · outbound

This paper cites TextGrad: Automatic "Differentiation" via Text.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents TextGrad: Automatic "Differentiation" via Text

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:48:49.521976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:41ca95f595f605c078a80b860dd9df65539c1a16bedbbfbf8be9292599a52cc7

Observation 7311449a-4606-4d17-8358-8fec45566b0d · outbound

This paper cites Unlocking long-horizon agentic search with large-scale end-to-end rl.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Unlocking long-horizon agentic search with large-scale end-to-end rl

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.187826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:b36a610f7a43f5bee65d2747a90f7a81b27d30d116e7372a9944940b3b7092cf

Observation 06870b1d-8722-4fea-8bb3-1141e75b3305 · outbound

This paper cites ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:48:49.538132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:e8eb1615fdef8246f6cb56bab91bbbc3dc8ae83980b3d6619f6926471f1f730f

Observation 7dcf6429-605c-42ca-925d-cafe84bc77f9 · outbound

This paper cites Optimizing {RLHF} training for large language models with stage fusion.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Optimizing {RLHF} training for large language models with stage fusion

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.204090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:d65e6c36ffd7b8adc817107297ab59866cb9bfecabf5735351ce5b204bbc42d1

Observation 9487bfb8-2fb1-45b3-a0a8-28adfd9e3a54 · outbound

This paper cites G-Core: A Simple, Scalable and Balanced RLHF Trainer.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents G-Core: A Simple, Scalable and Balanced RLHF Trainer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:48:49.524813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:f145433966e24cfe576b5407200064aeeee5d75e3d1bdb16beb33422f70098e1

Observation 79e78776-42d1-406d-8f8c-808815f3e29f · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Hybridflow: A flexible and efficient rlhf framework

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.212609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:e0058344df68aad16a4fc6081a8831f252370298ebfdc58c05e46b74da40220a

Observation 082f5c95-f61d-4dd9-8c06-0462846c29b6 · outbound

This paper cites StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:48:49.509444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:be36e8fd23e877781b008e187b348f93f261902299b105b3b0b1eb64b2d53dda

Observation 17173bbf-93a2-485d-9c69-c7d3c3888201 · outbound

This paper cites AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:48:49.549581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:b943fe56fd66e722dd966be3028f74c61c08fc7534fcd09b5b58a187ef62ce51

Observation 7c103216-7fca-42f0-8da6-95001fddfca1 · outbound

This paper cites Introducing the Model Context Protocol.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Introducing the Model Context Protocol

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.197808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:ea5fc6c25ba095e1487e46017642fc60c04e2452c78eed4170429dcda8278a0f

Observation 07fc8c61-7a8c-402a-a713-6ed78e051e15 · outbound

This paper cites Agent2agent (a2a) protocol.https://a2a-protocol.org/, 2025.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Agent2agent (a2a) protocol.https://a2a-protocol.org/, 2025

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.194269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:3b5d5cd087ef98649914efa22a506efce3791debf8a6e4c96b6d011bc772e982

Observation 79ec63d6-1e71-4b4b-b3e7-b1c6f0664c4b · outbound

This paper cites A Survey of AI Agent Protocols.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents A Survey of AI Agent Protocols

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:48:49.556107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:9030c8dd8b4b4adc84e1210af0aca569ab919102a672468936fa486c31726645

Observation aad236f2-dc72-4606-960e-cccab34bc2f6 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:48:49.553561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:221db249f2cab3ad8e610a7976de10738e7f4fff53db776d853c85f2dcbe9012

Observation 9092a0e7-8054-4066-95f6-fc2fc0a351e6 · outbound

This paper cites RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:48:49.559274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:8790f002179199603703ab5453fe8b904c61f12ce0c03f7612b6f7b9b7e896b4

Observation 3aa93142-cd4f-4de2-8eb6-38292f3c8ac6 · outbound

This paper cites Agent data protocol: Unifying datasets for diverse, effective fine-tuning of llm agents.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Agent data protocol: Unifying datasets for diverse, effective fine-tuning of llm agents

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:48:49.561726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:52ccf00a82815bdca1efb7f99a9b80def55ad035884c3427496a5e2e234f4f6f

Observation c8bf950d-8111-425e-9de1-2afa6e8be718 · outbound

This paper cites LangChain: The agent engineering platform.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents LangChain: The agent engineering platform

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.206668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:8b224a0a22bc53b82c92be700f86481f1af140faa3fc8c78ca8a0814b59a41aa

Observation 2db5f868-cab5-4c31-865f-3fed96a8be40 · outbound

This paper cites LangGraph: Build resilient language agents as graphs.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents LangGraph: Build resilient language agents as graphs

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.216644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:bb8b473cdc9eae537748f56b817ecb8c62db320028f9527a860b65d238507f96

Observation 41e9af60-311a-4abb-8ce8-1c387b0e3477 · outbound

This paper cites CrewAI: Framework for orchestrating role-playing, autonomous AI agents.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents CrewAI: Framework for orchestrating role-playing, autonomous AI agents

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.191761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:64d2fca2856da889d76cfd6228716d07bd86b5bc31e8f7c6b22adbc68a8dbb6f

Observation eeafd810-b69f-4294-8a62-ab7eb4ff6b8c · outbound

This paper cites OpenAI Agents SDK: A lightweight, powerful framework for multi-agent workflows.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents OpenAI Agents SDK: A lightweight, powerful framework for multi-agent workflows

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.201810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:f1940e9062d7cca69b2d01298888a9dc71e5f42312459d447714b7adf46f8343

Observation 639b0062-83c4-4c49-a225-30d408f222ee · outbound

This paper cites Claude Agent SDK.https://github.com/anthropics/claude-agent-sdk-python, 2025.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Claude Agent SDK.https://github.com/anthropics/claude-agent-sdk-python, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.210562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:2ee49a4307a4dc8da64209252056f3122a6b75cefca43e23a02fb479fea5c4c2

Observation 73883c96-3a86-45eb-b16f-3cc6483ff40b · outbound

This paper cites Agentprm: Process reward models for llm agents via step-wise promise and progress.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Agentprm: Process reward models for llm agents via step-wise promise and progress

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.214498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:7ddc5249dda329d16895b226829bbdda34bebe4429da0d0024286a3c83f5d9ab

Observation e4742587-7e2f-4d8b-99c3-607576e104d8 · outbound

This paper cites Rlanything: Forge environment, policy, and reward model in completely dynamic rl system.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Rlanything: Forge environment, policy, and reward model in completely dynamic rl system

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:48:49.512855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:a1da51d63418cb0ea99b2e6da9d54ce6ba78fc5466311a1654e65a96ca7db3a9

Observation 5e8abccd-d426-4fc2-8f39-021a16556997 · outbound

This paper cites Hermes agent: The self-improving ai agent built by nous research.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Hermes agent: The self-improving ai agent built by nous research

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:30:39.216467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:99f20d392998811cb6b83fd29d06b9eb52087ab49e0b104e4c4995a8ee4f3842

Pith citing papers

Observation 6a869fb7-4c00-430a-973c-afe7325b5047 · inbound

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning cites this paper.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:31.023088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:31.023088Z digest=sha256:0ac9eef7eda17fdc63dc842380213419e10c937dd1df0a29c62dc213f85247fa