Pith. sign in

Paper Citation Record · LEDGER

RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2504.20073.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.20073 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 107 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:01:24.792889Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 09feafe4-9c86-4aab-bd19-aa4064a5e969 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 189

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:32:01.417888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:840eda65857e3f1c4bff700509614b26e73216ab2c294f6777dff179cf9c397c

Observation 6f662b05-6ca3-490d-8994-c6cef5f2f129 · inbound

WebThinker: Empowering Large Reasoning Models with Deep Research Capability cites this paper.

WebThinker: Empowering Large Reasoning Models with Deep Research Capability RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:14:25.317621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T19:14:25.283645Z digest=sha256:7b8dd03df10ea59be9ca9cae53dd71ab02df5014d6a60d54ed37c1071ba01792

Observation 030595a1-8013-4c92-8910-92ee6bb51497 · inbound

Group-in-Group Policy Optimization for LLM Agent Training cites this paper.

Group-in-Group Policy Optimization for LLM Agent Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T09:15:08.193357Z digest=sha256:b2c83c7389d5ce87d7943b2b7329cd9691b820665386f25c577ab5efc3a43069

Observation 5ab7702e-0e98-4869-a4c7-6d183698e938 · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 75

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:55:40.342411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:20bdfd0f7b515afb12808e0a40b7ebeac128573e19459656a5a81b01239fd965

Observation 85946968-b8e4-49d9-a7fc-ae0c82e95606 · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:05:51.865184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:a0f17c173af630b74fd42aee39051bc50ee806a640437d2a3bfc6ce725c2c1c9

Observation a31baf8e-8ee3-419f-a2ea-3f0aba9b8c7a · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T22:23:15.493799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:582d4bab7698bf343c19859b310300e1fd5c1e8cece56b4754a7ad3bc4e944b8

Observation 691ecce4-10a8-4e84-81de-4c7146df54da · inbound

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems cites this paper.

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 101

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T23:21:42.174087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T23:21:42.029285Z digest=sha256:99250f5adcc73b27af6fe7bbeebc7ade05f8879e7eaa7edf5fd2005ab76eff3d

Observation c9593914-92a1-4285-8e4e-cfd5c89eafb6 · inbound

EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making cites this paper.

EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T21:01:24.792889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:01:24.792889Z digest=sha256:2c7fff8aafb0a2e12f435e41d78c79db2f5222dbf1cbf5c2a8f26b1ffcbbf7bc

Observation 8aaccf3a-20ba-495a-88f2-cf358bd7790b · inbound

Reinforced Language Models for Sequential Decision Making cites this paper.

Reinforced Language Models for Sequential Decision Making RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:56.119890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:56.119890Z digest=sha256:46ea5f2835dc9a3261ea5ef121c434aa984fdac80baa3cb8c55c721e164832bd

Observation ac04a7b9-7d04-41c4-be28-29de60856bbc · inbound

SSRL: Self-Search Reinforcement Learning cites this paper.

SSRL: Self-Search Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.728962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.728962Z digest=sha256:c3c8c59676a476c5227e15343eec398244bdb1638ca1ca3875d0490ebf95c338

Observation 324d735f-e02e-433e-9b4a-3469afc45f9a · inbound

BetaWeb: Towards a Blockchain-enabled Trustworthy Agentic Web cites this paper.

BetaWeb: Towards a Blockchain-enabled Trustworthy Agentic Web RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-05T18:55:33.722891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:55:33.722891Z digest=sha256:bd1bd9d98a2944488420deb3969a531b7ac5c98a7e247e2584cd63f6ab2ba206

Observation 20ea23a1-92a8-4a08-8f9b-c2c8a16bfe37 · inbound

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use cites this paper.

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T16:22:51.511495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:22:51.511495Z digest=sha256:19f025742a422dfaa46d56598b8f64a9a27fe249e5632562d853c7c85323f601

Observation 36916772-8f61-48ec-b5b7-bf8d65ed929e · inbound

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning cites this paper.

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:14.910639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:44:14.910639Z digest=sha256:b561432f955fc442cc0ffe820934210366095e3e3d2bb3081a99485cc467e7e3

Observation 2beff123-71bc-4bc3-bd4c-76dbac3f7860 · inbound

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning cites this paper.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:24.770094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:24.770094Z digest=sha256:fa026056b73b339f8ac82417577462d4a0ddd5a1410b1542d4a058dd8f607c87

Observation 4e4c5d38-576b-494f-b239-26b6df0b04bd · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.689674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:9496ddc5f3ef29b34ad7e9be59b677c9c06e3bf13da4e4fa2d21b91ec40932bf

Observation c0e8479f-58f0-426e-b168-7fa18833324b · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 191

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:43.345147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:43.345147Z digest=sha256:5de759e4ad9ef31c1cef252a47417341bad57682fcb758d5ed5cd2dbc5eee0ba

Observation 2f2a474c-08e4-4498-84cb-d20a478138f9 · inbound

Agent Learning via Early Experience cites this paper.

Agent Learning via Early Experience RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:04.269782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:04.269782Z digest=sha256:4fe5323d768a1eab01010571ddcb62e2455127d0666040b22a39ba7252e5280c

Observation 2bbf6ac5-1307-49ac-bbbe-a979c670fcf3 · inbound

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails cites this paper.

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-21T20:44:22.027160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T20:42:40.823721Z digest=sha256:87b60e0f4e9b6afb85c0370707f1ff6ea98811b1c108dc93a972907acec23162

Observation b552fccc-256f-4aa6-8cef-bdfd395e92f1 · inbound

The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination cites this paper.

The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:10:51.312197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T04:09:50.183494Z digest=sha256:84aabc404693b570d8ec1dfcd8c1835d6c4531da50be044715c4c923aa315faf

Observation 7c05bf3a-330b-487b-8dbb-4d062324d8c2 · inbound

Graph-Enhanced Policy Optimization in LLM Agent Training cites this paper.

Graph-Enhanced Policy Optimization in LLM Agent Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T07:17:44.658592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:17:44.658592Z digest=sha256:c867164a2fd3b90b825fac1c38654c0407953e567e184539aca3453f268727d1

Observation 59563790-729f-4fc6-a6d7-c7168d26fa22 · inbound

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory cites this paper.

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T18:03:38.550009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:03:38.550009Z digest=sha256:fb10a89138dce5ad78a2d2d76ff5eebb76169db7335815f7c77b44216d20cb28

Observation 9f29736f-0b7a-451f-8e22-4bbfdb37d26e · inbound

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning cites this paper.

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:42.182243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T16:30:42.182243Z digest=sha256:04b13db2d16b81320f0d801ee8ba0bff4ed7d9730e9d05c0a9ae549c58a0c235

Observation b9f191fc-e0ae-4a81-8328-6c2d48299049 · inbound

Reinforcement Learning for Self-Improving Agent with Skill Library cites this paper.

Reinforcement Learning for Self-Improving Agent with Skill Library RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:05:30.610948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:05:30.593804Z digest=sha256:4faa5040854f57ef7a5b217c668f59592d460c2575aabf8b1bafc83de02e7eed

Observation ec99d046-96b7-449d-9fbb-9947731020c9 · inbound

Beyond Experience Retrieval: Learning to Generate Utility-Optimized Structured Experience for Frozen LLMs cites this paper.

Beyond Experience Retrieval: Learning to Generate Utility-Optimized Structured Experience for Frozen LLMs RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:47:41.958456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T09:43:55.255948Z digest=sha256:8b67c3bc2dc16e53efebbd5433473b40bbc5e30d9bbfaf201bf3f2a87ff36f08

Observation 0ef3f9e4-a74f-43ac-b707-a85ec80bcedb · inbound

Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents cites this paper.

Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T12:40:08.627745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:4d7f786383ec1ada7867e408d638bc64ab199e097e66e979ebd1b3582aff5e97

Observation ece3d37d-2f0d-4447-80b0-1c334b7bf0f2 · inbound

HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents cites this paper.

HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T18:36:28.338027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:35:45.900606Z digest=sha256:112de293a1bba2a361cf832b4e0889a99feebb5c89adc538201998bb568379e2

Observation 44f760fe-5aac-4d76-8b02-ba402f612dc9 · inbound

Training LLMs for Multi-Step Tool Orchestration with Constrained Data Synthesis and Graduated Rewards cites this paper.

Training LLMs for Multi-Step Tool Orchestration with Constrained Data Synthesis and Graduated Rewards RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:18:22.746285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T00:15:56.194714Z digest=sha256:982baa399479b862ca586a9c61a44f78ddc3936ec2fdc047ea6a3f0eac521c03

Observation b31f1fa8-1c7f-41f1-8d3c-5765fff0547f · inbound

OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search cites this paper.

OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 31

Resolution
verified exact
orphan_title_repair, observed 2026-05-13T17:16:44.180135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T17:16:09.927267Z digest=sha256:1cb6b44a88285614ee1a494c2a1105e7d1bede3192c093aeaa28cecbb142066d

Observation 943c6789-6844-4526-b83a-c2c61459f484 · inbound

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization cites this paper.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:f04fafde1c28f922a201f7b44314e5c2f77abd1123268dd9578159eeeb6d9ab7

Observation 15d605c7-bc36-4672-b51c-0984e9909ea1 · inbound

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning cites this paper.

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T18:08:18.056525Z digest=sha256:cbcf7825b43f366990d26d2464f73a4180b8506bdcaae1a6d1dfcceaf6a3f66a

Observation 644bb98b-0882-48c1-ab52-9181e316eb54 · inbound

From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models cites this paper.

From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:09:36.341574Z digest=sha256:22bebb35f5043c5b0a86dac81bd1925d730463facea57048a65bfd476a7540e1

Observation 983c08d9-220c-45e6-9d95-2ae8b5048398 · inbound

Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents cites this paper.

Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T15:28:07.981488Z digest=sha256:b5a423a5480988e7818e037dab6c056169c22f3852f8b214b0ceed67cc313cfc

Observation 655764ac-ffe7-46c5-aabe-961de77ac732 · inbound

Mind DeepResearch Technical Report cites this paper.

Mind DeepResearch Technical Report RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:46:49.178896Z digest=sha256:ed0155430a582641c45a2f91decb0bb80ef4cc6c73241344c37f47a09ce77cac

Observation c5e44bd2-d50e-4964-b7d9-a0d1b125defe · inbound

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration cites this paper.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:cb06c59e432aeffe4692fe0b7d82aee864d5560a94a80bc0b8ce94bdb9d32f17

Observation f4cf9952-ad33-4e9f-8c30-89926b88acc9 · inbound

Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures cites this paper.

Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T04:31:28.242097Z digest=sha256:6edfd61bf1e9576d963dd958f3711cbaf3bc8a86402361f166e2c2d6fbe2520e

Observation 9bceaa4a-1cf9-40e0-898b-a2fd027ee004 · inbound

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning cites this paper.

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T04:26:59.781416Z digest=sha256:b939b16ee221f958514cf484a3540042fbcb71802989644a9ef03e333f6bc67b

Observation dcbea456-1568-4b82-bd7f-8fa34f17f7c9 · inbound

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning cites this paper.

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-05T12:30:59.971516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-05T12:23:00.487399Z digest=sha256:399cf28ee561976e0a62f1a9f57203171a9dc00d6c488af1a65cb2326a838077

Observation f2bb9c72-46d1-47e0-917c-46b24afd5490 · inbound

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training cites this paper.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:939811d0c338bf377eb852575ea9fa631bdb9e968b733f1dd10f4f2c2fdeddd1

Observation b28fdab9-0c1a-4183-8774-625d16ffb86d · inbound

Pause or Fabricate? Training Language Models for Grounded Reasoning cites this paper.

Pause or Fabricate? Training Language Models for Grounded Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T03:01:58.366028Z digest=sha256:c8b21b11bc9968e9f8c5c106b97e8cffeae5749a9d684f642c5bf0dc9828bc60

Observation ed0ef21a-ed4b-423a-8c7f-c380277a5ff0 · inbound

Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing Agents cites this paper.

Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T12:00:07.345611Z digest=sha256:77afccd6088ccc371a22fac3a622cd11f0810b4cb80df35ae5f5b74934c43862

Observation a1294068-88c4-4bce-9250-de59386412c3 · inbound

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training cites this paper.

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T06:36:33.914193Z digest=sha256:de2363cd2eef6aced4d9ec172e11db2444272b78687b9bf736a4a3468005ec94

Observation 921c004b-82bb-47f0-801d-b557f01cc9b5 · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:d62ca9ff27eef25fef17b5818abed101c0c02c46586cd7c0ae94e0fe32b275ee

Observation 7b56134a-cb9a-4da6-b448-033bfae798d1 · inbound

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning cites this paper.

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T19:36:00.359351Z digest=sha256:2c84113fd56ffbd22d8354da9e91e921dbdf1083ddf31917dc6ed8e954005c25

Observation ab820f4d-5ba5-451b-ba60-7dbe9186081e · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:dc151c0f4683a3e7a9275df44b093eca8d654e3ccbb7fa4df84f5c594167a734

Observation b18275cc-3954-4200-9074-b5d16f1b970f · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T10:23:52.522238Z digest=sha256:5ba4743ae04c8cb841a3927f2e3c1ac33f470ae38edd703ba44813c0a1853b1f

Observation 4a4a1fc4-50df-43db-a8bb-60dfe1286238 · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T02:00:00.663355Z digest=sha256:3e639927f9e4827b35d4befc881fbc006e1ac4cd8c8dcbfd56953a7d817c7f70

Observation 14c0c2fd-2909-42dc-9d5b-49ec524756e9 · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 83

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T07:17:28.448889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T07:17:13.708752Z digest=sha256:b544d730c78b2c35c05fb2362a026c136e19046c4fdb1891bb185adffdc61a1b

Observation ecd14ea2-58eb-438f-b41d-36cb35e4c4bb · inbound

A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping cites this paper.

A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T10:41:46.675257Z digest=sha256:b5055badad509b18bac471a333b5baf1b386ef9d7516fc9d4860f6c55e416ba4

Observation 4138b0f1-d58b-4d2e-90fd-a13ed32749f8 · inbound

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction cites this paper.

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T09:59:48.604813Z digest=sha256:7fffb14e7970054989eefb1f785275a37b0a316c8522108caf2a1a8613ab40e3

Observation f50d1723-9254-49d8-906a-0436e958d562 · inbound

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair cites this paper.

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:21:43.166304Z digest=sha256:b02d260f0ca4225a6b309441436078684138866b4fd89fa487e625edb57163b3

Observation 2764f7de-4e64-4785-b512-06afa90c7071 · inbound

AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design cites this paper.

AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:14:39.643357Z digest=sha256:13eb72da06091c5992d7911da4916a1ac6486c71df788970f63f048813104ba1

Observation c20f7f7b-8d3d-4f3d-ac8c-c47facd157b1 · inbound

MAGE: Multi-Agent Self-Evolution with Co-Evolutionary Knowledge Graphs cites this paper.

MAGE: Multi-Agent Self-Evolution with Co-Evolutionary Knowledge Graphs RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T05:13:28.089038Z digest=sha256:fd897d61c74aabdc9de870bb4220dc9708cf7dc15406265eaa5abb6c80763e13

Observation 3cfd89e0-f144-4799-a8e1-f62b5c3e7e6a · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T06:44:30.069444Z digest=sha256:460f4d5437be769e063fc43a1285eb315226e31cb288b6731bd97e3eeac10af1

Observation 9192f7c5-7de3-4233-b063-9daa01956329 · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:59:48.694854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T05:59:23.949124Z digest=sha256:f09057d4074fb263cb55f4bd3e54db5df2b3e9a0800e4d01425f6a69dcb79b57

Observation 7711b20f-5074-4cc0-b9e5-82086d47af10 · inbound

Learning Agentic Policy from Action Guidance cites this paper.

Learning Agentic Policy from Action Guidance RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:f8314963e516979090ce5fee3ea08293c8a022016f93cd5cb3b991fec7cfc6ac

Observation 75fde33f-3688-4568-9140-b260bc16a4f5 · inbound

Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy cites this paper.

Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:48:28.567125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:46:24.724553Z digest=sha256:e2ac84e4a0d1d77d450be71ed25ccbec1cb24098bbea3062e70acc4f16c81843

Observation c69755d3-72f3-4d54-85b5-1e857cbae7c9 · inbound

Look Before You Leap: Autonomous Exploration for LLM Agents cites this paper.

Look Before You Leap: Autonomous Exploration for LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-20T17:43:36.378048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T17:42:31.830822Z digest=sha256:cbc6e3ff0b3ce362d08729e6f4d1225cfd5881e386a3903046887d3111025de8

Observation 75a2db13-0db5-450d-8454-1655c57c8a90 · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 144

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:43:17.165237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:5e90ac5db0bb3187fb32f389ff160507d7be223e0938fed3b53a5960b8bb341d

Observation a7e320ea-5775-48f6-8350-82f09467bb73 · inbound

HydroAgent: Closing the Gap Between Frontier LLMs and Human Experts in Hydrologic Model Calibration via Simulator-Grounded RL cites this paper.

HydroAgent: Closing the Gap Between Frontier LLMs and Human Experts in Hydrologic Model Calibration via Simulator-Grounded RL RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:28:18.828024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T13:28:03.965754Z digest=sha256:e18b5e52c5b7806a54feecbf30bdd33b98037043d218c0c31689cedaf2c07988

Observation 92cb3926-1cd7-462c-acf9-e00d7658f33f · inbound

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents cites this paper.

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:43:05.690848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T05:41:23.712146Z digest=sha256:fa63ec9c1ab0a170d61631e7a07a1ee479129c32ac35330cc94abdd134b00b6f

Observation 8f6170ba-ec94-4679-9880-94e7e6a78c23 · inbound

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents cites this paper.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.463079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:fff7f5ab307cddd9234a68d4158a971bc237cddf1e628f2eb879ad8e73599f87

Observation 04ec941d-8ecb-4db9-9539-8cceb2e5f738 · inbound

SEAL: Synergistic Co-Evolution of Agents and Learning Environments cites this paper.

SEAL: Synergistic Co-Evolution of Agents and Learning Environments RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-30T13:44:40.937292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T13:38:10.466713Z digest=sha256:7f3ff73679f13506ca3a0b41267c7a0550da0395dfe17c8e85582293b6364ea8

Observation 3b40b939-7bbe-4ec7-b924-ae2b29d3c4df · inbound

Test-Time Deep Thinking to Explore Implicit Rules cites this paper.

Test-Time Deep Thinking to Explore Implicit Rules RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-30T11:54:38.222745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T11:52:15.163893Z digest=sha256:d7d78646dbada2885ccbcc7f6a4c6ae7b118a6287c23622b9b9371d27961bb69

Observation 893c9511-cc09-442e-8e2c-52d730b24c45 · inbound

"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise cites this paper.

"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:56:20.765294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T14:51:06.448678Z digest=sha256:fd71480d5265fda9a8dc1c3594cc7283a047c9699d27b24d2887779da181a1c8

Observation 30281e7c-4687-4ea5-a483-c266c1f14a32 · inbound

OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents cites this paper.

OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:16.416446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:46:50.587684Z digest=sha256:79a7cd30050093e8917969ad228cdd50c07ac844e5482c239017dba2aafb0425

Observation 09b01bac-cd26-4552-ad37-23c1c4b9fa62 · inbound

Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes cites this paper.

Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:26:22.814833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:20:13.991921Z digest=sha256:18b53f3841a677149f0cd5a4cb88758046bdead4b8a1094fee701311842913c1

Observation 9995315c-f420-4a84-b9bd-b3cf5ecc40e3 · inbound

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training cites this paper.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.892126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:9770543e1770f24b6574519be59509f6acbbcf37369c71f521972987c532d751

Observation a4d95838-ab3e-4dc1-994b-19eb79d3053f · inbound

Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection cites this paper.

Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T14:22:17.935687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T14:20:06.381334Z digest=sha256:f5a885d3baae7a158e3331e6bf43d6f9bed89ed260a0be554bd61eed668c5803

Observation 14f83664-ee98-4aa8-bae0-4d0bc5d1a688 · inbound

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning cites this paper.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.172159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:11e506c2bd22c6b483d3845ca59785ec23c2ec6b258877201aba1f866d5a66fe

Observation a7cf2cfc-85ae-4ad3-88ed-7106b7592131 · inbound

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning cites this paper.

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:26.241307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:19:31.702516Z digest=sha256:d56dc202aac8fb2439497c914232f34ce92fef1eea0b55f24c7804f468b36fde

Observation 4206eb12-62b0-4082-b686-b923468fc54e · inbound

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents cites this paper.

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T07:56:47.799866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T06:29:11.398007Z digest=sha256:3fa856b08eeff36dd81a090cf1d63d12c2a13ed6a7c7cfe93507acb75c0b9d98

Observation cdc6b572-e970-4eb5-abe7-2e63e2b901dd · inbound

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training cites this paper.

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:16:56.809012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T02:17:32.324432Z digest=sha256:b97022beeade295496b41f2f46ee3bbfd3e9883e563be088c8a82403a5989ab7

Observation 055d9650-f016-4f7e-b339-1092a20aeb76 · inbound

Signal-Driven Observation for Long-Horizon Web Agents cites this paper.

Signal-Driven Observation for Long-Horizon Web Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T01:31:29.240542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T01:25:11.149977Z digest=sha256:7a3bb512637fab5341c0b83410dc340bf51deb499c3c2ae8ba24be4550a60918

Observation 9216267e-9761-4ba9-affe-0bbfad21c897 · inbound

Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning cites this paper.

Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:07:28.253101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T17:28:58.574865Z digest=sha256:8fa1a9aa529495b31d244d72babdb9732ea3bcf1ac6e7a1b2498c85536a2ddf8

Observation 302ddb98-7377-47cd-bdea-3f923e19dafc · inbound

Escaping the KL Agreement Trap in On-Policy Distillation cites this paper.

Escaping the KL Agreement Trap in On-Policy Distillation RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:57:29.474501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T16:58:57.106367Z digest=sha256:608ce834720bf79225adb2ec2eff8848c77cbe1a2ebdf9b7c1b2ed6ca2930044

Observation 1535bad0-1d2e-4f17-9685-f9de617b5991 · inbound

HIPIF: Hierarchical Planning and Information Folding for Long-Horizon LLM Agent Learning cites this paper.

HIPIF: Hierarchical Planning and Information Folding for Long-Horizon LLM Agent Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:47:41.643997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:04:28.758614Z digest=sha256:9c21c5ddfca49e79d0a872de731cfd3ba5bd43bba03165cb8e22926c3b4fa209

Observation a415e6f4-e2c8-47ed-95f3-0ca19c67e792 · inbound

From Verdict to Process: Agentic Reinforcement Learning for Multi-Stage Fact Verification cites this paper.

From Verdict to Process: Agentic Reinforcement Learning for Multi-Stage Fact Verification RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:58:33.390469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:47:09.681690Z digest=sha256:cfb47ca8a70e42c67cef882f8895146427548e5d0d64e6c3f01898fc0c56ea77

Observation 89e55f14-1000-4b12-bcaa-986a3b883492 · inbound

MagicSim: A Unified Infrastructure for Executable Embodied Interaction cites this paper.

MagicSim: A Unified Infrastructure for Executable Embodied Interaction RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:58:58.092132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:00:07.465292Z digest=sha256:993c5bedc9c0703555ee4a4b67a8d9a320696018fb56da21762d96c8585c5500

Observation 4a7ab9b8-0470-45ee-a2e4-f46d6579b3a5 · inbound

Marginal Advantage Accumulation for Memory-Driven Agent Self-Evolution cites this paper.

Marginal Advantage Accumulation for Memory-Driven Agent Self-Evolution RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:39:30.791534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T17:45:49.272070Z digest=sha256:6c5086ebca53786e26c33a0b39fb7884cbe5bcf59a8b7192fa79b7584360b1dd

Observation 20bee695-c39e-41b3-b1bf-92354649a765 · inbound

Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning cites this paper.

Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:45.753940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T08:44:23.085858Z digest=sha256:dcc7d2122c50d6edd24b1d38b91dd1fa729e9692ceb74ee38a915e326033d725

Observation 4743eec0-3698-43dc-9ecd-e76ab5fefb42 · inbound

Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies cites this paper.

Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:59:47.008995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T08:11:30.091859Z digest=sha256:84c151a14bfd42c74c212bb6cd6964b9828d825ba73bd76ac2f3a8452b10e925

Observation f31d3551-2d7b-4670-9e87-00d8d3dcb77e · inbound

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It cites this paper.

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T20:50:12.373539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T19:28:36.499352Z digest=sha256:564bc060f35f0f47b55819b340b6d4576201fcc83d0eceff8c3fa877b1e8647a

Observation 5c01eefa-bda7-4e24-85e9-3336d39a822b · inbound

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation cites this paper.

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:24:22.387998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T07:15:44.387347Z digest=sha256:cde9022317f3d884d4476188dc619f31c1365c1c26ea9002709970b61ca71601

Observation 42fbcc21-3a76-4e7c-9dc0-234b73e2a021 · inbound

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry cites this paper.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:7d6caa1c5e5cbe831c0d1aae17593d0cc51bbc7c49456b8fd02f4333f4ad0b57

Observation 81905e31-a397-4083-8ada-1d9aa26f3c55 · inbound

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning cites this paper.

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T20:42:48.585661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:42:48.585661Z digest=sha256:14532d585f9ceda78bbadf4f6396417dcbba844d7fc7deaf3961d134f0ab4582

Observation a2f0c4dd-5b3b-4455-be53-92fa6e44ec70 · inbound

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents cites this paper.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:78f95a9d026a0f7730f9f70b9d08e6fdb3d7b572d1071cb06d32dc1b05717d6f

Observation a49d9804-1e64-4d2b-b360-bef7c5926a61 · inbound

CurateEvo: Data-Curation Evolving for Agentic Post-Training cites this paper.

CurateEvo: Data-Curation Evolving for Agentic Post-Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-08T15:35:07.149976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-08T15:33:01.138366Z digest=sha256:3282ba04b3e9b18776a9bb2d37daac2713e1c6d208ff28d18a739e6081734c51

Observation 9366d44b-ed90-4935-8687-061518498d23 · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-14T12:26:27.446079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:26:27.446079Z digest=sha256:040abc42ae4f68e77066cc985d78c05e925bb776d5e0194fd3d48a36ab72e7bf

Observation 4f621dae-7032-4da8-b946-fa806174ddbf · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T07:21:32.935264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:21:32.935264Z digest=sha256:fab754d79734ddce216aae634e90bebb6cc7a3ad87a2752afc338be25fc1b750

Observation 072b71ac-60e3-40ac-b0ad-f30414dc02fe · inbound

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories cites this paper.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:080a2e02f55d98428a17988f4e622dfd635e052b82f776e851fbe3ac03cfbe95

Observation 39265298-cdb6-48ae-8ec7-38b6236a36ec · inbound

ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning cites this paper.

ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T04:59:08.940069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T04:59:08.940069Z digest=sha256:54a1751fcdcd9a3b895ebdad9a764ae903cf06e81aba36d6450c116d63c720ae

Observation ee512434-d41c-4c64-b9a2-d401a3209b85 · inbound

Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning cites this paper.

Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T04:40:39.171769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:40:39.171769Z digest=sha256:7e6af0f41226164923a1675c9f4159488fe33d166c436e5e3d5858f66b59c977

Observation be013677-665f-491e-bb22-7ae8ab1adb42 · inbound

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning cites this paper.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.134716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.134716Z digest=sha256:3e23cf7a8895f3ca472545d28704b10c620d651022289566129c80484f35dd70

Observation b6a23d0b-ff2c-4156-9c67-44fa9a8e2822 · inbound

From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training cites this paper.

From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-02T09:45:49.605488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:45:49.605488Z digest=sha256:268234efff38b2dc6338e51bbc961c4b1687bda4a7b681b4d3263754f2e4b65a

Observation c31ec5f3-eb14-4508-9995-269e8aed0c6f · inbound

Interactive Task Alignment as a POMDP cites this paper.

Interactive Task Alignment as a POMDP RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T21:04:42.348283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:04:42.348283Z digest=sha256:bfeae0fd859e7e667cc3778a0e8c0331dd983cbe1940451ef4bbfba167d6394e

Observation 37782cd0-dcd7-4986-a590-a9cd9e2ec741 · inbound

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation cites this paper.

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T08:56:36.117089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:56:36.117089Z digest=sha256:c4bd0878dac5c92988860259691be7ca506324303ed02ca883d780792065c203

Observation b7cbb21b-0bbc-453e-9235-562f0b29ee0c · inbound

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD cites this paper.

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T10:42:41.407957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:42:41.407957Z digest=sha256:8eaabc9c01c433dd2a9dd67effdec6bfcb2f6c282c5147002507d463b86b01b7

Observation a7f59ec8-ad52-45a0-86d2-be21b378a15f · inbound

EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL cites this paper.

EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T12:21:14.541478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:21:14.541478Z digest=sha256:87b476ade6052af2d44af7cc523671c12364698dc1b375c5b5ac508c0e18b0a8

Observation 28f538ce-fa89-4896-a254-2995afafdb39 · inbound

AREX: Towards a Recursively Self-Improving Agent for Deep Research cites this paper.

AREX: Towards a Recursively Self-Improving Agent for Deep Research RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 16

Resolution
malformed identifier
no resolver link, observed 2026-08-01T07:26:20.954574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:26:20.954574Z digest=sha256:435060134f746dfc070e9e936533c1c6e8f3559bdc71da39d5374ead81196e87

Observation 424ee099-0126-4d97-9e0b-a3d9bc69968b · inbound

AREX: Towards a Recursively Self-Improving Agent for Deep Research cites this paper.

AREX: Towards a Recursively Self-Improving Agent for Deep Research RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T07:26:21.088294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:26:21.088294Z digest=sha256:72ff6b53831bf9b1f10a88873aa45df8855a64bb8030a8f0dc39b9a01f2158e1