Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2504.20073.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:01:24.792889Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 09feafe4-9c86-4aab-bd19-aa4064a5e969 · inbound
Reinforcement Learning from Human Feedback RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 189
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6f662b05-6ca3-490d-8994-c6cef5f2f129 · inbound
WebThinker: Empowering Large Reasoning Models with Deep Research Capability RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 030595a1-8013-4c92-8910-92ee6bb51497 · inbound
Group-in-Group Policy Optimization for LLM Agent Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5ab7702e-0e98-4869-a4c7-6d183698e938 · inbound
VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 85946968-b8e4-49d9-a7fc-ae0c82e95606 · inbound
Grounded Reinforcement Learning for Visual Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a31baf8e-8ee3-419f-a2ea-3f0aba9b8c7a · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 691ecce4-10a8-4e84-81de-4c7146df54da · inbound
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c9593914-92a1-4285-8e4e-cfd5c89eafb6 · inbound
EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aaccf3a-20ba-495a-88f2-cf358bd7790b · inbound
Reinforced Language Models for Sequential Decision Making RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac04a7b9-7d04-41c4-be28-29de60856bbc · inbound
SSRL: Self-Search Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 324d735f-e02e-433e-9b4a-3469afc45f9a · inbound
BetaWeb: Towards a Blockchain-enabled Trustworthy Agentic Web RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20ea23a1-92a8-4a08-8f9b-c2c8a16bfe37 · inbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36916772-8f61-48ec-b5b7-bf8d65ed929e · inbound
Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2beff123-71bc-4bc3-bd4c-76dbac3f7860 · inbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e4c5d38-576b-494f-b239-26b6df0b04bd · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c0e8479f-58f0-426e-b168-7fa18833324b · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 191
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f2a474c-08e4-4498-84cb-d20a478138f9 · inbound
Agent Learning via Early Experience RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bbf6ac5-1307-49ac-bbbe-a979c670fcf3 · inbound
From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b552fccc-256f-4aa6-8cef-bdfd395e92f1 · inbound
The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7c05bf3a-330b-487b-8dbb-4d062324d8c2 · inbound
Graph-Enhanced Policy Optimization in LLM Agent Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59563790-729f-4fc6-a6d7-c7168d26fa22 · inbound
Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f29736f-0b7a-451f-8e22-4bbfdb37d26e · inbound
AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9f191fc-e0ae-4a81-8328-6c2d48299049 · inbound
Reinforcement Learning for Self-Improving Agent with Skill Library RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ec99d046-96b7-449d-9fbb-9947731020c9 · inbound
Beyond Experience Retrieval: Learning to Generate Utility-Optimized Structured Experience for Frozen LLMs RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0ef3f9e4-a74f-43ac-b707-a85ec80bcedb · inbound
Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ece3d37d-2f0d-4447-80b0-1c334b7bf0f2 · inbound
HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 44f760fe-5aac-4d76-8b02-ba402f612dc9 · inbound
Training LLMs for Multi-Step Tool Orchestration with Constrained Data Synthesis and Graduated Rewards RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b31f1fa8-1c7f-41f1-8d3c-5765fff0547f · inbound
OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 943c6789-6844-4526-b83a-c2c61459f484 · inbound
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 15d605c7-bc36-4672-b51c-0984e9909ea1 · inbound
E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 644bb98b-0882-48c1-ab52-9181e316eb54 · inbound
From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 983c08d9-220c-45e6-9d95-2ae8b5048398 · inbound
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 655764ac-ffe7-46c5-aabe-961de77ac732 · inbound
Mind DeepResearch Technical Report RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5e44bd2-d50e-4964-b7d9-a0d1b125defe · inbound
Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f4cf9952-ad33-4e9f-8c30-89926b88acc9 · inbound
Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9bceaa4a-1cf9-40e0-898b-a2fd027ee004 · inbound
StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dcbea456-1568-4b82-bd7f-8fa34f17f7c9 · inbound
StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f2bb9c72-46d1-47e0-917c-46b24afd5490 · inbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b28fdab9-0c1a-4183-8774-625d16ffb86d · inbound
Pause or Fabricate? Training Language Models for Grounded Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed0ef21a-ed4b-423a-8c7f-c380277a5ff0 · inbound
Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a1294068-88c4-4bce-9250-de59386412c3 · inbound
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 921c004b-82bb-47f0-801d-b557f01cc9b5 · inbound
Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7b56134a-cb9a-4da6-b448-033bfae798d1 · inbound
T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ab820f4d-5ba5-451b-ba60-7dbe9186081e · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 126
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b18275cc-3954-4200-9074-b5d16f1b970f · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4a4a1fc4-50df-43db-a8bb-60dfe1286238 · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 14c0c2fd-2909-42dc-9d5b-49ec524756e9 · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ecd14ea2-58eb-438f-b41d-36cb35e4c4bb · inbound
A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4138b0f1-d58b-4d2e-90fd-a13ed32749f8 · inbound
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f50d1723-9254-49d8-906a-0436e958d562 · inbound
Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2764f7de-4e64-4785-b512-06afa90c7071 · inbound
AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c20f7f7b-8d3d-4f3d-ac8c-c47facd157b1 · inbound
MAGE: Multi-Agent Self-Evolution with Co-Evolutionary Knowledge Graphs RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3cfd89e0-f144-4799-a8e1-f62b5c3e7e6a · inbound
GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9192f7c5-7de3-4233-b063-9daa01956329 · inbound
GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7711b20f-5074-4cc0-b9e5-82086d47af10 · inbound
Learning Agentic Policy from Action Guidance RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 75fde33f-3688-4568-9140-b260bc16a4f5 · inbound
Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c69755d3-72f3-4d54-85b5-1e857cbae7c9 · inbound
Look Before You Leap: Autonomous Exploration for LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 75a2db13-0db5-450d-8454-1655c57c8a90 · inbound
DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 144
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a7e320ea-5775-48f6-8350-82f09467bb73 · inbound
HydroAgent: Closing the Gap Between Frontier LLMs and Human Experts in Hydrologic Model Calibration via Simulator-Grounded RL RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 92cb3926-1cd7-462c-acf9-e00d7658f33f · inbound
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8f6170ba-ec94-4679-9880-94e7e6a78c23 · inbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 04ec941d-8ecb-4db9-9539-8cceb2e5f738 · inbound
SEAL: Synergistic Co-Evolution of Agents and Learning Environments RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3b40b939-7bbe-4ec7-b924-ae2b29d3c4df · inbound
Test-Time Deep Thinking to Explore Implicit Rules RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 893c9511-cc09-442e-8e2c-52d730b24c45 · inbound
"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 30281e7c-4687-4ea5-a483-c266c1f14a32 · inbound
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 09b01bac-cd26-4552-ad37-23c1c4b9fa62 · inbound
Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9995315c-f420-4a84-b9bd-b3cf5ecc40e3 · inbound
SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a4d95838-ab3e-4dc1-994b-19eb79d3053f · inbound
Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 14f83664-ee98-4aa8-bae0-4d0bc5d1a688 · inbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a7cf2cfc-85ae-4ad3-88ed-7106b7592131 · inbound
Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4206eb12-62b0-4082-b686-b923468fc54e · inbound
Rethinking Continual Experience Internalization for Self-Evolving LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cdc6b572-e970-4eb5-abe7-2e63e2b901dd · inbound
When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 055d9650-f016-4f7e-b339-1092a20aeb76 · inbound
Signal-Driven Observation for Long-Horizon Web Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9216267e-9761-4ba9-affe-0bbfad21c897 · inbound
Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 302ddb98-7377-47cd-bdea-3f923e19dafc · inbound
Escaping the KL Agreement Trap in On-Policy Distillation RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1535bad0-1d2e-4f17-9685-f9de617b5991 · inbound
HIPIF: Hierarchical Planning and Information Folding for Long-Horizon LLM Agent Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a415e6f4-e2c8-47ed-95f3-0ca19c67e792 · inbound
From Verdict to Process: Agentic Reinforcement Learning for Multi-Stage Fact Verification RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 89e55f14-1000-4b12-bcaa-986a3b883492 · inbound
MagicSim: A Unified Infrastructure for Executable Embodied Interaction RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4a7ab9b8-0470-45ee-a2e4-f46d6579b3a5 · inbound
Marginal Advantage Accumulation for Memory-Driven Agent Self-Evolution RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 20bee695-c39e-41b3-b1bf-92354649a765 · inbound
Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4743eec0-3698-43dc-9ecd-e76ab5fefb42 · inbound
Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f31d3551-2d7b-4670-9e87-00d8d3dcb77e · inbound
Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c01eefa-bda7-4e24-85e9-3336d39a822b · inbound
UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 42fbcc21-3a76-4e7c-9dc0-234b73e2a021 · inbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81905e31-a397-4083-8ada-1d9aa26f3c55 · inbound
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2f0c4dd-5b3b-4455-be53-92fa6e44ec70 · inbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a49d9804-1e64-4d2b-b360-bef7c5926a61 · inbound
CurateEvo: Data-Curation Evolving for Agentic Post-Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9366d44b-ed90-4935-8687-061518498d23 · inbound
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f621dae-7032-4da8-b946-fa806174ddbf · inbound
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 072b71ac-60e3-40ac-b0ad-f30414dc02fe · inbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39265298-cdb6-48ae-8ec7-38b6236a36ec · inbound
ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee512434-d41c-4c64-b9a2-d401a3209b85 · inbound
Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be013677-665f-491e-bb22-7ae8ab1adb42 · inbound
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6a23d0b-ff2c-4156-9c67-44fa9a8e2822 · inbound
From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 2008
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c31ec5f3-eb14-4508-9995-269e8aed0c6f · inbound
Interactive Task Alignment as a POMDP RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37782cd0-dcd7-4986-a590-a9cd9e2ec741 · inbound
From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7cbb21b-0bbc-453e-9235-562f0b29ee0c · inbound
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7f59ec8-ad52-45a0-86d2-be21b378a15f · inbound
EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28f538ce-fa89-4896-a254-2995afafdb39 · inbound
AREX: Towards a Recursively Self-Improving Agent for Deep Research RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 424ee099-0126-4d97-9e0b-a3d9bc69968b · inbound
AREX: Towards a Recursively Self-Improving Agent for Deep Research RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.