Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T19:19:36.427337Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 100 of 299 outbound references and 64 inbound Pith citation observations for arXiv:2509.02547.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T19:19:36.427337Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T13:31:24.753889Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T12:15:01.137692Z
100 of 299 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2396945c-eabb-4b66-8703-39135f8b1aa8 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Proximal Policy Optimization Algorithms
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4bcdb4dc-b734-4cba-943e-6a221710b7ae · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f7ec9216-8168-4138-9a66-00fde475356e · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Playing Atari with Deep Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 51a38e8c-fdb4-48a4-8ea7-c7a11221d4f4 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A Technical Survey of Reinforcement Learning Techniques for Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c802a854-f512-4e56-88f1-ed39a82c88db · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reinforcement Learning Enhanced LLMs: A Survey
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 445e3597-d01a-4a9f-a365-cb6b197d9ec1 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dd1c919c-db40-48e9-bcb8-04eb7bfcfba3 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods.IEEE Transactions on Neural Networks and Learning Systems, 36(6):9737–9757, June 2025
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 974982fd-1f54-4684-bc3f-518d10cee338 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Synthetic Data RL: Task Definition Is All You Need
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8f97174b-0ea6-4fa5-8474-fd875a94a3fb · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Let large language models find the data to train themselves
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation df887682-effa-49cc-926f-47e4e88480b0 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reinforcement Pre-Training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 82991100-669e-4fa7-8439-456414022799 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Inference-aware fine-tuning for best- of-n sampling in large language models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 520c1980-21ef-4703-a46c-c65d147549b2 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A survey of reinforcement learning in large language models: From data generation to test-time inference.Available at SSRN 5128927
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 38cb5da6-4859-467e-ba09-a95700d73bad · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Training language models to follow instructions with human feedback
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fd962b78-169e-4c63-8873-44cb2333d49d · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reinforcement Learning for LLM Post-Training: A Survey
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 53a7ab6b-16a0-4246-bd5b-83e320dda34c · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 30828b37-65b1-424b-8307-cf32756ce4e1 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6fd84143-8792-4eaa-9dc3-3467e62e2971 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Position: Agent Should Invoke External Tools ONLY When Epistemically Necessary
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation be03dfa9-3240-45d1-952c-b28b3cf9a009 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey From System 1 to System 2: A Survey of Reasoning Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cc284fdc-decd-4222-9ecd-6c6fd5b7dd91 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Large Language Model Agent: A Survey on Methodology, Applications and Challenges
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 22dfb8e0-5607-4f3f-b478-cba99185768d · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey van Duijn, Niki Stein, Mike Preuss, Peter van der Putten, and Kees Joost Batenburg
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7908cbc2-eb3d-4f6d-9e2a-f0bbce367f94 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A review of prominent paradigms for LLM-based agents: Tool use, planning (including RAG), and feedback learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 96eb9ba1-e214-48e8-b858-4b0597943900 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey What are tools anyway? a survey from the language model perspective
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a7d74a82-9516-472f-be0d-43028de24e20 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8de79810-8874-492d-9ce7-aa077060f9c8 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 01aa4e0f-97c6-45a0-87d8-ecb1612155ec · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A Survey on Self-Evolution of Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b2532011-c579-4698-82fe-57b4dcb69540 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey LLMs Working in Harmony: A Survey on the Technological Aspects of Building Effective LLM-Based Multi Agent Systems
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9ec266e8-7f25-4a82-b1e8-27c581965ef1 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Agent AI: Surveying the Horizons of Multimodal Interaction
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4876a34b-7f0a-4df9-9596-17a9f86843cb · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4fcfd900-328a-435f-a8a1-afc240977055 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Multi-agent reinforcement learning: A comprehensive survey
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7a445999-f65f-4c74-8227-913d327376e2 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Multi-agent Reinforcement Learning: A Comprehensive Survey
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 01685371-3625-47d6-a637-18bed9d705c0 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Direct preference optimization: Your language model is secretly a reward model
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2c337282-0282-4e35-b81b-aebb6370eff7 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey OpenAI o1 System Card
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 58e06a1a-2b09-4152-9ce7-7630c00b9215 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dc1a6192-e18c-40a2-95a6-49ece5ccadf2 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Openai o3 and o4-mini: Next-generation reasoning models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4f5f303e-8ec3-4e6a-8cec-2d4640f9a9f7 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ee604cd5-9a49-4432-86e8-bc5da3c09b6d · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A general theoretical paradigm to understand learning from human preferences
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8ec5c89e-c66c-46a1-86db-9b1020d8e098 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Model alignment as prospect theoretic optimization
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b2ab2c70-a02f-4218-9207-3bb352e10aca · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d47cb896-c4a9-49d3-91dc-6ff91af02719 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Part i: Tricks or traps? a deep dive into rl for llm reasoning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9ad6a137-df59-4412-8227-3f7838253bb7 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Policy filtration for RLHF to mitigate noise in reward models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6179bf92-2916-48d4-8521-503ca9af2743 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 518b9d2a-53a1-498b-8f60-db0e3e94ec91 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Process supervision-guided policy optimization for code generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 85ca1492-1eb5-4891-b8f5-cfaa50b4e617 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey $\beta$-DPO: Direct preference optimization with dynamic $\beta$
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b69bf1a4-b249-410c-a560-13219291986a · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SimPO: Simple preference optimization with a reference- free reward
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 79879293-fddf-4ffa-b97a-0261ab07e005 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey ORPO: Monolithic Preference Optimization without Reference Model
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 808fbd4d-c3e0-4f1f-88bc-b8bff91c6cc2 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9c6eb917-b493-4053-8f5d-207c5f29a10f · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dc87ac1d-53ef-498d-85f7-32e9f5e14888 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 63add13d-d4b5-4056-b866-47cd0302d06f · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Group Sequence Policy Optimization
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3a6d2884-cd12-4c1c-9b70-725a57238a5d · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Geometric-mean policy optimization
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c5a2fb91-77c4-41df-a9a1-7264959a1673 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 122ae140-4eb5-4fc2-a6dc-cb60cc1aead6 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 181db150-9ff2-45ac-a448-0f9a114bf0df · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Understanding R1-Zero-Like Training: A Critical Perspective
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0a41b68b-c538-473c-99e2-959846fa58ca · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5ff6a99d-6525-439b-bee4-924540413cb4 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e3703e59-508a-46ab-96f4-cc98e1583768 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Bartoldson, Bhavya Kailkhura, Fan Lai, Jiawei Zhao, and Beidi Chen
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b4d40336-1471-4269-b24f-77da40690d0f · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4e4c5d38-576b-494f-b239-26b6df0b04bd · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 005cfd71-2387-45e5-adae-73e25cfbd840 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b2df9d6a-be8e-4b2f-ba4f-e856480b8e1c · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 81bbebef-4da3-4fce-b002-9c69f4ab4b61 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Understanding Tool-Integrated Reasoning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6e85e4c9-95eb-4636-b522-21e93dec0dec · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cf2cc2bc-117b-4cd6-847a-ba838efe361b · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 28aa91b4-ef3e-4a99-9638-7cb0a8cb5ec7 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d29ee98b-ad09-460f-b317-c504c15b29d7 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey arXiv preprint arXiv:2508.11408 , year=
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e7d8987a-1e34-4486-b289-3db04d99566c · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Perception-Aware Policy Optimization for Multimodal Reasoning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 22d80a05-e4f4-4ebe-913a-986afc61037c · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Pass@k training for adaptively balancing exploration and exploitation of large reasoning models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 799af097-4862-4a75-989d-849b0b503152 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c150adfb-e10f-4dfa-9a3d-fbff1e0149e7 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Llm-powered autonomous agents.lilianweng.github.io, Jun 2023
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3cb555e9-4f99-4d75-9994-d037a04b337d · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Agentsquare: Automatic LLM agent search in modular design space
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fddb92de-a7ef-4a3f-9b9a-e1b5541c8ca1 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey ISBN 979-8-89176-251-0
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a71b2933-cd04-489e-8ebe-cc12767fbab4 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 261db82e-e2be-4f29-8e14-7279eddd3777 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey LLM as a mastermind: A survey of strategic reasoning with large language models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 127adaf1-3e28-45c7-8f98-b287df11f931 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Tool learning with foundation models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 69b943a9-d04f-449b-8703-96f8c92b8312 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Elements of a theory of human problem solving.Psychological review, 65(3):151
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation abbb2d01-8e70-491f-9ce1-4fc5e002bb0b · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Understanding the planning of LLM agents: A survey
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 66cab004-858c-4dc3-8d18-86affe20f3c6 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Understanding the planning of LLM agents: A survey
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 42eeca06-3ad2-4c69-ba04-d4fd3d2760cc · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey React: Synergizing reasoning and acting in language models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 42575813-c981-482d-9242-1614ec963728 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reasoning with language model is planning with world model
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b9098cef-aa05-4468-ad87-3384385ee767 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Language agent tree search unifies reasoning, acting, and planning in language models
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ba991bd1-c692-43b5-b6ea-ba1d1a6dbab7 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Planning without search: Refining frontier llms with offline goal-conditioned rl
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a92bf813-e794-45f0-9ab9-3863594c978d · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Learning when to plan: Efficiently allocating test-time compute for llm agents
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 71995afe-1318-4193-8be6-7f97b73df15b · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Deshmukh
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ee4b2b14-200b-4ffe-979c-a74267c12c86 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Trial and error: Exploration- based trajectory optimization of LLM agents
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1fe0a0f9-2855-47fc-98c1-e0c7a53c56ae · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Voyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3f518fd1-643b-4a88-9340-2a1a77aafeae · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Dynamic speculative agent planning
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4c2ea35a-3944-4662-935b-63a1c61f2ed2 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5e4d2d28-a510-4f85-b061-dcabb6b549b8 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 68e79a83-d91e-47bf-8ebe-3ea74e9f02cf · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Planner-r1: Reward shaping enables efficient agentic rl with smaller llms
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d2762523-ea31-441f-93f4-1145ba12a75d · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Planner-r1: Reward shaping enables efficient agentic rl with smaller llms
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 63e3f720-3a8b-4398-ba5c-98ae87fe806d · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A brain-inspired agentic architec- ture to improve planning with llms.Nature Communications, 16(1):8633
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 08887fbe-0b1f-413f-9d99-f07452820a75 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Toolformer: Language models can teach them- selves to use tools
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d149853-178e-4757-a71c-fdf5f94f7ea8 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey FireAct: Toward Language Agent Fine-tuning
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 30507396-f3d7-4e55-9e8e-68627b5ccdd0 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Knowledge-Centric Hallucination Detection
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d5de3d6c-f9bf-4fc6-be28-3eff2b34e99f · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Agent-FLAN: Designing data and methods of effective agent tuning for large language models
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 035db9a4-bf6e-4035-89f2-2e35b5aca205 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey doi: 10.18653/v1/2024.findings-acl.557
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ae81e564-ef8e-4cbc-9baa-ae7934baf2ff · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Agentbank: Towards generalized llm agents via fine-tuning on 50000+ interaction trajectories
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9893c56f-167f-4db3-88cd-898bf89c4dfd · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey API-Bank : A comprehensive benchmark for tool-augmented LLMs
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 97d5aa68-d662-47c0-b740-2d8553e714df · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey ToolRL: Reward is All Tool Learning Needs
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d4db3474-c124-4d39-aa05-13d3579a4cd9 · outbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Acting less is reasoning more! teaching model to act efficiently
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d0843f26-55b0-4b87-80d1-6ab94b23f010 · inbound
What Factors Affect LLMs and RLLMs in Financial Question Answering? The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c2a11b12-6c43-47a2-a4f2-f527bcfd31f6 · inbound
HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8bdd12ed-b601-491c-adce-b6b0c8e02f79 · inbound
Graph-Enhanced Policy Optimization in LLM Agent Training The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23833c0c-2daa-4777-bb4e-3bbb44933c9d · inbound
Agentic Learner with Grow-and-Refine Multimodal Semantic Memory The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 92811bb6-2bb7-4dc7-a1f5-a1d8d3323824 · inbound
Training Multi-Image Vision Agents via End2End Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c1ee493e-32f3-4ecb-b41d-38c9b4c1c565 · inbound
Agentic Reasoning for Large Language Models The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation beb3b442-90aa-4ef1-9b55-b1494c07e1e6 · inbound
S$^2$GR: Stepwise Semantic-Guided Reasoning in Latent Space for Generative Recommendation The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d1b03946-a403-41ed-a954-c463b59934b3 · inbound
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a3a7ae48-876a-4096-9503-ade0391bc9f1 · inbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adcc7487-9bed-42e4-b1cd-e7fca77e05a9 · inbound
AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8edec5fe-a6ba-4813-ad89-71f80131a018 · inbound
Towards Knowledgeable Deep Research: Framework and Benchmark The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0fdc4e7b-aceb-41ba-a193-4fd9624d5aea · inbound
StaRPO: Stability-Augmented Reinforcement Policy Optimization The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4a6fd09a-7f0f-4c66-888e-bfd364676651 · inbound
E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e20370d8-dd87-4dfd-a222-7e82f415063a · inbound
From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c0add783-059e-4a77-b22d-44111e535a31 · inbound
AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2e0677ef-c984-4342-91ac-3cc70a947423 · inbound
Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1a0a8b86-ff0b-471a-bd3b-9e16118db1ff · inbound
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 130
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 307ac012-d529-4d07-9023-23f6f40fd6f5 · inbound
StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6001819a-e57f-4286-a0be-10e7bb80ca33 · inbound
StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 29c2dcb9-8346-4ef4-accb-b9041e761ea6 · inbound
SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ad6aa5a1-a4cd-48c8-9a86-ba8445237e5e · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6c2b1663-d5d3-46db-ac4d-2db269f4bc6b · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9000084b-0722-4fd9-a29f-22b7c57b3e1e · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c0cece03-ef9f-4792-972c-2e16d5fd7362 · inbound
Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 178
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e92c6fc1-7400-4258-afee-285011440b5e · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 165
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f0194b7a-1933-48bb-86ef-b2fecd5e8553 · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 32f7ce1b-c029-47aa-bd4a-5b64f5302d08 · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8917f524-9daa-4a54-8631-31c83540d360 · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6705439e-9314-460b-ae08-4e24b5a2f304 · inbound
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6411ce64-b1ed-4acf-a284-d52677013cc2 · inbound
Personalizing LLMs with Binary Feedback: A Preference-Corrected Optimization Framework The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ff787be5-ad14-4556-bcbd-6f07b151209e · inbound
Learning Agentic Policy from Action Guidance The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 592ecb5c-3926-42e0-9019-97984e70e84f · inbound
Reinforced Collaboration in Multi-Agent Flow Networks The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1c3ca816-0b82-4621-8982-83c43c5c1428 · inbound
NEWTON: Agentic Planning for Physically Grounded Video Generation The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e6690489-b4a0-4ba0-a2f6-f1c670648b1f · inbound
Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 279c5188-7d12-4d1c-8ec6-32364f4b2214 · inbound
Echo: Learning from Experience Data via User-Driven Refinement The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5a28c917-b334-4e04-b1b9-1c306c65e98e · inbound
Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bca960fa-fcc1-4ee5-9170-aad48d307215 · inbound
Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c9994a9a-d795-4eca-88b9-8282b8dc8293 · inbound
Libra: Efficient Resource Management for Agentic RL Post-Training The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a486bd6a-57e1-454c-b174-7cc35e28dbef · inbound
EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f5b1c40d-9b05-4f54-b95d-ebeb54d0a35e · inbound
Co-Evolving Skill Generation and Policy Optimization The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ad6b1f5b-7f82-42e4-b3b8-027497a35524 · inbound
Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ffe115d7-9925-48de-b56e-50e4915e100e · inbound
APPO: Agentic Procedural Policy Optimization The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8e5d4caa-6992-4673-bb21-e69135855cd3 · inbound
APPO: Agentic Procedural Policy Optimization The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cd33109-5bc7-416d-907d-560d5cd9e59c · inbound
PDAGENT-BENCH: Characterizing, Grounding, and Architecting LLM Agents for VLSI Physical Design The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eb877cd0-e9f2-47e4-85a9-9a7e727956d7 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 260
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 13eb0052-5713-4b22-84b9-572d6b235405 · inbound
AIR: Adaptive Interleaved Reasoning with Code in MLLMs The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 94344ad7-f12f-49de-9845-958e838cddb3 · inbound
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cc852ecd-9b9d-426e-bbbb-4bc1701b7652 · inbound
Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9879578b-5bcb-4dc6-ae8e-5dfd175261ba · inbound
Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 33657ff9-1549-4848-9780-49c8acfa3dd4 · inbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 76af6b6d-0be3-4282-b302-959eb43ce550 · inbound
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 59fbeac2-527b-467f-85be-ae71a3a1a923 · inbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3840bc46-ce83-427a-8711-a6671b20ec41 · inbound
Mathematical methods of reinforcement learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fd6df471-3b24-4d47-9423-2a9a6012f5b4 · inbound
Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ddc527bd-8339-4e66-ade0-586f28941835 · inbound
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 300
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac28d4f5-fac1-44ee-887f-b466e24e46be · inbound
Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93df52ba-fc98-4581-b12a-702612842bfc · inbound
ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdcfbaf5-e0f4-40e0-ac83-ece8e809376b · inbound
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 119
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6477b146-22e0-4b29-84b7-91020de15fca · inbound
From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64f52e8b-3915-41ba-a444-7e1ec7a8b90b · inbound
AgentOmnia: Scaling Agentic Models for Full-Scenario Applications The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 108
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4478aa35-c581-4ddd-9bf5-6208b2bf3672 · inbound
OmniQEC: discovering practical quantum error-correcting codes by an AI scientist The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a145b9c-9181-430b-acfa-2b13de815d79 · inbound
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48e765bb-755c-4a4d-b144-67fca8c8c5da · inbound
Deep Reinforcement Learning: From First Principles to Reasoning Models The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27ca895f-cc55-44a1-8c62-003964f5102b · inbound
From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.