Pith. sign in

Paper Citation Record · LEDGER

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

As of 5 August 2026, this Paper Citation Record lists 100 of 299 outbound references and 64 inbound Pith citation observations for arXiv:2509.02547.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.02547 v5

Coverage vector

measured 100 of 299 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T19:19:36.427337Z

measured 164 of 164 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 64 of 64 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:31:24.753889Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

100 of 299 outbound references displayed

  • verified exact55
  • verified fuzzy30
  • unresolved2
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch11

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2396945c-eabb-4b66-8703-39135f8b1aa8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Proximal Policy Optimization Algorithms

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.870824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:35662e04c717a15c4590c054065496815cc12cbdfbdff59b5404407e2929f7f6

Observation 4bcdb4dc-b734-4cba-943e-6a221710b7ae · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.884172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:893d4ce40092c4edf506c1c75a9fc7c0f7796a061ec6113ff0415194e6bfa8f8

Observation f7ec9216-8168-4138-9a66-00fde475356e · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Playing Atari with Deep Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.866403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:0e3c4ad732ada41f80d50bb9c906b32218d01ea07e97188f96c7b2ef54c02d46

Observation 51a38e8c-fdb4-48a4-8ea7-c7a11221d4f4 · outbound

This paper cites A Technical Survey of Reinforcement Learning Techniques for Large Language Models.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A Technical Survey of Reinforcement Learning Techniques for Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:18:25.553364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:af355a37f5bfd4ae135f2ef8e3f40570c3b6239497e6229b26c934bc503f191a

Observation c802a854-f512-4e56-88f1-ed39a82c88db · outbound

This paper cites Reinforcement Learning Enhanced LLMs: A Survey.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reinforcement Learning Enhanced LLMs: A Survey

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.875324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:6cdc220ad3b83eefc7dd1b38052f7e8bd64372e8fb8041cfd062877b40e46935

Observation 445e3597-d01a-4a9f-a365-cb6b197d9ec1 · outbound

This paper cites Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.880055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:d49a2678113426351c2b2ea85dd3f09669966d82dac43a8ff60c649d32abbccc

Observation dd1c919c-db40-48e9-bcb8-04eb7bfcfba3 · outbound

This paper cites Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods.IEEE Transactions on Neural Networks and Learning Systems, 36(6):9737–9757, June 2025.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods.IEEE Transactions on Neural Networks and Learning Systems, 36(6):9737–9757, June 2025

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:46.766902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:225ce71c1478c402a26e0c13feacd41982282ed739f94e7b3733286ee4bbece1

Observation 974982fd-1f54-4684-bc3f-518d10cee338 · outbound

This paper cites Synthetic Data RL: Task Definition Is All You Need.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Synthetic Data RL: Task Definition Is All You Need

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.858237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:31838cd952dd9e682939bd074d6bdfc578bdb864e717dbddc49be5ba6134ef81

Observation 8f97174b-0ea6-4fa5-8474-fd875a94a3fb · outbound

This paper cites Let large language models find the data to train themselves.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Let large language models find the data to train themselves

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.164052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:51ca8984b467558dbe60e4ed899e383339b2685066e4aa8838120cfdce05328f

Observation df887682-effa-49cc-926f-47e4e88480b0 · outbound

This paper cites Reinforcement Pre-Training.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reinforcement Pre-Training

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T19:21:48.478904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:772c509f4c5b52caa271c49642beb7a7e407521bb793e0fbd1be09a68ac7573a

Observation 82991100-669e-4fa7-8439-456414022799 · outbound

This paper cites Inference-aware fine-tuning for best- of-n sampling in large language models.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Inference-aware fine-tuning for best- of-n sampling in large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.439072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:3821ac95ae95e493ee16320e5d37673fff71c408aaa6dbade0373a6a4452155c

Observation 520c1980-21ef-4703-a46c-c65d147549b2 · outbound

This paper cites A survey of reinforcement learning in large language models: From data generation to test-time inference.Available at SSRN 5128927.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A survey of reinforcement learning in large language models: From data generation to test-time inference.Available at SSRN 5128927

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.442376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:f506699c97af0953800530547a3e7decac3f75cc6cb29754ef0a8d9388cb0022

Observation 38cb5da6-4859-467e-ba09-a95700d73bad · outbound

This paper cites Training language models to follow instructions with human feedback.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Training language models to follow instructions with human feedback

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.445626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:67a3982f17d6c2bad8b6cc9baec370f2214188f1a78e8b100bf0dff3027c8336

Observation fd962b78-169e-4c63-8873-44cb2333d49d · outbound

This paper cites Reinforcement Learning for LLM Post-Training: A Survey.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reinforcement Learning for LLM Post-Training: A Survey

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T19:21:48.376249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:6f244015c4630852eabd9602289bf072f19607525a44fa6f5d9b0aebda72ff5d

Observation 53a7ab6b-16a0-4246-bd5b-83e320dda34c · outbound

This paper cites A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:1cc4a95421d184146d01bc61c98ef319f5cf436c3a6b9bd2a9d4d18a7619a7a8

Observation 30828b37-65b1-424b-8307-cf32756ce4e1 · outbound

This paper cites Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment

Reference 16

Resolution
malformed identifier
doi_truncated, observed 2026-05-18T19:21:46.662487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:fe953d979df4511eb842957547f879436185320d991c964d0f810e9277b31558

Observation 6fd84143-8792-4eaa-9dc3-3467e62e2971 · outbound

This paper cites Position: Agent Should Invoke External Tools ONLY When Epistemically Necessary.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Position: Agent Should Invoke External Tools ONLY When Epistemically Necessary

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.390112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:db0fbd745139d44a283e688b28fb7645f56851942cf4caf191ad20bc2e081e84

Observation be03dfa9-3240-45d1-952c-b28b3cf9a009 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.426898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:9064a5af328b0939d8a161ebc7ffbc69f4a72ea96c8c0d3b4ead91aa3dec46ed

Observation cc284fdc-decd-4222-9ecd-6c6fd5b7dd91 · outbound

This paper cites Large Language Model Agent: A Survey on Methodology, Applications and Challenges.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Large Language Model Agent: A Survey on Methodology, Applications and Challenges

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T19:21:46.673527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:f98553e279b8ee4cd5c5859732771ed12a25637b886ca199f9c0774f47367c9b

Observation 22dfb8e0-5607-4f3f-b478-cba99185768d · outbound

This paper cites van Duijn, Niki Stein, Mike Preuss, Peter van der Putten, and Kees Joost Batenburg.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey van Duijn, Niki Stein, Mike Preuss, Peter van der Putten, and Kees Joost Batenburg

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:46.743137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:27983c10c8de2cfd9cadadac5513186ce6d0b80ea6a382c96f900de21187329b

Observation 7908cbc2-eb3d-4f6d-9e2a-f0bbce367f94 · outbound

This paper cites A review of prominent paradigms for LLM-based agents: Tool use, planning (including RAG), and feedback learning.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A review of prominent paradigms for LLM-based agents: Tool use, planning (including RAG), and feedback learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.429687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:05de42bcb47a8283546280ebc2c3b16c9f8462f8e07ad0c8beef1e9fd1841803

Observation 96eb9ba1-e214-48e8-b858-4b0597943900 · outbound

This paper cites What are tools anyway? a survey from the language model perspective.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey What are tools anyway? a survey from the language model perspective

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.423529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:f965aee65939f7260da064bca4da674aba49fee98426c3629a2267b5fd0989ae

Observation a7d74a82-9516-472f-be0d-43028de24e20 · outbound

This paper cites The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.613347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:278a200cfa9377a898c1545618e3814af22f2748f26cd50e585ddfe506096cc6

Observation 8de79810-8874-492d-9ce7-aa077060f9c8 · outbound

This paper cites an unresolved cited work.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-18T19:22:50.420467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:79f3cf59c8a863501e89a113b24d4d5684cbc60b0f694f324a15b461974e56d1

Observation 01aa4e0f-97c6-45a0-87d8-ecb1612155ec · outbound

This paper cites A Survey on Self-Evolution of Large Language Models.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A Survey on Self-Evolution of Large Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T19:21:48.350622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:82a2f2883edfa5871a37164363b41b0e798a604f0dc9f8a65dac9198269cb1a2

Observation b2532011-c579-4698-82fe-57b4dcb69540 · outbound

This paper cites LLMs Working in Harmony: A Survey on the Technological Aspects of Building Effective LLM-Based Multi Agent Systems.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey LLMs Working in Harmony: A Survey on the Technological Aspects of Building Effective LLM-Based Multi Agent Systems

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.246164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:6dea719f3c9001b2823f843eb5c3e17f2b1ebeada5183a3c09c3d5379d69747e

Observation 9ec266e8-7f25-4a82-b1e8-27c581965ef1 · outbound

This paper cites Agent AI: Surveying the Horizons of Multimodal Interaction.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Agent AI: Surveying the Horizons of Multimodal Interaction

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:46.719040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:33969a3d960326fc1a46baafc5dabfe80c47b7ad818c620ce4c5ff7bbd137b22

Observation 4876a34b-7f0a-4df9-9596-17a9f86843cb · outbound

This paper cites Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T19:21:46.618753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:75c2a89a4b556bb27bc18bb8ee01d38896afe77cb5deb6d4b517122f3d22ac9c

Observation 4fcfd900-328a-435f-a8a1-afc240977055 · outbound

This paper cites Multi-agent reinforcement learning: A comprehensive survey.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Multi-agent reinforcement learning: A comprehensive survey

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.426645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:9cb1c54dc2ba1a2b3f06efe66dd82c635a500791c31112e01791ae35a93c3d50

Observation 7a445999-f65f-4c74-8227-913d327376e2 · outbound

This paper cites Multi-agent Reinforcement Learning: A Comprehensive Survey.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Multi-agent Reinforcement Learning: A Comprehensive Survey

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.282484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:fd4a14a123de627ee7a71928d8c208b69ab6bf52f9a0532520de73ff968f11fb

Observation 01685371-3625-47d6-a637-18bed9d705c0 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Direct preference optimization: Your language model is secretly a reward model

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.432658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:797ef5073275105b1debe26012dd5fff57f8d720e16fd879f5673a36bdcab466

Observation 2c337282-0282-4e35-b81b-aebb6370eff7 · outbound

This paper cites OpenAI o1 System Card.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey OpenAI o1 System Card

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.266422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:a0aea0b690fb463b3a0c4292d57982d65ff1deb7a7921d894364abf59fb717d1

Observation 58e06a1a-2b09-4152-9ce7-7630c00b9215 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.138145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:d292367966a08a97709f2236a1d7033d78d244aecb9f59b44be8a790955f41af

Observation dc1a6192-e18c-40a2-95a6-49ece5ccadf2 · outbound

This paper cites Openai o3 and o4-mini: Next-generation reasoning models.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Openai o3 and o4-mini: Next-generation reasoning models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.405181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:7690c095c7321896396ae8b259aebe475183afeb2069b355c2c636f3bda79979

Observation 4f5f303e-8ec3-4e6a-8cec-2d4640f9a9f7 · outbound

This paper cites A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.183065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:a60bddf07ce96b1436650268ad52289559344ad1c622be742b4510538968ae55

Observation ee604cd5-9a49-4432-86e8-bc5da3c09b6d · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A general theoretical paradigm to understand learning from human preferences

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.402099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:428ad1a9a814ef0c287c31089e4575c7065664e16b3e6d6d615244af1daf240a

Observation 8ec5c89e-c66c-46a1-86db-9b1020d8e098 · outbound

This paper cites Model alignment as prospect theoretic optimization.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Model alignment as prospect theoretic optimization

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.408414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:16469afbee3d7ff3dbcb433b1451155cbdfe56ba1f0accc9ae7ff5e55ad6a352

Observation b2ab2c70-a02f-4218-9207-3bb352e10aca · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.013620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:8e02d0ac85241c08fb2e934018948a6af6434187eaddc65567aec178d943dc69

Observation d47cb896-c4a9-49d3-91dc-6ff91af02719 · outbound

This paper cites Part i: Tricks or traps? a deep dive into rl for llm reasoning.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Part i: Tricks or traps? a deep dive into rl for llm reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.411470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:ccf47b47995f1e4f2ad8532ca275db0118b708412ba5b36d4ac59e7925d85cae

Observation 9ad6a137-df59-4412-8227-3f7838253bb7 · outbound

This paper cites Policy filtration for RLHF to mitigate noise in reward models.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Policy filtration for RLHF to mitigate noise in reward models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.417568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:a6c2159d85a0af35020e29661ccd6b394480fb6a29f166edffc939c47d95571c

Observation 6179bf92-2916-48d4-8521-503ca9af2743 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T19:21:48.734746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:4583edaf8086579b0e742f341fa9d21ded62612a3167712cf33d5a0fb4579865

Observation 518b9d2a-53a1-498b-8f60-db0e3e94ec91 · outbound

This paper cites Process supervision-guided policy optimization for code generation.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Process supervision-guided policy optimization for code generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.435823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:c0e34225c7459d2adfe3c459d66b5eb3bfd8e710b283718ed04b7400c0969371

Observation 85ca1492-1eb5-4891-b8f5-cfaa50b4e617 · outbound

This paper cites $\beta$-DPO: Direct preference optimization with dynamic $\beta$.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey $\beta$-DPO: Direct preference optimization with dynamic $\beta$

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.414636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:c04a1d84a48865670213415889c696b1c5ebf431cfbc5a2f8767b06f1ca3be97

Observation b69bf1a4-b249-410c-a560-13219291986a · outbound

This paper cites SimPO: Simple preference optimization with a reference- free reward.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SimPO: Simple preference optimization with a reference- free reward

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.451664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:b7a1c93cdc6e4e62706a5a9128ff1547aed5cdc0e5747b797d8643edf062033e

Observation 79879293-fddf-4ffa-b97a-0261ab07e005 · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey ORPO: Monolithic Preference Optimization without Reference Model

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.143700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:6406e6b721537aff873802acd630b36cb52f2c3dc2182e7dd9a00bd518ad36ad

Observation 808fbd4d-c3e0-4f1f-88bc-b8bff91c6cc2 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:20f0730f9345123b30600f78e046ef05b37c6bd73b19e2db95d7c00556d71167

Observation 9c6eb917-b493-4053-8f5d-207c5f29a10f · outbound

This paper cites Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.068374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:1e1faef76040b7b42eacd8499fd8ddf9ec090ea5dec7a4f64f9d9492a9599f33

Observation dc87ac1d-53ef-498d-85f7-32e9f5e14888 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.409093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:633068edc5f1ee0622ac8f61f3b31b034d18371ee31ef72a61838430ead989aa

Observation 63add13d-d4b5-4056-b866-47cd0302d06f · outbound

This paper cites Group Sequence Policy Optimization.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Group Sequence Policy Optimization

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.052038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:82f80f3bdae909fc6f2a2474ddc0b03d54c192ed078efa719d94433a2809bbbc

Observation 3a6d2884-cd12-4c1c-9b70-725a57238a5d · outbound

This paper cites Geometric-mean policy optimization.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Geometric-mean policy optimization

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.512545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:b0123108e135848460b4ec4be42b30ca6f8a4e809d111a1073950cc98ac18ab4

Observation c5a2fb91-77c4-41df-a9a1-7264959a1673 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:cd84d7221c55c380e15e0339c1c2606b45950ab008c8aa15144f765ed4db0927

Observation 122ae140-4eb5-4fc2-a6dc-cb60cc1aead6 · outbound

This paper cites ReCode: Reinforcing Code Generation with Reasoning-Process Rewards.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey ReCode: Reinforcing Code Generation with Reasoning-Process Rewards

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T19:21:48.003226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:9f522d973a29da0b5d50462216ab879616c936eef468f628a515244f5329f5c8

Observation 181db150-9ff2-45ac-a448-0f9a114bf0df · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Understanding R1-Zero-Like Training: A Critical Perspective

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.073138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:5304836a74be3cd550307f487eb898211b436d8727eca0eb68a2fe3fd48667ae

Observation 0a41b68b-c538-473c-99e2-959846fa58ca · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.310900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:576a26679b82ab91cbe658231cd7ec8aab8912aa4b7739e47c1378c4499b0fca

Observation 5ff6a99d-6525-439b-bee4-924540413cb4 · outbound

This paper cites SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.803557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:30a03fa7fb827e508d48a101338557126d8ccff1cc98c4acadadd7602cad0003

Observation e3703e59-508a-46ab-96f4-cc98e1583768 · outbound

This paper cites Bartoldson, Bhavya Kailkhura, Fan Lai, Jiawei Zhao, and Beidi Chen.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Bartoldson, Bhavya Kailkhura, Fan Lai, Jiawei Zhao, and Beidi Chen

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.395864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:edea4219989caa1ff1a39e8ce7295a547e04423b20e6f15fef0f7f0d6755efcc

Observation b4d40336-1471-4269-b24f-77da40690d0f · outbound

This paper cites Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.638499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:a544d6e82d3b8826859661d609c5d48fa5f6ad133bb4ba0d57af234483bb6014

Observation 4e4c5d38-576b-494f-b239-26b6df0b04bd · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.689674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:b41a5148ade71663a59c6259984b522082e2c05f5ca297c237764d28820d0761

Observation 005cfd71-2387-45e5-adae-73e25cfbd840 · outbound

This paper cites GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T19:21:48.567762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:f3512faf0d32049b0fc7fba6782cd2e186935b43d96d0406c5de4cccc695d0f5

Observation b2df9d6a-be8e-4b2f-ba4f-e856480b8e1c · outbound

This paper cites Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.436079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:471ca93e370bd137ff8a80d620bebf9fa2f0b31ca502f4245b046f2864cc7a93

Observation 81bbebef-4da3-4fce-b002-9c69f4ab4b61 · outbound

This paper cites Understanding Tool-Integrated Reasoning.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Understanding Tool-Integrated Reasoning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.422826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:5475d10ab83d61a4cc42053b557bbc4603f5af90b9610eec8c83384ca13cba34

Observation 6e85e4c9-95eb-4636-b522-21e93dec0dec · outbound

This paper cites TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.770589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:c783013c4e893b24aa42ddd610eab3e30ec432b04859d95c1db3fc6558f4cdeb

Observation cf2cc2bc-117b-4cd6-847a-ba838efe361b · outbound

This paper cites EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.764302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:c4c0fb5ad3e2e3dccd9a890f068fe0a5e18a0e5d7f56e7c95483ff5dbc1c2698

Observation 28aa91b4-ef3e-4a99-9638-7cb0a8cb5ec7 · outbound

This paper cites Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.758177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:f1872d2642d3943334ccb2a90422b3d52def7c724cf1416dc7d57c88ae7b6f08

Observation d29ee98b-ad09-460f-b317-c504c15b29d7 · outbound

This paper cites arXiv preprint arXiv:2508.11408 , year=.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey arXiv preprint arXiv:2508.11408 , year=

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T19:21:48.441115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:3413c8adbf38374c750aa0340de878df3361f93ba6452221185d9b65509618df

Observation e7d8987a-1e34-4486-b289-3db04d99566c · outbound

This paper cites Perception-Aware Policy Optimization for Multimodal Reasoning.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.458569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:2154afe825ad69db97d9dda80f7ccb0f00a1eb93d782482f087a96791bd41e53

Observation 22d80a05-e4f4-4ebe-913a-986afc61037c · outbound

This paper cites Pass@k training for adaptively balancing exploration and exploitation of large reasoning models.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Pass@k training for adaptively balancing exploration and exploitation of large reasoning models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.385396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:f609d684338005a5a84d5923839e11d478c342e497d46027864e9af2517a2289

Observation 799af097-4862-4a75-989d-849b0b503152 · outbound

This paper cites Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.591192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:4773ea8882d9a8063e43a4b388131f46f2376a441a35d46345ecef80991a40fb

Observation c150adfb-e10f-4dfa-9a3d-fbff1e0149e7 · outbound

This paper cites Llm-powered autonomous agents.lilianweng.github.io, Jun 2023.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Llm-powered autonomous agents.lilianweng.github.io, Jun 2023

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.399038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:cec9b6082115014f37aedd3e91182d7a38c7159b68569e6ac10ff992a6cd7840

Observation 3cb555e9-4f99-4d75-9994-d037a04b337d · outbound

This paper cites Agentsquare: Automatic LLM agent search in modular design space.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Agentsquare: Automatic LLM agent search in modular design space

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.381883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:f3824d584b2c8c6db15c145a2e52fcfa91c206dd57930c5faa5d25d263dd0b66

Observation fddb92de-a7ef-4a3f-9b9a-e1b5541c8ca1 · outbound

This paper cites ISBN 979-8-89176-251-0.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey ISBN 979-8-89176-251-0

Reference 71

Resolution
metadata mismatch
doi, observed 2026-05-18T19:21:46.625626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:8b842e01f4464ed3c4af9d6c7a5bbfbaf2252b8b791d3e116d6dc27d8c0d12ad

Observation a71b2933-cd04-489e-8ebe-cc12767fbab4 · outbound

This paper cites an unresolved cited work.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-05-18T19:22:50.388581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:549248a70adcbbf00286aab48e645330d08930185bd57381dc3702720150af8a

Observation 261db82e-e2be-4f29-8e14-7279eddd3777 · outbound

This paper cites LLM as a mastermind: A survey of strategic reasoning with large language models.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey LLM as a mastermind: A survey of strategic reasoning with large language models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.366987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:c34b12330a3c59982d0a85b6c019336900517062e737a846b04ebf5ab6ecfcf0

Observation 127adaf1-3e28-45c7-8f98-b287df11f931 · outbound

This paper cites Tool learning with foundation models.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Tool learning with foundation models

Reference 74

Resolution
verified exact
doi, observed 2026-05-18T19:21:46.736820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:9299cb02d44e3eee980094398a074061b5d068600eb277a74d97c81267dabb54

Observation 69b943a9-d04f-449b-8703-96f8c92b8312 · outbound

This paper cites Elements of a theory of human problem solving.Psychological review, 65(3):151.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Elements of a theory of human problem solving.Psychological review, 65(3):151

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.363554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:b0037ce2eea141f911e482da1299898889dfc4389ab467c083ca1923a2683df2

Observation abbb2d01-8e70-491f-9ce1-4fc5e002bb0b · outbound

This paper cites Understanding the planning of LLM agents: A survey.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Understanding the planning of LLM agents: A survey

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.239818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:a8ef54ebff2d04441539fa4a0b36339525c134b8f6f38c8c1bcd4580a0c4dd74

Observation 66cab004-858c-4dc3-8d18-86affe20f3c6 · outbound

This paper cites Understanding the planning of LLM agents: A survey.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Understanding the planning of LLM agents: A survey

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:46.759838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:e99c0ec9832aee0538cbda08d7bff049f3a7e52e9277b289b66ffce5f0f8cfde

Observation 42eeca06-3ad2-4c69-ba04-d4fd3d2760cc · outbound

This paper cites React: Synergizing reasoning and acting in language models.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey React: Synergizing reasoning and acting in language models

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.370819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:4b346a2b946060c3d415c197384c0c739aa94e266c80bbe7168ca8a537527f27

Observation 42575813-c981-482d-9242-1614ec963728 · outbound

This paper cites Reasoning with language model is planning with world model.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reasoning with language model is planning with world model

Reference 79

Resolution
malformed identifier
raw_fallback, observed 2026-05-18T19:22:50.378508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:a6e9b52a52444ba878110639c1545df4152af054c45c82549be064b4c3ba5f46

Observation b9098cef-aa05-4468-ad87-3384385ee767 · outbound

This paper cites Language agent tree search unifies reasoning, acting, and planning in language models.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Language agent tree search unifies reasoning, acting, and planning in language models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.351629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:1f3397b40615f318baffd8b62815bf2753e608e2bd9acf509c7f761cb7d66cd7

Observation ba991bd1-c692-43b5-b6ea-ba1d1a6dbab7 · outbound

This paper cites Planning without search: Refining frontier llms with offline goal-conditioned rl.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Planning without search: Refining frontier llms with offline goal-conditioned rl

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.111021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:3b877f5b732feee495b1f59fbc6f836e02c28219502a28d72a06013d1b89f6d3

Observation a92bf813-e794-45f0-9ab9-3863594c978d · outbound

This paper cites Learning when to plan: Efficiently allocating test-time compute for llm agents.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Learning when to plan: Efficiently allocating test-time compute for llm agents

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.628450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:d1d7b3a4aa2a020f7e9f8c78143e75fb73ba4f7aa21c9130318b79d681fe992d

Observation 71995afe-1318-4193-8be6-7f97b73df15b · outbound

This paper cites Deshmukh.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Deshmukh

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.160251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:58f102070d7a060ad219bdea56e70c9cb9ec1c4c27cfbbbd398d1bca4dbd8314

Observation ee4b2b14-200b-4ffe-979c-a74267c12c86 · outbound

This paper cites Trial and error: Exploration- based trajectory optimization of LLM agents.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Trial and error: Exploration- based trajectory optimization of LLM agents

Reference 84

Resolution
verified exact
doi, observed 2026-05-18T19:21:46.667234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:c030fddd0023357d42da6a404c07c3294a03528a6a5a175ecf46bea7590dceba

Observation 1fe0a0f9-2855-47fc-98c1-e0c7a53c56ae · outbound

This paper cites Voyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Voyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.358704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:e2ee977fc3ac6bf13322742b5b0c1c0e3a0dfc023dc8a3fe653db80586f0ea46

Observation 3f518fd1-643b-4a88-9340-2a1a77aafeae · outbound

This paper cites Dynamic speculative agent planning.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Dynamic speculative agent planning

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.708756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:3baa8239f8625a90f3c9aca96d1acbaaedfd4ac10cf6533f26e5cfe2e246a570

Observation 4c2ea35a-3944-4662-935b-63a1c61f2ed2 · outbound

This paper cites Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.698956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:5717da55502f24697e1c5643798f3d317b574ad13db57d999500052b61427df6

Observation 5e4d2d28-a510-4f85-b061-dcabb6b549b8 · outbound

This paper cites PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-07-29T02:24:11.040454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:df6dca613001b4eed86fdb05959d6dcb8519f2b22807d509a671f1cbe15a659d

Observation 68e79a83-d91e-47bf-8ebe-3ea74e9f02cf · outbound

This paper cites Planner-r1: Reward shaping enables efficient agentic rl with smaller llms.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Planner-r1: Reward shaping enables efficient agentic rl with smaller llms

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.374625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:0f533928ae58a5c15bef05881ae1a3bd541e2790042b95cd1494a320a004a1eb

Observation d2762523-ea31-441f-93f4-1145ba12a75d · outbound

This paper cites Planner-r1: Reward shaping enables efficient agentic rl with smaller llms.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Planner-r1: Reward shaping enables efficient agentic rl with smaller llms

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.703490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:2ff6b130642fee018e2e8a3285830ec205b1bdc66f1c6a4b367d4ff90569a13c

Observation 63e3f720-3a8b-4398-ba5c-98ae87fe806d · outbound

This paper cites A brain-inspired agentic architec- ture to improve planning with llms.Nature Communications, 16(1):8633.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A brain-inspired agentic architec- ture to improve planning with llms.Nature Communications, 16(1):8633

Reference 91

Resolution
verified exact
doi, observed 2026-05-18T19:21:46.680300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:95df723ae6a5872f847c1bff5fe5402b55f2f38fdeac0ce0a70c2088404d12bb

Observation 08887fbe-0b1f-413f-9d99-f07452820a75 · outbound

This paper cites Toolformer: Language models can teach them- selves to use tools.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Toolformer: Language models can teach them- selves to use tools

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.348198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:30e115ed24836faf5674d7da717c3fa217557d99235b54208c01b58ba964fd5f

Observation 7d149853-178e-4757-a71c-fdf5f94f7ea8 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey FireAct: Toward Language Agent Fine-tuning

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.752794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:38b154e17eea2d93686cf4c19a095669dda5f86e03519dc098c4e0fd2704fbe2

Observation 30507396-f3d7-4e55-9e8e-68627b5ccdd0 · outbound

This paper cites Knowledge-Centric Hallucination Detection.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Knowledge-Centric Hallucination Detection

Reference 94

Resolution
metadata mismatch
doi, observed 2026-05-18T19:21:46.686021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:e9dda87b7568bf8224ac9fa11085d7e850b831af4a1478dd292f09f0aace7226

Observation d5de3d6c-f9bf-4fc6-be28-3eff2b34e99f · outbound

This paper cites Agent-FLAN: Designing data and methods of effective agent tuning for large language models.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Agent-FLAN: Designing data and methods of effective agent tuning for large language models

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.343624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:d7a4092ff74d24dc809260fa607e5780c05ddec6c8cbd7abce57b5da2fac947f

Observation 035db9a4-bf6e-4035-89f2-2e35b5aca205 · outbound

This paper cites doi: 10.18653/v1/2024.findings-acl.557.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey doi: 10.18653/v1/2024.findings-acl.557

Reference 96

Resolution
verified exact
doi, observed 2026-05-18T19:21:46.641108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:af9b14a30ee22aca6db298c3f87c0d196ff7d5cde3501a7c931a64d65674ecad

Observation ae81e564-ef8e-4cbc-9baa-ae7934baf2ff · outbound

This paper cites Agentbank: Towards generalized llm agents via fine-tuning on 50000+ interaction trajectories.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Agentbank: Towards generalized llm agents via fine-tuning on 50000+ interaction trajectories

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.354885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:df80e904be66680a3488ac9e6fb29db05e5b27887d6e968edb578eb3e558172f

Observation 9893c56f-167f-4db3-88cd-898bf89c4dfd · outbound

This paper cites API-Bank : A comprehensive benchmark for tool-augmented LLMs.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey API-Bank : A comprehensive benchmark for tool-augmented LLMs

Reference 98

Resolution
verified exact
doi, observed 2026-05-18T19:21:46.696153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:d7f1c06a39033648a5dd461864c0ffbe59f330e5b9da44c8a29de641ac176acd

Observation 97d5aa68-d662-47c0-b740-2d8553e714df · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey ToolRL: Reward is All Tool Learning Needs

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.675756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:9e4172ce2c5fbff9271e74e536556fd9d49ff3f288648b85f38ea1935d4f4d57

Observation d4db3474-c124-4d39-aa05-13d3579a4cd9 · outbound

This paper cites Acting less is reasoning more! teaching model to act efficiently.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Acting less is reasoning more! teaching model to act efficiently

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:22:50.392161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:7db5d1e851bc174b13580158da1b62145cc1c2cc40c6c6c3af2b79b9905df9d0

Pith citing papers

Observation d0843f26-55b0-4b87-80d1-6ab94b23f010 · inbound

What Factors Affect LLMs and RLLMs in Financial Question Answering? cites this paper.

What Factors Affect LLMs and RLLMs in Financial Question Answering? The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T05:07:04.507042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T05:06:43.406540Z digest=sha256:80f3e14eb7241e2b1dc0293aff2b8392a880c117916a9d43ed5a53141b8b2ace

Observation c2a11b12-6c43-47a2-a4f2-f527bcfd31f6 · inbound

HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation cites this paper.

HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T09:31:11.965415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T09:27:17.760078Z digest=sha256:250da8457824e94878a88550fc080f98125e60550d5135bfd2292ef3c77cb564

Observation 8bdd12ed-b601-491c-adce-b6b0c8e02f79 · inbound

Graph-Enhanced Policy Optimization in LLM Agent Training cites this paper.

Graph-Enhanced Policy Optimization in LLM Agent Training The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T07:17:44.680792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:17:44.680792Z digest=sha256:6b8f3468788d22e175c6c84c24d2ef4fe4aeeb7f237b214cd60beed024811118

Observation 23833c0c-2daa-4777-bb4e-3bbb44933c9d · inbound

Agentic Learner with Grow-and-Refine Multimodal Semantic Memory cites this paper.

Agentic Learner with Grow-and-Refine Multimodal Semantic Memory The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-17T04:29:01.636424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T04:27:40.232015Z digest=sha256:1c2f798a8336cbe5baa194f9857a2f7465649e6d07e33155be5717e31d97ffbb

Observation 92811bb6-2bb7-4dc7-a1f5-a1d8d3323824 · inbound

Training Multi-Image Vision Agents via End2End Reinforcement Learning cites this paper.

Training Multi-Image Vision Agents via End2End Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-17T01:01:24.314958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T00:59:28.618477Z digest=sha256:1bbf8405d75b6458370db9c1ab376ab2ff597bf5b9c3db7084ab55e7a49c3266

Observation c1ee493e-32f3-4ecb-b41d-38c9b4c1c565 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:14:25.827351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:b2c5d08710e9a1dc9937319bdc8f51900d58ac94e29457257592de8d8d8d73db

Observation beb3b442-90aa-4ef1-9b55-b1494c07e1e6 · inbound

S$^2$GR: Stepwise Semantic-Guided Reasoning in Latent Space for Generative Recommendation cites this paper.

S$^2$GR: Stepwise Semantic-Guided Reasoning in Latent Space for Generative Recommendation The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:24:12.918784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T14:23:44.856958Z digest=sha256:51e15f3107e06e8dde2e36dc1a491a96dca3877eedf49547706bb93f8ee0b209

Observation d1b03946-a403-41ed-a954-c463b59934b3 · inbound

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies cites this paper.

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:57:24.560026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T05:53:29.860037Z digest=sha256:80048ef09e7ca6c860a6eea9d0521eca75eaac8e9b2f1adfe3ad103ad33001e9

Observation a3a7ae48-876a-4096-9503-ade0391bc9f1 · inbound

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents cites this paper.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:45.930272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:45.930272Z digest=sha256:6f42fafac17940be3231c85856e6b0d90249c8273f8ab01a01c1fbc97745acac

Observation adcc7487-9bed-42e4-b1cd-e7fca77e05a9 · inbound

AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments cites this paper.

AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:30:49.538465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T19:07:46.077831Z digest=sha256:45e5fe160b869a58ad6ada91f2149a97238336f50d672a9611ac037f872d4312

Observation 8edec5fe-a6ba-4813-ad89-71f80131a018 · inbound

Towards Knowledgeable Deep Research: Framework and Benchmark cites this paper.

Towards Knowledgeable Deep Research: Framework and Benchmark The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:10:54.620416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:17:35.879705Z digest=sha256:1802bbf3ac2979912ffef3a8ef4db943a1682f75ee5614ee88684b32ded1aadb

Observation 0fdc4e7b-aceb-41ba-a193-4fd9624d5aea · inbound

StaRPO: Stability-Augmented Reinforcement Policy Optimization cites this paper.

StaRPO: Stability-Augmented Reinforcement Policy Optimization The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:21:00.050350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:11:49.056805Z digest=sha256:a41d6cc7843e1ba5fb45a2e2050c22199147d15e1861c4b7c70f1608fdf3cbf7

Observation 4a6fd09a-7f0f-4c66-888e-bfd364676651 · inbound

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning cites this paper.

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:25:59.858348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T18:08:18.056525Z digest=sha256:16958002196fd95d34de19719f39dc3545e14ed06ec0ef5490252bd3345af8fa

Observation e20370d8-dd87-4dfd-a222-7e82f415063a · inbound

From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models cites this paper.

From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:30:58.534608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:09:36.341574Z digest=sha256:578e02da8a690192f67324d5751991f383d995c6e5c89fbe9dba11079af9dafd

Observation c0add783-059e-4a77-b22d-44111e535a31 · inbound

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum cites this paper.

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-10T09:43:49.907062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T05:08:54.560648Z digest=sha256:41f8dd673bf4e83f22614e6f1c847cd2a50b205fd65b4ae80bc5364a56392737

Observation 2e0677ef-c984-4342-91ac-3cc70a947423 · inbound

Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures cites this paper.

Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:51:03.843942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:31:28.242097Z digest=sha256:18e456975e114ed868fc744b82eb415bb0a10583f3bdd7ab3d4bf1db072b35ab

Observation 1a0a8b86-ff0b-471a-bd3b-9e16118db1ff · inbound

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence cites this paper.

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 130

Resolution
verified exact
local_arxiv, observed 2026-05-10T05:25:54.680633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T05:24:00.503836Z digest=sha256:5419e6832ac6b83801faf41811ef173240a770f3779940e342e49fcb23e901f6

Observation 307ac012-d529-4d07-9023-23f6f40fd6f5 · inbound

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning cites this paper.

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:56:29.963355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:26:59.781416Z digest=sha256:d142b6156bbacfbaa9caf4534c24eb0457e4e4a6435785339605a66250d9d866

Observation 6001819a-e57f-4286-a0be-10e7bb80ca33 · inbound

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning cites this paper.

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-05T12:30:59.985134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-05T12:23:00.487399Z digest=sha256:85c1ad94bd9715c35d0dffa33b737d5356c54a0ced80467f1e2766f103e38ab9

Observation 29c2dcb9-8346-4ef4-accb-b9041e761ea6 · inbound

SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation cites this paper.

SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:31:06.624691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:22:30.012610Z digest=sha256:15a8281158d3c141363e4177ab9c1e9d400c94789985298ce81b5f93cc2e2469

Observation ad6aa5a1-a4cd-48c8-9a86-ba8445237e5e · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 123

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:21:29.931588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:55faa09af55a6a9a150214a263358544f10a811ada7d8dd6cb2be8a0e5b18818

Observation 6c2b1663-d5d3-46db-ac4d-2db269f4bc6b · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 123

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:11:17.477480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:7ae236a6b6f03638c61c03a892bc3c2453752372ca216cc6f827396d4faeb6d9

Observation 9000084b-0722-4fd9-a29f-22b7c57b3e1e · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 123

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:02:40.867075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:b70e45f5641e52cb2c94f681dd86eb5e0485f4da51d81c8c8745eab916fb2544

Observation c0cece03-ef9f-4792-972c-2e16d5fd7362 · inbound

Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory cites this paper.

Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 178

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:34.365503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T19:10:30.849963Z digest=sha256:f8a80fbb182e6c6d82f5297c30d1f124f83ec93e7eab4467f4434c8597c10db7

Observation e92c6fc1-7400-4258-afee-285011440b5e · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 165

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:15:49.622476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:8f4e77937a60040dab5ebc4cf9aeb0516f1d709627d75284607ce613d50c4980

Observation f0194b7a-1933-48bb-86ef-b2fecd5e8553 · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 72

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:01:12.713998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T10:23:52.522238Z digest=sha256:d7561f8397b68a22f8054dae5619e68f18c838428a81ce050dc96a744b73b000

Observation 32f7ce1b-c029-47aa-bd4a-5b64f5302d08 · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 72

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T04:00:56.870208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T02:00:00.663355Z digest=sha256:08444d72a1abe290886e602195d9fd8a526ff17ed08a836f6a8397e6cce4329c

Observation 8917f524-9daa-4a54-8631-31c83540d360 · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 72

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T07:17:28.188774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T07:17:13.708752Z digest=sha256:68211a7d1e64f068d7cb772be298f0cf543c96c5ddded7268a1be247444030df

Observation 6705439e-9314-460b-ae08-4e24b5a2f304 · inbound

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction cites this paper.

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:11:12.262776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T09:59:48.604813Z digest=sha256:a643a18ca5444d37b73efb7c276bd120a622c42f876bec8fb188ce2350ce88e9

Observation 6411ce64-b1ed-4acf-a284-d52677013cc2 · inbound

Personalizing LLMs with Binary Feedback: A Preference-Corrected Optimization Framework cites this paper.

Personalizing LLMs with Binary Feedback: A Preference-Corrected Optimization Framework The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T02:41:17.572918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T02:38:55.639568Z digest=sha256:833f6414de950a4194344f7ff927877726835c6ebcf24b35831f255d3684db09

Observation ff787be5-ad14-4556-bcbd-6f07b151209e · inbound

Learning Agentic Policy from Action Guidance cites this paper.

Learning Agentic Policy from Action Guidance The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.608879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:d87b5906fc240cd8129f7f2c4cc02238375a3a824bb78a1b041cf34320cb2d2b

Observation 592ecb5c-3926-42e0-9019-97984e70e84f · inbound

Reinforced Collaboration in Multi-Agent Flow Networks cites this paper.

Reinforced Collaboration in Multi-Agent Flow Networks The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:42:57.113425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T20:42:48.438057Z digest=sha256:95cbe222dc06ec5ee8ed0513cf96cbf7201397f4c6294404705f4d95a9bdd6be

Observation 1c3ca816-0b82-4621-8982-83c43c5c1428 · inbound

NEWTON: Agentic Planning for Physically Grounded Video Generation cites this paper.

NEWTON: Agentic Planning for Physically Grounded Video Generation The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:58:13.729421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T10:56:22.343333Z digest=sha256:eba5ce218e7ab905ed93850d1805893a1064a381c1a9faee5dfda37c55930e24

Observation e6690489-b4a0-4ba0-a2f6-f1c670648b1f · inbound

Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On cites this paper.

Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:33:12.352568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T10:31:16.368065Z digest=sha256:390fff69e5bcdddb91d235c738662d09d8fac0668772f5a0e00cc2dcfaa9002d

Observation 279c5188-7d12-4d1c-8ec6-32364f4b2214 · inbound

Echo: Learning from Experience Data via User-Driven Refinement cites this paper.

Echo: Learning from Experience Data via User-Driven Refinement The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:41:10.745267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T06:37:34.840129Z digest=sha256:179ed1f36e1bf073d41b9734213dfbbdeaed8d4d110a871e2651e6e7809cc989

Observation 5a28c917-b334-4e04-b1b9-1c306c65e98e · inbound

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments cites this paper.

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T16:53:40.551897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T16:51:36.524194Z digest=sha256:0cdee4d76a82f1db554a2c56ca1775cc5d51a3dfc6bea202bb036200e768292d

Observation bca960fa-fcc1-4ee5-9170-aad48d307215 · inbound

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning cites this paper.

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:43:28.975560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T13:38:01.819121Z digest=sha256:0ddcf66d620a7e32529fd0d722c95fe06b5a89641cd712ef581c7e77f8cc4726

Observation c9994a9a-d795-4eca-88b9-8282b8dc8293 · inbound

Libra: Efficient Resource Management for Agentic RL Post-Training cites this paper.

Libra: Efficient Resource Management for Agentic RL Post-Training The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:27.165537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:02:00.385932Z digest=sha256:e1f7f2c4a98e60d9384f274218d58973383d9b5d4a4696f8e3dbca510d996b31

Observation a486bd6a-57e1-454c-b174-7cc35e28dbef · inbound

EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management cites this paper.

EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:56:35.113117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T09:32:51.765633Z digest=sha256:1e5feccd63e2b5632c475cd7649a4bf0aa1ac5cabf909e7083cc471ea6ccf126

Observation f5b1c40d-9b05-4f54-b95d-ebeb54d0a35e · inbound

Co-Evolving Skill Generation and Policy Optimization cites this paper.

Co-Evolving Skill Generation and Policy Optimization The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:26.427426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:6af23310e034a05adec44d8332daf77add0de07c05707ca797e147369e6120f4

Observation ad6b1f5b-7f82-42e4-b3b8-027497a35524 · inbound

Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning cites this paper.

Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:07:28.226428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T17:28:58.574865Z digest=sha256:306d240dd0361dd7e2ee5ee4567b9762cd84e20fa50b8d7ccb9d054b3f7158ff

Observation ffe115d7-9925-48de-b56e-50e4915e100e · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-07-03T09:37:49.371227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T10:21:55.485624Z digest=sha256:451ba6e9415ba99ed3cd787a341477dd39d4c9d3489c3faf3b8b21d2cb615c55

Observation 8e5d4caa-6992-4673-bb21-e69135855cd3 · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:35.232674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:35.232674Z digest=sha256:35bc24b0e4cac1abf3d778d6c8a29756b4d5ce9a8973ff1c4d14e073171bb51a

Observation 0cd33109-5bc7-416d-907d-560d5cd9e59c · inbound

PDAGENT-BENCH: Characterizing, Grounding, and Architecting LLM Agents for VLSI Physical Design cites this paper.

PDAGENT-BENCH: Characterizing, Grounding, and Architecting LLM Agents for VLSI Physical Design The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-03T19:08:49.836672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T02:11:34.174504Z digest=sha256:c49a8561ae15fbd10d8530efad24478f0add7dff2fbdc868b2b30f98acb81509

Observation eb877cd0-e9f2-47e4-85a9-9a7e727956d7 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 260

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:09:40.470466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:ab16203274e9e465103fda2d94f0e37d5ac38c45254c7506081d3b48b29da15a

Observation 13eb0052-5713-4b22-84b9-572d6b235405 · inbound

AIR: Adaptive Interleaved Reasoning with Code in MLLMs cites this paper.

AIR: Adaptive Interleaved Reasoning with Code in MLLMs The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:09:45.005873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T09:06:38.001604Z digest=sha256:2d47bd8d9fe12c0315652d9940c7ef58d235f61a82f529a993fb61b85f6c33c3

Observation 94344ad7-f12f-49de-9845-958e838cddb3 · inbound

OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning cites this paper.

OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:29:52.131453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-26T05:07:16.926615Z digest=sha256:80de6ae8729addc19702daefc8a9ae00ec9c2d76716cd6d817ae906389b85f3c

Observation cc852ecd-9b9d-426e-bbbb-4bc1701b7652 · inbound

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning cites this paper.

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T14:29:53.373940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-26T03:51:51.827622Z digest=sha256:eda9dfed58ae38bcb3094bf86ac2279c04619f0d67761c1ca6fa94ac05bb2e86

Observation 9879578b-5bcb-4dc6-ae8e-5dfd175261ba · inbound

Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting cites this paper.

Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:47:22.214005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T20:40:11.059493Z digest=sha256:232ebbdd34a5dd07e2042b5dce98a65dae525813810bbd0b0c72729c57157adb

Observation 33657ff9-1549-4848-9780-49c8acfa3dd4 · inbound

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents cites this paper.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-02T06:06:40.599975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T05:58:39.107614Z digest=sha256:fe77b121a1b12deb4c48601aa5ab95ab8035f009fbc6be94f7a348fa8a5f3cc4

Observation 76af6b6d-0be3-4282-b302-959eb43ce550 · inbound

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents cites this paper.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:48:49.516258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:abb9e77722775a6a2ece40fefdddab86294d8dbe27f61c2f30f896483e64924c

Observation 59fbeac2-527b-467f-85be-ae71a3a1a923 · inbound

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents cites this paper.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.742129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:ca451c6627f56715e6662f1110ebe2ca86593943568554acd193a3cef5b6d079

Observation 3840bc46-ce83-427a-8711-a6671b20ec41 · inbound

Mathematical methods of reinforcement learning cites this paper.

Mathematical methods of reinforcement learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.774476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:9a5eabd40e6134236965d4860084c3fd915ff0c5cc13d06cb01fbbe4d72b1f93

Observation fd6df471-3b24-4d47-9423-2a9a6012f5b4 · inbound

Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning cites this paper.

Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-09T18:26:26.279667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T18:19:12.180499Z digest=sha256:8fdb2f84921074e39b9dfee8d850fe5b08320ff6a37a2c3b8e563e541d93a44c

Observation ddc527bd-8339-4e66-ade0-586f28941835 · inbound

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy cites this paper.

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 300

Resolution
unresolved
no resolver link, observed 2026-07-14T06:30:16.612345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:30:16.612345Z digest=sha256:163a51d399c33ab63908bfc2c2fed7fdbbf1a772b9c2393764c66788ddfdfaf0

Observation ac28d4f5-fac1-44ee-887f-b466e24e46be · inbound

Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling cites this paper.

Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T04:54:14.134572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:54:14.134572Z digest=sha256:9a41ceec0ff9043aa937e8c8323950a7e64053a043d6d9877ebac94aae122c8e

Observation 93df52ba-fc98-4581-b12a-702612842bfc · inbound

ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability cites this paper.

ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T06:36:47.629823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:36:47.629823Z digest=sha256:78287a82892f0f35dd683c4280c009b62f7fc34fdcc955a32af445c14a24c9a2

Observation fdcfbaf5-e0f4-40e0-ac83-ece8e809376b · inbound

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL cites this paper.

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-02T14:50:18.593672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:50:18.593672Z digest=sha256:4cb1fff6453dd02033cc4bfe4657b878a979d78d660be8fd2d591a360c1f312f

Observation 6477b146-22e0-4b29-84b7-91020de15fca · inbound

From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training cites this paper.

From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T09:45:49.621635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:45:49.621635Z digest=sha256:fca35b3bf929e1c2dd26d938ae59cc528d0d112519c9f817aea2b1fe312b356b

Observation 64f52e8b-3915-41ba-a444-7e1ec7a8b90b · inbound

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications cites this paper.

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-01T03:38:25.752951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:38:25.752951Z digest=sha256:ece2e24a97e65c13cf360afb6d1caf9708ae8ff6aa6f3b3f0f6ad5fb6bc57ac8

Observation 4478aa35-c581-4ddd-9bf5-6208b2bf3672 · inbound

OmniQEC: discovering practical quantum error-correcting codes by an AI scientist cites this paper.

OmniQEC: discovering practical quantum error-correcting codes by an AI scientist The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T01:20:49.101238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:20:49.101238Z digest=sha256:e1f5a003ea576ed182d99632e8984deacfafc988473c79e332afcd030db2fbe8

Observation 7a145b9c-9181-430b-acfa-2b13de815d79 · inbound

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution cites this paper.

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-30T21:12:59.446937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:12:59.446937Z digest=sha256:775f5c40057c9e3f822d5849733bdd8c5af05fe5dcda92cab25b65b5212d66a2

Observation 48e765bb-755c-4a4d-b144-67fca8c8c5da · inbound

Deep Reinforcement Learning: From First Principles to Reasoning Models cites this paper.

Deep Reinforcement Learning: From First Principles to Reasoning Models The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:05.118218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:05.118218Z digest=sha256:62d80412e919fb4a7664ffadf372825c109130d54533acda388d15a51c5c6c4e

Observation 27ca895f-cc55-44a1-8c62-003964f5102b · inbound

From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution cites this paper.

From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T13:31:24.753889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:31:24.753889Z digest=sha256:ef86c0d3e1cb0595980c4db64b6496d6e5ab580fb428b9dbaeff1a07fa056eea