Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:07:18.380838Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 6 inbound Pith citation observations for arXiv:2509.06980.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:07:18.380838Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T18:00:17.278040Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T12:16:57.747266Z
11 of 11 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6c0e1a07-587d-4e63-88ab-a40d9ff44466 · outbound
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Introducing gpt 5
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4e916ce5-dffe-4c26-8060-2fc22b9698e7 · outbound
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Agent rl scaling law: Agent rl with spontaneous code execution for mathematical problem solving, 2025
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a0e9739-c96a-49fe-9c7d-c669497d230e · outbound
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Agentic reinforced policy optimization, 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3357afd6-2c30-499e-b906-b4dd26856800 · outbound
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Agentic reasoning and tool integration for llms via reinforcement learning, 2025
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 536182d1-6685-4491-9134-d11dd6372bee · outbound
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Search-o1: Agentic search-enhanced large reasoning models, 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 525d9313-2543-4522-9255-460b0002a9f0 · outbound
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Search-r1: Training llms to reason and leverage search engines with reinforcement learning, 2025
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8bdf368b-a736-4ff5-b227-6c47ad6c23f7 · outbound
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Mmsearch-r1: Incentivizing lmms to search, 2025
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c4d907dc-e03e-42c8-a09e-407c739803a9 · outbound
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Deepresearcher: Scaling deep research via reinforcement learning in real-world environments, 2025
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b161d047-2c1b-404d-8a62-0ce2e56eafd8 · outbound
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f37fafe-1ae4-47f6-9956-5a28792d3deb · outbound
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17e17ad6-cbbf-4845-a144-4485a5f87b7a · outbound
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use HybridFlow: A Flexible and Efficient RLHF Framework
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06872784-d724-4494-9912-f89a2a61ba75 · inbound
LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bddbbd6a-8510-4f17-aa96-93e1f8c41bf3 · inbound
SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1142874d-a039-4f24-b1a4-1b215bfacb85 · inbound
VistaHop: Benchmarking Long-Horizon Visual DeepSearch RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ab8021cf-9c68-49ae-902d-b902bba21c4e · inbound
VistaHop: Benchmarking Long-Horizon Visual DeepSearch RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa4b8412-441a-466e-9849-001daf2e9e7e · inbound
TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f73945fa-ece9-4df3-9ccd-86ed0e883e3a · inbound
From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.