Pith. sign in

Paper Citation Record · LEDGER

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

As of 8 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 6 inbound Pith citation observations for arXiv:2509.06980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06980 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:07:18.380838Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:00:17.278040Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:16:57.747266Z

Reference resolution

11 of 11 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6c0e1a07-587d-4e63-88ab-a40d9ff44466 · outbound

This paper cites Introducing gpt 5.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Introducing gpt 5

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.485775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T13:07:18.350381Z digest=sha256:7f4ad04abdfc3017cbdb34fe80cf92a2de43e31c94769dcd713c153a68354ed3

Observation 4e916ce5-dffe-4c26-8060-2fc22b9698e7 · outbound

This paper cites Agent rl scaling law: Agent rl with spontaneous code execution for mathematical problem solving, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Agent rl scaling law: Agent rl with spontaneous code execution for mathematical problem solving, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.477442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T13:07:18.354025Z digest=sha256:dcb8362a7508d68230de7343ff3bc7131abaa4645465c48b5ddb0779c1206c24

Observation 4a0e9739-c96a-49fe-9c7d-c669497d230e · outbound

This paper cites Agentic reinforced policy optimization, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Agentic reinforced policy optimization, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T13:07:18.357017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:07:18.357017Z digest=sha256:8b6f7467e4323184b14b19f0723ebc2bc19dec2881e523ebbca188f6d2b3cae5

Observation 3357afd6-2c30-499e-b906-b4dd26856800 · outbound

This paper cites Agentic reasoning and tool integration for llms via reinforcement learning, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Agentic reasoning and tool integration for llms via reinforcement learning, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.463675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T13:07:18.359940Z digest=sha256:78dfc152300b0564afe73428c859345048cbe51ab56b676194285ccd8fde3c99

Observation 536182d1-6685-4491-9134-d11dd6372bee · outbound

This paper cites Search-o1: Agentic search-enhanced large reasoning models, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Search-o1: Agentic search-enhanced large reasoning models, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T13:07:18.362922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:07:18.362922Z digest=sha256:ace66e92463e8cc35d57de4fa084a628cd638c96f2f77908d49baa9bee879d15

Observation 525d9313-2543-4522-9255-460b0002a9f0 · outbound

This paper cites Search-r1: Training llms to reason and leverage search engines with reinforcement learning, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Search-r1: Training llms to reason and leverage search engines with reinforcement learning, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.450757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T13:07:18.365922Z digest=sha256:3ada49e4fde48bf8f9ab00953a935eec3f0bfc76ae65e6db378b67015e6f944b

Observation 8bdf368b-a736-4ff5-b227-6c47ad6c23f7 · outbound

This paper cites Mmsearch-r1: Incentivizing lmms to search, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Mmsearch-r1: Incentivizing lmms to search, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.442344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T13:07:18.368961Z digest=sha256:f6d4e8857b12037df8fd8382c4020605fa0375108b4ea3c4d6e747489fc10603

Observation c4d907dc-e03e-42c8-a09e-407c739803a9 · outbound

This paper cites Deepresearcher: Scaling deep research via reinforcement learning in real-world environments, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Deepresearcher: Scaling deep research via reinforcement learning in real-world environments, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.433534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T13:07:18.372033Z digest=sha256:cea6a458665772bc7ddaf23bef12b7c4802ac41a7579aeaf43ee29b5b8ca1522

Observation b161d047-2c1b-404d-8a62-0ce2e56eafd8 · outbound

This paper cites Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T13:07:18.374868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:07:18.374868Z digest=sha256:67c66fcbfce7bc8283798d7550c8b168bb7209a109087a93c8ff09356f988415

Observation 8f37fafe-1ae4-47f6-9956-5a28792d3deb · outbound

This paper cites TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T13:07:18.377992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:07:18.377992Z digest=sha256:60d5f0b40860032778028ed1d5676deac223ed32353e6ec9cc8bc6277edd6a3b

Observation 17e17ad6-cbbf-4845-a144-4485a5f87b7a · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use HybridFlow: A Flexible and Efficient RLHF Framework

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T13:07:18.380838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:07:18.380838Z digest=sha256:f9235ac1a8cc39ff267d006e05bc886968e0996858b35fab67d26303f86e4d12

Pith citing papers

Observation 06872784-d724-4494-9912-f89a2a61ba75 · inbound

LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services cites this paper.

LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T18:00:17.278040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:00:17.278040Z digest=sha256:914b9bded77ae5c97b99006d731ee61d080c9dc5a79cc89095e22a4a826c3d09

Observation bddbbd6a-8510-4f17-aa96-93e1f8c41bf3 · inbound

SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents cites this paper.

SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:46:48.471379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:26:57.084201Z digest=sha256:059f6cf12a24d1659c97ebc4e9305ea3094600187e2209e78c7c201e26dcb491

Observation 1142874d-a039-4f24-b1a4-1b215bfacb85 · inbound

VistaHop: Benchmarking Long-Horizon Visual DeepSearch cites this paper.

VistaHop: Benchmarking Long-Horizon Visual DeepSearch RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:29.261024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T10:35:45.449334Z digest=sha256:b503079016d2816bd743f8b885fa1cf7665b56c5aa81d952ecea754913cfbb24

Observation ab8021cf-9c68-49ae-902d-b902bba21c4e · inbound

VistaHop: Benchmarking Long-Horizon Visual DeepSearch cites this paper.

VistaHop: Benchmarking Long-Horizon Visual DeepSearch RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T12:32:59.871356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:32:59.871356Z digest=sha256:151a929ba5f6636aad75c9e3a81be169240f11fb7a45259cf1cf41c3ec2cc04c

Observation aa4b8412-441a-466e-9849-001daf2e9e7e · inbound

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents cites this paper.

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:57.748936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T02:11:11.638029Z digest=sha256:35a9b3683240584f246e8e7a6aca6534147eee92bf412ca1dab731a767682b12

Observation f73945fa-ece9-4df3-9ccd-86ed0e883e3a · inbound

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation cites this paper.

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T08:56:36.010374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:56:36.010374Z digest=sha256:d4b6647ad3eb9594b5edb018ded06b620537053e6a8308f4edac08d9b01d5335