Pith. sign in

Paper Citation Record · LEDGER

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

As of 8 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 6 inbound Pith citation observations for arXiv:2509.06980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06980 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:07:18.380838Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:00:17.278040Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:16:57.747266Z

Reference resolution

11 of 11 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6c0e1a07-587d-4e63-88ab-a40d9ff44466 · outbound

This paper cites Introducing gpt 5.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Introducing gpt 5

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.485775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T13:07:18.350381Z digest=sha256:a0145b7a1b8e6791d78df40c6750be1331487504b1dd4441b4bda42f1e3e2477

Observation 4e916ce5-dffe-4c26-8060-2fc22b9698e7 · outbound

This paper cites Agent rl scaling law: Agent rl with spontaneous code execution for mathematical problem solving, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Agent rl scaling law: Agent rl with spontaneous code execution for mathematical problem solving, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.477442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T13:07:18.354025Z digest=sha256:404a9498c8b3d01cfe7189cb67a3e3cfc6036f14cee53722238645bd81f4d9d5

Observation 4a0e9739-c96a-49fe-9c7d-c669497d230e · outbound

This paper cites Agentic reinforced policy optimization, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Agentic reinforced policy optimization, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T13:07:18.357017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:07:18.357017Z digest=sha256:8b6f7467e4323184b14b19f0723ebc2bc19dec2881e523ebbca188f6d2b3cae5

Observation 3357afd6-2c30-499e-b906-b4dd26856800 · outbound

This paper cites Agentic reasoning and tool integration for llms via reinforcement learning, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Agentic reasoning and tool integration for llms via reinforcement learning, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.463675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T13:07:18.359940Z digest=sha256:008627959fadd18009fb545feb81ef6820768f1a64acea75b517c9b5e4d37775

Observation 536182d1-6685-4491-9134-d11dd6372bee · outbound

This paper cites Search-o1: Agentic search-enhanced large reasoning models, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Search-o1: Agentic search-enhanced large reasoning models, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T13:07:18.362922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:07:18.362922Z digest=sha256:ace66e92463e8cc35d57de4fa084a628cd638c96f2f77908d49baa9bee879d15

Observation 525d9313-2543-4522-9255-460b0002a9f0 · outbound

This paper cites Search-r1: Training llms to reason and leverage search engines with reinforcement learning, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Search-r1: Training llms to reason and leverage search engines with reinforcement learning, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.450757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T13:07:18.365922Z digest=sha256:776823f432b17507f9f46987bd5d1e0d9fc50ab3dd1b04ebb2cb2c58f178dad4

Observation 8bdf368b-a736-4ff5-b227-6c47ad6c23f7 · outbound

This paper cites Mmsearch-r1: Incentivizing lmms to search, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Mmsearch-r1: Incentivizing lmms to search, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.442344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T13:07:18.368961Z digest=sha256:36c70798c9c4a0b88636ecf2bfaf93ba5f77fdee43038b34078cc34542cfb3f8

Observation c4d907dc-e03e-42c8-a09e-407c739803a9 · outbound

This paper cites Deepresearcher: Scaling deep research via reinforcement learning in real-world environments, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Deepresearcher: Scaling deep research via reinforcement learning in real-world environments, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.433534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T13:07:18.372033Z digest=sha256:c819747ede37ac4ab726fc8c2a3e6d8d6474a986c6a79e9f43bb8fef7253e10f

Observation b161d047-2c1b-404d-8a62-0ce2e56eafd8 · outbound

This paper cites Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T13:07:18.374868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:07:18.374868Z digest=sha256:67c66fcbfce7bc8283798d7550c8b168bb7209a109087a93c8ff09356f988415

Observation 8f37fafe-1ae4-47f6-9956-5a28792d3deb · outbound

This paper cites TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T13:07:18.377992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:07:18.377992Z digest=sha256:60d5f0b40860032778028ed1d5676deac223ed32353e6ec9cc8bc6277edd6a3b

Observation 17e17ad6-cbbf-4845-a144-4485a5f87b7a · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use HybridFlow: A Flexible and Efficient RLHF Framework

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T13:07:18.380838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:07:18.380838Z digest=sha256:f9235ac1a8cc39ff267d006e05bc886968e0996858b35fab67d26303f86e4d12

Pith citing papers

Observation 06872784-d724-4494-9912-f89a2a61ba75 · inbound

LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services cites this paper.

LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T18:00:17.278040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:00:17.278040Z digest=sha256:914b9bded77ae5c97b99006d731ee61d080c9dc5a79cc89095e22a4a826c3d09

Observation bddbbd6a-8510-4f17-aa96-93e1f8c41bf3 · inbound

SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents cites this paper.

SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:46:48.471379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:26:57.084201Z digest=sha256:ac0ee156e06febbdefad25461c1286384de8ec2f83ea86e036732ddb5fab887b

Observation 1142874d-a039-4f24-b1a4-1b215bfacb85 · inbound

VistaHop: Benchmarking Long-Horizon Visual DeepSearch cites this paper.

VistaHop: Benchmarking Long-Horizon Visual DeepSearch RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:29.261024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:35:45.449334Z digest=sha256:b98842b3684edcafbce7218cbe3cded25fd64fee09848695422eae73687db47d

Observation ab8021cf-9c68-49ae-902d-b902bba21c4e · inbound

VistaHop: Benchmarking Long-Horizon Visual DeepSearch cites this paper.

VistaHop: Benchmarking Long-Horizon Visual DeepSearch RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T12:32:59.871356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:32:59.871356Z digest=sha256:151a929ba5f6636aad75c9e3a81be169240f11fb7a45259cf1cf41c3ec2cc04c

Observation aa4b8412-441a-466e-9849-001daf2e9e7e · inbound

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents cites this paper.

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:57.748936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T02:11:11.638029Z digest=sha256:638af5e728135594e1fe674b6821d701119c1cab94d09ee8e2d29ab5abebcc8c

Observation f73945fa-ece9-4df3-9ccd-86ed0e883e3a · inbound

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation cites this paper.

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T08:56:36.010374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:56:36.010374Z digest=sha256:d4b6647ad3eb9594b5edb018ded06b620537053e6a8308f4edac08d9b01d5335