Pith. sign in

Paper Citation Record · LEDGER

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning

As of 10 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2601.11960.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.11960 v3

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T10:01:26.786756Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved35
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de93cd5b-a255-4ae4-b4af-eb7e298cce7c · outbound

This paper cites Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and R \' e mi Munos.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and R \' e mi Munos

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:23.124399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:23.124399Z digest=sha256:546adfaef983aa73fdcd0b98015cb61d479d7c065b596066cf456350afee97ea

Observation df6e4778-5269-45c0-88f6-44e35b543106 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Reasoning with Exploration: An Entropy Perspective

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:23.184597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:23.184597Z digest=sha256:affcff7d09937a20770608c8cdbb026d78a3f6d4ba9cd181c565daa4be532332

Observation 83810eff-5631-467e-9a1c-b3d5faa660be · outbound

This paper cites Christiano, Jan Leike, Tom B.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Christiano, Jan Leike, Tom B

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:23.323839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:23.323839Z digest=sha256:c9d5fcfdb93428cc5dcd981cd57061b0d7846335e327143d79319bc2c1f36a0b

Observation e4f750a5-f8d4-4b08-a41a-cae44bc892b2 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:23.465502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:23.465502Z digest=sha256:853a1dbcebfa244fc6000eefe54899c43878841f5883490e8d74fe1aab523b98

Observation 2a4df2cc-5786-4c62-8279-14453832d4d8 · outbound

This paper cites an unresolved cited work.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:23.598386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:23.598386Z digest=sha256:fbc6ede7d428f208ed368a8a632323bb2363348b93845a153107d6b864ccced2

Observation 08bdac19-a887-43f9-8756-62dddc72b42e · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:23.711845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:23.711845Z digest=sha256:4b5c82ba180204c3e32491652a0de7cfc131d85b6ce12af75ba18c0a2b9b3455

Observation fe7394cb-3bf2-42cc-95bf-a360fdece666 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:23.793563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:23.793563Z digest=sha256:ce5cc9467f23a7f05bfd1ba5d6b015306d3787f621499309f33718e17710749a

Observation a3718b81-1ec8-407e-83b2-b33a2b6ebd14 · outbound

This paper cites DeepSeek-V3 Technical Report.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning DeepSeek-V3 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:23.977634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:23.977634Z digest=sha256:eb0a5774e986cb16df48ec680774d8b8e26388f4a223c89f1e83492c3a05e8cb

Observation 4359e8d4-22d2-4f64-ae2b-4eb44359f6ea · outbound

This paper cites an unresolved cited work.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:24.108061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:24.108061Z digest=sha256:839e4dcf79d00fe123d0e29d3056124ddc6bdba8533a745e2c399dfa376f1b62

Observation 6ddb480e-8dd8-4ffb-9799-f6aab7625168 · outbound

This paper cites an unresolved cited work.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:24.165526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:24.165526Z digest=sha256:70371a0d481d5e325663e7aeef514b360aa2c8c1c0634cfd9f19605fdb721010

Observation 896f061b-00f4-4a0e-95e1-d87943c4a162 · outbound

This paper cites an unresolved cited work.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:24.235859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:24.235859Z digest=sha256:afb912696681a1e2ec7c6028000722b7a2a75e6cd6cdff2f44788044c9eb5ae0

Observation ecc0eea7-c38b-4c39-be4b-8c2c7c7efa47 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:24.289678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:24.289678Z digest=sha256:1a5e97885023f7b09b57a430bf8456c2f464c467fbc88fcb77096c9ba46f37d9

Observation 7cefa712-b75c-4fbc-ab67-0b56b131c00d · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:24.361961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:24.361961Z digest=sha256:549b56c5f27af07e282e9a5be66ffeeef6521b7e9eb41ca50593cce19e5ce380

Observation 77d9dbf8-a841-4e7d-ba04-9821845fcff6 · outbound

This paper cites an unresolved cited work.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:24.418088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:24.418088Z digest=sha256:bdf2d89d641d4ad84266d55683db20748a4e3bab0a6f0a9185cb9d5126d95696

Observation 1e5abbef-0fb4-472b-920e-6c71c43a1b5f · outbound

This paper cites an unresolved cited work.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work

Reference 15

Resolution
malformed identifier
no resolver link, observed 2026-08-03T10:01:24.482177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:24.482177Z digest=sha256:a53d978d702712997200dd95f1d9f0990d5eedfac33119faf568b38d9c7af68c

Observation 26f60728-75b5-4ae3-aa3c-b533ba7f0d97 · outbound

This paper cites an unresolved cited work.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:24.583677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:24.583677Z digest=sha256:1a4ae4193978e87ade19dc74e926393e46107f82a256da08388020830d5c1935

Observation e03b6bc5-cfe7-43a8-b400-aeb42e093ad6 · outbound

This paper cites an unresolved cited work.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:24.703473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:24.703473Z digest=sha256:f5acb14d783c5e1a8697e1b5a091a84628cfbb0a158f0177a66ccda48918ce28

Observation 31f5a61a-9a6b-4c37-9878-502cfd63b686 · outbound

This paper cites Park, Junsu Kim, Gyeongman Kim, Jinyoung Jo, Sean Choi, Jaewoong Cho, and Ernest K.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Park, Junsu Kim, Gyeongman Kim, Jinyoung Jo, Sean Choi, Jaewoong Cho, and Ernest K

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:24.827479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:24.827479Z digest=sha256:7d55594e74a48a1a519b740f16867df2a14fc05459a5f0aede04ebe5cef0d837

Observation 072e22b8-610e-4e70-9e52-1140402e5bdd · outbound

This paper cites Qwen2.5 Technical Report.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Qwen2.5 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:24.966307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:24.966307Z digest=sha256:4470f73884cb50f03b041169502b1ab79fba16d68c7b9af1e42c3705a573032d

Observation 73f7241f-5bfc-422f-b655-5e7e3e9525d8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:25.049961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:25.049961Z digest=sha256:053c874b59711454b2bd1e5646475407524bbc85e4ef9c10573545544d5b9e4a

Observation 0e73d007-56cc-40a1-ab2c-bab8052de6f1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:25.107146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:25.107146Z digest=sha256:42d801c70539102ab5fb3ee8117010d8019badbbac27f43111d6889fd3127bb4

Observation 123152b8-da22-41d5-a73c-f9116eb03e7f · outbound

This paper cites an unresolved cited work.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:25.159166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:25.159166Z digest=sha256:2785bf0c11a7f2329bd8d66dcaf437043763e3a409d45d216de4c0592333e25f

Observation f259c5e4-0b05-431d-b06d-860e347777f2 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:25.225411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:25.225411Z digest=sha256:cc4efda0e33960ba7214e2c4f8ba3ddb01adbac08d23a6d44f5d0ecf56e497ec

Observation 507e3a6b-9f24-490b-b33b-6d11de912508 · outbound

This paper cites Sutton and Andrew G.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Sutton and Andrew G

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:25.350439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:25.350439Z digest=sha256:3021f93a7c7d15a68a724d2d22f6e859583721a027eaae1297201cb878a8ebe5

Observation 6e54eb95-f327-4220-9e32-1e2f4eafbb16 · outbound

This paper cites an unresolved cited work.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:25.453248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:25.453248Z digest=sha256:4a7c9692ae1d670a5f77c7d8db54645f6d6866f61778a09332d8ad4965a874bf

Observation 03613c47-f757-4d02-85a3-20eb907510f6 · outbound

This paper cites ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:25.490489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:25.490489Z digest=sha256:dd1a952b352ae8f26877c7fbcda6ff84b5c63e211752d4110fab2f4707219c77

Observation 31b0c102-880e-4755-bbaa-8d84bca1bb8a · outbound

This paper cites Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:25.614727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:25.614727Z digest=sha256:5ad213f62f6071f695144f4b6f6bd44cf9f4a18ed446298e98ff08fa579ed2e9

Observation dd69d07a-9e81-4deb-8017-5450a211bc63 · outbound

This paper cites an unresolved cited work.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:25.745607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:25.745607Z digest=sha256:f8a97d1a52b2751c07779d38ef80c3ab1f08a2cc3a61b58366665c16b22e2097

Observation 68a64716-ac48-493d-98da-ceb42c2f96f6 · outbound

This paper cites an unresolved cited work.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:25.860954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:25.860954Z digest=sha256:4e1ffc216db5c34656253233a944090f507c1bd371c657e3cb2a9980a2619bdd

Observation d1989704-4b62-4235-ab9d-3fb24ae7f43d · outbound

This paper cites an unresolved cited work.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:25.980371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:25.980371Z digest=sha256:bcb58fc227e808a4f5f4cfcae9399ea13315666db70745bd88eea46b1da63215

Observation 924bd161-e955-4f97-9a89-d39d01e7d3bb · outbound

This paper cites an unresolved cited work.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:26.120049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:26.120049Z digest=sha256:9f40ddb8871008a178528db36d3e9bc1300bdfdf2d2ce93b351e3fb0505b0f38

Observation ef88384c-f0aa-4c12-8df2-2224cc34eff7 · outbound

This paper cites Qwen3 Technical Report.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Qwen3 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:26.286249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:26.286249Z digest=sha256:4869611b08070fde8fd452f6c0dab11813ee31a5fe8598c228f36b83589d7fa7

Observation eb28cef4-4463-444b-999f-a9bd72234041 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:26.427526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:26.427526Z digest=sha256:58ad0b9bcfab56109cacbca615c5d3174c4ffa3b4c0dbb19a9d72a96f88b496c

Observation 4b43a749-ca41-4c1a-841d-7204f0597a61 · outbound

This paper cites an unresolved cited work.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:26.544678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:26.544678Z digest=sha256:92e8d3856424697f9e24da0d34ade20c65d2e8576c6c092901f80a96a33f96be

Observation 69cc710c-9a75-43bd-8724-32f641e3ce03 · outbound

This paper cites online" 'onlinestring :=.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning online" 'onlinestring :=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:26.631117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:26.631117Z digest=sha256:ec77bf8c4d2bef7ca75bf14ae43c0de7fcd741edc4674efb9128e9091c1a0aa3

Observation fe36e1f0-98b3-4da3-a948-14f4ff5733e9 · outbound

This paper cites write newline.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning write newline

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:26.786756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:26.786756Z digest=sha256:a45fd4136fa6aff97293bf993fa3dba006af7a985097db77c39b22cfc4915542

Pith citing papers

No inbound Pith citation observations are available.