Pith. sign in

Paper Citation Record · LEDGER

Process Reward Informed Tree Rollout for Effective Multi-Turn RL

As of 23 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2607.15610.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15610 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T22:51:02.115809Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc654b2b-d399-4ac3-bb76-dc3e5a7b49ff · outbound

This paper cites Agentic Reinforced Policy Optimization.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Agentic Reinforced Policy Optimization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.855827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.855827Z digest=sha256:68d3153281b0fe4d7cf4cf2c0c0f0bf25023bdbdce623eeba9b7d8b7c2d13d89

Observation 1af78faa-aa71-4e41-a3a5-aa37f897a72f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.862640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.862640Z digest=sha256:390df72ff7de94caa6e31ea159fa34b441da15ef31ddce1551fb92b87073edd2

Observation d266afe9-452f-42fb-9f2e-54c0854fbfac · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.867466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.867466Z digest=sha256:b2513fcf1a6b990528d979d3d7ad36413cc24f92e6ba9141dddae240f24746e3

Observation ceda4a33-923c-4f12-abc3-ca1097d64285 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in Neural Information Processing Systems , volume=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.872812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.872812Z digest=sha256:becac3089961fb76a30ac729399bea43fd6377a8d9dfff6e708d84ad610b6623

Observation 11173a2b-ceb1-4a1b-944e-6b0beb524a81 · outbound

This paper cites arXiv preprint arXiv:2602.22817 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2602.22817 , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.878075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.878075Z digest=sha256:e8edab1fdaef8e7183d4c29c84fe6422f6ed9dcbc7360e4d28fdf932c17f3c04

Observation 64c0fd4b-8384-4fa9-8a1b-84f79f30dfb6 · outbound

This paper cites arXiv preprint arXiv:2509.06040 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2509.06040 , year=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.882565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.882565Z digest=sha256:24b8b33dc421a0cccc35d9558d44e22e96552da5b96923ae601d31c6db221805

Observation b8af38c2-5456-4162-b519-7ed240284fe8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in Neural Information Processing Systems , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.887952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.887952Z digest=sha256:5b522a04dcc736b52eb986a878ab718e7eaae757018964d31e3ac6b6aad9439a

Observation d980537e-6949-4b51-a9d2-f9032e0bee0b · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in Neural Information Processing Systems , volume=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.892992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.892992Z digest=sha256:eeba0d3fd02f37a7b02561575dea883be6e8db8c14d34720d8649103a6a5fd7b

Observation c17c63ae-e387-4181-9083-421ea1a67109 · outbound

This paper cites arXiv preprint arXiv:2603.03078 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2603.03078 , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.897762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.897762Z digest=sha256:de7ab4e4745786c540a36b0473ca58ace2e7b34f9aa2962627a6a9b185a2b194

Observation a1c6a846-5789-41fe-b98c-d59c96edafdf · outbound

This paper cites arXiv preprint arXiv:2511.16108 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2511.16108 , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.902422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.902422Z digest=sha256:33f5c372a09f24a30c1327f51efc493c14f6c18061dcdc9c3b0acb4f82e4b31e

Observation bf37dc69-9ff4-49e7-966e-590967662b13 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.907282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.907282Z digest=sha256:5e1ed773ec4d84d4d0cf26ae044f5c4c34d6c82f10803c99f6ad30475bbf9dc5

Observation c3369599-5db8-429f-bf20-1b03f9917d68 · outbound

This paper cites Advances in neural information processing systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in neural information processing systems , volume=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.912213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.912213Z digest=sha256:54d635367b5f94a99d6f88651518be7c827e0653538bae6c213f30c10c5648e2

Observation 3a644d5f-e4da-4824-9785-b8729c3db3eb · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.916601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.916601Z digest=sha256:03db7325494cd486bcb35dc7d45cc6a775957b29e1ddff273f7c2532842b9834

Observation 4a037f56-8f47-4f3f-ad54-cee3b2c3fba0 · outbound

This paper cites arXiv preprint arXiv:2407.01476 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2407.01476 , year=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.921572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.921572Z digest=sha256:3dadf7b900935739c88d1a5448c7e4d81b0a5d2529612d591a9a293462fadf23

Observation 9cb62683-61db-48f2-b63c-25c1b718a0cd · outbound

This paper cites TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.926088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.926088Z digest=sha256:b88cd970c632b34beb59a2494038ceaac7434476d8eedd7283a69b4317d0ae7e

Observation 76d710ab-ab6d-47ec-90a7-b16436448d20 · outbound

This paper cites arXiv preprint arXiv:2603.02216 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2603.02216 , year=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.931380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.931380Z digest=sha256:4514c733865d52fc11cb2e6c63c8b73e7561fd0f6556810ae1b7a0ba4e734d24

Observation 0595bac0-09dc-473d-97ec-77345758cebb · outbound

This paper cites arXiv preprint arXiv:2510.14545 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2510.14545 , year=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.936445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.936445Z digest=sha256:149c922dfec79b82af1b510d2ad039748fe81a6fefc1c633ffaea5cf66337e4a

Observation 31f50e61-ee74-448b-b5e9-749f725fe73a · outbound

This paper cites an unresolved cited work.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.942075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.942075Z digest=sha256:5afb58fd6c549aa512761de7294a22812c7032fd17563469ba4d43d43dec9f88

Observation fc1243b5-7dda-455d-a4a6-c4b49a9be0ef · outbound

This paper cites TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.946686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.946686Z digest=sha256:30248fd3b9f39d9d6c7792004ac5a4559f91b447c01be379217bb43c03ab6ae9

Observation 053ca571-7a11-480d-a77e-2b38be2ce1ee · outbound

This paper cites arXiv preprint arXiv:2512.08153 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2512.08153 , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.952267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.952267Z digest=sha256:d2b8447e645ccffec3983f0606c41b84f3dee44858d40eacb2be3e353b31aacb

Observation 69dbb63a-ebc6-4007-9b44-5d513f8f1081 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.957238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.957238Z digest=sha256:65d4074d86677664d103727c402093cb8c6a3a19319766f5cdb3b8909e817f95

Observation 88e9ea15-985d-4c67-a444-d926277ba6d5 · outbound

This paper cites arXiv preprint arXiv:2603.00296 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2603.00296 , year=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.961785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.961785Z digest=sha256:8a384c6b84d8287c3cafb9540a14c6d5ac0130089be59b732a549a1427177829

Observation ca1c3646-9076-402a-b257-8b84b5901d4b · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.966286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.966286Z digest=sha256:3dcb15b340f7a70baac100d7ccece30d779115ca29dccba694e358585b5fe8aa

Observation 9687b349-3f62-4dc5-a68f-30c48f4e2a2a · outbound

This paper cites International Conference on Learning Representations , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL International Conference on Learning Representations , volume=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.970877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.970877Z digest=sha256:279f2051a820ad32e928685fe66af5ee7a4e3335442f7a9c94b0099ce979a0b7

Observation 409d0468-5740-4bbc-bd5c-738da42d02ef · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.975760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.975760Z digest=sha256:192a5ed84419596c1067ad024fef54084e5b924e8043197844110e04b1bdc211

Observation bd2f6b3c-1eda-4706-b9e2-96700058a548 · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Skywork Open Reasoner 1 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.981544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.981544Z digest=sha256:b76130feed4afb4e5ca1c08bc0991451bda8c35f857d805c5ad4fca29a76ffa6

Observation 41784db7-cd49-4d7e-a3a2-c2d3b8546a4d · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Process Reinforcement through Implicit Rewards

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.986493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.986493Z digest=sha256:36d4dcb7525516817f9dc231c00b8d9119bd71a88155f8e7ae139b0b22a7e1c1

Observation fe98da61-d19b-4833-80e0-609915bfd183 · outbound

This paper cites arXiv preprint arXiv:2509.19199 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2509.19199 , year=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.991250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.991250Z digest=sha256:7adfaab3d44bdcfd559d3ec43a36919ee7cb79ec707dc54491b82af8bf06deb3

Observation 238ddf24-c58b-4aeb-9108-74853ddc1fc8 · outbound

This paper cites arXiv preprint arXiv:2602.17025 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2602.17025 , year=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.996019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.996019Z digest=sha256:c6f3bc2dbcf50b12ea4e322efb7e8262ed517c7c6d6845a63460f0aede30fa70

Observation a0dbd9bd-7ce5-405e-9bd5-b21bbd184ace · outbound

This paper cites Advances in neural information processing systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in neural information processing systems , volume=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.000591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.000591Z digest=sha256:7cc5e52dfd8ad07e2475f924658ea104776fef97ad91529980610bdb011061ae

Observation e27a9318-3394-4ff3-b71a-92ec7e67eb52 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Training Verifiers to Solve Math Word Problems

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.005624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.005624Z digest=sha256:cdf359ca9351bede3f698e9d84f43b418923c5944afef01a588fe49f090e8bd7

Observation c10ddbf0-bbf3-4b6a-9c03-05f6d232a9a9 · outbound

This paper cites International Conference on Learning Representations , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL International Conference on Learning Representations , volume=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.011033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.011033Z digest=sha256:1d113256dbc8b64b45c2376e749603a89e9d6d5503a44c80e9f273e23d89ac26

Observation 9adaac36-5103-476c-a731-d4ea3d76b883 · outbound

This paper cites nature , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL nature , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.017289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.017289Z digest=sha256:0aeaf17b4e88ceaa20015a3f86e4931616a7549c19f90f8dcfa56d0697c0a700

Observation bd08aaba-361d-4a92-8b93-176bc1499bb7 · outbound

This paper cites From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.022603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.022603Z digest=sha256:fe2f674ff6c5192a8887634e567e7fcebbbbe7073e65fef731b8dac543cae86f

Observation ea0a0d05-9fda-4bad-b6ae-6be651077ec0 · outbound

This paper cites OpenAI Gym.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL OpenAI Gym

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.027774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.027774Z digest=sha256:4a671669cdd97969c75e20187ef65fc9b79549d758dbf9e1041ff27bae824728

Observation abd75e59-6196-4652-8578-6a365dae4207 · outbound

This paper cites an unresolved cited work.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.033389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.033389Z digest=sha256:b6a524679406c880a9719a7052d66b7b6094b0d5e872eebbcb003ba207074f7c

Observation 4af383bc-1e81-4ed7-9655-97be86f4b61d · outbound

This paper cites R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.038204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.038204Z digest=sha256:17e5f25553c613fb1e177f268a8d47a4cba74331cd8283666eb2fe665be94e4c

Observation 9420448c-6ad6-4882-a529-04e6680b56b1 · outbound

This paper cites International Conference on Learning Representations , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL International Conference on Learning Representations , volume=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.043466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.043466Z digest=sha256:0434e010f498b93d61f71c02d5f6a92948eff5e971f8c4f4a1f86e570d6d54d8

Observation e74c68b8-fb29-4fe6-8ef7-7daac7c193c5 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Qwen2.5-Coder Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.048732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.048732Z digest=sha256:e6da97c3dae1d191dacffd9a1c049053f127e0a7fec1ff0563fb3cd17938b314

Observation 3dde4b1b-2356-4576-b3ab-1c4318d0a4e3 · outbound

This paper cites 2024 , eprint=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL 2024 , eprint=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.054044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.054044Z digest=sha256:6f9aa8ab4018291ba98fd8b565d00349dc02fbebd2f7d20c935e295a2684302b

Observation bd4a0d4e-9a80-47d5-8b47-16f9224ae776 · outbound

This paper cites Qwen3 Technical Report.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Qwen3 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.060253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.060253Z digest=sha256:ee06ef2c5bbb1209138099e3b4b7e6bf354eaf670287ee45d4159f304973cdd2

Observation 275992c1-7290-4e6d-b0ca-31e35f9d3b1c · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in Neural Information Processing Systems , volume=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.065858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.065858Z digest=sha256:95b69859b7172f4196a19001174e26369ccdcadfbc44eb5aa32e7985874fa34e

Observation c977a3cf-2263-4106-a43a-a5e1645285b2 · outbound

This paper cites European conference on machine learning , pages=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL European conference on machine learning , pages=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.070415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.070415Z digest=sha256:5045eef72785f1576df5fcb009c7dab52841a3a0042a98b27c9a1c9819701608

Observation 09524048-5675-44be-8290-323d18595cf0 · outbound

This paper cites 2016 , eprint=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL 2016 , eprint=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.075513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.075513Z digest=sha256:a12cbda97001b9bf4d8c7d9e167305114e95ad551d4fc6ab666061a6d6c5954d

Observation 8d089323-bf1f-4a43-af2c-bc1f7304286b · outbound

This paper cites 2018 , eprint=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL 2018 , eprint=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.080472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.080472Z digest=sha256:376b8657dedd7b502b0af6214af98f7847bcd8ed6f827a2ad9163e7686a3be6d

Observation 10d9657f-9991-4954-9eb6-098bd0580983 · outbound

This paper cites 2026 , eprint=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL 2026 , eprint=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.086794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.086794Z digest=sha256:8933893f78610e200516583fceee325e70b56a26f0fc1eae99464c78111a618c

Observation 7381c2bd-ed72-49b0-a58e-91b207009301 · outbound

This paper cites arXiv preprint arXiv:2510.24302 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2510.24302 , year=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.092156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.092156Z digest=sha256:f07761781941cfd555e6971b4baecfa4baf227025143da076eb59e5fe8012c4f

Observation 2a0dc8ef-ad46-42da-8a6a-bbb4f8074e52 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL ReAct: Synergizing Reasoning and Acting in Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.096738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.096738Z digest=sha256:14bbfbb927b609e3e32715522478ba9ad7e140beac2ceb22b89acba95eac6341

Observation 256e16d3-4bb1-4f15-94ee-d41e4d810f61 · outbound

This paper cites Advances in neural information processing systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in neural information processing systems , volume=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.101875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.101875Z digest=sha256:8bcf405de50d710e9ddbdbf87e035ff5ae38fd8566c7c9ceb178407520b59da0

Observation ea78f280-d06d-45e1-8314-7e893df889d2 · outbound

This paper cites Advances in neural information processing systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in neural information processing systems , volume=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.106487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.106487Z digest=sha256:57c9261f55bb2f65ded3e093e4af70a60fd19b3089c3622c8439b1df3e440d44

Observation f6dd1b67-63af-4ea8-b0ae-c701f91f2835 · outbound

This paper cites Findings of the Association for Computational Linguistics: EACL 2026 , pages=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Findings of the Association for Computational Linguistics: EACL 2026 , pages=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.111074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.111074Z digest=sha256:22cedaa2392dd55fde319b6e79aad7d73b5e847518131e5caab6f9e408a2dd4d

Observation 9ff0fa1f-36fc-4df7-9af0-d5877a854285 · outbound

This paper cites SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.115809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.115809Z digest=sha256:cc35a874a36648041ed71704366bad348507c188393a7591bcca28226f31d091

Pith citing papers

No inbound Pith citation observations are available.