Pith. sign in

Paper Citation Record · LEDGER

Process Reward Informed Tree Rollout for Effective Multi-Turn RL

As of 8 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2607.15610.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15610 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T22:51:02.115809Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc654b2b-d399-4ac3-bb76-dc3e5a7b49ff · outbound

This paper cites Agentic Reinforced Policy Optimization.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Agentic Reinforced Policy Optimization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.855827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.855827Z digest=sha256:c9e3aae82ce76722f211b3f73f3bb27533eb898d984d280d735179dde1cd93fb

Observation 1af78faa-aa71-4e41-a3a5-aa37f897a72f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.862640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.862640Z digest=sha256:c5ddbc2a4016d3bc997d98121b636f12a722b8153a61f6afe254df0a0817e2a3

Observation d266afe9-452f-42fb-9f2e-54c0854fbfac · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.867466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.867466Z digest=sha256:5b65906c73d0abf9a8ee84bea3f79a140659464b554c7e12f185d2e586cb4105

Observation ceda4a33-923c-4f12-abc3-ca1097d64285 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in Neural Information Processing Systems , volume=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.872812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.872812Z digest=sha256:d28e606f7dafa79442b9e2abc6d82f8fb95fd9e01c828be12c8847fe0f5a21e6

Observation 11173a2b-ceb1-4a1b-944e-6b0beb524a81 · outbound

This paper cites arXiv preprint arXiv:2602.22817 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2602.22817 , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.878075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.878075Z digest=sha256:30ffb2b7d99ea874860ea80910f7ce427daaac98a22ce67ccfb1ba54cb565161

Observation 64c0fd4b-8384-4fa9-8a1b-84f79f30dfb6 · outbound

This paper cites arXiv preprint arXiv:2509.06040 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2509.06040 , year=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.882565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.882565Z digest=sha256:c99c284411f31c7138426214966a3653aedc616030fb3bc471469dee8bb840d5

Observation b8af38c2-5456-4162-b519-7ed240284fe8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in Neural Information Processing Systems , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.887952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.887952Z digest=sha256:c308d7f48aca6fae693f5896bc1cd0a2088343a4009b21cdbeb88fb9c1da7401

Observation d980537e-6949-4b51-a9d2-f9032e0bee0b · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in Neural Information Processing Systems , volume=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.892992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.892992Z digest=sha256:5c713ccb7173b2c9c4d10095dd029564d1918b6cf2bfa09715b731f024cb716b

Observation c17c63ae-e387-4181-9083-421ea1a67109 · outbound

This paper cites arXiv preprint arXiv:2603.03078 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2603.03078 , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.897762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.897762Z digest=sha256:86c8f0ecdc5607014faf4912cf9207743ea357d4bad87a3928d3cfe0f0cfef66

Observation a1c6a846-5789-41fe-b98c-d59c96edafdf · outbound

This paper cites arXiv preprint arXiv:2511.16108 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2511.16108 , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.902422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.902422Z digest=sha256:033ad91ae55afef804e4bae4e7a160483de999c4d70bc479e3b6b33c7d113ec1

Observation bf37dc69-9ff4-49e7-966e-590967662b13 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.907282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.907282Z digest=sha256:44f32c69fa74b75694d02f699d1561bf8d58b91e5aec7a402eb107b336ec732e

Observation c3369599-5db8-429f-bf20-1b03f9917d68 · outbound

This paper cites Advances in neural information processing systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in neural information processing systems , volume=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.912213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.912213Z digest=sha256:0b2cdacc6c943902484a7c828ee089ca8bdda07a6daf8711d26fb2245ccad7ac

Observation 3a644d5f-e4da-4824-9785-b8729c3db3eb · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.916601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.916601Z digest=sha256:0dc6651d654ee0f1161f64fa9fba5227ada04b32b07ab9ce6be1f78c5820112c

Observation 4a037f56-8f47-4f3f-ad54-cee3b2c3fba0 · outbound

This paper cites arXiv preprint arXiv:2407.01476 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2407.01476 , year=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.921572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.921572Z digest=sha256:73b2d638d3546c88d1b5a446606a01d038c244465d5d5f63275d821d1fc9baf3

Observation 9cb62683-61db-48f2-b63c-25c1b718a0cd · outbound

This paper cites TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.926088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.926088Z digest=sha256:8e3155edf6cb064b77eae9b9a7e9f142d116cf69cc308abff24bde158c107325

Observation 76d710ab-ab6d-47ec-90a7-b16436448d20 · outbound

This paper cites arXiv preprint arXiv:2603.02216 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2603.02216 , year=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.931380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.931380Z digest=sha256:22f81b458990137546610d8d8d04b40dc9afe03c310f66ff998c9ca33a1e37e3

Observation 0595bac0-09dc-473d-97ec-77345758cebb · outbound

This paper cites arXiv preprint arXiv:2510.14545 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2510.14545 , year=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.936445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.936445Z digest=sha256:81437c6375f6afa4a947161b34ba9908fc652ae02552a9b33982e04145223caa

Observation 31f50e61-ee74-448b-b5e9-749f725fe73a · outbound

This paper cites an unresolved cited work.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.942075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.942075Z digest=sha256:a499c9c2ed93d555a4b6e463b227e0a513ae93bf6e0a3787ec551de070517a56

Observation fc1243b5-7dda-455d-a4a6-c4b49a9be0ef · outbound

This paper cites TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.946686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.946686Z digest=sha256:e41eb69fb3e57a908235f33d312f1f493dbf05d322568f5168791e212cf97206

Observation 053ca571-7a11-480d-a77e-2b38be2ce1ee · outbound

This paper cites arXiv preprint arXiv:2512.08153 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2512.08153 , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.952267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.952267Z digest=sha256:412bc5f41441f6c1dd31caa80ac00ff854a4aea134443a25026de6a013b34b5c

Observation 69dbb63a-ebc6-4007-9b44-5d513f8f1081 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.957238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.957238Z digest=sha256:a9f1ec4e904a63362e9a4eaed2dcc54df51d9b127e6b31dbc8b9677210863963

Observation 88e9ea15-985d-4c67-a444-d926277ba6d5 · outbound

This paper cites arXiv preprint arXiv:2603.00296 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2603.00296 , year=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.961785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.961785Z digest=sha256:efa7a4b1bd2ff24bd56cef59256c3cfdbe773974515c5cd56da66adb42fba692

Observation ca1c3646-9076-402a-b257-8b84b5901d4b · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.966286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.966286Z digest=sha256:5dc93f517f9be07e0329a6bbe0d3ad9d27709c97ab2b901e59abc6d22476e208

Observation 9687b349-3f62-4dc5-a68f-30c48f4e2a2a · outbound

This paper cites International Conference on Learning Representations , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL International Conference on Learning Representations , volume=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.970877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.970877Z digest=sha256:b2903b3a449b1782dd07d8b250217cbd1d8f374105ad01fbb2caee6c268d496f

Observation 409d0468-5740-4bbc-bd5c-738da42d02ef · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.975760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.975760Z digest=sha256:58b8f188888f09565aed927836e04e555245318ccb1fe91aa00135a9d081f8c8

Observation bd2f6b3c-1eda-4706-b9e2-96700058a548 · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Skywork Open Reasoner 1 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.981544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.981544Z digest=sha256:ffeabb3a38b7980c9427c45b3c89b2ad8ea6de46c80699dc55b7ebc4718d7857

Observation 41784db7-cd49-4d7e-a3a2-c2d3b8546a4d · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Process Reinforcement through Implicit Rewards

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.986493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.986493Z digest=sha256:f521099d76a266eb729f9327b534e882c1bb82cc439474543abc26793c8154b8

Observation fe98da61-d19b-4833-80e0-609915bfd183 · outbound

This paper cites arXiv preprint arXiv:2509.19199 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2509.19199 , year=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.991250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.991250Z digest=sha256:c1c6d48facd153c19e86522d359191b3cf2c696f82cf77c4763017efe94030f4

Observation 238ddf24-c58b-4aeb-9108-74853ddc1fc8 · outbound

This paper cites arXiv preprint arXiv:2602.17025 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2602.17025 , year=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:01.996019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:01.996019Z digest=sha256:1316e8858f3a971a241efba43d41c2c2f321a2f3be53f530e5e0a15ed219b717

Observation a0dbd9bd-7ce5-405e-9bd5-b21bbd184ace · outbound

This paper cites Advances in neural information processing systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in neural information processing systems , volume=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.000591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.000591Z digest=sha256:dbdb59c7da5d4c1b10416819f15d1fe2b0856579aeee4dec87b68b536d90749e

Observation e27a9318-3394-4ff3-b71a-92ec7e67eb52 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Training Verifiers to Solve Math Word Problems

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.005624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.005624Z digest=sha256:0674115d02ceb3cffcc0cd0296816c9372721b3580be9f729ce2cdcaa70df120

Observation c10ddbf0-bbf3-4b6a-9c03-05f6d232a9a9 · outbound

This paper cites International Conference on Learning Representations , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL International Conference on Learning Representations , volume=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.011033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.011033Z digest=sha256:06e2b3c3abd4daa6dacea6f241831883d0d4d72388b5ea53f2fb416cc48a512b

Observation 9adaac36-5103-476c-a731-d4ea3d76b883 · outbound

This paper cites nature , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL nature , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.017289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.017289Z digest=sha256:f307ff8cec105f9c2aedbb9798893294089ae5105c2e4ccc531f5d2641397172

Observation bd08aaba-361d-4a92-8b93-176bc1499bb7 · outbound

This paper cites From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.022603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.022603Z digest=sha256:a6a0b3d0963d424a45ea19235d7d85924c0fb2893bc612ead10a17603a1f474a

Observation ea0a0d05-9fda-4bad-b6ae-6be651077ec0 · outbound

This paper cites OpenAI Gym.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL OpenAI Gym

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.027774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.027774Z digest=sha256:90b0adbc8f25a9d504c93756621a320af846258de41ca576059930e756ed7557

Observation abd75e59-6196-4652-8578-6a365dae4207 · outbound

This paper cites an unresolved cited work.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.033389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.033389Z digest=sha256:cf4ebca17f5af74b13bdb2576110c4966b9a8104a2051150ebf31f218f246a67

Observation 4af383bc-1e81-4ed7-9655-97be86f4b61d · outbound

This paper cites R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.038204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.038204Z digest=sha256:252ce429cbc7f1d0cd1aee39b2ea680bf986d7dfa591b81c69000005f60aa037

Observation 9420448c-6ad6-4882-a529-04e6680b56b1 · outbound

This paper cites International Conference on Learning Representations , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL International Conference on Learning Representations , volume=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.043466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.043466Z digest=sha256:b2e1998568ea2040f3f63f8d0d778ec8375be988d9dbd06cd982311bc53541b5

Observation e74c68b8-fb29-4fe6-8ef7-7daac7c193c5 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Qwen2.5-Coder Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.048732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.048732Z digest=sha256:56f00e4a18107e6133ce1c42e9e7a39440a70e2f17682b2c6e73828cca5ea842

Observation 3dde4b1b-2356-4576-b3ab-1c4318d0a4e3 · outbound

This paper cites 2024 , eprint=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL 2024 , eprint=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.054044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.054044Z digest=sha256:d1380496881a1b275579eeabfb481901c57f0b0ac7e2326c9ea384ed38cb8259

Observation bd4a0d4e-9a80-47d5-8b47-16f9224ae776 · outbound

This paper cites Qwen3 Technical Report.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Qwen3 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.060253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.060253Z digest=sha256:674c06637cdb631ada6c3ecbbb87fd5c6bf779be09f32ac170474fcb1fec2a20

Observation 275992c1-7290-4e6d-b0ca-31e35f9d3b1c · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in Neural Information Processing Systems , volume=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.065858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.065858Z digest=sha256:1cebe1c43d335d6ab9dfcc3e7fc3838f77bcbab800d8530075846df8dac92149

Observation c977a3cf-2263-4106-a43a-a5e1645285b2 · outbound

This paper cites European conference on machine learning , pages=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL European conference on machine learning , pages=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.070415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.070415Z digest=sha256:61b346f81de7a5fa803ca521e6e1a7d45eba0d0f705a9de447a9781a641c6833

Observation 09524048-5675-44be-8290-323d18595cf0 · outbound

This paper cites 2016 , eprint=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL 2016 , eprint=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.075513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.075513Z digest=sha256:6511251665c76ec8a34c4e15508a3dd53f89e5203c92926571cdc5ce53a0b298

Observation 8d089323-bf1f-4a43-af2c-bc1f7304286b · outbound

This paper cites 2018 , eprint=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL 2018 , eprint=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.080472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.080472Z digest=sha256:cbbc7521ef24a7f740f879c1592165951c77ec0b484a76250e33213cb82eac68

Observation 10d9657f-9991-4954-9eb6-098bd0580983 · outbound

This paper cites 2026 , eprint=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL 2026 , eprint=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.086794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.086794Z digest=sha256:25e3124110542d0a6c76e957c488cf452a6f6b38fd03c5d35fa0b3c07c18e039

Observation 7381c2bd-ed72-49b0-a58e-91b207009301 · outbound

This paper cites arXiv preprint arXiv:2510.24302 , year=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL arXiv preprint arXiv:2510.24302 , year=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.092156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.092156Z digest=sha256:ce6d485247ccbcc0dcc5103376843cc5d9e13d7fe9184a14ad882fa13195ed7a

Observation 2a0dc8ef-ad46-42da-8a6a-bbb4f8074e52 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL ReAct: Synergizing Reasoning and Acting in Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.096738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.096738Z digest=sha256:f628945f4000b726da355a04cf114ff2e66e032cca55d84fbad735f1ad935f0d

Observation 256e16d3-4bb1-4f15-94ee-d41e4d810f61 · outbound

This paper cites Advances in neural information processing systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in neural information processing systems , volume=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.101875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.101875Z digest=sha256:a4738b211d66218012cc4826945805dbeefadf1a96124d7fbd666cee7ff59f29

Observation ea78f280-d06d-45e1-8314-7e893df889d2 · outbound

This paper cites Advances in neural information processing systems , volume=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Advances in neural information processing systems , volume=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.106487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.106487Z digest=sha256:2a03ed0187e7fe494531c42fb5b178cc2ac2a4a188a8a48c8c13d9468049a0de

Observation f6dd1b67-63af-4ea8-b0ae-c701f91f2835 · outbound

This paper cites Findings of the Association for Computational Linguistics: EACL 2026 , pages=.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL Findings of the Association for Computational Linguistics: EACL 2026 , pages=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.111074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.111074Z digest=sha256:664e446f1beacaced9226f89d41cf5d04ce77592f87c9360292cb2a922ad9442

Observation 9ff0fa1f-36fc-4df7-9af0-d5877a854285 · outbound

This paper cites SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T22:51:02.115809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:51:02.115809Z digest=sha256:d9956fe0348d2a14c5a5bd024e4ae497170667b9436d41a1f6708f5b6d02eb50

Pith citing papers

No inbound Pith citation observations are available.