Pith. sign in

Paper Citation Record · LEDGER

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

As of 21 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 17 inbound Pith citation observations for arXiv:2506.11902.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11902 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:08:25.628815Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:07:48.796217Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9000ba0-c45d-4c5e-bc4e-749e5da8624e · outbound

This paper cites GPT-4 Technical Report.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.099844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.099844Z digest=sha256:f2b4d207b4bbf2560ebde16178c57702873a2d85aa2c0aa009096eed90f84bae

Observation 35c08fa4-6399-477d-bc3e-b7597da039ab · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:28.054101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T01:08:20.171800Z digest=sha256:211b16372514cdcf314861846b4187474ef6396a9a85b31f10987d5d9142aa40

Observation 548c0e0d-4b53-40d9-8dc8-97d8b7c4fd0f · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.268467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.268467Z digest=sha256:2602d448ec6c3adb3f4c64c322bcddc69ae6c9f7743ec2655e043cdb38122ba8

Observation 37859dfb-bf4d-40f2-b6a5-eede338bbcf6 · outbound

This paper cites AlphaMath Almost Zero: Process Supervision without Process.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search AlphaMath Almost Zero: Process Supervision without Process

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.439050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.439050Z digest=sha256:c06dcaf577a1ad54293ca9d79322228eb154782004e4603dbcbd5e9ff166fb7a

Observation 0474d1d2-34ee-482b-892b-0c3854886c24 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:27.863940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T01:08:20.554819Z digest=sha256:913f7669546d85f6955842dfa0e340866e25871aae2591365d7eb9404694df97

Observation d0808632-c636-405e-bf7f-551e4a35ad44 · outbound

This paper cites The Llama 3 Herd of Models.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.672429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.672429Z digest=sha256:ea8ef12986794214efab99f82ec7139bfa9ade797dcf0b6c71d8bc86f99b16c8

Observation c54a79fa-7e5e-42a8-85a3-f7797b28d389 · outbound

This paper cites Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.750335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.750335Z digest=sha256:f395801aeea75c70b629faebbde815b6c0a119e86f979c95e3e9d7cb4dfc8e9a

Observation d7d0511f-bc7d-4b3f-9910-c310f4ad58b5 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.862537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.862537Z digest=sha256:4c8f030bc479998bb61d437f5021c59a879558a49a6ed48aa894308635b077ba

Observation 108809f1-b85c-46a5-b31e-70b41564c3ba · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.018108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.018108Z digest=sha256:7914bc497639e7008ffc15f1fb22a4b070cf9514afe1537e8884f5061f6ab495

Observation db28df58-6259-4ddf-81cc-9f7145447025 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.132363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.132363Z digest=sha256:7ce6cfbb242876a5d820255fc08fe572472c033bcdb98f2a17532918db5819b1

Observation 59be97ce-94dc-4111-8a38-12fe2946a9f4 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.246828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.246828Z digest=sha256:dfb5687ea2351b3ba4f0e01528f585039bc2ea060458cd64333d4e2e14480a8b

Observation e23d2773-23fa-4156-87aa-225ec2af809b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Measuring Massive Multitask Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.368779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.368779Z digest=sha256:a34e8efc5133b146ca983abdd95fc9b36dd1cc7d4e968c84938ac0e65d1e4c3e

Observation 2a69b891-6729-46bb-b446-d107a225d97c · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.489039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.489039Z digest=sha256:fe0ba4cc4fc5ecc488700806d176bbe345335766bc9a5ce146d29d3af66e1331

Observation 3abbd6aa-3f04-4587-8587-14516d4153ce · outbound

This paper cites T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.609464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.609464Z digest=sha256:08bd45a4e20494dbcba87c931263be522b14bbacd0d383ca3350322b17fc9f40

Observation b173984c-824c-4026-bc23-9654b25d5a29 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.704002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.704002Z digest=sha256:1a319dc89034200a696ebc2e0383688b97ee88b017c8e493c8416fabfa4325f7

Observation ebe55ede-7678-4f2c-8c1e-c65490ac9874 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.838257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.838257Z digest=sha256:80c88d3fb744ca8786e0e8436690eb8e947840d141cd4bf5c73fe6ca5ff3694e

Observation 2a630102-fced-4a9e-8dd4-3ff21196ca73 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Gonzalez, Hao Zhang, and Ion Stoica

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.937967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.937967Z digest=sha256:4eb04b916d69969d3b2e2813dbb5360516d91a3c8af7514b6e8e38269d02605b

Observation 4e3ceeda-aeb0-40d4-8a56-5a201fa75cd1 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:27.647972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T01:08:22.107828Z digest=sha256:6146f72eb8a06d25baffe52632eae2ef0ad32343f810908f235b8a2ee1e7f420

Observation 5f105cbc-8e7c-4f4e-8a0e-b94a0bd87094 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:22.249034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:22.249034Z digest=sha256:c6797c9849732eba86d86856977b5400599b7b5525517ba67359bb5a1dd4d72e

Observation bd96f29a-1019-4ae5-b74b-1fc215bc35c9 · outbound

This paper cites Let's Verify Step by Step.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Let's Verify Step by Step

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:22.406148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:22.406148Z digest=sha256:6e78ee2c973605da68999d36d871b13704dbe50c026b3da9c1fa19032e68589a

Observation b3a8ffeb-dcb2-4621-ab00-4827b7660a87 · outbound

This paper cites Let's verify step by step.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Let's verify step by step

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:08:27.514928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T01:08:22.548614Z digest=sha256:92e836ba93851dc9be1f86803d9b4d95e9f5e55866c37f170dcbd3abdd1f9015

Observation 17994dfb-b49c-4538-9ba7-4905d912477d · outbound

This paper cites Large Language Model Guided Tree-of-Thought.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Large Language Model Guided Tree-of-Thought

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:22.626559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:22.626559Z digest=sha256:68bbe8e932f0181630657a7c17f1c1e3eedf26a5b6d23766e00a074ac5bbc939

Observation 5acb4d0d-a41c-4560-b514-bfffe22778e6 · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search StarCoder 2 and The Stack v2: The Next Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:22.745145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:22.745145Z digest=sha256:b5d5b5bdcfcc47abbd1a47776f0554dac049f53e0f6cd8f9361cc83dbc37e9c5

Observation 0a1ef31d-1ec5-4bb9-be91-38cada151174 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:22.851720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:22.851720Z digest=sha256:26b094f6ac1e19a2dddc65975f503bc35314bb9bdad12c0e9e7851e8afcdecec

Observation 360f76b9-4154-482c-8917-9fd6eac0f072 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:27.334489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T01:08:23.007752Z digest=sha256:6dd2858bc069a22e6ad6663372fff9e71d87642c1091770eaa4ab1bc8514afe1

Observation de9385a3-4caf-4419-9694-4bcc5477f941 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:27.175293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T01:08:23.112005Z digest=sha256:171c616e9cd917c5895c12fc8223ce234a2461e6f115d2439ba1d9298bf27c7c

Observation 01541c63-caf3-4805-886b-c1a41b8a9ab9 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:26.901684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T01:08:23.214596Z digest=sha256:1fc1a6cefac87a911858e7dc7b2c219f091f9521b5d8aba3003599fdb6b3ada5

Observation 68c28502-06e4-489e-a638-0b5bf15d708a · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:26.772305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T01:08:23.430414Z digest=sha256:138842872e28e9875cef69d1b4bbefa51e6df23b4ed957e81b9588fff8b453e3

Observation e8512c18-f4aa-4c11-a56c-5ce891616b30 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:23.607958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:23.607958Z digest=sha256:b385eca0a8de5bcd63a77eba392e4628df12c1d4cc9f5f0c81d09cc544cf2e48

Observation ba6f07d1-6cbf-4aca-af99-7d7ed57b24b3 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:26.643901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T01:08:23.744564Z digest=sha256:f20b10bac17739966632b75973aa7bea06f4aa8b220f75bdc487ed0d31db0888

Observation b2ed575e-2c68-4be3-bc27-733b2bc5dc5a · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:23.919121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:23.919121Z digest=sha256:fa276f5a47a83966f7541b4f22d725c966fc8d587dc6df16e02bbb372ae9052d

Observation 0d289415-6c0c-42d4-8ffd-a3037d2b62e4 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.026164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.026164Z digest=sha256:2bc2670fb51a16467080b2e63370aeaad6b49a4d2b0dc53c44b95c26ddf4b30d

Observation 6a7f969b-4cb2-4938-ad60-c9dd3bfae858 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.164924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.164924Z digest=sha256:bab988faaa5b9fff04ab523a1a65171cd950dbdc281511f66221ac8bf2c61997

Observation c8f2eee2-1bb1-44aa-bead-6ad9de2b6459 · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.323361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.323361Z digest=sha256:60252759a0bc1b8dce3c634792f4b049410a644887d11cdb714eb35c865ef498

Observation c42a747e-1287-4068-9841-9947be8584b2 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:26.497716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T01:08:24.412522Z digest=sha256:5d75e4a0122ef90462fdb6dde745618ef3bd967a845b830cea587e19091cac93

Observation cd4fc2bf-5f38-4144-b847-aa1bc12bed1f · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.493167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.493167Z digest=sha256:bf70bdfa3f27eecb8f06f69ec06a662d26e4384bf02b90ccf4a3bd5ee2129c7b

Observation 6ecd0436-4be8-4c61-a716-91e27d955e8d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Gemini: A Family of Highly Capable Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.572387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.572387Z digest=sha256:d6f5a6624fb8b72d3f3046c8f8fef6e50314731b3d68b8a2bb36a12b29b3c5c4

Observation b6ef6b9a-dc9e-4d59-87dd-d15515afa047 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.708005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.708005Z digest=sha256:579c0c7c963362765afb7932d29f950e5cba81e45ce3414d935e89d03a27562f

Observation 536beee4-e060-470b-a045-4f6a9fec0f9d · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.775505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.775505Z digest=sha256:1c92799d586700f4a6a47b8a5606d00d7c5adde88ed11c0e79124e2289791ef3

Observation a76676cf-d4fe-4126-96ea-2073f65e973b · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:26.349520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T01:08:24.889023Z digest=sha256:232b0145ecbd83c7cfa759c44467d4e77443521d8d347f09e6f7fa19ee4cdf63

Observation f6e930f5-f5ae-4d25-b27e-4f1f5de066af · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.959999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.959999Z digest=sha256:8c2a2895d35b5a4cb0c250b48d8641996681e17448150f83a70d0469f76944fa

Observation 23e05b3c-5c67-4de9-af7d-a2f197b49a0f · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.027204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.027204Z digest=sha256:0525fbdebfbc2e33fb9c2bb97828dbd6b239b8dcaa6f79390a80954a1ce7d251

Observation fb7408bf-4e5e-4c70-91be-f632880a6d68 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.129553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.129553Z digest=sha256:0db9286b0bb43ea2618461a81d779f801c149cb95d6e4a72c728f33efaab698f

Observation 835d2202-c34e-47ce-8dd7-cbb7280c706c · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.265647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.265647Z digest=sha256:784323c58fe76bb25f93630f11333db1aa8985dfdd9da8e350343c0087454c70

Observation 78f5cf1a-7238-481a-aaac-6960c8b4c934 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Instruction-Following Evaluation for Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.355982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.355982Z digest=sha256:4ff2547a2dc6b709f5023a98e869034063c66e80161ea18daaa0511cf272a35b

Observation 6eb98aae-9e09-4c8b-a53c-e63acfe06a87 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.428290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.428290Z digest=sha256:ace67a4c3faf568a81e319965c998d2afe72989ef34b32203115186c0ddf7705

Observation 377648b2-6d7e-47a0-a253-9d73c9182140 · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.510840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.510840Z digest=sha256:332d743e2e6c6cabcb565d754de6b4058487e96221b275485790c7e35663ef36

Observation ee4802b8-704c-4ce6-ba17-55dc5ff8dbb2 · outbound

This paper cites online" 'onlinestring :=.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search online" 'onlinestring :=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.577547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.577547Z digest=sha256:82b9d0f4f18d17d8edca466e11de3a4265a33a86179636f2ed7f6b0ef3eef1ba

Observation b5960a35-e0d5-46f2-9186-5756badda292 · outbound

This paper cites write newline.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search write newline

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.628815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.628815Z digest=sha256:f18a2fa6bacb13f2a8e042085a3d12511937adcb0bde7380730e36f7471181d6

Pith citing papers

Observation d4741241-e8f2-4fa6-b6c2-25f2c4592963 · inbound

ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism cites this paper.

ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:01.975680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:01.975680Z digest=sha256:e02ae56822f1e17664a45b9bacc03b4988963d25d1bcbf2102b0f595fbc9b459

Observation 8bcab73e-01f1-4882-91c9-cc739c0130bd · inbound

TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling cites this paper.

TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:48.796217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:48.796217Z digest=sha256:21bd56b0ce8756c22dcfbb3bf27420ae3b9c9b37fececc0419ebd2f7a9532afd

Observation 41286303-b7db-49f1-89ba-016153a2aaca · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 198

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:05:31.659995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:440d4f70cb0233ce0a2dbc9c97663998fcf8bd889ff601f2f1ef958f9aae2f14

Observation 55afe15c-6588-48c6-871e-1df5cc9c14b4 · inbound

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation cites this paper.

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T11:11:11.329062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:11:11.329062Z digest=sha256:68128d14e2cb205d390edfae45acf70f67d7cfc20135915290ff4b97a7ca7a03

Observation bc205017-7caf-4ec6-b1fc-0c19b7228b78 · inbound

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse cites this paper.

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:05:39.076738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T02:04:46.559881Z digest=sha256:3039a94ef431387a5aa74376646dbe60d93ed8dbb17b2cb09f4d296224daa0d1

Observation 223c20ca-91da-4c58-9739-b077b2209246 · inbound

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning cites this paper.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:27.901205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:27.901205Z digest=sha256:906c839d6cc445b8329faa210a0b36aa4ce62da32452a31da63b806edee7ed48

Observation f59c23a7-3b89-4b5d-97a9-95dbc550f765 · inbound

Your Model Diversity, Not Method, Determines Reasoning Strategy cites this paper.

Your Model Diversity, Not Method, Determines Reasoning Strategy TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:56:02.900863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T15:16:30.448400Z digest=sha256:982f6bffa828885aa0c5f3a36de3eca6628985788f93fcc1aa443636b41a54dd

Observation 871d0549-81eb-41e7-bc1d-f3d15e446620 · inbound

Mind DeepResearch Technical Report cites this paper.

Mind DeepResearch Technical Report TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:50:20.510822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T11:46:49.178896Z digest=sha256:eeb9f063b350cda2e65f263216928467185c20e6a299ec5e3668cebb04e38956

Observation d84b49d2-2575-42a7-b3b7-89b17e034bd9 · inbound

A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping cites this paper.

A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:56:08.732519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T10:41:46.675257Z digest=sha256:943d22f49a2056f4b5e333b06275533e9c47f1ff7ca55b1ef3653d2c58441510

Observation 9dae7e03-dc10-483b-afb3-bfb4057ee305 · inbound

Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias cites this paper.

Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:41:24.001119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T00:51:58.315109Z digest=sha256:ad90425c2415b36fa8c9ef405aa1c67ec62140e664e68640dba6e74847276a93

Observation d66b4fc2-8352-4580-91ae-06aaccc2e7c2 · inbound

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling cites this paper.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.272895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:966c21216fe1481b5238e6148ee8dc2e22cae169eddd070faa3e127db6b355ed

Observation 6bdbc50d-42f7-4a2e-9bfc-6124572ea550 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 130

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.465370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:7a1338ddade0d218adcd18c4f8dccd376c3291fa02b7fa0aecee0acc5eb2bb83

Observation 42992776-2f8f-409a-aa5d-dd9f01562b47 · inbound

Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs cites this paper.

Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:46:33.124385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T09:40:35.258818Z digest=sha256:5f0943f11bc84c28e27b08dc01ce3bb007c743a426464c017e8a716d8a913a39

Observation 54e78ec6-2080-409f-a345-eaf32746f106 · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 128

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:56.077665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:fa7c33c0f9274e61891133dda0e31758605856bde91d099dd30abe9ecbcde4ba

Observation 0e55e482-269d-4501-83d7-1fdd41060ce4 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.748310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:a2e7ba7dd2e3617bf01f6008877eb14f763c7bc935d5dc4fc33de810f44be3b9

Observation 56d2bf61-7fb1-45cb-9a2a-53bbd9ab9731 · inbound

Optimizing CUDA like a Human: Micro-Profiling Tools as Expert Surrogates for LLM-Based GPU Kernel Optimization cites this paper.

Optimizing CUDA like a Human: Micro-Profiling Tools as Expert Surrogates for LLM-Based GPU Kernel Optimization TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:59:56.473477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T01:12:04.486570Z digest=sha256:9269563399e7165197bf564df70f53f97a9766ea5ac96130e6f60036a6cd070c

Observation e6123da0-4012-48b2-8d19-fb1ac55d7222 · inbound

CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents cites this paper.

CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T14:03:48.459093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:352a7e2f86d17c176d3acfcf512b111d4980eaf2c6419ba36a7fc6d5556277f8