Pith. sign in

Paper Citation Record · LEDGER

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 16 inbound Pith citation observations for arXiv:2506.11902.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11902 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:08:25.628815Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:04:01.975680Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9000ba0-c45d-4c5e-bc4e-749e5da8624e · outbound

This paper cites GPT-4 Technical Report.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.099844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.099844Z digest=sha256:6f7fedf47f6713faceeddc5dd9bf91f9abe9370694e0ccac3dda17765cbfeba9

Observation 35c08fa4-6399-477d-bc3e-b7597da039ab · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:28.054101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T01:08:20.171800Z digest=sha256:2123f048cb50700742f3f50d35534297d6ae4cc5a48e761159c3a5879fa0526e

Observation 548c0e0d-4b53-40d9-8dc8-97d8b7c4fd0f · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.268467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.268467Z digest=sha256:cab4a604de7c113de30fcc3c6d8184d8c63cf57384078128941c5d83d4cf85d1

Observation 37859dfb-bf4d-40f2-b6a5-eede338bbcf6 · outbound

This paper cites AlphaMath Almost Zero: Process Supervision without Process.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search AlphaMath Almost Zero: Process Supervision without Process

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.439050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.439050Z digest=sha256:1a663dc68666c420b469fac7c9f1ed9234fad9c84d76b1b7cc5af1071b1b22e5

Observation 0474d1d2-34ee-482b-892b-0c3854886c24 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:27.863940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T01:08:20.554819Z digest=sha256:660e8fe01da0ceb3d9400363a08139fedde565f1c74dc510cd60f602900822fb

Observation d0808632-c636-405e-bf7f-551e4a35ad44 · outbound

This paper cites The Llama 3 Herd of Models.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.672429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.672429Z digest=sha256:766cc6dd9d32c22b3ff1eafc642347ee2fa62e34ebf37d741e53a7136c3fb59d

Observation c54a79fa-7e5e-42a8-85a3-f7797b28d389 · outbound

This paper cites Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.750335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.750335Z digest=sha256:bd8140e026c5ffd6f8e577089cc459016194871a7e88dcb398c101666d0674dc

Observation d7d0511f-bc7d-4b3f-9910-c310f4ad58b5 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.862537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.862537Z digest=sha256:9f8b7f7daba017d5312951e6c8f769c37736eb64a91afe0e629b2855da164fed

Observation 108809f1-b85c-46a5-b31e-70b41564c3ba · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.018108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.018108Z digest=sha256:848ce5ddcc44743bc330faed468b147cc56d8af4110fd0e0c150341b059c599d

Observation db28df58-6259-4ddf-81cc-9f7145447025 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.132363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.132363Z digest=sha256:c1e4bb6871c6a7c47f6b32b1b10a638825644165f07cf65ca7066d931e90cc0c

Observation 59be97ce-94dc-4111-8a38-12fe2946a9f4 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.246828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.246828Z digest=sha256:3b062103b35cc58a52a994e9f0836d32b76460a20785fc6845820c64c9ab4536

Observation e23d2773-23fa-4156-87aa-225ec2af809b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Measuring Massive Multitask Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.368779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.368779Z digest=sha256:e46b8a4f45e3765beea47458b09feab070414185d03978f3d9b23927160b6d39

Observation 2a69b891-6729-46bb-b446-d107a225d97c · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.489039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.489039Z digest=sha256:cf667bcdc87c7103830589474fbe66d2c16b8cbdd72cdb41a3cbcf5683bccf84

Observation 3abbd6aa-3f04-4587-8587-14516d4153ce · outbound

This paper cites T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.609464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.609464Z digest=sha256:0abcff25b04b3beed4384c632bf856e26930393990527af28656626456f42930

Observation b173984c-824c-4026-bc23-9654b25d5a29 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.704002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.704002Z digest=sha256:846a635e4e9afb7e2ea8de6e55a7d0b8621852ed86ffb0da51d274d89b6b8cd2

Observation ebe55ede-7678-4f2c-8c1e-c65490ac9874 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.838257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.838257Z digest=sha256:46430f8b14dff8f6dfff8f0bf5f39ed1b8bc15cd4afdb759a3b9d7966d94041e

Observation 2a630102-fced-4a9e-8dd4-3ff21196ca73 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Gonzalez, Hao Zhang, and Ion Stoica

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.937967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.937967Z digest=sha256:f1c6be2499708d6df0dca28766c1fe59e122c298de6c36e8c781a9ca0732f5d0

Observation 4e3ceeda-aeb0-40d4-8a56-5a201fa75cd1 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:27.647972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T01:08:22.107828Z digest=sha256:acc5a3f828afccc2be3c8987b54390f9b9a4da88ff9c6146aa9ef0bf71208ee5

Observation 5f105cbc-8e7c-4f4e-8a0e-b94a0bd87094 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:22.249034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:22.249034Z digest=sha256:fa519946fd75b0b0d28a19673db14e6c10d2a54b2081a1f453fe349ed1454b6f

Observation bd96f29a-1019-4ae5-b74b-1fc215bc35c9 · outbound

This paper cites Let's Verify Step by Step.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Let's Verify Step by Step

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:22.406148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:22.406148Z digest=sha256:4f612fe3ea07c2d907161d50b34b08dc2b27cac9bdfc7788c0ff17c32c588691

Observation b3a8ffeb-dcb2-4621-ab00-4827b7660a87 · outbound

This paper cites Let's verify step by step.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Let's verify step by step

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:08:27.514928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T01:08:22.548614Z digest=sha256:5bacefb3359c48092ee53a511189ca33fbb0e600f6700bef76dca0d585a08993

Observation 17994dfb-b49c-4538-9ba7-4905d912477d · outbound

This paper cites Large Language Model Guided Tree-of-Thought.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Large Language Model Guided Tree-of-Thought

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:22.626559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:22.626559Z digest=sha256:382aece7b00c33179c16b719c06861525622acacc3c1616b39bd580882099e20

Observation 5acb4d0d-a41c-4560-b514-bfffe22778e6 · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search StarCoder 2 and The Stack v2: The Next Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:22.745145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:22.745145Z digest=sha256:b366167a9feb1863919d3dcf2d1212e6259e0ba6686cdbceb8c90807b07cab5e

Observation 0a1ef31d-1ec5-4bb9-be91-38cada151174 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:22.851720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:22.851720Z digest=sha256:815c6dcb97d18f1e27f450bb6c7ebec21b808ac9d6aec82b7558df8e128357ab

Observation 360f76b9-4154-482c-8917-9fd6eac0f072 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:27.334489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T01:08:23.007752Z digest=sha256:809dafc44c9a90fa43533c29485d59663f1f1cfccdab4107cdbf0073f6880793

Observation de9385a3-4caf-4419-9694-4bcc5477f941 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:27.175293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T01:08:23.112005Z digest=sha256:debbceb35978396bbf4ba283a76c53fdbba77a8ed823454954fc76720a3ee1d6

Observation 01541c63-caf3-4805-886b-c1a41b8a9ab9 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:26.901684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T01:08:23.214596Z digest=sha256:b4bda675b6b1a4da5ee5b99d4983c3676b591babd72883f05d796517e2cf82e6

Observation 68c28502-06e4-489e-a638-0b5bf15d708a · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:26.772305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T01:08:23.430414Z digest=sha256:e46d93021f83b0b051a1d7bb9bb6070c8085e768c73b0e22aa9c416c249b72b4

Observation e8512c18-f4aa-4c11-a56c-5ce891616b30 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:23.607958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:23.607958Z digest=sha256:c327c87c801458aa3533ce51ec10757343a339064264d5e39dd4207739415bb9

Observation ba6f07d1-6cbf-4aca-af99-7d7ed57b24b3 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:26.643901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T01:08:23.744564Z digest=sha256:5670151c504cf8f8c9768f93f942b04519582b01554587ec7064bb0db64316c6

Observation b2ed575e-2c68-4be3-bc27-733b2bc5dc5a · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:23.919121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:23.919121Z digest=sha256:ab86ab0103581ac0063cde1a1782a2809b470ceec129f99bcf1013f52c7f5407

Observation 0d289415-6c0c-42d4-8ffd-a3037d2b62e4 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.026164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.026164Z digest=sha256:bb7e606220d5730cfeab69389ecbaa64bdcab74a37e551940d9e63fbc6714e9a

Observation 6a7f969b-4cb2-4938-ad60-c9dd3bfae858 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.164924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.164924Z digest=sha256:5eec04c3d4bbceb7da0c3bcdb02658f8969effd429fb065ab05db21934ff2185

Observation c8f2eee2-1bb1-44aa-bead-6ad9de2b6459 · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.323361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.323361Z digest=sha256:6a7f56bbfd8fc25a39c663bb8f746e2e608789025f6946e43c47b7efa566cd25

Observation c42a747e-1287-4068-9841-9947be8584b2 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:26.497716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T01:08:24.412522Z digest=sha256:60f34c6e0bdb919662013a3a9eac457634d7451951d8d138488369e6ca27ec3e

Observation cd4fc2bf-5f38-4144-b847-aa1bc12bed1f · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.493167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.493167Z digest=sha256:2115a8c82ab145a00bc22001d501f1df9052eda3b26bfc49fdd9ba381633633f

Observation 6ecd0436-4be8-4c61-a716-91e27d955e8d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Gemini: A Family of Highly Capable Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.572387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.572387Z digest=sha256:3bb07aa513c3f7aa341a3858e28dd32f36ef9907ca7367f22a7e587054142b81

Observation b6ef6b9a-dc9e-4d59-87dd-d15515afa047 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.708005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.708005Z digest=sha256:9335ab4194c9b1fa32b73f7a4580b67b4385ac09daa278eea5288ff5f7429072

Observation 536beee4-e060-470b-a045-4f6a9fec0f9d · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.775505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.775505Z digest=sha256:ee6ee658e858bd51319252458a67692a13d9d902d0e4d0c9dc5df7e2b39e4f19

Observation a76676cf-d4fe-4126-96ea-2073f65e973b · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:08:26.349520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T01:08:24.889023Z digest=sha256:2cb2fb1ab2b0f058a693c01a9df3bc5439d19bfda240ee35971faa2e981589ec

Observation f6e930f5-f5ae-4d25-b27e-4f1f5de066af · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.959999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.959999Z digest=sha256:2fe44fa25eea237a533fae59d956dbe601496985380ce81932b8c198b6f90465

Observation 23e05b3c-5c67-4de9-af7d-a2f197b49a0f · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.027204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.027204Z digest=sha256:0cda5988a1a13764e5a0c9ed2eb464b3ad508cd0e86bfc4234dbe105204b5a43

Observation fb7408bf-4e5e-4c70-91be-f632880a6d68 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.129553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.129553Z digest=sha256:68a3f350981cc9eb42140c24e3b07dde58c7c688b0dcadba300664c97329e738

Observation 835d2202-c34e-47ce-8dd7-cbb7280c706c · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.265647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.265647Z digest=sha256:67d1bbd775359bce853ca0a250ee054551007e8dbc72aa05134ae729e516b31a

Observation 78f5cf1a-7238-481a-aaac-6960c8b4c934 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Instruction-Following Evaluation for Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.355982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.355982Z digest=sha256:311a6b1b5f6ff4be72cd44fb5959de78cae23851cfc90c66d2e52edc291e756c

Observation 6eb98aae-9e09-4c8b-a53c-e63acfe06a87 · outbound

This paper cites an unresolved cited work.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.428290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.428290Z digest=sha256:6a3178925e281318a8aa9ac5177c1de50be5e2124c0b2ec33b8bdf7dcf578737

Observation 377648b2-6d7e-47a0-a253-9d73c9182140 · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.510840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.510840Z digest=sha256:0013c0c60af75754c4a854fe657a7516702ad223864ee08def7a6d18c39b6b81

Observation ee4802b8-704c-4ce6-ba17-55dc5ff8dbb2 · outbound

This paper cites online" 'onlinestring :=.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search online" 'onlinestring :=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.577547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.577547Z digest=sha256:f828f626d2b29ce4dcba24c98974331065fb63455006bae3a16a71324bae952f

Observation b5960a35-e0d5-46f2-9186-5756badda292 · outbound

This paper cites write newline.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search write newline

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.628815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.628815Z digest=sha256:916e9c5c0feac16f180f4fd19edd52d9293ecc0aa6fbab450a82946e73ebe49e

Pith citing papers

Observation d4741241-e8f2-4fa6-b6c2-25f2c4592963 · inbound

ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism cites this paper.

ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:01.975680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:01.975680Z digest=sha256:8972bc09046ec3cc6deefe234bcc8538956950531465db3c8240244f12c66d18

Observation 41286303-b7db-49f1-89ba-016153a2aaca · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 198

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:05:31.659995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:39ba69f69d306d686714499394e97b7eb3dec00bf91cbf48280b33098bc0aed9

Observation 55afe15c-6588-48c6-871e-1df5cc9c14b4 · inbound

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation cites this paper.

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T11:11:11.329062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:11:11.329062Z digest=sha256:9b82466dac6c3b514b1422e3ae4abd31006a8974015770897f536e8a0cd47b7d

Observation bc205017-7caf-4ec6-b1fc-0c19b7228b78 · inbound

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse cites this paper.

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:05:39.076738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T02:04:46.559881Z digest=sha256:1644dc2140378cbf2f8e82b5e881fefc5e512796717654543418451048f7eb4d

Observation 223c20ca-91da-4c58-9739-b077b2209246 · inbound

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning cites this paper.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:27.901205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:27.901205Z digest=sha256:019cbd32ba434ef7c039a43c77ef7437e63f8d269b875f9954971c48ecf72661

Observation f59c23a7-3b89-4b5d-97a9-95dbc550f765 · inbound

Your Model Diversity, Not Method, Determines Reasoning Strategy cites this paper.

Your Model Diversity, Not Method, Determines Reasoning Strategy TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:56:02.900863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T15:16:30.448400Z digest=sha256:bb636c5b7d85e7944330e5d8d3d2d9174e8ffeaa6d8db841fc2f9207c63044bb

Observation 871d0549-81eb-41e7-bc1d-f3d15e446620 · inbound

Mind DeepResearch Technical Report cites this paper.

Mind DeepResearch Technical Report TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:50:20.510822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:46:49.178896Z digest=sha256:1e9b7a053ed1a44b9d9f80bf63f5dceacc7fa7f7b0b896e0ecbd1d9b80dc75a2

Observation d84b49d2-2575-42a7-b3b7-89b17e034bd9 · inbound

A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping cites this paper.

A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:56:08.732519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T10:41:46.675257Z digest=sha256:af3e424d01b873793f8f1c09856b6375ef1dd9cf17d5e705cbcd65bbd73be4aa

Observation 9dae7e03-dc10-483b-afb3-bfb4057ee305 · inbound

Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias cites this paper.

Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:41:24.001119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T00:51:58.315109Z digest=sha256:c9bbc9c1215d6ad0fabb9b926ffa3a760d18151b993da64251ad1e5c2af869e6

Observation d66b4fc2-8352-4580-91ae-06aaccc2e7c2 · inbound

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling cites this paper.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.272895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:83931d7b9deed3751ed4e5deb12cffd08649db58ec07c75210b3737259ddbd5f

Observation 6bdbc50d-42f7-4a2e-9bfc-6124572ea550 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 130

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.465370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:c169451dc84757f94d92758b777464479f08128917303cce21c27643e17a31ae

Observation 42992776-2f8f-409a-aa5d-dd9f01562b47 · inbound

Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs cites this paper.

Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:46:33.124385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T09:40:35.258818Z digest=sha256:19afeed0fd787d91eff48ebe3e0c1ba54f87e2168e83ecc030b2fc6bc4991162

Observation 54e78ec6-2080-409f-a345-eaf32746f106 · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 128

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:56.077665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:745631442269ff8f975da91d0c21d911a2a8ae031d2aa7581259e27652c197e4

Observation 0e55e482-269d-4501-83d7-1fdd41060ce4 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.748310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:b78a9db4f859977108771ad9694e2d355eb40c38a3d4c0a0099dcf7049660f81

Observation 56d2bf61-7fb1-45cb-9a2a-53bbd9ab9731 · inbound

Optimizing CUDA like a Human: Micro-Profiling Tools as Expert Surrogates for LLM-Based GPU Kernel Optimization cites this paper.

Optimizing CUDA like a Human: Micro-Profiling Tools as Expert Surrogates for LLM-Based GPU Kernel Optimization TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:59:56.473477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T01:12:04.486570Z digest=sha256:534e33c7f28debe5ab20b2bba213e68c022e1c90e8e7752372bc12dc66730277

Observation e6123da0-4012-48b2-8d19-fb1ac55d7222 · inbound

CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents cites this paper.

CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T14:03:48.459093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:03121cd8e2045ef44333600e932264552466f4c57f56f868538414b77f42d935