Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T01:08:25.628815Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 16 inbound Pith citation observations for arXiv:2506.11902.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T01:08:25.628815Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:04:01.975680Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T12:15:01.137692Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f9000ba0-c45d-4c5e-bc4e-749e5da8624e · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35c08fa4-6399-477d-bc3e-b7597da039ab · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 548c0e0d-4b53-40d9-8dc8-97d8b7c4fd0f · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37859dfb-bf4d-40f2-b6a5-eede338bbcf6 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search AlphaMath Almost Zero: Process Supervision without Process
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0474d1d2-34ee-482b-892b-0c3854886c24 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0808632-c636-405e-bf7f-551e4a35ad44 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search The Llama 3 Herd of Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c54a79fa-7e5e-42a8-85a3-f7797b28d389 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7d0511f-bc7d-4b3f-9910-c310f4ad58b5 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 108809f1-b85c-46a5-b31e-70b41564c3ba · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db28df58-6259-4ddf-81cc-9f7145447025 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59be97ce-94dc-4111-8a38-12fe2946a9f4 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e23d2773-23fa-4156-87aa-225ec2af809b · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Measuring Massive Multitask Language Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a69b891-6729-46bb-b446-d107a225d97c · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3abbd6aa-3f04-4587-8587-14516d4153ce · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b173984c-824c-4026-bc23-9654b25d5a29 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebe55ede-7678-4f2c-8c1e-c65490ac9874 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a630102-fced-4a9e-8dd4-3ff21196ca73 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Gonzalez, Hao Zhang, and Ion Stoica
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e3ceeda-aeb0-40d4-8a56-5a201fa75cd1 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f105cbc-8e7c-4f4e-8a0e-b94a0bd87094 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd96f29a-1019-4ae5-b74b-1fc215bc35c9 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Let's Verify Step by Step
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3a8ffeb-dcb2-4621-ab00-4827b7660a87 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Let's verify step by step
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 17994dfb-b49c-4538-9ba7-4905d912477d · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Large Language Model Guided Tree-of-Thought
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5acb4d0d-a41c-4560-b514-bfffe22778e6 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search StarCoder 2 and The Stack v2: The Next Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a1ef31d-1ec5-4bb9-be91-38cada151174 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 360f76b9-4154-482c-8917-9fd6eac0f072 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation de9385a3-4caf-4419-9694-4bcc5477f941 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01541c63-caf3-4805-886b-c1a41b8a9ab9 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68c28502-06e4-489e-a638-0b5bf15d708a · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e8512c18-f4aa-4c11-a56c-5ce891616b30 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba6f07d1-6cbf-4aca-af99-7d7ed57b24b3 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b2ed575e-2c68-4be3-bc27-733b2bc5dc5a · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d289415-6c0c-42d4-8ffd-a3037d2b62e4 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a7f969b-4cb2-4938-ad60-c9dd3bfae858 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f2eee2-1bb1-44aa-bead-6ad9de2b6459 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c42a747e-1287-4068-9841-9947be8584b2 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd4fc2bf-5f38-4144-b847-aa1bc12bed1f · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ecd0436-4be8-4c61-a716-91e27d955e8d · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Gemini: A Family of Highly Capable Multimodal Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6ef6b9a-dc9e-4d59-87dd-d15515afa047 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 536beee4-e060-470b-a045-4f6a9fec0f9d · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a76676cf-d4fe-4126-96ea-2073f65e973b · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6e930f5-f5ae-4d25-b27e-4f1f5de066af · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23e05b3c-5c67-4de9-af7d-a2f197b49a0f · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb7408bf-4e5e-4c70-91be-f632880a6d68 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 835d2202-c34e-47ce-8dd7-cbb7280c706c · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f5cf1a-7238-481a-aaac-6960c8b4c934 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Instruction-Following Evaluation for Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eb98aae-9e09-4c8b-a53c-e63acfe06a87 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 377648b2-6d7e-47a0-a253-9d73c9182140 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee4802b8-704c-4ce6-ba17-55dc5ff8dbb2 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search online" 'onlinestring :=
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5960a35-e0d5-46f2-9186-5756badda292 · outbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search write newline
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4741241-e8f2-4fa6-b6c2-25f2c4592963 · inbound
ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41286303-b7db-49f1-89ba-016153a2aaca · inbound
A Survey of Reinforcement Learning for Large Reasoning Models TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 198
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55afe15c-6588-48c6-871e-1df5cc9c14b4 · inbound
XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc205017-7caf-4ec6-b1fc-0c19b7228b78 · inbound
Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 223c20ca-91da-4c58-9739-b077b2209246 · inbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f59c23a7-3b89-4b5d-97a9-95dbc550f765 · inbound
Your Model Diversity, Not Method, Determines Reasoning Strategy TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 871d0549-81eb-41e7-bc1d-f3d15e446620 · inbound
Mind DeepResearch Technical Report TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d84b49d2-2575-42a7-b3b7-89b17e034bd9 · inbound
A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9dae7e03-dc10-483b-afb3-bfb4057ee305 · inbound
Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d66b4fc2-8352-4580-91ae-06aaccc2e7c2 · inbound
Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6bdbc50d-42f7-4a2e-9bfc-6124572ea550 · inbound
Trust Region On-Policy Distillation TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 130
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42992776-2f8f-409a-aa5d-dd9f01562b47 · inbound
Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54e78ec6-2080-409f-a345-eaf32746f106 · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e55e482-269d-4501-83d7-1fdd41060ce4 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56d2bf61-7fb1-45cb-9a2a-53bbd9ab9731 · inbound
Optimizing CUDA like a Human: Micro-Profiling Tools as Expert Surrogates for LLM-Based GPU Kernel Optimization TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6123da0-4012-48b2-8d19-fb1ac55d7222 · inbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.