Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2508.17445.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T13:44:42.721995Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:00:08.411000Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 6e85e4c9-95eb-4636-b522-21e93dec0dec · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b5dedc86-dbeb-4988-8b31-c267152942ee · inbound
A Survey of Reinforcement Learning for Large Reasoning Models TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 291
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 22201236-da65-4a03-9c75-7c4e2ebdd27b · inbound
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8880798-37da-44a6-9a62-48596e83e9bc · inbound
Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dfc14eb6-8026-4391-ba5f-de4300779ffe · inbound
On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2f3b64a6-9fd2-43a0-8be4-fa6899118e7c · inbound
MARS$^2$: Scaling Multi-Agent Tree Search via Reinforcement Learning for Code Generation TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 92c7244c-7220-47d0-971d-44ce66a8e495 · inbound
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1649c2d8-7890-4b00-94aa-ccdd4798102d · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 747d6ebb-d3ee-4917-93d5-e6774cf4a896 · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f8863bac-0668-4a09-84a7-0acb625f2df0 · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1acda47f-3864-43ec-8e71-e8f981bbeee4 · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3ac7030b-4d37-4549-9f82-002235bb6dec · inbound
Tree-based Credit Assignment for Multi-Agent Memory System TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 50b649c1-6e75-4318-9f10-f13359b616de · inbound
Confidence-Aware Alignment Makes Reasoning LLMs More Reliable TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a9ae4831-4b5b-4f29-a2a6-f36d7dab1cf8 · inbound
StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4e57ca9d-bf71-4157-9800-7d59dd00135d · inbound
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d7ea33ff-bead-4edc-88cc-e6aaa813b95e · inbound
Trust Region On-Policy Distillation TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 137
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fedcaa51-2c99-452b-bd1a-73b58a5ec6ec · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 154
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 26158168-54cf-450a-91a7-ab672c20e535 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8b3c164f-9214-446d-b6ae-aea487e3649a · inbound
Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routing TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4ecadc22-68e6-4932-8c2f-95fbb06ea81b · inbound
Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routing TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e7b180d2-1e1e-452d-9600-1f82833cda0d · inbound
Learning with a Single Rollout via Monte Carlo Pass@k Critic TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fc1243b5-7dda-455d-a4a6-c4b49a9be0ef · inbound
Process Reward Informed Tree Rollout for Effective Multi-Turn RL TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d2a8488-18fd-4831-9eac-02fb84f50a6f · inbound
Distilled Reinforcement Learning for LLM Post-training TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76a670ab-125a-49ab-bcf2-60f6d7ef615a · inbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f465618d-2642-4c03-99f4-85b6e8c5f630 · inbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.