Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 64 inbound Pith citation observations for arXiv:2406.03816.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:10:15.354011Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T03:45:55.622776Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 2d718454-7d82-4fec-81f0-200f2633fc22 · inbound
Improve Mathematical Reasoning in Language Models by Automated Process Supervision ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8653292a-1be5-41ef-928d-61de44463806 · inbound
Large Language Models Can Self-Improve in Long-context Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6289d1b0-481a-484b-99e3-60ccfa79acc8 · inbound
SRA-MCTS: Self-driven Reasoning Augmentation with Monte Carlo Tree Search for Code Generation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 834bdc76-c0b8-48e9-bf5d-50a86808bc3f · inbound
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 153
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e065f485-159f-40ea-b431-68dd1d3c3f96 · inbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b3640ac-c4ef-4313-8320-1b7cd2f390a1 · inbound
Enhancing LLM Reasoning with Reward-guided Tree Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d756a0-ec1d-41f9-bc42-cadb0297302a · inbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5352d337-a3b6-45dc-8f9f-5af1dd93fa91 · inbound
RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b01f65d-f688-4e31-b814-775cf5a28ef1 · inbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 321cbda7-ef41-4c1e-a222-db8c3246af63 · inbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19a858cb-1c7b-49a1-bf42-23ffd7a464a8 · inbound
Seed-CTS: Unleashing the Power of Tree Search for Superior Performance in Competitive Coding Tasks ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53413c37-0369-418d-90f3-b140edbcdde2 · inbound
Progressive Multimodal Reasoning via Active Retrieval ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 121
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44145f61-6d57-4e46-b63e-bd490181b2cc · inbound
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4615a65b-8a81-4837-b8ef-f944ea4ab645 · inbound
Language Models as Continuous Self-Evolving Data Engineers ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da23cc91-d9c0-4101-9133-778167ea3036 · inbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bdef6c5-b27f-42c2-b52d-610c5aedda5b · inbound
Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f2e1174-af77-4c1d-9cf2-9f6d4ef505a5 · inbound
Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0052e1b-d013-423c-afb1-66ecb64e6ba5 · inbound
Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98c184f5-fd52-4b32-b907-79930e83b2ee · inbound
Reasoning-Enhanced Self-Training for Long-Form Personalized Text Generation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deab5b4d-2433-4df5-8191-4c55e8a57584 · inbound
Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdc81c39-977f-4985-87d5-629e578dcf47 · inbound
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 185
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 09ecfb35-f421-46bf-b2b2-220425071ba2 · inbound
RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems? ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c147c86a-1c30-4cc9-bce2-c402dc638100 · inbound
T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5044d472-b652-4884-99f1-01dd711cd724 · inbound
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e3b6c22-85b4-4e11-8a79-ed7b9756ce08 · inbound
Parameter-Efficient Fine-Tuning for Foundation Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73b2a888-cbde-4897-8090-6905b21c5e3a · inbound
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bc6ab0c-4925-41b2-aa0f-7f7ec2aca181 · inbound
Locality-aware Fair Scheduling in LLM Serving ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed1a7084-8618-462c-b28c-14a43dfe3562 · inbound
Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42aaf597-5211-412e-b112-53e20c5aebc7 · inbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ec4b393-0e04-4406-a39f-d5adf40228bc · inbound
Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19cb2dbc-c86d-4fa6-aa64-4abbd030d2d6 · inbound
Holistically Guided Monte Carlo Tree Search for Intricate Information Seeking ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c915cf3-30ac-4f5a-a29b-5b1223555aed · inbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4199f7f-f4b0-424c-8252-8677147d40fd · inbound
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9fb5c6d-7ef9-4cc8-9766-d999c3b3bc77 · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 172
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation db777363-2f3f-434a-b00f-222e6d3c1778 · inbound
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eece3f41-da3b-462b-b495-83bc24ef556c · inbound
SplitReason: Learning To Offload Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8286826-da15-48d6-ad40-2d08c7019740 · inbound
Token-Efficient RL for LLM Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3141de9-49b4-40cb-bd22-61bbde2108b8 · inbound
Accelerating Large Language Model Reasoning via Speculative Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4819ef61-d895-46a2-aa92-b62eefbc383a · inbound
Enhancing Large Language Models with Reward-guided Tree Search for Knowledge Graph Question and Answering ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ab26c4c-3faf-40a7-a2b6-f62068ae0116 · inbound
Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 657dc109-a0cd-412b-860d-200919d6f8d1 · inbound
O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 595069d6-b431-4584-8492-333a1d28070e · inbound
Generalizing Large Language Model Usability Across Resource-Constrained ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 170
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f62cfd8-7bdb-4bc1-b40d-e42e352344dc · inbound
Fostering Video Reasoning via Next-Event Prediction ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b468fc3c-b25a-4c5b-9425-e2a8afe5b8c3 · inbound
Structured Pruning for Diverse Best-of-N Reasoning Optimization ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc91fd04-9f02-4dc1-80dc-a1c9320020f8 · inbound
CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4a3a9e7-5c69-4fc8-bffa-56f9115de103 · inbound
VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 835d2202-c34e-47ce-8dd7-cbb7280c706c · inbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9bd3714-6d04-4bae-b32b-264eec0dba3e · inbound
Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5be8b77b-eb10-4589-9779-57bf1ac7d3d4 · inbound
Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60d23114-b4cc-4d04-a3e8-356b4f0a3d6f · inbound
EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7bd3656-d3c4-4d7e-bdca-adc3cee26df6 · inbound
Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4404c5fd-2fa1-423a-a3b1-419b3d1c9fc1 · inbound
Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55537287-1e32-4ea5-97b8-a6533ee96552 · inbound
Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96cc5f10-7a1c-42e7-95eb-358b94a1a8c9 · inbound
Your Model Diversity, Not Method, Determines Reasoning Strategy ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6856d430-8a02-4ba5-a41e-b32e1038547c · inbound
PARM: Pipeline-Adapted Reward Model ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 77e634fa-d469-4914-96f5-b82c1ad7af7e · inbound
Self-Improvement for Fast, High-Quality Plan Generation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 96f5121e-c87d-416a-abe8-776e950c0f9c · inbound
Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3141565e-11ba-4c49-99e7-97baa43cf0d9 · inbound
Efficient Test-time Inference for Generative Planning Models with OCL Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c8b26145-bd85-472d-ac7c-7dd1757adcc6 · inbound
Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fc6e7769-7fc5-47cd-a782-be6c4bdcb884 · inbound
Theoria: Rewrite-Acceptability Verification over Informal Reasoning States ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d5ea40c1-a14f-41a6-8463-a32c16913dd9 · inbound
Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 98df491d-4780-4f9e-9cc5-a0fea5e01a3a · inbound
UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02923fb9-c267-4ac0-a44d-b3571225cc57 · inbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76aac002-9dd0-447e-aa6e-1327ef9e72ab · inbound
LeAct: Learning to Reason from Expert Actions ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.