Pith. sign in

Paper Citation Record · LEDGER

ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 64 inbound Pith citation observations for arXiv:2406.03816.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.03816 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 64 of 64 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:10:15.354011Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T03:45:55.622776Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2d718454-7d82-4fec-81f0-200f2633fc22 · inbound

Improve Mathematical Reasoning in Language Models by Automated Process Supervision cites this paper.

Improve Mathematical Reasoning in Language Models by Automated Process Supervision ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:45.987128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T20:53:45.878221Z digest=sha256:24bb95566f6dd1f5f195c07d1ad83fbf47161ec97b78d33ac7ddfd2bfb4e6132

Observation 8653292a-1be5-41ef-928d-61de44463806 · inbound

Large Language Models Can Self-Improve in Long-context Reasoning cites this paper.

Large Language Models Can Self-Improve in Long-context Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-12T21:59:11.624138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:59:11.624138Z digest=sha256:ab8ae890a6938cc67ea583f6dd46cef0e1da625e3870fb933fdfe42435d5c064

Observation 6289d1b0-481a-484b-99e3-60ccfa79acc8 · inbound

SRA-MCTS: Self-driven Reasoning Augmentation with Monte Carlo Tree Search for Code Generation cites this paper.

SRA-MCTS: Self-driven Reasoning Augmentation with Monte Carlo Tree Search for Code Generation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T19:05:29.094385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:05:29.094385Z digest=sha256:3e63d401ff3c219834b0ed359d48039bbc73496b9a788e4e09af4bf3cbc12179

Observation 834bdc76-c0b8-48e9-bf5d-50a86808bc3f · inbound

Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering cites this paper.

Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:49.569025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:28:49.569025Z digest=sha256:a43f721d18de72fb239d56a9d5e91f4d05b3ed61d3c49490f42931c7f75a4ba0

Observation e065f485-159f-40ea-b431-68dd1d3c3f96 · inbound

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment cites this paper.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.061179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.061179Z digest=sha256:4995f3b5b5f8027ae856889e32657d7cb2749f0c4067cc3bce2e56cb17fcba3f

Observation 4b3640ac-c4ef-4313-8320-1b7cd2f390a1 · inbound

Enhancing LLM Reasoning with Reward-guided Tree Search cites this paper.

Enhancing LLM Reasoning with Reward-guided Tree Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T18:18:28.721532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:18:28.721532Z digest=sha256:025ab2ef7fcd2a3244fca261d24fbef4fbdcbcc811b958a999082f4da7c78e75

Observation 88d756a0-ec1d-41f9-bc42-cadb0297302a · inbound

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision cites this paper.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.208906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.208906Z digest=sha256:dff1a3defaa49e54258148aea390688a043cd1f82d7d45528c49e57d486abc1d

Observation 5352d337-a3b6-45dc-8f9f-5af1dd93fa91 · inbound

RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models cites this paper.

RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:16.547714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:08:16.547714Z digest=sha256:fd201badf9144b9b1e7793f90bb4d0d76565e33ee013236bf974aff74af60094

Observation 6b01f65d-f688-4e31-b814-775cf5a28ef1 · inbound

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension cites this paper.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.727434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.727434Z digest=sha256:94ab36becf1c187cc6e766f9cb8ad2ddb846d8e7f8536d5099c96ebfec4cb582

Observation 321cbda7-ef41-4c1e-a222-db8c3246af63 · inbound

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models cites this paper.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.695790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.695790Z digest=sha256:95afcf7a9e3724214a8a43e38f4cbf4ee1a56f2f4e7f45ad2318b5b61188a9d4

Observation 19a858cb-1c7b-49a1-bf42-23ffd7a464a8 · inbound

Seed-CTS: Unleashing the Power of Tree Search for Superior Performance in Competitive Coding Tasks cites this paper.

Seed-CTS: Unleashing the Power of Tree Search for Superior Performance in Competitive Coding Tasks ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T14:02:49.098915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:02:49.098915Z digest=sha256:e5aaca9b85d5f6b1cde63800d75f5056b6025db8860b3fa42fa23e04145025f2

Observation 53413c37-0369-418d-90f3-b140edbcdde2 · inbound

Progressive Multimodal Reasoning via Active Retrieval cites this paper.

Progressive Multimodal Reasoning via Active Retrieval ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-11T11:55:10.620968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:55:10.620968Z digest=sha256:dc36cfe38af3cec730c1f8de600b106022181782249f0ec8f5cd929ffe2a58be

Observation 44145f61-6d57-4e46-b63e-bd490181b2cc · inbound

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation cites this paper.

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T11:40:13.912864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:40:13.912864Z digest=sha256:009c9e393c0ba55104a2fd195aa636ed622d05b4827ed472e055bc9b37efaa40

Observation 4615a65b-8a81-4837-b8ef-f944ea4ab645 · inbound

Language Models as Continuous Self-Evolving Data Engineers cites this paper.

Language Models as Continuous Self-Evolving Data Engineers ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T11:40:52.627913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:40:52.627913Z digest=sha256:316ba3c3b2370ff0c30a111635ce2c1c98c32e136cca17468fa5961a907daa51

Observation da23cc91-d9c0-4101-9133-778167ea3036 · inbound

A Systematic Examination of Preference Learning through the Lens of Instruction-Following cites this paper.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.737718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.737718Z digest=sha256:55a9c977cfe0b27b0206048d058fec8f8b87780ab1319e01e8b9d75e6894b60f

Observation 7bdef6c5-b27f-42c2-b52d-610c5aedda5b · inbound

Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning cites this paper.

Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T11:08:56.611535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:08:56.611535Z digest=sha256:c01a08752499311f7a3ed447e46222bd26bcd340c1af3d0cbfd67342adb8f610

Observation 6f2e1174-af77-4c1d-9cf2-9f6d4ef505a5 · inbound

Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search cites this paper.

Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:53:11.873185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:53:11.873185Z digest=sha256:3117d1a4bc956a703d513ce04f8b0ddc78f7f95edf6b0245331bd651154b94bb

Observation e0052e1b-d013-423c-afb1-66ecb64e6ba5 · inbound

Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search cites this paper.

Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:07.277408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:37:07.277408Z digest=sha256:ee53460fede4535e275d5ebfd0242c5bb875071b407013605d908432fc4b8ef1

Observation 98c184f5-fd52-4b32-b907-79930e83b2ee · inbound

Reasoning-Enhanced Self-Training for Long-Form Personalized Text Generation cites this paper.

Reasoning-Enhanced Self-Training for Long-Form Personalized Text Generation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:44.174187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:44.174187Z digest=sha256:ded43be9afe50d612c13b7acbb951e3996c7da2b59366230718fc4c1070644fc

Observation deab5b4d-2433-4df5-8191-4c55e8a57584 · inbound

Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design cites this paper.

Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:29.534104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:26:29.534104Z digest=sha256:b61305fe8ec5f5f0f9cc2c700a9f98b7cad016e2a94a9dd2e09c0f4af27aab58

Observation fdc81c39-977f-4985-87d5-629e578dcf47 · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 185

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:20:59.421827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:ba2a9b88798531acb308b6db487b0c504e580e9d49c94f8819d82e41fa6d8034

Observation 09ecfb35-f421-46bf-b2b2-220425071ba2 · inbound

RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems? cites this paper.

RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems? ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T18:30:56.608249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:30:56.608249Z digest=sha256:a0d2cbd1c699e764365927eaa6392294af1abf3147874af5ad88ef5987e1f38f

Observation c147c86a-1c30-4cc9-bce2-c402dc638100 · inbound

T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling cites this paper.

T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T18:05:43.206830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:05:43.206830Z digest=sha256:bc19f0cc858b09735e5dc4c06873cbb23f933445b0cd969d859b24615cf75a4b

Observation 5044d472-b652-4884-99f1-01dd711cd724 · inbound

Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning cites this paper.

Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T15:49:44.048289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:49:44.048289Z digest=sha256:60187b9e0f7a07c045f038a646520b2b454c287171cc753a325488f690f5e8b5

Observation 0e3b6c22-85b4-4e11-8a79-ed7b9756ce08 · inbound

Parameter-Efficient Fine-Tuning for Foundation Models cites this paper.

Parameter-Efficient Fine-Tuning for Foundation Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T15:38:02.891230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:38:02.891230Z digest=sha256:b6ac4560cf580efef7312e593293190d141a1573dde9df6d9ed3a192ca53d0b5

Observation 73b2a888-cbde-4897-8090-6905b21c5e3a · inbound

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step cites this paper.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.345905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.345905Z digest=sha256:804e4da97eac0d18eb955e71ac56ce64eb186610acf3ebf775e4315105819dbb

Observation 1bc6ab0c-4925-41b2-aa0f-7f7ec2aca181 · inbound

Locality-aware Fair Scheduling in LLM Serving cites this paper.

Locality-aware Fair Scheduling in LLM Serving ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T15:25:03.583593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:25:03.583593Z digest=sha256:88c2e4742d87d732a3e6a45e707eb5db8e6084033dc73a3c9282d67f8c4ec417

Observation ed1a7084-8618-462c-b28c-14a43dfe3562 · inbound

Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning cites this paper.

Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T14:15:46.160179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:15:46.160179Z digest=sha256:7efbcbe8321d078497e391baa1e13efa6ed922b8910aa45dadf65df1aee4eb31

Observation 42aaf597-5211-412e-b112-53e20c5aebc7 · inbound

On Almost Surely Safe Alignment of Large Language Models at Inference-Time cites this paper.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.771963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.771963Z digest=sha256:a672aad349df7433d496d4e92ae2227e6bacfd913ac53c9537ecc5e173fe6f24

Observation 1ec4b393-0e04-4406-a39f-d5adf40228bc · inbound

Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models cites this paper.

Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T00:27:17.680587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:27:17.680587Z digest=sha256:49d47014c16fbb5f34f04fce47dbad0e74a481ce0f74730b19d633b5caceb64d

Observation 19cb2dbc-c86d-4fa6-aa64-4abbd030d2d6 · inbound

Holistically Guided Monte Carlo Tree Search for Intricate Information Seeking cites this paper.

Holistically Guided Monte Carlo Tree Search for Intricate Information Seeking ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T21:40:26.993871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:40:26.993871Z digest=sha256:372ee70a0cb117cbb2b266a78c5a51c057aa6c19236b5cc27bccc8b32a98d342

Observation 2c915cf3-30ac-4f5a-a29b-5b1223555aed · inbound

PIPA: Preference Alignment as Prior-Informed Statistical Estimation cites this paper.

PIPA: Preference Alignment as Prior-Informed Statistical Estimation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:53.446974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:53.446974Z digest=sha256:0f8e4e34fbe78b821d39c4bbf5fa0a0010a275bb6bb5608ebe803111e1fdac71

Observation e4199f7f-f4b0-424c-8252-8677147d40fd · inbound

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition cites this paper.

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:53.635679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:25:53.635679Z digest=sha256:4be4a232d42be6fb7951888ca8090c2a9e33628ba1a74d12db813272d0d9c81b

Observation d9fb5c6d-7ef9-4cc8-9766-d999c3b3bc77 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 172

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.208820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:d532d723c4c2b88f6c2a23b2acac953d316ecbec1542aed993bf8b7ee1bf3154

Observation db777363-2f3f-434a-b00f-222e6d3c1778 · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.788835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:4ee95fda7980e7909b8d858518f39a92585d43a9424e34bb21ab3a63275dd91d

Observation eece3f41-da3b-462b-b495-83bc24ef556c · inbound

SplitReason: Learning To Offload Reasoning cites this paper.

SplitReason: Learning To Offload Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:10:15.354011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:10:15.354011Z digest=sha256:659699bbae64f6498af172242876cca2743003fc3004db906537164b846874b5

Observation b8286826-da15-48d6-ad40-2d08c7019740 · inbound

Token-Efficient RL for LLM Reasoning cites this paper.

Token-Efficient RL for LLM Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:09.328405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:24:09.328405Z digest=sha256:a424163eb7be2059a2fe2f8cd2cab99d1117237738d669f0b6af48c7976a9bff

Observation c3141de9-49b4-40cb-bd22-61bbde2108b8 · inbound

Accelerating Large Language Model Reasoning via Speculative Search cites this paper.

Accelerating Large Language Model Reasoning via Speculative Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:17:36.670733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:17:36.670733Z digest=sha256:d72dc08f86157e7c2e88e6ef9e8c85689c53c625b27b400e68e6c657f7624c2c

Observation 4819ef61-d895-46a2-aa92-b62eefbc383a · inbound

Enhancing Large Language Models with Reward-guided Tree Search for Knowledge Graph Question and Answering cites this paper.

Enhancing Large Language Models with Reward-guided Tree Search for Knowledge Graph Question and Answering ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:09.431662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:09.431662Z digest=sha256:b277bd257e18e05cb6c3adbfeb7862881b834620ecf9ba6409a4b83b9bc7f030

Observation 5ab26c4c-3faf-40a7-a2b6-f62068ae0116 · inbound

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning cites this paper.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.099971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:50.099971Z digest=sha256:dabf1a0197e3b7030dce252a7ecdd262a8ef5fb393351a928d97afe128b61f9c

Observation 657dc109-a0cd-412b-860d-200919d6f8d1 · inbound

O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering cites this paper.

O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:53.041020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:53.041020Z digest=sha256:170dab734f883627644dc3bfad92b4dfb646a36b3ee368a9c801f1176cdd2755

Observation 595069d6-b431-4584-8492-333a1d28070e · inbound

Generalizing Large Language Model Usability Across Resource-Constrained cites this paper.

Generalizing Large Language Model Usability Across Resource-Constrained ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 170

Resolution
unresolved
no resolver link, observed 2026-08-15T22:08:56.095573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:08:56.095573Z digest=sha256:fc6d0276ee2a2e60d302148cd21bfb68d90de56a21420fe6479b3130f23c0c42

Observation 8f62cfd8-7bdb-4bc1-b40d-e42e352344dc · inbound

Fostering Video Reasoning via Next-Event Prediction cites this paper.

Fostering Video Reasoning via Next-Event Prediction ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:45.635718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:45.635718Z digest=sha256:73491cc36ce22c31ca5fbeb8b9b29f87caa998987a051a3b53b7f7e46d58ba4a

Observation b468fc3c-b25a-4c5b-9425-e2a8afe5b8c3 · inbound

Structured Pruning for Diverse Best-of-N Reasoning Optimization cites this paper.

Structured Pruning for Diverse Best-of-N Reasoning Optimization ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:23.724916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:23.724916Z digest=sha256:d6c1dcfe20f533d83e236fde07ee4ecf364413527533eeb97110bf84fd52347f

Observation cc91fd04-9f02-4dc1-80dc-a1c9320020f8 · inbound

CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning cites this paper.

CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:38.545811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:36:38.545811Z digest=sha256:36e9073ebd925b580d40203d217298af5047ca6cb7f072ecb88ab61f6a3f24b6

Observation f4a3a9e7-5c69-4fc8-bffa-56f9115de103 · inbound

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism cites this paper.

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:23.240182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:23.240182Z digest=sha256:58c41802c15f1a923d3725f1b27516f79ef634b347cc6c4069021ce335c4e2ce

Observation 835d2202-c34e-47ce-8dd7-cbb7280c706c · inbound

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search cites this paper.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.265647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.265647Z digest=sha256:b572dcfa16ddcdde3e6c83c509f0d1e988a7ed6b188c7110c4a1f8d280ef02ca

Observation e9bd3714-6d04-4bae-b32b-264eec0dba3e · inbound

Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs cites this paper.

Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:54.490103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:03:54.490103Z digest=sha256:51d743abab7a82f9c58db6e9e935a89d23d18ec7230529345415c3a0461d852c

Observation 5be8b77b-eb10-4589-9779-57bf1ac7d3d4 · inbound

Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments cites this paper.

Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:16.296100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:16.296100Z digest=sha256:771420cda99b428faefa693da84aab23030dbb5b6b5a5b1eff8015b0d53c13ac

Observation 60d23114-b4cc-4d04-a3e8-356b4f0a3d6f · inbound

EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique cites this paper.

EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:04:55.263407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:04:55.263407Z digest=sha256:e7c63064315c76b8c0d9fc16f54b9f28c61cc5ef4fa86e137c06221089dfadd1

Observation d7bd3656-d3c4-4d7e-bdca-adc3cee26df6 · inbound

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation cites this paper.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.307933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.307933Z digest=sha256:26f3452f8d495844556bdcbfcc83eea50c19cdf17abedecf0705ac5f647079bb

Observation 4404c5fd-2fa1-423a-a3b1-419b3d1c9fc1 · inbound

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL cites this paper.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.478842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.478842Z digest=sha256:061d259af10822d61617311715f20614af0f18ba2402f3a2f78b4c58a2795e22

Observation 55537287-1e32-4ea5-97b8-a6533ee96552 · inbound

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving cites this paper.

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T06:40:27.771393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:40:27.771393Z digest=sha256:7a3593bd1b5793d664e3ddbda017049f4cfe72cae1a1d6f66d1a266afea208e3

Observation 96cc5f10-7a1c-42e7-95eb-358b94a1a8c9 · inbound

Your Model Diversity, Not Method, Determines Reasoning Strategy cites this paper.

Your Model Diversity, Not Method, Determines Reasoning Strategy ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:02.754586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T15:16:30.448400Z digest=sha256:160ddac0fbbbaeb7874097b6552985be3eead4cc73dfcaef46001546607dba54

Observation 6856d430-8a02-4ba5-a41e-b32e1038547c · inbound

PARM: Pipeline-Adapted Reward Model cites this paper.

PARM: Pipeline-Adapted Reward Model ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:39.512713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T05:15:26.015817Z digest=sha256:5cb982aedf8cb9bbe6b06215c44383802ca3de71580faab58dc24f9b03709baa

Observation 77e634fa-d469-4914-96f5-b82c1ad7af7e · inbound

Self-Improvement for Fast, High-Quality Plan Generation cites this paper.

Self-Improvement for Fast, High-Quality Plan Generation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:41:30.662509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-07T04:13:17.612428Z digest=sha256:accf0a81f1a872e11187c346798fad6c2db7ea32d206e1bcea2147cd853df954

Observation 96f5121e-c87d-416a-abe8-776e950c0f9c · inbound

Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning cites this paper.

Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:06.893478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T17:16:00.499779Z digest=sha256:0df70572a775eb86d6c516ff4b5cb2d3453137c90b4e1026a13fb156bd13eac0

Observation 3141565e-11ba-4c49-99e7-97baa43cf0d9 · inbound

Efficient Test-time Inference for Generative Planning Models with OCL Search cites this paper.

Efficient Test-time Inference for Generative Planning Models with OCL Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:52:35.134640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T18:53:08.784890Z digest=sha256:b712a71ec2ec2cc73c4348f213e17942cb5cc23f2e774226bf1f69b526381341

Observation c8b26145-bd85-472d-ac7c-7dd1757adcc6 · inbound

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces cites this paper.

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.915321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T05:46:26.938277Z digest=sha256:91034dba7849033e981db46eed556d59ef5fc591b6cc86497cba87971cd48f90

Observation fc6e7769-7fc5-47cd-a782-be6c4bdcb884 · inbound

Theoria: Rewrite-Acceptability Verification over Informal Reasoning States cites this paper.

Theoria: Rewrite-Acceptability Verification over Informal Reasoning States ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:56.432836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-02T12:12:20.879377Z digest=sha256:5b8865d24a02c38aeb3cbbc73a8e79cc08d3804975e75d220ac869ec6262caba

Observation d5ea40c1-a14f-41a6-8463-a32c16913dd9 · inbound

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops cites this paper.

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:45:55.624038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-09T03:36:57.168246Z digest=sha256:3be528c975b2685e19c0383e25f056834cd0d8a537a790d8b269c16c92754dab

Observation 98df491d-4780-4f9e-9cc5-a0fea5e01a3a · inbound

UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp cites this paper.

UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T10:51:16.019022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:51:16.019022Z digest=sha256:188fa34ed34cbdc6e14796b1964f369d30e32a908a29b0a953e4bbf7edb78df7

Observation 02923fb9-c267-4ac0-a44d-b3571225cc57 · inbound

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents cites this paper.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.726132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.726132Z digest=sha256:216de08ca72686641561cba6764ed6378ca409b9e403d220d5985290260b9c20

Observation 76aac002-9dd0-447e-aa6e-1327ef9e72ab · inbound

LeAct: Learning to Reason from Expert Actions cites this paper.

LeAct: Learning to Reason from Expert Actions ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T06:32:24.768710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:32:24.768710Z digest=sha256:20481eb67c2f2b3813e3e259bfeabed75e6cf74bad114bed37155431ea674e01