Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:32:05.187906Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2412.17397.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:32:05.187906Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-13T01:36:23.845366Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-13T01:36:24.154635Z
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6139c537-2c34-45e4-8e74-2e0cbe47a633 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning , " * write output.state after.block = add.period write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef7f8bf7-0f0d-4e82-9a03-493d51278cba · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c59b594-4057-4559-a966-d31d539f04ed · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 716b68ea-4806-4bab-b082-434262358aff · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae8fbef-08ce-4121-a41a-c5cc708d1804 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Training Verifiers to Solve Math Word Problems
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66c77d0e-e133-40d8-a43f-febd697f0b96 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41634b16-a98d-4af0-8a17-582fc290f46a · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning The Llama 3 Herd of Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6724a79-e296-4dd7-a083-36322016908b · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f014e3-d931-424a-a8a8-12e468bd6d2a · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3c0ff638-9bb8-4f13-b778-5e620ca44c38 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Reasoning with Language Model is Planning with World Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbf429ce-cf85-401b-9596-23268c5f20e6 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcb1c06b-7139-42f8-81a7-08aa29e10fe8 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1842be97-7bc8-460b-a9ba-ba85042970e7 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 65d61a2c-4590-4299-a8fe-b9d85fd0e639 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Mistral 7B
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7b1a721-e5af-48b7-b9d8-048a594e227d · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 74cf42fa-1b7a-4a0f-bf3a-6bb4590f3474 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68af0466-d675-43b6-a933-312d41ea6f2e · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Training Language Models to Self-Correct via Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2b6c74d-03e9-4bf5-8291-ca00818a1fc7 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 87c9b8d4-ace4-4bca-b25c-80f786c7afcd · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Let's Verify Step by Step
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e460728-aeb0-42c6-b628-d7877ce4171a · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c3b8cd5-0cd1-4746-8667-0f53affc4735 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 10f51c8c-cbf8-4e23-b0af-e854b271ee2a · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 087d3004-cef3-48f6-96e5-55305776a496 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning REFINER: Reasoning Feedback on Intermediate Representations
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c47a144d-ef43-4a36-b39b-70530ebf90d4 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Recursive Introspection: Teaching Language Model Agents How to Self-Improve
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6fe4b0d-cfa8-4917-ba07-97a6fb580142 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning D.; Ermon, S.; and Finn, C
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 590eabcc-daa2-4879-8a24-6f44dd336a21 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8f2afc7e-e1c1-42a0-adc0-1045b8ccbb5b · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Multi-turn Reinforcement Learning from Preference Human Feedback
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc4032c3-1d91-45f8-9527-a34ed4c73115 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2dd6853-cff5-4612-8f08-88757ad4fca9 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Reflexion: Language Agents with Verbal Reinforcement Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe3f1d92-ba3d-40d6-be20-3812083b914e · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2f7f02e-3b18-485d-b1cd-f459a8631c03 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Offline RL for Natural Language Generation with Implicit Language Q Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a40f0ab-ea8f-4ca9-8c65-e908bf7b7388 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4fd0710-afef-4a15-b087-da2858a9caff · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aa358fe-ae80-4232-ac82-68d297a01d65 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Solving math word problems with process- and outcome-based feedback
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 174d3ed4-ae9a-4391-a886-334acd0be3ba · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Generating Sequences by Learning to Self-Correct
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1af8582-daa7-4077-a3b8-63c3cd2b011f · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning A.; Ostendorf, M.; and Hajishirzi, H
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8af40692-2384-4830-b52c-7e7b4d258317 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68b738bf-6638-44f4-8d49-efedbaff11b0 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Self-Evaluation Guided Beam Search for Reasoning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87c2715b-ed7d-4e77-8b7e-84faff3d4faa · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Building Math Agents with Multi-Turn Iterative Preference Learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7adbc43-b337-4f3d-98d4-cdbd1ffc4ac7 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b01707f-56f0-460b-bddd-0eb8ccca49de · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1be35d6f-3f50-4936-88ef-3fd03c599e9a · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Small Language Models Need Strong Verifiers to Self-Correct Reasoning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c14088f-63e6-409b-82e9-e14c9c121d84 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eda91110-2467-45dd-8607-670cd82758d6 · outbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Solving Math Word Problems via Cooperative Reasoning induced Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20d65286-dcfe-4a98-bb6c-fa62cf15d1cf · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning
Reference 155
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.