Pith. sign in

Paper Citation Record · LEDGER

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning

As of 12 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2412.17397.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17397 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:32:05.187906Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T01:36:23.845366Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T01:36:24.154635Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6139c537-2c34-45e4-8e74-2e0cbe47a633 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.007464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.007464Z digest=sha256:df61c853457d88afab27fb436b5e1e7729e53edbc485836022a0c999833ec60e

Observation ef7f8bf7-0f0d-4e82-9a03-493d51278cba · outbound

This paper cites write newline.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.012085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.012085Z digest=sha256:d0a0d608e9c018b0bb19e58e4511088e048ab43071756d982a19935e456a0e35

Observation 6c59b594-4057-4559-a966-d31d539f04ed · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.016894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.016894Z digest=sha256:cff2d4285d667e9e575449067f1cad54d100f42020c122305498430924be3ba0

Observation 716b68ea-4806-4bab-b082-434262358aff · outbound

This paper cites RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.021572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.021572Z digest=sha256:2bf6a73f26984ca6bf7a6035a83dd1a9e95ac031843bd85de4ef1fdfcc6ce15c

Observation 9ae8fbef-08ce-4121-a41a-c5cc708d1804 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.025822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.025822Z digest=sha256:7a959110ee924286be637a3d85697ed658506bf34eff70344572cf08251cf12d

Observation 66c77d0e-e133-40d8-a43f-febd697f0b96 · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.030152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.030152Z digest=sha256:b6af8222d854258788169c35f6432f8903aa7f1404e3e431082dc36b9214b14e

Observation 41634b16-a98d-4af0-8a17-582fc290f46a · outbound

This paper cites The Llama 3 Herd of Models.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.033716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.033716Z digest=sha256:d8771ec4a48be4b4af6468cb9688f2d9d3e2ef1d6af484064699f730f8c8ed4e

Observation e6724a79-e296-4dd7-a083-36322016908b · outbound

This paper cites Stop Regressing: Training Value Functions via Classification for Scalable Deep RL.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.037290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.037290Z digest=sha256:e3866bc8b76eb135f82c3b016906feea43f10fbc74e22238fdce9d31f0b13fd6

Observation 81f014e3-d931-424a-a8a8-12e468bd6d2a · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:32:05.866211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:32:05.041677Z digest=sha256:686ebea636cfa1ebee91060f34fc0bbdb9712474e64a96f1a9d2ddc409088c20

Observation 3c0ff638-9bb8-4f13-b778-5e620ca44c38 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Reasoning with Language Model is Planning with World Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.045671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.045671Z digest=sha256:d36f09a3904398cbf9d5f5412a76a9aa64ed7a7ff296e7133112602dab56b3ec

Observation dbf429ce-cf85-401b-9596-23268c5f20e6 · outbound

This paper cites GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.050096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.050096Z digest=sha256:71f52b60c8ae214ce1a8277fcf816385aa4ce9495a176505a832bc5ab758ee7f

Observation bcb1c06b-7139-42f8-81a7-08aa29e10fe8 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.054517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.054517Z digest=sha256:b91fc8ca68006c6a6c7d3060ff6164c7305175d532fc0ebda2ecb4242da1ba9b

Observation 1842be97-7bc8-460b-a9ba-ba85042970e7 · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:32:05.854172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:32:05.058941Z digest=sha256:9d1e92764052f73086037a515c3dbeef9a9b3b616f549225af34a7b84b04c37b

Observation 65d61a2c-4590-4299-a8fe-b9d85fd0e639 · outbound

This paper cites Mistral 7B.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.062776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.062776Z digest=sha256:120f37ca7884e54b8cb859e98a13ac31b8c970323e4aa9f561bcc87379b8b639

Observation e7b1a721-e5af-48b7-b9d8-048a594e227d · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:32:05.842298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:32:05.066919Z digest=sha256:f27658a502b8a1816a0c48b51f4525606e7e42032e9d4d24639ae34d90b9831f

Observation 74cf42fa-1b7a-4a0f-bf3a-6bb4590f3474 · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.071278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.071278Z digest=sha256:6a451637b31a05b8a759115aa3b42293ee3a118dafcad8434b68b66ba30c5e9e

Observation 68af0466-d675-43b6-a933-312d41ea6f2e · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Training Language Models to Self-Correct via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.075477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.075477Z digest=sha256:97279e184d697fd86132b82fd0304315d6359136b117568ef423b590f8d7af82

Observation c2b6c74d-03e9-4bf5-8291-ca00818a1fc7 · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:32:05.822646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:32:05.079473Z digest=sha256:e0a1a10f7def2ff5cde76c6b448d49230947f47664407c36ec97f3dbfe8929f3

Observation 87c9b8d4-ace4-4bca-b25c-80f786c7afcd · outbound

This paper cites Let's Verify Step by Step.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Let's Verify Step by Step

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.083498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.083498Z digest=sha256:27eda6d490db63ec08d96db0c1be6010ca8dd3feb0bc745385944c3aeb28541a

Observation 6e460728-aeb0-42c6-b628-d7877ce4171a · outbound

This paper cites Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.087359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.087359Z digest=sha256:22ee0589657112b801f66379c3c816b9ef4eee7a369cdd226f17ae3b129d7429

Observation 1c3b8cd5-0cd1-4746-8667-0f53affc4735 · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:32:05.811023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:32:05.091033Z digest=sha256:19765a532fa3ca9db4320279a1fc472cdbfaa14195a2b1b5ac47e2f1d76a8a37

Observation 10f51c8c-cbf8-4e23-b0af-e854b271ee2a · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.094613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.094613Z digest=sha256:84d70c6de419920d4dddefb0fb355cf602029b01920d611325f9808045fe2e60

Observation 087d3004-cef3-48f6-96e5-55305776a496 · outbound

This paper cites REFINER: Reasoning Feedback on Intermediate Representations.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning REFINER: Reasoning Feedback on Intermediate Representations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.098268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.098268Z digest=sha256:3c81aa8117faed8c7f115f06d6d9a334962c43e83c133ca82ca70618e2e39170

Observation c47a144d-ef43-4a36-b39b-70530ebf90d4 · outbound

This paper cites Recursive Introspection: Teaching Language Model Agents How to Self-Improve.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Recursive Introspection: Teaching Language Model Agents How to Self-Improve

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.102208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.102208Z digest=sha256:890081a2027d73cd5c116a4f6f4d50d7ff500f29b8a112645d27955f58ccb2cd

Observation d6fe4b0d-cfa8-4917-ba07-97a6fb580142 · outbound

This paper cites D.; Ermon, S.; and Finn, C.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning D.; Ermon, S.; and Finn, C

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.106426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.106426Z digest=sha256:e9c0aff663a1293b5122976991769f24eef3151512d23ba29d849433d27c19c2

Observation 590eabcc-daa2-4879-8a24-6f44dd336a21 · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:32:05.785053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:32:05.109858Z digest=sha256:a4a910048f7935d50130a55e55241e574f546426e6c99f698a13ce80402c0452

Observation 8f2afc7e-e1c1-42a0-adc0-1045b8ccbb5b · outbound

This paper cites Multi-turn Reinforcement Learning from Preference Human Feedback.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.113680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.113680Z digest=sha256:dd13dca6edcf36c0a26bfe4ffd384fa7c3970f9467b12f9bccd2c5b48d2b0976

Observation fc4032c3-1d91-45f8-9527-a34ed4c73115 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.117598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.117598Z digest=sha256:608ea4be87ccb1ffd9a920a6313ddbbbdd35c5a5957bb1897c1e6dc4aab74db1

Observation e2dd6853-cff5-4612-8f08-88757ad4fca9 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.121465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.121465Z digest=sha256:c0d9d3326010e250cb7bd396f325a0746ff3abd845adbf4a946e447c13ab6f04

Observation fe3f1d92-ba3d-40d6-be20-3812083b914e · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.125316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.125316Z digest=sha256:c6adc711517cce215844f9bfb7da52f092434bd948829dee72903d63c52dd346

Observation d2f7f02e-3b18-485d-b1cd-f459a8631c03 · outbound

This paper cites Offline RL for Natural Language Generation with Implicit Language Q Learning.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Offline RL for Natural Language Generation with Implicit Language Q Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.128971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.128971Z digest=sha256:653a044470826b11318fbabade59765f813f045ae2edd8c5792f640c3b5e97c2

Observation 8a40f0ab-ea8f-4ca9-8c65-e908bf7b7388 · outbound

This paper cites DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.132531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.132531Z digest=sha256:2be0ced64eacb9802efc5eaf398a9ed8177a693a3d9f1e378d6235bdb790952a

Observation c4fd0710-afef-4a15-b087-da2858a9caff · outbound

This paper cites OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.140190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.140190Z digest=sha256:1363322e7670d0c6dfca39f266a6acff9b7b5f7e93d1e23b92be3f919ab2d518

Observation 2aa358fe-ae80-4232-ac82-68d297a01d65 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Solving math word problems with process- and outcome-based feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.144130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.144130Z digest=sha256:4b27531299655141a83a056284621a41c6e12bbf56214e209173eb2ea1e5710f

Observation 174d3ed4-ae9a-4391-a886-334acd0be3ba · outbound

This paper cites Generating Sequences by Learning to Self-Correct.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Generating Sequences by Learning to Self-Correct

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.148451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.148451Z digest=sha256:1befc6fcc3e9adae1e24c3f432bab004bbc3b4edf98bf38f98bde5ebef1aa04c

Observation f1af8582-daa7-4077-a3b8-63c3cd2b011f · outbound

This paper cites A.; Ostendorf, M.; and Hajishirzi, H.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning A.; Ostendorf, M.; and Hajishirzi, H

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:32:05.774600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T05:32:05.153337Z digest=sha256:590bcefb2bdea968e5c1d3a81512a8958ac29eb1a31f6a01f330b3b9a3481ce3

Observation 8af40692-2384-4830-b52c-7e7b4d258317 · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.157765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.157765Z digest=sha256:9aeff597a5837f0ddd535b0c8fed8ef54530aaec6efadac1450db0a807bad2aa

Observation 68b738bf-6638-44f4-8d49-efedbaff11b0 · outbound

This paper cites Self-Evaluation Guided Beam Search for Reasoning.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Self-Evaluation Guided Beam Search for Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.161940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.161940Z digest=sha256:303ca31320499c95bdfbb555df1a26a162a938ec8b57e75b2578fb80e08b848d

Observation 87c2715b-ed7d-4e77-8b7e-84faff3d4faa · outbound

This paper cites Building Math Agents with Multi-Turn Iterative Preference Learning.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Building Math Agents with Multi-Turn Iterative Preference Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.166120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.166120Z digest=sha256:1dc735e8a96409da89f2c792aefdba9743486dd7defe95fab35e284aa8b0b773

Observation c7adbc43-b337-4f3d-98d4-cdbd1ffc4ac7 · outbound

This paper cites an unresolved cited work.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.170329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.170329Z digest=sha256:19d6b8872cc8798c4e10ce6bc962f330e65965844d5007d8d226e0552071fce1

Observation 5b01707f-56f0-460b-bddd-0eb8ccca49de · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.174422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.174422Z digest=sha256:fb7af4565f3b701f12e16ac1584d77fbb1a82093e3a5cb57a3dcc62212c34391

Observation 1be35d6f-3f50-4936-88ef-3fd03c599e9a · outbound

This paper cites Small Language Models Need Strong Verifiers to Self-Correct Reasoning.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Small Language Models Need Strong Verifiers to Self-Correct Reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.179010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.179010Z digest=sha256:e6e76746a90d507d8b737ee9d78223b9313bac83aaea93aaed2f26a131655afb

Observation 9c14088f-63e6-409b-82e9-e14c9c121d84 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.183607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.183607Z digest=sha256:2124e4a2ead884537e0315a2185efde2240c62efe3fa3866a540e7a119c10367

Observation eda91110-2467-45dd-8607-670cd82758d6 · outbound

This paper cites Solving Math Word Problems via Cooperative Reasoning induced Language Models.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Solving Math Word Problems via Cooperative Reasoning induced Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.187906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.187906Z digest=sha256:8a187afae8137b52e48b1408d3fbfaa4491d757de62a597fd74a5f2308e9783a

Pith citing papers

Observation 20d65286-dcfe-4a98-bb6c-fa62cf15d1cf · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.157579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:5b606e2d094fb68b0affead733d0182798075dec7fb7896e58cf211df5c13add