Pith. sign in

Paper Citation Record · LEDGER

Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 91 inbound Pith citation observations for arXiv:2405.00451.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.00451 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 91 of 91 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:16:31.845707Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 80d9cd6b-36de-46fc-9d5f-4d915d320ab6 · inbound

Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering cites this paper.

Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 135

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:49.296695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:28:49.296695Z digest=sha256:9ecd3b4e1d8ce1f5c5803a587b8b90d6b5bd048e4888cbef3ed15e47cff4cd6b

Observation 03f1aaf6-a946-4a12-8d6b-55ed8d4c73d0 · inbound

Towards Adaptive Mechanism Activation in Language Agent cites this paper.

Towards Adaptive Mechanism Activation in Language Agent Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T05:09:47.944328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:09:47.944328Z digest=sha256:4789d31474f21d07aee89196b418db567ffec2379f8a4aa569d7769598ecce29

Observation 492cfe2a-9964-483a-9eb5-c6b866a3df65 · inbound

From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons cites this paper.

From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T17:55:13.631034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:55:13.631034Z digest=sha256:b1976185d3364bce5aa695f9417b71322d023597fd91b5b40df19d6bd5dd51f6

Observation 9e51b3f5-5634-463f-bc0e-2fa5f74376bd · inbound

A Systematic Examination of Preference Learning through the Lens of Instruction-Following cites this paper.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.732265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.732265Z digest=sha256:ef719248243a3e14c69b5a3f9e70b48791e42ee093e5b8b823732dba4d4158da

Observation 5d1afb34-b79c-41d5-a450-4c01cd1a1143 · inbound

Formal Mathematical Reasoning: A New Frontier in AI cites this paper.

Formal Mathematical Reasoning: A New Frontier in AI Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 204

Resolution
unresolved
no resolver link, observed 2026-08-11T10:51:30.375984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:51:30.375984Z digest=sha256:8514dcccd19a84d83ccea17ad412cae0053e81193ff65fa53dc36b0ede7e268b

Observation f562407d-f349-4e75-ab5c-e7e219e36a82 · inbound

System-2 Mathematical Reasoning via Enriched Instruction Tuning cites this paper.

System-2 Mathematical Reasoning via Enriched Instruction Tuning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:55.177931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:55.177931Z digest=sha256:7b2af8e3ea7255724a1dead4c40e4807a551940ab2d29692034414e91637ecfb

Observation 8af40692-2384-4830-b52c-7e7b4d258317 · inbound

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning cites this paper.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.157765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.157765Z digest=sha256:3b749a1b1247e22a98360a06dcf3fbbd882cf65115d02f24a37937a4285b694e

Observation 88b8f79a-7ea5-42ed-aba9-38b42bd20d11 · inbound

Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search cites this paper.

Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:53:11.851150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:53:11.851150Z digest=sha256:5dcd59210999ace50e9a5429cab53fb53b505f17f7f41e733b3fa46607ca6088

Observation 454a3d22-e998-4677-8a83-aa9d8470165a · inbound

Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search cites this paper.

Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:07.251016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:37:07.251016Z digest=sha256:420a0b795c2c66af8080914ae00f3a0ae8b4d3a83acf5cfc161f87b89a0a74c0

Observation fc1a4bc7-b4e5-41a8-8d5c-28877f071ae2 · inbound

ReARTeR: Retrieval-Augmented Reasoning with Trustworthy Process Rewarding cites this paper.

ReARTeR: Retrieval-Augmented Reasoning with Trustworthy Process Rewarding Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:59.671117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:59.671117Z digest=sha256:64e90eac1ebc5343873a959fb3093b90eb6df54c6362a70a1364df1c65974dcb

Observation ce697286-0ca5-40b7-afc7-b01a640d9510 · inbound

Aligning Instruction Tuning with Pre-training cites this paper.

Aligning Instruction Tuning with Pre-training Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T20:10:34.939845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:10:34.939845Z digest=sha256:9f39bb197bc876b3aa248d97e0c4bf1179a7eac00735e3b8c795ab4260e99285

Observation 86487eec-c682-434f-a033-314558ab702d · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 167

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:20:59.372426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:41f6cc68f88b9ccd30c5d8c501b72a2d7c060c87c59cf39c14f7927b2d7c0e2c

Observation ae83e03f-ef37-4ed0-b88f-b93fdb94aeee · inbound

Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps cites this paper.

Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:45:17.625495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T11:45:17.473970Z digest=sha256:11f5ff2cb35177e84c99f502d0662cd253f8f712e0112699edd4b96fb065eae5

Observation 7f15cafc-3d7b-4b6b-ba01-2abf8ce793f2 · inbound

Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search cites this paper.

Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:07:30.543382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T04:06:23.521344Z digest=sha256:ec6f4479e6716a4639c883e3c30c115ac82d3a2a5009259e42283b11fc6a14ca

Observation ca48cf8c-0bac-4244-8be4-0fafc950c87c · inbound

Adversarial Reasoning at Jailbreaking Time cites this paper.

Adversarial Reasoning at Jailbreaking Time Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T14:49:08.851523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:49:08.851523Z digest=sha256:c04698f6f84ff4e6bb8f8fadfdb79cba9bb9476cd8e3f36a72f06de6552e2f0b

Observation 77bd2389-0d2d-41bf-8f95-2dc1e2df1484 · inbound

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation cites this paper.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.802953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.802953Z digest=sha256:a29876a59a03dfa619a4649739a44de2599f9d1ce894a5631b43bcf9da7854c8

Observation 7c5120fd-060a-47a6-a75e-7e2775365954 · inbound

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information cites this paper.

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T13:25:52.121741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:25:52.121741Z digest=sha256:bae2cd4722854a6fb346a3ed096d57cb405cb875550d559a81900f02dd0dd0bb

Observation 54b0c245-e256-405e-8ae9-67f8c9c35f16 · inbound

Policy Guided Tree Search for Enhanced LLM Reasoning cites this paper.

Policy Guided Tree Search for Enhanced LLM Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:31.925154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:20:31.925154Z digest=sha256:a9cfd719f206a86d6b47b00064ea1e9ee7dd9513513389f24825efcb3138e8ff

Observation f123c1d2-b171-4acc-94cb-4186e2f1a67d · inbound

The Process of Categorical Clipping at the Core of the Genesis of Concepts in Synthetic Neural Cognition cites this paper.

The Process of Categorical Clipping at the Core of the Genesis of Concepts in Synthetic Neural Cognition Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 148

Resolution
unresolved
no resolver link, observed 2026-08-10T17:37:46.149210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:37:46.149210Z digest=sha256:a26882a44f947ef2f5f6295b2728690f09999bb9c679447445038d3e4040d9f2

Observation e7b5c070-3cd7-411a-8e25-612bae7b00b3 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 150

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:36:24.137693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:1b7be457cf29c0ef04c2ab26c2611b85e2f84eca8084c46dadea2489ecd5a6c7

Observation f0e36665-2446-416b-a711-e2d2184c22d7 · inbound

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL cites this paper.

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T15:15:46.329142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T15:15:46.255296Z digest=sha256:ddf3b397cd701dc29ecaef273f16f9b8cc182774f2c7617560e9fdb17766af4d

Observation d344b008-9029-4da1-81dc-eeaff6a905f7 · inbound

A Survey of Scaling in Large Language Model Reasoning cites this paper.

A Survey of Scaling in Large Language Model Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 231

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:22:09.274394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T21:20:07.238992Z digest=sha256:4f58509ae2312f6cdc1db686ac9dcf4d8103980e6e923dd04625de521c8550be

Observation 1c7c246d-3af2-4f84-bb1f-362cefc559c0 · inbound

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs cites this paper.

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:42:39.110987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T18:42:39.023650Z digest=sha256:04694079203e59bf65b7f98ec97c5f50495e4597dce2b232434bf2a5423e052c

Observation 805bb766-d5da-4d30-91a4-13e94bd898b0 · inbound

$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark cites this paper.

$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:16:31.845707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:16:31.845707Z digest=sha256:9e98cfbda3023712d5a54d8dd658b941b75c2660ca69085dfd6dec8bd3cff69e

Observation deaae569-f1a2-4a1f-9d4a-da39a1702705 · inbound

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning cites this paper.

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:46:57.117383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T18:46:11.571831Z digest=sha256:08f4349f31baa9d8be7ca27019e4822ace8f66a144cd92fe945ad9bca1a85b3a

Observation 804c29ac-bac8-4d13-83af-ec00d9c83727 · inbound

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models cites this paper.

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-16T05:12:18.622009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:12:18.622009Z digest=sha256:b362763424d027e91e537d598a1e7491e0cff2b10e0f054c3d1333279196d340

Observation ae813dbc-8b2a-44ac-b6b5-de339479dfe9 · inbound

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control cites this paper.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.115772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.115772Z digest=sha256:52c6b240ff6dc96345ca157daf4849c6e97020115a873bbf28054f65e3f8c0ed

Observation 4f91cc1d-8841-40dd-9af8-9bcdc9cc5114 · inbound

CEC-Zero: Chinese Error Correction Solution Based on LLM cites this paper.

CEC-Zero: Chinese Error Correction Solution Based on LLM Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T21:44:05.867538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:44:05.867538Z digest=sha256:80ec9be56eb2892618f3bde3ff8359a9782930538cdf6255cef5903557004249

Observation 41751754-982f-4da2-806c-b400fd83650d · inbound

Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents cites this paper.

Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 451

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:49.367478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:49.367478Z digest=sha256:0046c9e621410a4e35e128bf4ec82abe0161d6beb3a56a08a1c294a2a83495c6

Observation 3996829b-ccc4-498e-994f-8efd91710467 · inbound

EquivPruner: Boosting Efficiency and Quality in LLM-Based Search via Action Pruning cites this paper.

EquivPruner: Boosting Efficiency and Quality in LLM-Based Search via Action Pruning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:01.832895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:06:01.832895Z digest=sha256:b6efd865380d443086298c7a463e20114c4e6b726cb6d55eb348fa74a373aaf1

Observation b0e8a27f-7ab7-4583-9ee6-bbe4a596520d · inbound

Reward Model Generalization for Compute-Aware Test-Time Reasoning cites this paper.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.461851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.461851Z digest=sha256:4480019fed5dd0f38e787bb360238f631d8c8ea92bf3e0d81634a99942f30705

Observation aab6bc26-c491-4c70-a393-298260d74354 · inbound

First Finish Search: Efficient Test-Time Scaling in Large Language Models cites this paper.

First Finish Search: Efficient Test-Time Scaling in Large Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:42.491053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:42.491053Z digest=sha256:195bb96acf818dc955c5862ee4d67a0734a1c7a4fd363fc89cceb146da8b1088

Observation 6ec0619b-14a4-4f90-832b-942fe44fa2fa · inbound

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning cites this paper.

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:59.566653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:22:59.566653Z digest=sha256:0f581fd0140c3eed3dde5d3d5fb65abc5330e2ab2cd9efe10304d0098c0f09c4

Observation d28e2eda-1c8d-4a34-b6ab-382e6809f35a · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 288

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:10.962028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:10.962028Z digest=sha256:1c7e005458505def7176df0be174e07f7a1eeb4b5200dcbe064028fa9ee4e6cb

Observation 9bf720c4-d531-484a-9cea-ec5a7b6a8677 · inbound

Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence cites this paper.

Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:08.638944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:08.638944Z digest=sha256:1a425438ef1e3fc71697dd6a2cd5eb0a29dbbe04645d8396890a8fd1a1f72b7d

Observation 53db4d41-a0ef-45ab-a84d-865bbc38514c · inbound

Control-R: Towards controllable test-time scaling cites this paper.

Control-R: Towards controllable test-time scaling Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:14:21.204522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:14:21.204522Z digest=sha256:f00a834a46919fb0c55bb51679480b97c51007a13706a77b7b6dcb462ef2b803

Observation a09a2f29-2492-42e0-b9f1-000907bd094a · inbound

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning cites this paper.

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:41.085938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:42:41.085938Z digest=sha256:b25a52623957b117667255a5141eb1dc7d510d8f79816a8c968ee08048566bb2

Observation f65d5de8-86c7-4087-8689-c8d1d2d00d61 · inbound

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning cites this paper.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:31.403312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:31.403312Z digest=sha256:74f4096dcf326540e69e971cd6abc6d6ab6a3506236b15872e4c2e7cc9bebe2d

Observation 8f17ab33-77d9-4642-8147-261668aec8db · inbound

AI Agent Behavioral Science cites this paper.

AI Agent Behavioral Science Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 172

Resolution
unresolved
no resolver link, observed 2026-08-07T11:00:54.079028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:00:54.079028Z digest=sha256:1e5bebb946ba7685d96d6c7c64d58053cc83fea700241b2384aebd0cd9661a3a

Observation 3f1b2dee-12e0-4eea-ac87-090838360a58 · inbound

SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation cites this paper.

SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:49.657996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:49.657996Z digest=sha256:e089f107e9fcfc06cb1c34a34e45907fd50c406eef18e3b81823406c29266908

Observation 52378c17-6432-4034-b1cb-652c1758abaf · inbound

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs cites this paper.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.319293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.319293Z digest=sha256:fbd61df20b9096a72f44bd3e643785d5b011dc13403c610b66203f58bef23ee6

Observation 23e05b3c-5c67-4de9-af7d-a2f197b49a0f · inbound

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search cites this paper.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.027204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.027204Z digest=sha256:0525fbdebfbc2e33fb9c2bb97828dbd6b239b8dcaa6f79390a80954a1ce7d251

Observation 5dfb5fc4-6623-405c-9e18-0c96c5b66eea · inbound

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach cites this paper.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.502275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.502275Z digest=sha256:6ce8d1d37213c5eb2cac2ffce064fa1fb06fdf86fb28b3ba1a4e6df22828a0b2

Observation e070fe5c-7e87-41b8-8a10-5bf39df4fad6 · inbound

PLoP: Precise LoRA Placement for Efficient Finetuning of Large Models cites this paper.

PLoP: Precise LoRA Placement for Efficient Finetuning of Large Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:50:14.576341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:50:14.576341Z digest=sha256:2799b91415f42c8b1d1038a50b2f8c132b2cefdc28a43da49bb25c2c3df831dd

Observation adf48413-f128-401a-9f20-b1603e517bd6 · inbound

Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments cites this paper.

Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:15.833229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:15.833229Z digest=sha256:ac7b6ab6623c7ee99f5cdb6be2749a0b1f8a739f6eca3956e011d416734d761b

Observation 11e6cd9d-3b1c-4b70-ac25-4c9c5de4798b · inbound

Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models cites this paper.

Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:56.996116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:56.996116Z digest=sha256:58657828daa9c706c3faae1dbea064d5987e47e72c00aab1ffff794dc3f1af97

Observation e5e80796-9eed-4796-bcae-9ef353d7f30d · inbound

CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning cites this paper.

CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T18:39:18.997563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:39:18.997563Z digest=sha256:a312fe8a6b24843851b453fc2835ae2f6877069920d2fc618e692311df3121ba

Observation 1c03327a-0a26-410a-9f02-e0d0947c75a1 · inbound

ChipSeek: Optimizing Verilog Generation via EDA-Integrated Reinforcement Learning cites this paper.

ChipSeek: Optimizing Verilog Generation via EDA-Integrated Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T06:52:08.359154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T06:48:29.015759Z digest=sha256:8963140adf85af55adf7d3780b6b0b4c0f6d7d1830b166ed0f379340dcd8ec44

Observation da8eb2be-757c-408d-9572-55982b3900b1 · inbound

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning cites this paper.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:05.025670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:05.025670Z digest=sha256:a5c4f58fe7ac393ae4c37e6dceff2d279a85018b43db038d054dee01a2cd073d

Observation 06bec0d1-228a-4eb2-8616-7e5318d67b2e · inbound

Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality cites this paper.

Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 198

Resolution
unresolved
no resolver link, observed 2026-08-05T15:38:55.207976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:38:55.207976Z digest=sha256:2bfa553bd5494d757c84ca6d5ce543e13a78fb85850124a2d8c06dee2c4aeed1

Observation 3ed8bf8c-fce5-4454-a622-c5bef9720374 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.898650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.898650Z digest=sha256:5a08dccece6c1daa136278c6754315335d7035bc9b0a8a94614d9339742e9f14

Observation aa814ac9-bc2e-474d-9163-cfea9b35fa5e · inbound

Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts cites this paper.

Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:42:38.694348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T13:42:07.883909Z digest=sha256:f3e1eaae579312603b9b561d4b5e9e915e2ddff6f4c503fc8605892cc1c99b13

Observation 576f190c-e985-4613-8133-602f0e0edbac · inbound

Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm cites this paper.

Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:39.878932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:44:39.878932Z digest=sha256:ef1c059c7089f100721e531d2b791108ed685b645a4d0fb3d3101ce802be97ec

Observation fbf825b4-52e4-4b83-8952-4c4bae846dfd · inbound

Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models cites this paper.

Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T09:56:13.094934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T09:53:57.765473Z digest=sha256:3de6003a9e8d82ebc37fb931e8039401e5392014df1ab1c87832650cd1ba0988

Observation 0f043711-7720-401a-a758-a72b39b21b18 · inbound

The Art of Scaling Reinforcement Learning Compute for LLMs cites this paper.

The Art of Scaling Reinforcement Learning Compute for LLMs Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.038199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:8a65c86dbe2f97b63c00b583edb068eee3f8a2f546f605c339d70e605ceb1fd5

Observation b56b40c4-68a7-4348-99d7-7eb4484f2f73 · inbound

When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs cites this paper.

When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T08:11:52.628741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:11:52.628741Z digest=sha256:d08a3c62bfbbeb581aaf4287fe9b14be251856f491d56bb3e665466a350c5386

Observation 1336d905-1baf-45fd-9289-70be4433e537 · inbound

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning cites this paper.

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:13:47.957447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T01:11:26.411893Z digest=sha256:b034a3ba76a168a2a775799baf9fef3a34d1ae2c5b4fc7601f41855f15413747

Observation f41c982a-01ce-450e-8365-cea84e37ddb8 · inbound

Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention cites this paper.

Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T17:15:17.701284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:15:17.701284Z digest=sha256:ae00a757903050f5005eb7e0c8c0b723a1fbb50e24af03eef7984e7db3aa9edc

Observation 1427f7f5-f06b-4302-885a-0aeb295ecd74 · inbound

Can I Have Your Order? Monte-Carlo Tree Search for Slot Filling Ordering in Diffusion Language Models cites this paper.

Can I Have Your Order? Monte-Carlo Tree Search for Slot Filling Ordering in Diffusion Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T23:50:37.479610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:50:37.479610Z digest=sha256:979f9c659808502e856f624d6c4b31460d99ea9d448d19fba7daa60c4bf59e6c

Observation 8b0a4382-6d8c-49de-bae7-f6361f8e2a36 · inbound

Online Self-Calibration Against Hallucination in Vision-Language Models cites this paper.

Online Self-Calibration Against Hallucination in Vision-Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:10.758067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T20:20:27.931679Z digest=sha256:8449318960a1fdad033441e196b74827bab74c8a6905716d26bcad09d21782f8

Observation 68714fb5-68aa-476d-9c59-4a3b6fff7f38 · inbound

StoryAlign: Evaluating and Training Reward Models for Story Generation cites this paper.

StoryAlign: Evaluating and Training Reward Models for Story Generation Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:31:07.625809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T17:29:13.549559Z digest=sha256:8c866de34e1e2626e4e8bdc58c17a6e643fbf68692b6871edbb4e4b662981acc

Observation c8ef7e11-e612-4411-85d7-be0817289dcd · inbound

A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping cites this paper.

A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:56:08.803918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T10:41:46.675257Z digest=sha256:8044c7e35eba672a6360396792c958dcfebdd23d730c4d452fa0b63f350d6574

Observation a3040011-7ea3-44ec-a821-3ec480e3f0aa · inbound

CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation cites this paper.

CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:25:58.162248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T02:28:20.366674Z digest=sha256:d79720e87dac041d20eb60b03d89ce2495d2ff0c9db8582a3e7d901a0204d467

Observation 68c07757-b090-4bc8-8122-a6503e8fca4c · inbound

APCD: Adaptive Path-Contrastive Decoding for Reliable Large Language Model Generation cites this paper.

APCD: Adaptive Path-Contrastive Decoding for Reliable Large Language Model Generation Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:26:25.114868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T05:22:25.475956Z digest=sha256:affb85c5827a8db695c57367d3566315808af959f00faca40090323352be5822

Observation c55c9fe3-ad48-4d46-99ac-8f5548f67e17 · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:51:30.308824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:51:52.375703Z digest=sha256:27e050c3fe51ea2f1b57ffbef5e55738e7f00eb316f9c8d435b34ade19b18909

Observation 57f9ef26-61e7-4a7b-9c54-95760f60ccb7 · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:15:03.377542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T05:11:32.053440Z digest=sha256:5900d1933d7e725b7b1f11d0de20ca4609999ffb4758bfcf1800b8185c40d263

Observation 15b827cc-a7dd-4a11-9f5c-b797760a1816 · inbound

Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces cites this paper.

Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:30.885704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:29:33.497561Z digest=sha256:dc1a51956b34e34ca2fd53b2a4669b8cc4a5bcfbca7e7efe7bc4bdf424bb0ffe

Observation 48f329f6-4104-4082-8909-01fa1b90e315 · inbound

Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces cites this paper.

Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:25:06.761279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T22:24:26.528532Z digest=sha256:07e734861106faf40ab9b654e7d8035a35cace4a0cf025859dfd943632aae7f0

Observation 64f8b47c-59d7-4d39-a4ba-3ba70e8c492b · inbound

Scalable Token-Level Hallucination Detection in Large Language Models cites this paper.

Scalable Token-Level Hallucination Detection in Large Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:52:22.455333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T05:49:23.534294Z digest=sha256:29fd5cb6783608c3a8fca5e2be985ac5d13eb2756bece780a1c9ca471cc41897

Observation 6e1d9e44-cb39-4d6f-94dd-d1797095a35d · inbound

ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training cites this paper.

ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:04:40.483302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T12:57:37.344369Z digest=sha256:189e8637f7e1025838c33d9a62f80c479dd0c7fc9e3f6aeee40fa4d65def4384

Observation bae10265-a6c0-4d2c-bc44-db096746d7aa · inbound

ATLAS: Agentic Test-time Learning-to-Allocate Scaling cites this paper.

ATLAS: Agentic Test-time Learning-to-Allocate Scaling Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.086263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T15:27:28.290178Z digest=sha256:235fac0886491e8ac2b615ba386f6320151712eac75f9d7854df93048b35280e

Observation 6e49eddc-22ef-4fce-a20d-44cbeb608b66 · inbound

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces cites this paper.

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:46:48.943212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T05:46:26.938277Z digest=sha256:73e8bf54cd1f9f4f4a9902a31249db40b88b90bb476679547445ecf32cc4aee5

Observation f6beadf6-5ee6-45d4-b3d0-40f50864e89c · inbound

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning cites this paper.

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:25.354455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T19:36:57.231932Z digest=sha256:23a351b70733f6cea2c50f3e8f64f4f923c9d743e8347c306b96be3c65cf38d7

Observation 2bc63d21-22b7-432f-b193-237cc353379a · inbound

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning cites this paper.

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.752547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T13:55:35.363377Z digest=sha256:015d853451382a9579db12b723fd2552aeb8c38a330f48ee5c713171070d74b2

Observation 28117a7f-5f6a-4c38-b65b-444a7c346444 · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:37:49.336402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T10:21:55.485624Z digest=sha256:54d4f24f4b63c6c61c106385dbc832447f3fa94b79dbc7febd296ed738477969

Observation e9977467-00f8-4151-8f7b-7bd97cb7a3cf · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:34.483141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:34.483141Z digest=sha256:3c2b5fb260736de525c344ef7f6cffa60314ea180741d41c55cb061633ed57f1

Observation be34a2a5-a3d3-41c9-9d02-0bf58118a20d · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:33.369931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T06:39:34.199607Z digest=sha256:d7e766a32fe1296d1709b257becb473c34ece71d0e0acc8439437578faabdd62

Observation b85efc86-f7b5-42fb-a897-3b8873fc3b52 · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:23.860218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:23.860218Z digest=sha256:c0accc3ab71be4b6b84663245444500774f40162b746aaa9bdddcb317797d256

Observation 824af795-821b-434e-bc1f-1cdba21dd220 · inbound

Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies cites this paper.

Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:44:59.791395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T18:38:48.512777Z digest=sha256:aaef70b64be7c6ebe0bec76cff01a2482a596efc001a74a6a6dc507a19fa5738

Observation 1410f6a1-38bf-4ab8-90ae-624ba2f850b4 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 236

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.727786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:766e91dfdd8114af5abccb0bb359bb18581aee4aee5c2f4f3f6010ddc76e68db

Observation bd4cebbd-1f7e-4e37-97ef-503b2241cf42 · inbound

Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routing cites this paper.

Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routing Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:20:06.708349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-25T21:29:40.564852Z digest=sha256:6f41a98ad6c55b613270031c1dcb44a210d7331f423c9a48229a5f43e0bb6904

Observation 0d3dd942-0e67-4127-8a6d-21c116e33c71 · inbound

Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routing cites this paper.

Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routing Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T06:35:29.541290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T06:33:45.824727Z digest=sha256:7c9a67bf023b0f56acc18fb0f8f5a9ab22e35b4a3da8e0045c254b55718c435f

Observation f526cda6-e761-4871-83c3-6b5e4f3a5928 · inbound

Addressing Over-Refusal in LLMs with Competing Rewards cites this paper.

Addressing Over-Refusal in LLMs with Competing Rewards Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:55:35.605736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-01T06:59:12.695984Z digest=sha256:7d2179dd5c8ceb09aac720c260188edc3fba2a281d361ea995c14d796dfdd550

Observation f5348cad-0dee-43c0-8c6e-a9a67b0c39fe · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:442a45f90e00b295b835fa01d42d133a6613093ff12eb173b66448edadbddac9

Observation 611ae85f-0e67-41c6-b984-776c7bf72877 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.787640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.787640Z digest=sha256:1ac10d2d93ce19753d796e394a93525796bc77b9fd97ca4b934e616e953918f2

Observation 9ab0484c-befc-4fdd-877f-07d23aa35314 · inbound

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL cites this paper.

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-02T14:50:17.718586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:50:17.718586Z digest=sha256:1ce4266dac903cfd21e20527a919c9a057d0b67998a4b8855853cb32f4f4384c

Observation aaaaf2ff-7240-4b6f-b5d1-0bd167f1baa2 · inbound

Theoretical Foundations of $\max$@$k$ Reinforcement Learning cites this paper.

Theoretical Foundations of $\max$@$k$ Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T15:43:00.157724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:43:00.157724Z digest=sha256:12e88573b4b2486f67925ff70a5aec02cacf91711a44136a1f1c3926f8d7da77

Observation ab2469e5-4b98-4ee1-8904-3346aacf9b3e · inbound

Thought-Level Beam Search for Reasoning cites this paper.

Thought-Level Beam Search for Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T00:38:25.264770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:38:25.264770Z digest=sha256:523681dd3a394a2e94f74c6f2b090f600f554d345aa735bb80e5851143bf0327

Observation 42eeda29-89e7-4ce3-853e-7adcccbc75ce · inbound

Thought-Level Beam Search for Reasoning cites this paper.

Thought-Level Beam Search for Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T14:30:50.755233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:30:50.755233Z digest=sha256:d8703ea0ef50bf0c7a7617aaaa936050dc377cc998d659aba32bac4cda9fbefe

Observation ac3caf52-92a7-4a2a-a78f-cce02a3d2cf7 · inbound

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling cites this paper.

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 166

Resolution
unresolved
no resolver link, observed 2026-08-12T14:10:45.699834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:10:45.699834Z digest=sha256:2504b5c874b91a9fa39350930e7d1dd39182c073a0c1dd80743a9557d7322c2a

Observation 42632a1c-3669-4b61-9f89-470f2fcefc82 · inbound

Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations cites this paper.

Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:21:59.252295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:21:59.252295Z digest=sha256:605380cab9d8c6d35adc2d378504bef5f812fb9abe4daf80ead77cf9b1f86708