Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 91 inbound Pith citation observations for arXiv:2405.00451.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:16:31.845707Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 80d9cd6b-36de-46fc-9d5f-4d915d320ab6 · inbound
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 135
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03f1aaf6-a946-4a12-8d6b-55ed8d4c73d0 · inbound
Towards Adaptive Mechanism Activation in Language Agent Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 492cfe2a-9964-483a-9eb5-c6b866a3df65 · inbound
From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e51b3f5-5634-463f-bc0e-2fa5f74376bd · inbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d1afb34-b79c-41d5-a450-4c01cd1a1143 · inbound
Formal Mathematical Reasoning: A New Frontier in AI Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 204
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f562407d-f349-4e75-ab5c-e7e219e36a82 · inbound
System-2 Mathematical Reasoning via Enriched Instruction Tuning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8af40692-2384-4830-b52c-7e7b4d258317 · inbound
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88b8f79a-7ea5-42ed-aba9-38b42bd20d11 · inbound
Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 454a3d22-e998-4677-8a83-aa9d8470165a · inbound
Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc1a4bc7-b4e5-41a8-8d5c-28877f071ae2 · inbound
ReARTeR: Retrieval-Augmented Reasoning with Trustworthy Process Rewarding Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce697286-0ca5-40b7-afc7-b01a640d9510 · inbound
Aligning Instruction Tuning with Pre-training Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86487eec-c682-434f-a033-314558ab702d · inbound
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 167
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ae83e03f-ef37-4ed0-b88f-b93fdb94aeee · inbound
Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7f15cafc-3d7b-4b6b-ba01-2abf8ce793f2 · inbound
Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ca48cf8c-0bac-4244-8be4-0fafc950c87c · inbound
Adversarial Reasoning at Jailbreaking Time Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77bd2389-0d2d-41bf-8f95-2dc1e2df1484 · inbound
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c5120fd-060a-47a6-a75e-7e2775365954 · inbound
LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54b0c245-e256-405e-8ae9-67f8c9c35f16 · inbound
Policy Guided Tree Search for Enhanced LLM Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f123c1d2-b171-4acc-94cb-4186e2f1a67d · inbound
The Process of Categorical Clipping at the Core of the Genesis of Concepts in Synthetic Neural Cognition Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 148
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7b5c070-3cd7-411a-8e25-612bae7b00b3 · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f0e36665-2446-416b-a711-e2d2184c22d7 · inbound
LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d344b008-9029-4da1-81dc-eeaff6a905f7 · inbound
A Survey of Scaling in Large Language Model Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 231
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1c7c246d-3af2-4f84-bb1f-362cefc559c0 · inbound
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 805bb766-d5da-4d30-91a4-13e94bd898b0 · inbound
$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deaae569-f1a2-4a1f-9d4a-da39a1702705 · inbound
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 804c29ac-bac8-4d13-83af-ec00d9c83727 · inbound
Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 122
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae813dbc-8b2a-44ac-b6b5-de339479dfe9 · inbound
Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f91cc1d-8841-40dd-9af8-9bcdc9cc5114 · inbound
CEC-Zero: Chinese Error Correction Solution Based on LLM Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41751754-982f-4da2-806c-b400fd83650d · inbound
Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 451
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3996829b-ccc4-498e-994f-8efd91710467 · inbound
EquivPruner: Boosting Efficiency and Quality in LLM-Based Search via Action Pruning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e8a27f-7ab7-4583-9ee6-bbe4a596520d · inbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aab6bc26-c491-4c70-a393-298260d74354 · inbound
First Finish Search: Efficient Test-Time Scaling in Large Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ec0619b-14a4-4f90-832b-942fe44fa2fa · inbound
ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d28e2eda-1c8d-4a34-b6ab-382e6809f35a · inbound
Large Language Models for Planning: A Comprehensive and Systematic Survey Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 288
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bf720c4-d531-484a-9cea-ec5a7b6a8677 · inbound
Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53db4d41-a0ef-45ab-a84d-865bbc38514c · inbound
Control-R: Towards controllable test-time scaling Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a09a2f29-2492-42e0-b9f1-000907bd094a · inbound
Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f65d5de8-86c7-4087-8689-c8d1d2d00d61 · inbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f17ab33-77d9-4642-8147-261668aec8db · inbound
AI Agent Behavioral Science Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 172
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f1b2dee-12e0-4eea-ac87-090838360a58 · inbound
SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52378c17-6432-4034-b1cb-652c1758abaf · inbound
ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23e05b3c-5c67-4de9-af7d-a2f197b49a0f · inbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dfb5fc4-6623-405c-9e18-0c96c5b66eea · inbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e070fe5c-7e87-41b8-8a10-5bf39df4fad6 · inbound
PLoP: Precise LoRA Placement for Efficient Finetuning of Large Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf48413-f128-401a-9f20-b1603e517bd6 · inbound
Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11e6cd9d-3b1c-4b70-ac25-4c9c5de4798b · inbound
Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5e80796-9eed-4796-bcae-9ef353d7f30d · inbound
CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c03327a-0a26-410a-9f02-e0d0947c75a1 · inbound
ChipSeek: Optimizing Verilog Generation via EDA-Integrated Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation da8eb2be-757c-408d-9572-55982b3900b1 · inbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06bec0d1-228a-4eb2-8616-7e5318d67b2e · inbound
Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 198
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ed8bf8c-fce5-4454-a622-c5bef9720374 · inbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa814ac9-bc2e-474d-9163-cfea9b35fa5e · inbound
Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 576f190c-e985-4613-8133-602f0e0edbac · inbound
Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbf825b4-52e4-4b83-8952-4c4bae846dfd · inbound
Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0f043711-7720-401a-a758-a72b39b21b18 · inbound
The Art of Scaling Reinforcement Learning Compute for LLMs Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b56b40c4-68a7-4348-99d7-7eb4484f2f73 · inbound
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1336d905-1baf-45fd-9289-70be4433e537 · inbound
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f41c982a-01ce-450e-8365-cea84e37ddb8 · inbound
Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1427f7f5-f06b-4302-885a-0aeb295ecd74 · inbound
Can I Have Your Order? Monte-Carlo Tree Search for Slot Filling Ordering in Diffusion Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b0a4382-6d8c-49de-bae7-f6361f8e2a36 · inbound
Online Self-Calibration Against Hallucination in Vision-Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 68714fb5-68aa-476d-9c59-4a3b6fff7f38 · inbound
StoryAlign: Evaluating and Training Reward Models for Story Generation Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c8ef7e11-e612-4411-85d7-be0817289dcd · inbound
A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a3040011-7ea3-44ec-a821-3ec480e3f0aa · inbound
CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 68c07757-b090-4bc8-8122-a6503e8fca4c · inbound
APCD: Adaptive Path-Contrastive Decoding for Reliable Large Language Model Generation Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c55c9fe3-ad48-4d46-99ac-8f5548f67e17 · inbound
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 57f9ef26-61e7-4a7b-9c54-95760f60ccb7 · inbound
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 15b827cc-a7dd-4a11-9f5c-b797760a1816 · inbound
Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 48f329f6-4104-4082-8909-01fa1b90e315 · inbound
Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 64f8b47c-59d7-4d39-a4ba-3ba70e8c492b · inbound
Scalable Token-Level Hallucination Detection in Large Language Models Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6e1d9e44-cb39-4d6f-94dd-d1797095a35d · inbound
ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bae10265-a6c0-4d2c-bc44-db096746d7aa · inbound
ATLAS: Agentic Test-time Learning-to-Allocate Scaling Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6e49eddc-22ef-4fce-a20d-44cbeb608b66 · inbound
Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f6beadf6-5ee6-45d4-b3d0-40f50864e89c · inbound
Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2bc63d21-22b7-432f-b193-237cc353379a · inbound
TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 28117a7f-5f6a-4c38-b65b-444a7c346444 · inbound
APPO: Agentic Procedural Policy Optimization Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e9977467-00f8-4151-8f7b-7bd97cb7a3cf · inbound
APPO: Agentic Procedural Policy Optimization Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be34a2a5-a3d3-41c9-9d02-0bf58118a20d · inbound
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b85efc86-f7b5-42fb-a897-3b8873fc3b52 · inbound
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 824af795-821b-434e-bc1f-1cdba21dd220 · inbound
Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1410f6a1-38bf-4ab8-90ae-624ba2f850b4 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 236
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bd4cebbd-1f7e-4e37-97ef-503b2241cf42 · inbound
Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routing Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0d3dd942-0e67-4127-8a6d-21c116e33c71 · inbound
Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routing Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f526cda6-e761-4871-83c3-6b5e4f3a5928 · inbound
Addressing Over-Refusal in LLMs with Competing Rewards Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f5348cad-0dee-43c0-8c6e-a9a67b0c39fe · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 611ae85f-0e67-41c6-b984-776c7bf72877 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ab0484c-befc-4fdd-877f-07d23aa35314 · inbound
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 111
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaaaf2ff-7240-4b6f-b5d1-0bd167f1baa2 · inbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab2469e5-4b98-4ee1-8904-3346aacf9b3e · inbound
Thought-Level Beam Search for Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42eeda29-89e7-4ce3-853e-7adcccbc75ce · inbound
Thought-Level Beam Search for Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac3caf52-92a7-4a2a-a78f-cce02a3d2cf7 · inbound
ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 166
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42632a1c-3669-4b61-9f89-470f2fcefc82 · inbound
Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.