Pith. sign in

Paper Citation Record · LEDGER

Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2309.17179.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.17179 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:47:31.242429Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aea7076a-c96b-45b0-84c0-e62d2b2c97e4 · inbound

Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations cites this paper.

Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:34:15.902550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T22:34:15.638114Z digest=sha256:30f0ccecfefec7114e4469d3f88799dd1ba0e1ad0a0aea3d0b816bde7024e2b4

Observation a9446a2e-c207-429b-a153-efdfed86db07 · inbound

Improve Mathematical Reasoning in Language Models by Automated Process Supervision cites this paper.

Improve Mathematical Reasoning in Language Models by Automated Process Supervision Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:45.903535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:53:45.878221Z digest=sha256:5c7bde9e7955aa7209273689f12fa61f6fc038d27e735fffa393ad075ad20903

Observation cfbc782a-e2b0-4875-8f83-4673f0a6e7f6 · inbound

QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search cites this paper.

QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:31.242429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:31.242429Z digest=sha256:ab4dbb4b433d3bdee4e3cff11938c6ce925426633cdd91318ffc266d902adb9b

Observation bdd185cc-189c-4585-a351-a8aeff4a8531 · inbound

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition cites this paper.

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:53.332278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:25:53.332278Z digest=sha256:fec3692da575786cb3e486ea49fc68f44c7c37100e94fef7d0d71f22920a5bf1

Observation 4ef57d7c-c201-481b-b7ab-80775d952152 · inbound

Policy Guided Tree Search for Enhanced LLM Reasoning cites this paper.

Policy Guided Tree Search for Enhanced LLM Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:31.753947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:20:31.753947Z digest=sha256:29c498a13f7f35d3b6bf4423a5919a1e2b816b613499bc0f9b258f5734af45d4

Observation e92021b1-0c6e-4369-8774-ec43f7b0ccf9 · inbound

Bag of Tricks for Inference-time Computation of LLM Reasoning cites this paper.

Bag of Tricks for Inference-time Computation of LLM Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:24.496163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:35:24.496163Z digest=sha256:c004068364a784e8ff9f30a87b1bcc062d672b9d3e4942b81ab6381cc548ba04

Observation fb68f2ca-0621-4020-bd19-b463ba31ccd1 · inbound

Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement cites this paper.

Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:31:43.547304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:31:43.494099Z digest=sha256:c8825d19b91450c9840c843ed68a90835d35c4e225307924c03e978a03dce9a4

Observation fb995c4c-c156-4409-a4c3-748bfb0c287a · inbound

QM-ToT: A Medical Tree of Thoughts Reasoning Framework for Quantized Model cites this paper.

QM-ToT: A Medical Tree of Thoughts Reasoning Framework for Quantized Model Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:45:04.025581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T19:44:16.630377Z digest=sha256:6d70b388578b4691935e93e678541aac6f22862600ceefcede40a4484909eaa3

Observation beb2dcd8-ce09-4a45-ac89-9df398fc6e18 · inbound

MMATH: A Multilingual Benchmark for Mathematical Reasoning cites this paper.

MMATH: A Multilingual Benchmark for Mathematical Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.912353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:23:07.912353Z digest=sha256:30a6b080cfc1c3f9b34fdc0a606f2d15de81635f486aeb72a1a68995779227a9

Observation db427559-1d5c-4206-92ed-026d48eec8af · inbound

Can Past Experience Accelerate LLM Reasoning? cites this paper.

Can Past Experience Accelerate LLM Reasoning? Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.952185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.952185Z digest=sha256:81a1aa149ba8418086365a2e4d2887bdb44c94175f4d25468f774883973e73fb

Observation 8769437d-e1a6-40e5-a546-f22ebb8d114b · inbound

How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning cites this paper.

How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:21.438917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:21.438917Z digest=sha256:fee81e28b6bc37038bbd9c23c8905595d8b75326d3eb6c7882cfb6babd31d707

Observation 9ad7cede-ca98-4c61-8f14-ecbbdd2bc5ab · inbound

Structured Pruning for Diverse Best-of-N Reasoning Optimization cites this paper.

Structured Pruning for Diverse Best-of-N Reasoning Optimization Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:23.561140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:23.561140Z digest=sha256:412ea4809d963c067e40941223bc12d01cfca48f9ec4d9caba88f699502af8e9

Observation 650cc220-635d-40a7-87eb-790fc13501dd · inbound

Kinetics: Rethinking Test-Time Scaling Laws cites this paper.

Kinetics: Rethinking Test-Time Scaling Laws Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:30:34.180416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:30:34.180416Z digest=sha256:4af3a44ad62d99b640cd27dbb202c164a8a50cef1cade54ee94e5d51422a07f5

Observation 4234ff1d-a7b8-47fb-b9ec-4092ced0f56f · inbound

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism cites this paper.

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:18.616592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:18.616592Z digest=sha256:5344e55be4c3307d1e05bccd6a966add876764986bca116000e504316d4ff8df

Observation e67df85b-d344-4f51-8833-ae66fee42643 · inbound

Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty cites this paper.

Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:38:09.465060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:38:09.465060Z digest=sha256:02ba45893b0cd4ca602a60a4cc11f7aeda10a84fc42e0e440617ba8bb35ee3ce

Observation c54a79fa-7e5e-42a8-85a3-f7797b28d389 · inbound

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search cites this paper.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.750335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.750335Z digest=sha256:bd8140e026c5ffd6f8e577089cc459016194871a7e88dcb398c101666d0674dc

Observation 2c47a8aa-1244-4b5e-bdc4-7b2deac56277 · inbound

KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality cites this paper.

KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T07:37:08.989689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T07:33:08.719028Z digest=sha256:7914c4f3949bb9f79dfa66a7e88aa5ebd15632c981c9551206b203d8ea82d32c

Observation 6b7d3b68-72f8-48e2-94ec-d7f9494fb526 · inbound

Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments cites this paper.

Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:14.714930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:14.714930Z digest=sha256:54c5928436336395ce289cc5a334c4dbb641099738cb0b2200c942aadc6e2765

Observation 5eeb45b6-3c6d-4900-93a7-03e4147598f1 · inbound

Reasoning in machine vision by learning fast and slow thinking cites this paper.

Reasoning in machine vision by learning fast and slow thinking Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:18.579086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:18.579086Z digest=sha256:ef3c6d10bc305e3230a177763ef5dc8dcc5217ce6227408e537162f132aae204

Observation 2a1f5f4c-bd4b-4007-af55-873c9248df71 · inbound

Data Diversification Methods In Alignment Enhance Math Performance In LLMs cites this paper.

Data Diversification Methods In Alignment Enhance Math Performance In LLMs Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:26.040552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:43:26.040552Z digest=sha256:727c33afe0b66670348ee5a072beb21280fd7eef07f3677b18189a68d747626c

Observation e1d90dd1-c705-42d0-baef-f00aca515744 · inbound

Enhancing Test-Time Scaling of Large Language Models with Hierarchical Retrieval-Augmented MCTS cites this paper.

Enhancing Test-Time Scaling of Large Language Models with Hierarchical Retrieval-Augmented MCTS Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:12.161024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:28:12.161024Z digest=sha256:43d195110ed4319ea8f4cd9c51eccf86e1ca3a72a69c979d7cab596935e46ba2

Observation 60d6e2b2-8357-4380-9b77-2c4b9d9fc6b8 · inbound

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models cites this paper.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.209380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.209380Z digest=sha256:922ab2f9a2d62b17ba9165e53d2127e3a80bde9f1ac95b7b3e81447b25b36a1f

Observation 2a39a040-6db1-4621-9b31-ed1e1b145f3a · inbound

It's Not That Simple. An Analysis of Simple Test-Time Scaling cites this paper.

It's Not That Simple. An Analysis of Simple Test-Time Scaling Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:09:04.281051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:09:04.281051Z digest=sha256:a6c45c96a3067dbc1b81d5ff1257d921fc938d47a59967e28ee4262a02e8a4f2

Observation 6bacaa48-b2da-445a-9bf2-4aa70ba0a491 · inbound

LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra cites this paper.

LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:28:42.237425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:28:42.237425Z digest=sha256:a4223fab284fa380fa22ec51bd40ed57041e1a0ee7bceb0d2cd335504176739e

Observation 84845823-462d-4b71-adb7-056d11f47575 · inbound

Let's Revise Step-by-Step: A Unified Local Search Framework for Code Generation with LLMs cites this paper.

Let's Revise Step-by-Step: A Unified Local Search Framework for Code Generation with LLMs Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:46.343818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:09:46.343818Z digest=sha256:0c6ff0d02702220dde485265656e7d117489699ee864369353105fb1569d2135

Observation 14285457-87e1-4d05-a671-243df4b6e228 · inbound

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time cites this paper.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T11:57:16.716582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:57:16.716582Z digest=sha256:522037f8398de52efcd93beeaf1fb73eff8a7278681d6bb0cc1d21b9c977da37

Observation 398b7170-af66-43e5-b8c6-f39c12d228aa · inbound

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models cites this paper.

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-04T18:53:02.654971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:53:02.654971Z digest=sha256:a8db8efbcd18188e3a00794f468815467eeb8b8a93985a385ca9eefce03a99cf

Observation 3f246477-bd08-4ba3-8922-04f6a7bc2519 · inbound

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning cites this paper.

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T12:52:26.719841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:52:26.719841Z digest=sha256:46acd5d0a6c1fad6317262eff16dd2bdb01c0ca27e5a1831bc09e5e9686929b5

Observation 028a9c3e-e00f-4664-bf40-d95bf6581261 · inbound

DiffCoT: Diffusion-styled Chain-of-Thought Reasoning in LLMs cites this paper.

DiffCoT: Diffusion-styled Chain-of-Thought Reasoning in LLMs Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:21:07.901256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T17:19:58.334098Z digest=sha256:4784514be6746e12ad7aaaa6bb9c5e69bd5d037c81ccf5aaacbd5976ff9f3c6a

Observation 788d1018-d4fc-4a17-8e50-9b4dfe58040a · inbound

Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models cites this paper.

Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:22:55.276086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T13:21:36.606855Z digest=sha256:4f7c3507c7ba271ec18ff39f3769b3992e3ee2d4678c14b46af1e02d52284a07

Observation 0754fcb6-bf24-483b-9a0b-0deb0bafddd8 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:26.052180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:aff34dbd393d3904d3e04528b6107cfe181db1e0e9296b18021f9b446c99b516

Observation a5728366-d70a-4ffb-9660-d96556437e88 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.542090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.542090Z digest=sha256:582fca25153097ea269e1321293d7f5ca6803f91b2cf0df1363f5625abd47754

Observation 2a2df38d-1ab1-4573-b6f5-5d192859296b · inbound

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning cites this paper.

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:52.722018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:40:41.642852Z digest=sha256:b7c0b7641632b36a572f907e85502dbed742f33695731ef8abe01268b4fffcfd

Observation 761dd4d6-29a3-4eb1-8972-b07c89f37a19 · inbound

Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models cites this paper.

Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:20:52.418936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:36:01.200412Z digest=sha256:c1902cee881068c371a9889313169f61e58585c2ba817448bba06b6e3e9a88af

Observation 66be9b4c-024e-473f-ae01-4a23529ff515 · inbound

Evaluation-driven Scaling for Scientific Discovery cites this paper.

Evaluation-driven Scaling for Scientific Discovery Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:26:05.589442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:39:52.204043Z digest=sha256:cb2d5a079806155a1908c571cdb773919d4af10ba3edcb55b7bdf9421ab6c453

Observation 93f3611e-068a-4ac9-b5a0-fb844a575fd3 · inbound

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding cites this paper.

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:31:08.919771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T16:29:05.186607Z digest=sha256:d83fe877594375ac44fe4b61ff29bb88a5822acb453d0515064748d805d13e75

Observation 64f389a3-ece9-4f54-8a47-654524236766 · inbound

SAT: Sequential Agent Tuning for Coordinator Free Plug and Play Multi-LLM Training with Monotonic Improvement Guarantees cites this paper.

SAT: Sequential Agent Tuning for Coordinator Free Plug and Play Multi-LLM Training with Monotonic Improvement Guarantees Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:38:42.293790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T09:34:33.271537Z digest=sha256:38069ec0f0a57b99b2257c2fe6dbd8a79b3d1da8aea299a0eeb7f481c1045b07

Observation 5210ee3d-d234-4f0a-a335-3a0f50e898ff · inbound

Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning cites this paper.

Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:06.956367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:16:00.499779Z digest=sha256:885ab17c64a59d7206b76f3ad7e9cd6128854cbda5095167a9109eedb95e588c

Observation ff54ab4a-a1f1-47c5-8e7f-8ebbb47b3abf · inbound

CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation cites this paper.

CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:25:58.264821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T02:28:20.366674Z digest=sha256:9603927f12513ad379ca0fa2e790318c69b9984d5378fa129979baa7ca7a34b5

Observation ada3a1b9-1026-41b1-9cad-17bf54fbcee6 · inbound

V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning cites this paper.

V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:41:42.171923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:03:01.612516Z digest=sha256:8a084a1856ddbfc9b381c53618574e5a6685d453a88ea2524cd50bb7957c20a7

Observation cb9cf998-12a0-426d-826e-0ab8d015ecb7 · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:30.023684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:51:52.375703Z digest=sha256:2e94d0b7e8107b1fad6cda061ac3e17992bb965f492c9d349fbb0ba82da07b71

Observation 4f34326e-528b-4d96-9ad3-1fc96a76fa5d · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:15:03.262296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T05:11:32.053440Z digest=sha256:9b309be22f041b2ac938f23f96f98f23e3fe222e9aff693636a3eaeee1ff3fa6

Observation 4077066b-ce29-43c3-ad50-fc741a014a0d · inbound

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling cites this paper.

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:56:30.042468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T10:25:10.559953Z digest=sha256:8c23ca08fe1d480299612772432b24e16a203d0e95b5ac3a424778cda66957a4

Observation 9c889228-de92-44ce-852a-8375a79b6547 · inbound

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces cites this paper.

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.956800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T05:46:26.938277Z digest=sha256:05a08012eb8d8b295e3a0da6d0b1d4ec48054a7348a1c404042b145aa45b26bd

Observation 1cc95fe5-48ec-40d7-879e-988120a46c37 · inbound

DICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination cites this paper.

DICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.900391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T19:58:32.016341Z digest=sha256:535902fd3f8679f56392839a0cadb326fdee5aafb23f2b0709e5fe59893eba50

Observation 3c60c9d5-bffe-4fc4-937f-e71ab1969215 · inbound

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning cites this paper.

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.750084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T13:55:35.363377Z digest=sha256:d489c8923709ed9d39a74caabdd2972de7e5a801bcbc3789d4716110e3fbca2a

Observation e84f6b6a-14d9-492c-acf4-7991eabb59ea · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:37:49.320578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T10:21:55.485624Z digest=sha256:a480743de85bd676063ad6bf7c6aebb967309cf9196c86338c51782ee5d29920

Observation 956695dc-6775-41b4-930a-08ea6cbae14b · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:26.571855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:26.571855Z digest=sha256:c8ca91a85934bafe9de63315ef1c5b852720959b776102789d7d2f38a277c4bc

Observation 26e4a92d-54d2-45ef-982a-7632f69be8e3 · inbound

Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies cites this paper.

Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:44:59.784865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T18:38:48.512777Z digest=sha256:249e7bbf88d2345d433698949409d9bdefe8799cf3fa97c56c98cc8e3da462d4

Observation c83a8d58-2ff5-4cae-ac64-129a2da31a2c · inbound

DecompRL: Solving Harder Problems by Learning Modular Code Generation cites this paper.

DecompRL: Solving Harder Problems by Learning Modular Code Generation Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:38:39.805075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-03T16:30:34.793328Z digest=sha256:b5e1e7f05c33c598f5799f7f89c2483afb7c45a7d377633f434f810a0ea30d22