REVIEW 48 references
TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems
T0 review · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A new benchmark parses LLM reasoning traces on Go life-and-death problems into search trees and shows that search organization, not token volume, distinguishes stronger reasoners.
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
The authors convert each model's written reasoning trace into a search tree: which moves were considered, in what order, how much effort was spent on each branch, and which branches were judged winning or losing. From these trees they compute Search Efficiency (SearchE), which combines three signals: effort wasted on wrong branches, whether the correct move was examined first, and how early the correct move appeared. They compare this with Token Efficiency, which is simply accuracy divided by the number of tokens used.
On 600 puzzles across easy, medium, and hard tiers, most LLMs solved only a fraction of problems. Stronger models such as Gemini-3.1-Pro did much better on easy puzzles but still collapsed on hard open-ended ones. The authors find that models with better SearchE also have higher accuracy, while models that produce long chains of thought do not necessarily search better. They contrast this with KataGo, a neural-guided Go program that solves most easy puzzles with the same budget, and with plain MCTS, which behaves like shallow undirected search. The paper argues that evaluating how models allocate reasoning effort across branches is a missing dimension in LLM benchmarking.
Extended reading notes
Core claim
Current LLMs remain far from stable tsumego solving, and Search Efficiency, a process-level score computed from wrong-branch waste, first-hit behavior, and correct-candidate rank, aligns with answer accuracy more consistently than token efficiency. As stated in §5.3: 'SearchEisastrongerprocesssignalthanTokenE.' If true, this establishes search organization as a measurable, distinct dimension of LLM reasoning beyond answer accuracy and token cost.
Load-bearing premise
The pipeline assumes that visible CoT traces faithfully externalize the model's internal search organization. The paper acknowledges in Appendix L: 'Search trees are extracted from observable CoT traces, so our process metrics describe the visible reasoning path rather than the model's latent computation.' If proprietary summaries or omitted branches misrepresent internal search, then SearchE comparisons measure text organization, not the claimed search ability. This premise enters at §4.1 and is load-bearing for every search metric in the paper.
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
free parameters (2)
- SearchE weights =
0.5/0.3/0.2
- TokenE scaling constant =
1000
assumptions (3)
- domain assumption Visible CoT traces faithfully externalize the model's internal search organization
- domain assumption Gemini-2.5-Pro extraction preserves metric-relevant tree structure with 93-98% agreement
- domain assumption First-move hit rate is a valid measure of tsumego solving
Cite this review
Pith. "Pith review of TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems." pith.science (2026). https://pith.science/paper/552DAJ6C
@misc{pith2026260813221,
author = {Pith},
title = {Pith review of: TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems},
year = {2026},
howpublished = {\url{https://pith.science/paper/552DAJ6C}},
note = {Machine review of arXiv:2608.13221}
}
read the original abstract
The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and allocate reasoning resources--that is, how they organize search. Prior process-level methods focus on the coherence and redundancy of chain-of-thought (CoT), and most benchmark tasks have a single objective solvable by static capabilities such as derivation and tool use, leaving search organization unmeasured. We introduce TsuGO, a process-level reasoning benchmark for evaluating Search Efficiency in LLM reasoning through Go life-and-death problems. These problems provide closed and verifiable solution spaces with an inherent adversarial structure, making candidate generation, response checking, branch comparison, and backtracking necessary parts of reasoning rather than incidental trace patterns. By constraining the solution space, TsuGO disentangles domain knowledge from search organization, parses CoT into a structured search tree, and reports Search Efficiency together with Token Efficiency and other diagnostic metrics and visualizations. Experiments show that current LLMs remain far from stable tsumego solving: stronger models succeed by finding the correct candidate earlier and sustaining effort on productive branches, but most models still behave much closer to unguided search algorithms than to neural-guided KataGo. Longer CoT or higher Token Efficiency does not necessarily imply better search. Our results identify search organization and reasoning-resource allocation as missing dimensions in LLM reasoning evaluation.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
arXiv preprint arXiv:2110.14168 , year=
Training verifiers to solve math word problems , author=. arXiv preprint arXiv:2110.14168 , year=
-
[2]
Measuring mathematical problem solving with the
Hendrycks, Dan and Burns, Collin and Kadavath, Saurav and Arora, Akul and Basart, Steven and Tang, Eric and Song, Dawn and Steinhardt, Jacob , journal=. Measuring mathematical problem solving with the
-
[3]
NeurIPS , year=
Chain-of-thought prompting elicits reasoning in large language models , author=. NeurIPS , year=
- [4]
- [5]
-
[6]
Li, Siyi and Shi, Jiajun and Ni, Shiwen and Zhang, Ge and Li, Shuaimin and Wang, Shijian and Wen, Zhoufutu and Li, Yizhi and Alinejad-Rokny, Hamid and Liu, Jiaheng and Yang, Min and Huang, Wenhao , journal=. CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in
-
[7]
ReEfBench: Quantifying the Reasoning Efficiency of
Fu, Zhizhang and Gu, Yuancheng and Hu, Chenkai and Liu, Hanmeng and Zhang, Yue , journal=. ReEfBench: Quantifying the Reasoning Efficiency of
-
[8]
Han, Simeng and Schoelkopf, Hailey and Zhao, Yilun and Qi, Zhenting and Riddell, Martin and others , booktitle=
Show all 48 references
-
[9]
International Conference on Learning Representations , year=
Language Models are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought , author=. International Conference on Learning Representations , year=
-
[10]
Lin, Bill Yuchen and Le Bras, Ronan and Richardson, Kyle and Sabharwal, Ashish and Poovendran, Radha and others , journal=
-
[11]
Liu, Yiwei and Li, Yucheng and Li, Xiao and Cheng, Gong , journal=
-
[12]
Golovneva, Olga and Chen, Moya Peng and Poff, Spencer and Corredor, Martin and Zettlemoyer, Luke and Fazel-Zarandi, Maryam and Celikyilmaz, Asli , booktitle=
-
[13]
2023 , doi=
Prasad, Archiki and Saha, Swarnadeep and Zhou, Xiang and Bansal, Mohit , booktitle=. 2023 , doi=
2023
-
[14]
Inference-Time Computations for
Parashar, Shubham and Olson, Blake and Khurana, Sambhav and Li, Eric and Ling, Hongyi and Caverlee, James and Ji, Shuiwang , journal=. Inference-Time Computations for
-
[15]
, journal=
Gandhi, Kanishk and Chakravarthy, Ayush and Singh, Anikait and Lile, Nathan and Goodman, Noah D. , journal=. Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective
-
[16]
Mastering the game of
Silver, David and Huang, Aja and Maddison, Chris J and Guez, Arthur and Sifre, Laurent and Van Den Driessche, George and Schrittwieser, Julian and Antonoglou, Ioannis and Panneershelvam, Veda and Lanctot, Marc and others , journal=. Mastering the game of
-
[17]
Mastering the game of
Silver, David and Schrittwieser, Julian and Simonyan, Karen and Antonoglou, Ioannis and Huang, Aja and Guez, Arthur and Hubert, Thomas and Baker, Lucas and Lai, Matthew and Bolton, Adrian and others , journal=. Mastering the game of
-
[18]
Bandit based
Kocsis, Levente and Szepesv. Bandit based. Machine Learning: ECML 2006 , pages=. 2006 , publisher=
2006
-
[19]
, journal=
Wu, David J. , journal=. Accelerating self-play learning in
-
[20]
arXiv preprint , year=
Do Not Think That Much for 2+3! Overthinking with Chain-of-Thought , author=. arXiv preprint , year=
-
[21]
arXiv preprint , year=
Qwen3 Technical Report , author=. arXiv preprint , year=
- [22]
-
[23]
arXiv preprint , year=
Gemini: A family of highly capable multimodal models , author=. arXiv preprint , year=
- [24]
- [25]
- [26]
-
[27]
2026 , howpublished=
2026
-
[28]
2026 , howpublished=
Gemini 3.1 Pro: A Smarter Model for Your Most Complex Tasks , author=. 2026 , howpublished=
2026
-
[29]
arXiv preprint arXiv:2302.13071 , year=
Chess as a Testbed for Language Model State Tracking , author=. arXiv preprint arXiv:2302.13071 , year=
-
[30]
ICLR , year=
Emergent world representations: Exploring a sequence model trained on a synthetic task , author=. ICLR , year=
-
[31]
Think you have solved question answering? Try
Clark, Peter and Cowhey, Isaac and Etzioni, Oren and Khot, Tushar and Sabharwal, Ashish and Schoenick, Carissa and Tafjord, Oyvind , journal=. Think you have solved question answering? Try
-
[32]
TMLR , year=
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models , author=. TMLR , year=
-
[33]
ICLR , year=
Let's verify step by step , author=. ICLR , year=
-
[34]
2025 , doi=
Song, Mingyang and Su, Zhaochen and Qu, Xiaoye and Zhou, Jiawei and Cheng, Yu , booktitle=. 2025 , doi=
2025
-
[35]
ICML , year=
AlphaZero-like tree-search can guide large language model decoding and training , author=. ICML , year=
-
[36]
Stream of search (
Gandhi, Kanishk and Lee, Denise and Grand, Gabriel and Liu, Muxin and Cheng, Winson and Suhr, Alane and Goodman, Noah D , journal=. Stream of search (
-
[37]
Advances in Neural Information Processing Systems , volume=
Tree of Thoughts: Deliberate Problem Solving with Large Language Models , author=. Advances in Neural Information Processing Systems , volume=
-
[38]
International Conference on Learning Representations , year=
Self-Consistency Improves Chain of Thought Reasoning in Language Models , author=. International Conference on Learning Representations , year=
-
[39]
Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Graph of Thoughts: Solving Elaborate Problems with Large Language Models , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=. 2024 , doi=
2024
-
[40]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=
Reasoning with Language Model is Planning with World Model , author=. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=
2023
-
[41]
arXiv preprint arXiv:2310.04406 , year=
Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models , author=. arXiv preprint arXiv:2310.04406 , year=
-
[42]
Zheng, Chujie and Zhang, Zhenru and Zhang, Beichen and Lin, Runji and Lu, Keming and Yu, Bowen and Liu, Dayiheng and Zhou, Jingren and Lin, Junyang , booktitle=
-
[43]
Luo, Haotian and Shen, Li and He, Haiying and Wang, Yibo and Liu, Shiwei and Li, Wei and Tan, Naiqiang and Cao, Xiaochun and Tao, Dacheng , journal=
-
[44]
Li, Zhiyuan and Chang, Yi and Wu, Yuan , journal=
-
[45]
Shen, Chenhui and others , journal=
-
[46]
arXiv preprint arXiv:2503.16419 , year=
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models , author=. arXiv preprint arXiv:2503.16419 , year=
-
[47]
arXiv preprint arXiv:2412.15797 , year=
Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning , author=. arXiv preprint arXiv:2412.15797 , year=
-
[48]
Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of
Ma, Yichuan and Li, Linyang and Chen, Yongkang and Li, Peiji and Ye, Jiasheng and Guo, Qipeng and Lin, Dahua and Chen, Kai , booktitle=. Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of. 2025 , doi=
2025
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.