Pith. sign in

REVIEW 4 major objections 4 minor 48 references

What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles

T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read TurtleSoup-Bench: LLMs trail humans in imaginative reasoning

desk verdict A plausible benchmark paper for Turtle Soup puzzles, but the abstract alone can't support the central claims about imaginative reasoning. read the letter →

arxiv 2508.10358 v1 pith:7HS36GOI submitted 2025-08-14 cs.AI

classification cs.AI
keywords imaginativereasoningTurtleSoupbenchmarklargelanguagemodelsinteractiveevaluationMosaic-Agenthypothesisrevisionbilingual
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to show that imaginative reasoning—the active process of forming, testing, and revising hypotheses when information is scarce—can be measured with Turtle Soup puzzles, and that today's large language models are far weaker at it than humans. To do this, it builds TurtleSoup-Bench, a set of 800 bilingual puzzles, pairs it with Mosaic-Agent, an agent that interacts with a puzzle by asking questions and making guesses, and scores answers on logical consistency, detail completion, and conclusion alignment. Experiments with leading LLMs report clear capability limits, repeating failure patterns, and a significant human–model gap. If the paper is right, static benchmarks have been missing a distinct reasoning ability, and interactive, exploratory evaluation belongs on the LLM testing agenda.

What carries the argument

The central objects are TurtleSoup-Bench, a collection of 800 Turtle Soup puzzles (a game in which one player knows a bizarre scenario and others reconstruct it through yes/no questions); Mosaic-Agent, an agent wrapper that lets a language model act as the questioner and guesser; and a three-part evaluation protocol that scores outputs for logical consistency, detail completion, and conclusion alignment. The machinery works as a unit: the puzzles supply information-sparse scenarios, the agent supplies the interactive loop, and the protocol turns open-ended puzzle-solving into comparable scores across models and humans.

What would settle it

Run the same 800 puzzles through a head-to-head study in which humans and leading LLMs use the identical Mosaic-Agent interaction budget; if any LLM matches or exceeds human scores on all three dimensions, the paper's claimed significant gap would be falsified. A second check: if human raters' independent judgments of imaginative reasoning do not track the protocol's scores, the operationalization collapses.

Watch

Extended reading notes

Core claim

The paper claims that the combination of a large interactive puzzle set and a structured evaluation protocol exposes a measurable deficit: leading large language models cannot match human performance when they must actively seek information before forming a final explanation. On the paper's own terms, the discovery is not just that models make mistakes, but that their mistakes cluster into identifiable patterns under the three evaluated dimensions—logical consistency, detail completion, and conclusion alignment. The benchmark, the agent, and the protocol together provide a large-scale bilingual instrument for observing this failure.

Load-bearing premise

The load-bearing premise is that Turtle Soup puzzles plus the three evaluation dimensions (logical consistency, detail completion, conclusion alignment) validly capture 'imaginative reasoning' in information-sparse environments; the abstract offers no independent validation of that mapping.

Editorial extensions

If this is right

  • If the gap is real, static question-answering benchmarks understate what LLMs cannot do; interactive information-seeking tasks must be part of capability testing.
  • Model behavior in the question-asking phase becomes a first-class object of study: how a model chooses questions likely predicts how well it solves the puzzle.
  • The bilingual 800-puzzle resource makes cross-lingual comparison of imaginative reasoning possible at scale.
  • The same protocol can be reused to measure progress: an LLM that closes the gap would have to improve at hypothesis revision, not just memorized knowledge.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The puzzle format could be recast as a controlled measure of information gain: one could score each question by how much it reduces the space of possible stories, connecting the benchmark to optimal-questioning theory.
  • The three evaluation dimensions are scored on final outputs; a natural extension is to model the revision process itself, such as whether a model updates its hypothesis after a 'no' answer, which the current protocol may only capture indirectly.
  • Because Turtle Soup puzzles are culturally embedded, the bilingual design may reveal whether imaginative reasoning transfers across languages or depends on the cultural content of the story.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript, as provided, consists solely of the abstract for arXiv:2508.10358. It announces TurtleSoup-Bench, a bilingual, interactive benchmark of 800 Turtle Soup puzzles; Mosaic-Agent, an agent for assessing LLM performance; and a multi-dimensional evaluation protocol measuring logical consistency, detail completion, and conclusion alignment. The abstract claims that experiments with leading LLMs reveal clear capability limits, common failure patterns, and a significant performance gap compared to humans. No methods, data, scoring rubrics, example puzzles, statistical analyses, or validation are presented in the available text.

Significance. If substantiated, the work could offer a valuable new evaluation paradigm for exploratory and imaginative reasoning in LLMs, with practical implications for benchmark design and for understanding LLM behavior in information-sparse settings. The bilingual and interactive features are potentially distinguishing contributions. However, none of these contributions are currently verifiable because the manuscript provides no supporting evidence or methodological detail. The significance assessment is therefore conditional on the missing content being supplied and validated.

major comments (4)
  1. [Abstract / overall manuscript] The central claims—introduction of the first large-scale bilingual interactive benchmark, the proposed Mosaic-Agent, and the reported experimental findings—are entirely unsupported in the available text. No benchmark construction process, puzzle sampling strategy, agent architecture, evaluation protocol details, or experimental results are provided. As a journal submission, this is insufficient to assess soundness. This is load-bearing because the abstract's conclusions rest entirely on this missing material.
  2. [Abstract / evaluation protocol] The three evaluation dimensions (logical consistency, detail completion, conclusion alignment) are asserted to measure 'imaginative reasoning' without any construct validity evidence. No correlation with human holistic judgments, established reasoning/creativity measures, or inter-rater reliability is reported. The reported 'significant performance gap compared to humans' could equally reflect the specific scoring rubric rather than a difference in imaginative reasoning. The manuscript needs to justify these dimensions as a valid operationalization and rule out alternative interpretations such as generic textual coherence or puzzle-solving skill.
  3. [Abstract / interactive and bilingual setup] The interactive and bilingual aspects introduce potential confounds that are not addressed in the abstract. For example, LLM performance could depend on the question-asking policy or on translation quality rather than on imaginative reasoning. The manuscript must describe how these confounds are controlled or analyzed (e.g., ablations, human baseline protocols, translation checks). Without such controls, the capability conclusions are ambiguous.
  4. [Abstract / Mosaic-Agent] Mosaic-Agent is introduced as a novel agent for assessing LLM performance, but no information is given about how it interacts with the puzzles, what information it receives, or how its behavior is scored. If the agent itself shapes the evaluation (e.g., by generating questions or interim hypotheses), then the reported performance may depend on the agent's design choices. This needs explicit description and sensitivity analysis before the results can be interpreted.
minor comments (4)
  1. [Overall] The paper contains no references to related benchmarks or prior work, making it impossible to verify the 'first' claim or situate the contribution. Please provide a proper related-work section.
  2. [Abstract] No example puzzles or answer rubrics are shown, which would be essential for assessing the benchmark's quality and face validity.
  3. [Overall] The term 'imaginative reasoning' is not defined precisely; a formal or operational definition is needed to link the benchmark to the intended construct.
  4. [Abstract] Statistical details are absent: number of models, repetitions, significance tests, effect sizes, and variance. Reporting these is necessary for evaluating the reliability of the claimed human–LLM gap.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the benchmark/agent/evaluation claims are empirical and descriptive; construct-validity concerns are not circularity.

full rationale

This paper is an empirical benchmark paper rather than a formal derivation. The abstract describes constructing a benchmark (TurtleSoup-Bench), an agent (Mosaic-Agent), and an evaluation protocol (logical consistency, detail completion, conclusion alignment), then reports experimental results on LLMs. There is no equation or fitted parameter that is later renamed as a prediction; no result is derived from the definition of the benchmark by construction. The claim that LLMs show capability limits and a performance gap versus humans is an empirical report on a bespoke evaluation setting, not a mathematical consequence of the benchmark's definition. The authors do design the benchmark, agent, and protocol themselves, but that alone is not circular: benchmark papers routinely use author-designed tasks, and circularity would require the conclusion to reduce to the design choice (e.g., defining 'imaginative reasoning' as whatever the protocol measures and then asserting the protocol measures imaginative reasoning). The abstract makes no such definitional closure or external-validity claim. The skeptical concern about construct validity—whether the three dimensions truly measure imaginative reasoning rather than generic puzzle-solving—is a legitimate scientific critique but falls under correctness/validity, not derivation circularity. With no equations, no fitted inputs, and no load-bearing self-citations visible in the abstract, the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 1 invented entities

No free parameters are fitted because the paper does not present a mathematical model. The main assumptions are that the puzzle format and the custom evaluation protocol validly measure the target construct. Mosaic-Agent is an introduced component rather than a discovered entity.

assumptions (3)
  • domain assumption Turtle Soup puzzles require and elicit imaginative reasoning.
    The benchmark's validity relies on the assumption that the puzzle format is a faithful probe of imaginative reasoning in information-sparse settings, as stated in the abstract's motivation.
  • domain assumption The three evaluation dimensions (logical consistency, detail completion, conclusion alignment) capture the quality of imaginative reasoning.
    The protocol is introduced without external validation, so it presumes these dimensions are the correct ones.
  • domain assumption Human performance on the benchmark serves as a valid reference for comparison.
    The abstract reports a 'significant performance gap compared to humans,' but does not describe how human ground truth was established or whether human raters were consulted.
invented entities (1)
  • Mosaic-Agent
    purpose: A novel agent designed to assess LLM performance on TurtleSoup-Bench.
    This is a new component introduced by the paper, described only at a high level in the abstract. No external validation or independent evaluation of the agent is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles." pith.science (2026). https://pith.science/paper/7HS36GOI

@misc{pith2026250810358,
  author       = {Pith},
  title        = {Pith review of: What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7HS36GOI}},
  note         = {Machine review of arXiv:2508.10358}
}
read the original abstract

We investigate the capacity of Large Language Models (LLMs) for imaginative reasoning--the proactive construction, testing, and revision of hypotheses in information-sparse environments. Existing benchmarks, often static or focused on social deduction, fail to capture the dynamic, exploratory nature of this reasoning process. To address this gap, we introduce a comprehensive research framework based on the classic "Turtle Soup" game, integrating a benchmark, an agent, and an evaluation protocol. We present TurtleSoup-Bench, the first large-scale, bilingual, interactive benchmark for imaginative reasoning, comprising 800 turtle soup puzzles sourced from both the Internet and expert authors. We also propose Mosaic-Agent, a novel agent designed to assess LLMs' performance in this setting. To evaluate reasoning quality, we develop a multi-dimensional protocol measuring logical consistency, detail completion, and conclusion alignment. Experiments with leading LLMs reveal clear capability limits, common failure patterns, and a significant performance gap compared to humans. Our work offers new insights into LLMs' imaginative reasoning and establishes a foundation for future research on exploratory agent behavior.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

48 extracted references · 36 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Anthropic, C. 2025. 3.7 sonnet and claude code

  4. [4]

    Bai, G.; Liu, J.; Bu, X.; He, Y.; Liu, J.; Zhou, Z.; Lin, Z.; Su, W.; Ge, T.; Zheng, B.; and Ouyang, W. 2024. MT -Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volum...

  5. [5]

    Binz, M.; and Schulz, E. 2023. Using cognitive psychology to understand GPT-3. Proceedings of the National Academy of Sciences, 120(6): e2218523120

  6. [6]

    T.; Li, Y.; Lundberg, S.; et al

    Bubeck, S.; Chadrasekaran, V.; Eldan, R.; Gehrke, J.; Horvitz, E.; Kamar, E.; Lee, P.; Lee, Y. T.; Li, Y.; Lundberg, S.; et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4

  7. [7]

    Cai, Y.; Gu, Z.; Du, Z.; Ye, Z.; Cao, S.; Xu, Y.; Feng, H.; and Chen, P. 2025. MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments. arXiv:2501.01652

  8. [8]

    Comanici, G.; Bieber, E.; Schaekermann, M.; Pasupat, I.; Sachdeva, N.; Dhillon, I.; Blistein, M.; Ram, O.; Zhang, D.; Rosen, E.; et al. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities

Show all 48 references
  1. [9]

    Curvo, P. M. P. 2025. The Traitors: Deception and Trust in Multi-Agent Language Model Simulations. arXiv:2505.12923

  2. [10]

    Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024. The llama 3 herd of models

  3. [11]

    Eriksson, M.; Purificato, E.; Noroozian, A.; Vinagre, J.; Chaslot, G.; Gomez, E.; and Fernandez-Llorca, D. 2025. Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation. arXiv:2502.06559

  4. [12]

    Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P.; Bi, X.; et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning

  5. [13]

    Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021. Measuring Massive Multitask Language Understanding. In International Conference on Learning Representations

  6. [14]

    Hong, S.; Zhuge, M.; Chen, J.; Zheng, X.; Cheng, Y.; Wang, J.; Zhang, C.; Wang, Z.; Yau, S. K. S.; Lin, Z.; Zhou, L.; Ran, C.; Xiao, L.; Wu, C.; and Schmidhuber, J. 2024. Meta GPT : Meta Programming for A Multi-Agent Collaborative Framework. In The Twelfth International Confer...

  7. [15]

    Hsia, J.; Pruthi, D.; Singh, A.; and Lipton, Z. 2024. Goodhart ' s Law Applies to NLP ' s Explanation Benchmarks. In Graham, Y.; and Purver, M., eds., Findings of the Association for Computational Linguistics: EACL 2024, 1322--1335. St. Julian ' s, Malta: Association for Compu...

  8. [16]

    P.; Perelman, A.; Ramesh, A.; Clark, A.; Ostrow, A.; Welihinda, A.; Hayes, A.; Radford, A.; et al

    Hurst, A.; Lerer, A.; Goucher, A. P.; Perelman, A.; Ramesh, A.; Clark, A.; Ostrow, A.; Welihinda, A.; Hayes, A.; Radford, A.; et al. 2024. Gpt-4o system card

  9. [17]

    Kidd, C.; and Hayden, B. Y. 2015. The psychology and neuroscience of curiosity. Neuron, 88(3): 449--460

  10. [18]

    Lan, Y.; Hu, Z.; Wang, L.; Wang, Y.; Ye, D.; Zhao, P.; Lim, E.-P.; Xiong, H.; and Wang, H. 2024. LLM -Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay. In Al-Onaizan, Y.; Bansal, M.; and Chen, Y.-N., eds., Proceedings of the 2024 Conference...

  11. [19]

    Li, X.; Bao, K.; Ma, Y.; Li, M.; Wang, W.; Men, R.; Zhang, Y.; Feng, F.; Liu, D.; and Lin, J. 2025. MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation. arXiv:2505.17123

  12. [20]

    Liang, T.; He, Z.; tse Huang, J.; Wang, W.; Jiao, W.; Wang, R.; Yang, Y.; Tu, Z.; Shi, S.; and Wang, X. 2023. Leveraging Word Guessing Games to Assess the Intelligence of Large Language Models. arXiv:2310.20499

  13. [21]

    Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. 2023. G -Eval: NLG Evaluation using Gpt-4 with Better Human Alignment. In Bouamor, H.; Pino, J.; and Bali, K., eds., Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2511--2522. Sin...

  14. [22]

    Ma, C.; Zhang, J.; Zhu, Z.; Yang, C.; Yang, Y.; Jin, Y.; Lan, Z.; Kong, L.; and He, J. 2024. AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents. In Globerson, A.; Mackey, L.; Belgrave, D.; Fan, A.; Paquet, U.; Tomczak, J.; and Zhang, C., eds., Advances in Neur...

  15. [23]

    S.; O'Brien, J.; Cai, C

    Park, J. S.; O'Brien, J.; Cai, C. J.; Morris, M. R.; Liang, P.; and Bernstein, M. S. 2023. Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, UIST '23. New York, NY, USA: Associ...

  16. [24]

    Press, O.; Zhang, M.; Min, S.; Schmidt, L.; Smith, N.; and Lewis, M. 2023. Measuring and Narrowing the Compositionality Gap in Language Models. In Bouamor, H.; Pino, J.; and Bali, K., eds., Findings of the Association for Computational Linguistics: EMNLP 2023, 5687--5711. Sing...

  17. [25]

    Sato, T.; Ozaki, S.; and Yokoyama, D. 2024. An Implementation of Werewolf Agent That does not Truly Trust LLM s. In Kano, Y., ed., Proceedings of the 2nd International AIWolfDial Workshop, 58--67. Tokyo, Japan: Association for Computational Linguistics

  18. [26]

    L.; Addis, D

    Schacter, D. L.; Addis, D. R.; and Buckner, R. L. 2007. Remembering the past to imagine the future: the prospective brain. Nature Reviews Neuroscience, 8(9): 657--661

  19. [27]

    R.; and Yao, S

    Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K. R.; and Yao, S. 2023. Reflexion: language agents with verbal reinforcement learning. In Thirty-seventh Conference on Neural Information Processing Systems

  20. [28]

    Tang, H. 2025. Henre Tang's Personal Website. https://tanghenre.com/. Accessed: 2025-07-28

  21. [29]

    Wang, G.; Xie, Y.; Jiang, Y.; Mandlekar, A.; Xiao, C.; Zhu, Y.; Fan, L.; and Anandkumar, A. 2024 a . Voyager: An Open-Ended Embodied Agent with Large Language Models. Transactions on Machine Learning Research

  22. [30]

    Wang, L.; Lian, J.; Huang, Y.; Dai, Y.; Li, H.; Chen, X.; Xie, X.; and Wen, J.-R. 2025. C haracter B ox: Evaluating the Role-Playing Capabilities of LLM s in Text-Based Virtual Worlds. In Chiruzzo, L.; Ritter, A.; and Wang, L., eds., Proceedings of the 2025 Conference of the N...

  23. [31]

    Wang, N.; Peng, Z.; Que, H.; Liu, J.; Zhou, W.; Wu, Y.; Guo, H.; Gan, R.; Ni, Z.; Yang, J.; Zhang, M.; Zhang, Z.; Ouyang, W.; Xu, K.; Huang, W.; Fu, J.; and Peng, J. 2024 b . R ole LLM : Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models. In...

  24. [32]

    Wei, C.; Chen, J.; and Xu, J. 2025. Exploring Large Language Models for Word Games:Who is the Spy? arXiv:2503.15235

  25. [33]

    Wikipedia contributors . 2024. Situation puzzle --- Wikipedia , The Free Encyclopedia. [Online; accessed 25-July-2024]

  26. [34]

    Wu, D.; Shi, H.; Sun, Z.; and Liu, B. 2024. Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Findings of the Association for Computational Linguistics: ACL 2024, 8225--...

  27. [35]

    K.; Pu, Y.; Willis, K.; and Liu, B

    Wu, S.; Khasahmadi, A.; Katz, M.; Jayaraman, P. K.; Pu, Y.; Willis, K.; and Liu, B. 2023. Cad-llm: Large language model for cad generation . In NeurIPS 2023 Workshop on Machine Learning for Creativity and Design

  28. [36]

    H.; Katz, M.; Jayaraman, P

    Wu, S.; Khasahmadi, A. H.; Katz, M.; Jayaraman, P. K.; Pu, Y.; Willis, K.; and Liu, B. 2025 a . CadVLM: Bridging Language and Vision in the Generation of Parametric CAD Sketches. In Leonardis, A.; Ricci, E.; Roth, S.; Russakovsky, O.; Sattler, T.; and Varol, G., eds., Computer...

  29. [37]

    Wu, S.; Zhang, H.; Li, Y.; Effaty, F.; Ataei, A.; and Liu, B. 2025 b . Seeing Beyond Words: MatVQA for Challenging Visual-Scientific Reasoning in Materials Science. arXiv:2505.18319

  30. [38]

    Xi, Z.; Chen, W.; Guo, X.; et al. 2025. The rise and potential of large language model based agents: a survey. Science China Information Sciences, 68: 121101

  31. [39]

    Xu, Z.; Gu, W.; Yu, C.; Wu, Y.; and Wang, Y. 2025. Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization. In Forty-second International Conference on Machine Learning

  32. [40]

    Xu, Z.; Yu, C.; Fang, F.; Wang, Y.; and Wu, Y. 2024. Language agents with reinforcement learning for strategic play in the Werewolf game. In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR.org

  33. [41]

    Yang, A.; Li, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Gao, C.; Huang, C.; Lv, C.; et al. 2025 a . Qwen3 technical report

  34. [42]

    Yang, D.; and Jin, Q. 2024. What Makes a Good Story and How Can We Measure It? A Comprehensive Survey of Story Evaluation. arXiv:2408.14622

  35. [43]

    V.; Movahedi, M.; Li, M.; Ji, H.; Zhang, H.; and Zhang, T

    Yang, R.; Chen, H.; Zhang, J.; Zhao, M.; Qian, C.; Wang, K.; Wang, Q.; Koripella, T. V.; Movahedi, M.; Li, M.; Ji, H.; Zhang, H.; and Zhang, T. 2025 b . EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents. In Forty-seco...

  36. [44]

    R.; and Cao, Y

    Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K. R.; and Cao, Y. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. In The Eleventh International Conference on Learning Representations

  37. [45]

    Y.; Xu, Z.; Zhu, Y.; Shi, X.; Li, M.; and Smola, A

    Yu, P.; Shen, D.; Meng, S.; Lee, J.; Yin, W.; Cui, A. Y.; Xu, Z.; Zhu, Y.; Shi, X.; Li, M.; and Smola, A. 2025. RPGBENCH: Evaluating Large Language Models as Role-Playing Game Engines. arXiv:2502.00595

  38. [46]

    Yu, Q.; Song, S.; Fang, K.; Shi, Y.; Zheng, Z.; Wang, H.; Niu, S.; and Li, Z. 2024. TurtleBench: Evaluating Top Language Models via Real-World Yes/No Puzzles. arXiv:2410.05262

  39. [47]

    Zhang, Z.; Lan, Y.; Chen, Y.; Wang, L.; Wang, X.; and Wang, H. 2025. DVM: Towards Controllable LLM Agents in Social Deduction Games. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE

  40. [48]

    Zhu, Q.; Zhao, R.; Du, J.; Gui, L.; and He, Y. 2024. PLAYER*: Enhancing LLM-based Multi-Agent Communication and Interaction in Murder Mystery Games. CoRR, abs/2404.17662

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.