Pith. sign in

REVIEW 4 major objections 5 minor 58 references

This paper claims that a social-deduction agent becomes more effective and more natural when it watches players' video and replies through an animated avatar, and when its reasoning is trained to depend on the evidence it actually cites.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 16:50 UTC pith:WCW25JIH

load-bearing objection CaM-Wolf is a solid systems paper for multimodal Werewolf agents; the causal faithfulness reward is the least supported piece, but it's a worthwhile integration with addressable issues. the 4 major comments →

arxiv 2607.26393 v1 pith:WCW25JIH submitted 2026-07-29 cs.AI

CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games

classification cs.AI
keywords social deduction gamesWerewolfmultimodal agentscausal reasoningreinforcement learningcounterfactual interventionhuman-AI interactionvideo generation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper introduces CaM-Wolf, a Werewolf-playing agent that goes beyond text: it processes video of other players, transcribes their speech and describes their gestures and expressions, reasons about hidden roles through explicit premise-deduction-conclusion chains, and answers through a talking avatar. The agent's reasoner is trained with a reward that checks, via counterfactual premise removal, whether each cited piece of evidence truly changes the conclusion; if removing a cited premise does not change the conclusion, or removing an uncited premise does, the agent is penalized. In agent-versus-agent games, CaM-Wolf reports higher win rates than text-based baselines, especially when playing the Village team, and in a 5-player human study it wins the most games and is voted against least often. The paper argues this is the first demonstration that integrating multimodal perception and generation into a social-deduction agent improves both gameplay and human trust.

Core claim

CaM-Wolf is the first social-deduction agent that both perceives video of other players and responds through generated avatar video. Its causal-aware Reasoner is trained with reinforcement learning using a counterfactual premise intervention: for each correct role conclusion, the training removes a cited premise (or an uncited premise) from the game log and re-infers the role with a separate judge model. If removing a cited premise leaves the conclusion unchanged, the agent is penalized for citing irrelevant evidence; if removing an uncited premise changes the conclusion, it is penalized for omitting critical evidence. This steers the agent to ground its role judgments in the behavioral obse

What carries the argument

The load-bearing mechanism is the counterfactual premise-intervention check used to define a faithfulness reward. For every correctly predicted role, the training removes one premise the agent cited, or one premise it did not cite, from the game log and asks a lightweight language model whether the conclusion flips. A cited premise that does not change the conclusion incurs a penalty; an uncited premise that does change it incurs a penalty. These penalties are combined into a reward that drives a relative-policy-optimization update, teaching the policy to cite exactly the evidence on which its reasoning genuinely depends.

Load-bearing premise

The faithfulness reward assumes that a separate judge model's conclusion after removing a premise reliably indicates whether the original model's conclusion truly depended on that premise, rather than reflecting the judge's own biases or random variation.

What would settle it

Take a fixed set of game logs, compute the faithfulness reward with several different judge models, and also with randomly deleted premises instead of the paper's targeted deletion; if the reward labels differ sharply across judges, or if random deletion flips conclusions as often as deleting cited premises does, then the reward is not measuring genuine causal dependence.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the reported results hold, adding video perception and avatar output to a social-deduction agent changes gameplay outcomes, not just interface: CaM-Wolf reaches 60-70% win rates against strong text-based baselines depending on team assignment.
  • The causal-aware reward raises role-identification accuracy to 32.6%, up from the 22-28% range of baseline agents, suggesting that rewarded evidence-grounded reasoning transfers to better in-game judgments.
  • In mixed human-AI games, the video-input, avatar-output agent is voted against least often (0.21 average votes versus 3.45 for a text agent) and wins 46.8% of games, the highest of all agents in the study.
  • Users rank the full video-to-video interaction paradigm highest on naturalness, engagement, and satisfaction, implying that the multimodal interface is a measurable component of an agent's social performance.
  • The counterfactual intervention procedure is a reusable training signal: the same remove-and-reinfer logic could be applied to any reasoning task where faithfulness to cited evidence matters.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The counterfactual faithfulness check is a general diagnostic: applying it to any chain-of-thought output could reveal unsupported claims, independent of the Werewolf setting, whenever ground-truth roles or outcomes are available.
  • The paper does not isolate whether the win-rate gains come from the causal-aware training or from the extra multimodal information; an experiment that feeds the same visual description as text to the reasoner, or ablates video input, would separate these factors.
  • Because the faithfulness reward depends on a single judge model, its validity hinges on that judge's stability; a cheap, direct check would compare reward labels across several judge models and prompt phrasings, or against random premise deletion as a null baseline.
  • The causal premise that behavior is shaped by hidden roles is strong in Werewolf but may weaken in games with less role-behavior coupling; testing the method in games like Avalon or in negotiation tasks would reveal how far the approach generalizes.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. CaM-Wolf adds a multimodal perception/performer pipeline to a Werewolf agent: video inputs are transcribed and described as text, and outputs are rendered as talking avatars. The central technical novelty is a causal-aware Reasoner trained by GRPO with a reward that removes premises from the game log and asks a separate 14B judge to re-infer the conclusion; if the conclusion changes, the premise is deemed relevant. The paper reports agent-vs-agent win rates (Table 1), role identification accuracy (Table 2), ablations (Table 3), a human study (Table 4), and user-preference rankings (Figure 6).

Significance. If the central claim holds, CaM-Wolf is a meaningful step toward natural multimodal interaction in social deduction games and a potentially useful training signal for faithful, evidence-grounded reasoning. The system is described with enough detail to reproduce the pipeline, and the code is promised. However, the novelty is currently supported by three pieces of evidence that are not yet convincing: (1) the performance differences are within/near statistical noise at the reported sample sizes; (2) the counterfactual-judge reward is not validated as a measure of causal dependence; and (3) the human study lacks enough detail and power to support the interaction-quality claims. These are fixable with additional experiments and analysis, so the manuscript has merit but requires revision.

major comments (4)
  1. [§4.2, Table 1; §4.3, Table 3; §4.4, Table 4] All headline results are reported as point estimates without confidence intervals, standard errors, or significance tests. With 50 games per setting, the standard error of a win rate is about 7 percentage points, making a 10-point gap (e.g., 60% vs 50% against GPT-4o) statistically non-significant; even the 70% vs 50% gap against Qwen2.5-72B is only borderline after multiple comparisons. The human study has 16 participants playing 5 games each and no per-agent exposure counts. Please report intervals/tests, or run additional games, before claiming 'superior' and 'clear advantage.'
  2. [§3.3.2–3.3.3, Eq. (7)] The faithfulness reward is computed by removing a premise and asking a separate judge (Qwen2.5-14B-Instruct) to re-infer the conclusion, but the judge is never run on the un-intervened log, and no controls are reported for sampling stochasticity or judge/policy disagreement. A changed conclusion may therefore reflect the judge's own bias or variance, not genuine dependence on the removed premise. This is load-bearing: the ablation in Table 3 removes the reward as a whole, so it does not validate the proxy. Please provide (a) judge consistency on un-intervened logs, (b) agreement between the judge and the policy's own self-intervention, (c) a human spot-check of premise removals, and (d) at least one robustness check with a different judge family. Without this, Eq. (7) is an unvalidated noisy reward.
  3. [§3.3 and Conclusion] The text repeatedly describes the intervention procedure as 'causal discovery' and says the Reasoner learns 'causal relations between observable behaviors and hidden roles.' What is actually measured is whether a separate judge's conclusion changes when a premise is removed; no causal graph, structural equation, or exogenous variation is identified. The claim is stronger than the operationalization. Please either temper the causal language or provide a formal argument that the intervention-based δ is a valid causal effect estimate for the policy's own inference. This is a correctness-risk issue because 'causal-aware' is the paper's main novelty.
  4. [§4.4, Table 4] The human-AI mixed-game results are not adequately powered or described. It is unclear how many games each agent appeared in, how the random selection of four agents was balanced, whether participants were blinded to agent identity, and what the variance was. The claim that CaM-Wolf receives the 'fewest votes' (0.21 vs. 0.46 for LSPO) may be driven by a few games. Please provide per-agent game counts, distributions, and statistical tests (e.g., mixed-effects model with participant as random effect).
minor comments (5)
  1. [Table 4] The baseline is called 'LSA' in Table 4 but 'SLA' in the body and Table 1. Please unify the name.
  2. [§4.1.1] 'across 16 H20 GPUs' should be 'on 16 H20 GPUs'; also specify the VRAM/parallelism details if the training time is claimed to be 6 hours.
  3. [§4.2] The observation that Team Werewolf consistently wins more is explained by 'high-entropy environments'; a citation or additional analysis of role balance would help, since it is a notable asymmetry.
  4. [Figure 3] The diagram mixes reward values and examples ('+1', '−1/2') with the training loop; consider separating the reward-scheme schematic from the example log for readability.
  5. [Abstract/Introduction] The phrase 'the first SDG agent that integrates multimodal perception and generation' should be softened or supported with a more explicit comparison to the closest prior systems (§2.2), since authors later acknowledge that some prior work uses multimodal analysis.

Circularity Check

0 steps flagged

No significant circularity: central claims rest on external benchmarks and an empirical RL reward, not on a self-referential derivation.

full rationale

The paper's main claims—multimodal gameplay performance and human-AI interaction quality—are evaluated against external baselines (Tables 1, 2, and 4) using win rates and role-identification accuracy, which are independent of any fitted parameter or prior self-cited result. The causal-aware Reasoner is trained with a reward (Eq. 7) whose faithfulness component depends on a separate lightweight judge model re-inferring after premise removal. This is an empirical operationalization of 'causal evidence' rather than a definitional reduction: the reward does not by construction equal the reported performance, and the judge is not the policy itself. Any concern that the judge may measure disagreement or prompt sensitivity instead of genuine causal dependence is a validity threat to the training signal, not a circularity in the derivation chain. The paper's self-citations appear only in related-work context (e.g., refs. [15], [52], [54]) and are not load-bearing for the proposed method or its evaluation. No uniqueness theorem, ansatz-via-citation, or fitted-input-called-prediction pattern is present. Accordingly, no circular step meets the evidentiary standard required by the rubric.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

The paper introduces no new physical or conceptual entities. It relies on two load-bearing domain assumptions: the role-behavior causal link and the reliability of the intervention judge model. The reward coefficients and dataset size are hand-chosen free parameters. No formal proof is attempted; this is an empirical systems paper.

free parameters (3)
  • Reward coefficients = -1 (correct), -1/|P_i| and -0.5 (faithfulness), -1 (format), -1 (repeat)
    Hand-chosen coefficients in Equation (7); the central training signal is shaped by these manual weights, not fitted to a validation set.
  • Training data size = 500 self-play games, 3000 speaking turns
    Section 4.1.1: the number of generated game logs and sampled turns is a design choice that affects model quality.
  • GRPO hyperparameters = G=8, ε=0.2, β=0.04, batch 128, LoRA rank 16, lr=1e-6, 2 epochs
    Section 4.1.1: standard RL hyperparameters chosen by hand; included for completeness as they affect the optimization.
axioms (4)
  • domain assumption Players' observable behaviors (speech, gestures, facial expressions) are causally shaped by their hidden roles.
    Section 1: 'each player's behaviors in SDGs are inherently shaped by their hidden roles'. This premise motivates the entire causal-aware training.
  • domain assumption Counterfactual premise intervention with Qwen2.5-14B-Instruct reliably determines whether a conclusion depends on a premise.
    Section 3.3.2: the faithfulness reward assumes that a conclusion change after removing a premise correctly identifies causal relevance; the judge model's re-inference is treated as ground truth.
  • standard math GRPO update rule from DeepSeek-Math is a valid policy optimization for LLMs.
    Equations (9)-(11) cite Shao et al. [28]; the paper relies on this prior RL method without re-derivation.
  • domain assumption The Werewolf game rules (7-player and 5-player variants) are a valid testbed for social deduction.
    Section 4.1: 2 werewolves, 1 seer, 1 guardian, 3 villagers; the paper assumes these rules elicit the intended social reasoning.

pith-pipeline@v1.3.0-daily-deepseek · 15333 in / 12293 out tokens · 114652 ms · 2026-08-01T16:50:17.083768+00:00 · methodology

0 comments
read the original abstract

Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents. These games require complex social skills such as reasoning, deception, and collaboration. While recent advances in large language models (LLMs) have driven significant progress in SDG agents, current approaches are predominantly text-based, overlooking the multimodal nature that is fundamental to human social interaction. To bridge this gap, we introduce CaM-Wolf, the first SDG agent that integrates multimodal perception and generation. CaM-Wolf processes video inputs from other players, employs a causal-aware Reasoner trained via reinforcement learning to establish logical chains between observable behaviors and hidden roles, and presents itself through an animated avatar. Our experiments and user study show that CaM-Wolf achieves superior agent gameplay performance and enhances the quality of human-AI interaction. This work represents a significant advancement towards creating more human-like AI agents capable of participating in nuanced social dynamics. Our code is available at https://3dagentworld.github.io/avatar_wolf.

Figures

Figures reproduced from arXiv: 2607.26393 by Deheng Ye, Hao Wang, Jiarui He, Nanjie Yao, Peilin Zhao, Zheng Zhang.

Figure 1
Figure 1. Figure 1: Different interaction paradigms of SDG agents. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The overall framework of CaM-Wolf. The Perceiver processes video inputs from human players, extracting transcribed [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Overview of the Reasoner’s training process. The Reasoner first generates structured reasoning with explicit premise [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Examples of generated videos. The left side shows [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 6
Figure 6. Figure 6: Human ranking scores for different interaction [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

58 extracted references · 5 canonical work pages

  1. [1]

    Qi Chai, Zhang Zheng, Junlong Ren, Deheng Ye, Zichuan Lin, and Hao Wang. 2025. CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative Tasks. InFindings of the Association for Computational Linguistics: EMNLP 2025, Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (Eds.). Association for Computational Linguistics...

  2. [2]

    David Maxwell Chickering and Christopher Meek. 2015. Selective Greedy Equiv- alence Search: finding optimal Bayesian networks using a polynomial number of score evaluations. InProceedings of the Thirty-First Conference on Uncertainty in Artificial Intelligence. 211–219

  3. [3]

    Kai-Hendrik Cohrs, Gherardo Varando, Emiliano Diaz, Vasileios Sitokonstantinou, and Gustau Camps-Valls. 2024. Large Language Models for Constrained-Based Causal Discovery.arXiv preprint arXiv:2406.07378(2024)

  4. [4]

    Yang Dai, Oubo Ma, Longfei Zhang, Xingxing Liang, Shengchao Hu, Mengzhu Wang, Shouling Ji, Jincai Huang, and Li Shen. 2024. Is mamba compatible with trajectory optimization in offline reinforcement learning?. InProceedings of the 38th International Conference on Neural Information Processing Systems(Vancou- ver, BC, Canada)(NIPS ’24). Curran Associates In...

  5. [5]

    Honghao Fu, Junlong Ren, Qi Chai, Deheng Ye, Yujun Cai, and Hao Wang. 2025. VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft. arXiv:2508.18722 [cs.AI] https://arxiv.org/abs/2508.18722

  6. [6]

    Qijun Gan, Ruizi Yang, Jianke Zhu, Shaofei Xue, and Steven Hoi. 2025. Omni- Avatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation. arXiv:2506.18866 [cs.CV] https://arxiv.org/abs/2506.18866

  7. [7]

    Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrish- nan, Fiona Ryan, Jayant Sharma, Michael Wray, Mengmeng Xu, Eric Zhongcong Xu, Chen Zhao, Siddhant Bansal, Dhruv Batra, Vincent Car...

  8. [8]

    Longxiang He, Deheng Ye, Junbo Tan, Xueqian Wang, and Li Shen. 2025. Robust Policy Expansion for Offline-to-Online RL under Diverse Data Corruption. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=F7y7JMaTvj

  9. [9]

    Yuya Hirata, Michimasa Inaba, Kenichi Takahashi, Fujio Toriumi, Hirotaka Osawa, Daisuke Katagami, and Kousuke Shinoda. 2016. Werewolf Game Modeling Using Action Probabilities Based on Play Log Analysis. InComputers and Games. https://api.semanticscholar.org/CorpusID:37838481

  10. [10]

    Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. InInternational Conference on Learning Representations. https: //openreview.net/forum?id=nZeVKeeFYf9

  11. [11]

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2025. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.ACM Trans. Inf. Syst.43, 2, Article 42 (Jan. 2025), 55 pages. doi:10.1145/3703155

  12. [12]

    Xuanfa Jin, Ziyan Wang, Yali Du, Meng Fang, Haifeng Zhang, and Jun Wang. 2024. Learning to Discuss Strategically: A Case Study on One Night Ultimate Werewolf. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=1f82rnwCbl

  13. [13]

    Emre Kıcıman, Robert Ness, Amit Sharma, and Chenhao Tan. 2023. Causal reasoning and large language models: Opening a new frontier for causality.arXiv preprint arXiv:2305.00050(2023)

  14. [14]

    Bolin Lai, Hongxin Zhang, Miao Liu, Aryan Pariani, Fiona Ryan, Wenqi Jia, Shirley Anugrah Hayati, James Rehg, and Diyi Yang. 2023. Werewolf Among Us: Multimodal Resources for Modeling Persuasion Behaviors in Social Deduction Games. InFindings of the Association for Computational Linguistics: ACL 2023, Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (E...

  15. [15]

    Yihuai Lan, Zhiqiang Hu, Lei Wang, Yang Wang, Deheng Ye, Peilin Zhao, Ee-Peng Lim, Hui Xiong, and Hao Wang. 2024. LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Asso...

  16. [16]

    Sangmin Lee, Bolin Lai, Fiona Ryan, Bikram Boote, and James M. Rehg. 2024. Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations. In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 14585–14595. doi:10.1109/CVPR52733. 2024.01382

  17. [17]

    Toups Dugas, Gillian Smith, and Rose Bohrer

    Shano Liang, Max Chen, Phoebe O. Toups Dugas, Gillian Smith, and Rose Bohrer

  18. [18]

    Jonathan Light, Min Cai, Sheng Shen, and Ziniu Hu. 2023. From Text to Tactic: Evaluating LLMs Playing the Game of Avalon. InNeurIPS 2023 Foundation Models for Decision Making Workshop. https://openreview.net/forum?id=ltUrSryS0K

  19. [19]

    Stephanie Long, Alexandre Piché, Valentina Zantedeschi, Tibor Schuster, and Alexandre Drouin. 2023. Causal Discovery with Language Models as Imperfect Experts. InICML 2023 Workshop on Structured Probabilistic Inference & Generative Modeling. https://openreview.net/forum?id=RXlvYZAE49

  20. [20]

    Haohao Luo, Jiayi Kuang, Wei Liu, Ying Shen, Jian Luan, and Yang Deng. 2025. Browsing Like Human: A Multimodal Web Agent with Experiential Fast-and- Slow Thinking. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (E...

  21. [21]

    Weiyu Ma, Qirui Mi, Yongcheng Zeng, Xue Yan, Runji Lin, Yuqiao Wu, Jun Wang, and Haifeng Zhang. 2024. Large Language Models Play StarCraft II:Benchmarks and A Chain of Summarization Approach. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id= kEPpD7yETM

  22. [22]

    Noritsugu Nakamura, Michimasa Inaba, Kenichi Takahashi, Fujio Toriumi, Hi- rotaka Osawa, Daisuke Katagami, and Kousuke Shinoda. 2016. Constructing a Human-like agent for the Werewolf Game using a psychological model based multiple perspectives.2016 IEEE Symposium Series on Computational Intelligence (SSCI)(2016), 1–8. https://api.semanticscholar.org/Corpu...

  23. [23]

    Rodriguez, Montek Kalsi, Nicolas Chapados, M

    Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Juan A. Rodriguez, Montek Kalsi, Nicolas Chapados, M. Tamer Özsu, Aishwarya Agrawal, David Vazquez, Christopher Pal, Perouz Taslakian, Spandana Gella, and Sai Rajeswar. 2025. UI- Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction. InForty-second International Conference on Machine...

  24. [24]

    Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li,...

  25. [25]

    Joseph D Ramsey. 2015. Scaling up greedy causal search for continuous variables. arXiv preprint arXiv:1507.07749(2015)

  26. [26]

    Parkes, and Joshua B

    Jack Serrino, Max Kleiman-Weiner, David C. Parkes, and Joshua B. Tenenbaum. 2019.Finding friend and foe in multi-agent games. Curran Associates Inc., Red Hook, NY, USA

  27. [27]

    Xiao Shao, Weifu Jiang, Fei Zuo, and Mengqing Liu. 2024. SwarmBrain: Em- bodied agent for real-time strategy game StarCraft II via large language models. arXiv:2401.17749 [cs.AI] https://arxiv.org/abs/2401.17749

  28. [28]

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024. DeepSeek- Math: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300 [cs.CL] https://arxiv.org/abs/2402.03300

  29. [29]

    Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. 2025. HybridFlow: A Flexible and Efficient RLHF Framework. InProceedings of the Twentieth European Conference on Computer Systems(Rotterdam, Netherlands)(EuroSys ’25). Association for Computing Machinery, New York, NY, USA, 1279–1297. doi:10....

  30. [30]

    Zijing Shi, Meng Fang, Shunfeng Zheng, Shilong Deng, Ling Chen, and Yali Du

  31. [31]

    2001.Causation, prediction, and search

    Peter Spirtes, Clark Glymour, and Richard Scheines. 2001.Causation, prediction, and search. MIT press

  32. [32]

    Peter L Spirtes, Christopher Meek, and Thomas S Richardson. 2013. Causal inference in the presence of latent variables and selection bias.arXiv preprint MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Zheng Zhang et al. arXiv:1302.4983(2013)

  33. [33]

    Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2024. Voyager: An Open-Ended Embodied Agent with Large Language Models.Transactions on Machine Learning Research (2024). https://openreview.net/forum?id=ehfRiF0R3a

  34. [34]

    Shenzhi Wang, Chang Liu, Zilong Zheng, Siyuan Qi, Shuo Chen, Qisen Yang, Andrew Zhao, Chaofei Wang, Shiji Song, and Gao Huang. 2023. Avalon’s Game of Thoughts: Battle Against Deception through Recursive Contempla- tion. arXiv:2310.01320 [cs.AI] https://arxiv.org/abs/2310.01320

  35. [35]

    Shenzhi Wang, Chang Liu, Zilong Zheng, Siyuan Qi, Shuo Chen, Qisen Yang, Andrew Zhao, Chaofei Wang, Shiji Song, and Gao Huang. 2024. Boosting LLM Agents with Recursive Contemplation for Effective Deception Handling. InFind- ings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association fo...

  36. [36]

    Tianhe Wang and Tomoyuki Kaneko. 2018. Application of Deep Reinforcement Learning in Werewolf Game Agents.2018 Conference on Technologies and Appli- cations of Artificial Intelligence (TAAI)(2018), 28–33. https://api.semanticscholar. org/CorpusID:57191228

  37. [37]

    Xiaoqiang Wang and Bang Liu. 2025. OSCAR: Operating System Control via State- Aware Reasoning and Re-Planning. InThe Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=VuTrZzrPfn

  38. [38]

    Yiqin Wang, Haoji Zhang, Jingqi Tian, and Yansong Tang. 2025. Ponder & Press: Advancing Visual GUI Agent towards General Computer Control. InFindings of the Association for Computational Linguistics: ACL 2025, Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 146...

  39. [39]

    Dekun Wu, Haochen Shi, Zhiyuan Sun, and Bang Liu. 2024. Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mys- tery Games. InFindings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computa- tional Linguistics, Bangkok, Thailand, 822...

  40. [40]

    Junda Wu, Tong Yu, Xiang Chen, Haoliang Wang, Ryan Rossi, Sungchul Kim, Anup Rao, and Julian McAuley. 2024. DeCoT: Debiasing Chain-of-Thought for Knowledge-Intensive Tasks in Large Language Models via Causal Intervention. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Lun-Wei Ku, Andre Ma...

  41. [41]

    Shuang Wu, Liwen Zhu, Tao Yang, Shiwei Xu, Qiang Fu, Yang Wei, and Haobo Fu. 2024. Enhance Reasoning for Large Language Models in the Game Werewolf. arXiv:2402.02330 [cs.AI] https://arxiv.org/abs/2402.02330

  42. [42]

    Jing Xiang and Seyoung Kim. 2013. A* Lasso for learning a sparse Bayesian network structure for continuous variables.Advances in neural information processing systems26 (2013)

  43. [43]

    Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, Bin Zhang, Xiong Wang, Yunfei Chu, and Junyang Lin. 2025. Qwen2.5-Omni Technical Report. arXiv:2503.20215 [cs.CL] https://arxiv.org/abs/2503.20215

  44. [44]

    Yuzhuang Xu, Shuo Wang, Peng Li, Fuwen Luo, Xiaolong Wang, Weidong Liu, and Yang Liu. 2023. Exploring large language models for communication games: An empirical study on werewolf.arXiv preprint arXiv:2309.04658(2023)

  45. [45]

    Zelai Xu, Wanjun Gu, Chao Yu, Yi Wu, and Yu Wang. 2025. Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization. InForty-second International Conference on Machine Learning. https: //openreview.net/forum?id=N2mOBiSqhc

  46. [46]

    Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli. 2025. Hallucination is Inevitable: An Innate Limitation of Large Language Models. arXiv:2401.11817 [cs.CL] https: //arxiv.org/abs/2401.11817

  47. [47]

    Zelai Xu, Chao Yu, Fei Fang, Yu Wang, and Yi Wu. 2024. Language agents with reinforcement learning for strategic play in the Werewolf game. InProceedings of the 41st International Conference on Machine Learning(Vienna, Austria)(ICML’24). JMLR.org, Article 2285, 31 pages

  48. [48]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. InInternational Conference on Learning Representations (ICLR)

  49. [49]

    Junkun Yuan, Xu Ma, Ruoxuan Xiong, Mingming Gong, Xiangyu Liu, Fei Wu, Lanfen Lin, and Kun Kuang. 2023. Instrumental Variable-Driven Domain Gener- alization with Unobserved Confounders.ACM Trans. Knowl. Discov. Data17, 8, Article 118 (June 2023), 21 pages. doi:10.1145/3595380

  50. [50]

    Matej Zečević, Moritz Willig, Devendra Singh Dhami, and Kristian Kersting. 2023. Causal Parrots: Large Language Models May Talk Causality But Are Not Causal. Transactions on Machine Learning Research(2023)

  51. [51]

    Yuzhe Zhang, Yipeng Zhang, Yidong Gan, Lina Yao, and Chen Wang. 2024. Causal graph discovery with retrieval-augmented generation based large language mod- els.arXiv preprint arXiv:2402.15301(2024)

  52. [52]

    Zheng Zhang, Yihuai Lan, Yangsen Chen, Lei Wang, Xiang Wang, and Hao Wang

  53. [53]

    Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. 2024. GPT-4V(ision) is a generalist web agent, if grounded. InProceedings of the 41st International Conference on Machine Learning(Vienna, Austria)(ICML’24). JMLR.org, Article 2538, 37 pages

  54. [54]

    Zhang Zheng, Deheng Ye, Peilin Zhao, and Hao Wang. 2026. The Stackelberg Speaker: Optimizing Persuasive Communication in Social Deduction Games. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Maria Liakata, Viviane P. Moreira, Jiajun Zhang, and David Jurgens (Eds.). Association for Comput...

  55. [55]

    In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

    DVM: Towards Controllable LLM Agents in Social Deduction Games. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1–5. doi:10.1109/ICASSP49660.2025.10888525

  56. [58]

    Qiyang Zhou, Xu Ruihang, Peng Wang, WenJie Lu, Xiaochun Cao, Naiqiang Tan, and Li Shen. 2026. HTAC: Hierarchical Task-Aware Composition for Contin- ual Offline Reinforcement Learning. InForty-third International Conference on Machine Learning. https://openreview.net/forum?id=akfJfpUEBj

  57. [2023]

    arXiv:2312.17515 [cs.CL] https://arxiv.org/abs/2312.17515

    Cooperation on the Fly: Exploring Language Agents for Ad Hoc Teamwork in the Avalon Game. arXiv:2312.17515 [cs.CL] https://arxiv.org/abs/2312.17515

  58. [2025]

    doi:10.1145/3721121

    The Collaborative Sensemaking Play of Jubensha Games: A Deconstruction, Taxonomy, and Analysis.ACM Games3, 1, Article 6 (March 2025), 34 pages. doi:10.1145/3721121