REVIEW 4 major objections 5 minor 58 references
This paper claims that a social-deduction agent becomes more effective and more natural when it watches players' video and replies through an animated avatar, and when its reasoning is trained to depend on the evidence it actually cites.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 16:50 UTC pith:WCW25JIH
load-bearing objection CaM-Wolf is a solid systems paper for multimodal Werewolf agents; the causal faithfulness reward is the least supported piece, but it's a worthwhile integration with addressable issues. the 4 major comments →
CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
CaM-Wolf is the first social-deduction agent that both perceives video of other players and responds through generated avatar video. Its causal-aware Reasoner is trained with reinforcement learning using a counterfactual premise intervention: for each correct role conclusion, the training removes a cited premise (or an uncited premise) from the game log and re-infers the role with a separate judge model. If removing a cited premise leaves the conclusion unchanged, the agent is penalized for citing irrelevant evidence; if removing an uncited premise changes the conclusion, it is penalized for omitting critical evidence. This steers the agent to ground its role judgments in the behavioral obse
What carries the argument
The load-bearing mechanism is the counterfactual premise-intervention check used to define a faithfulness reward. For every correctly predicted role, the training removes one premise the agent cited, or one premise it did not cite, from the game log and asks a lightweight language model whether the conclusion flips. A cited premise that does not change the conclusion incurs a penalty; an uncited premise that does change it incurs a penalty. These penalties are combined into a reward that drives a relative-policy-optimization update, teaching the policy to cite exactly the evidence on which its reasoning genuinely depends.
Load-bearing premise
The faithfulness reward assumes that a separate judge model's conclusion after removing a premise reliably indicates whether the original model's conclusion truly depended on that premise, rather than reflecting the judge's own biases or random variation.
What would settle it
Take a fixed set of game logs, compute the faithfulness reward with several different judge models, and also with randomly deleted premises instead of the paper's targeted deletion; if the reward labels differ sharply across judges, or if random deletion flips conclusions as often as deleting cited premises does, then the reward is not measuring genuine causal dependence.
If this is right
- If the reported results hold, adding video perception and avatar output to a social-deduction agent changes gameplay outcomes, not just interface: CaM-Wolf reaches 60-70% win rates against strong text-based baselines depending on team assignment.
- The causal-aware reward raises role-identification accuracy to 32.6%, up from the 22-28% range of baseline agents, suggesting that rewarded evidence-grounded reasoning transfers to better in-game judgments.
- In mixed human-AI games, the video-input, avatar-output agent is voted against least often (0.21 average votes versus 3.45 for a text agent) and wins 46.8% of games, the highest of all agents in the study.
- Users rank the full video-to-video interaction paradigm highest on naturalness, engagement, and satisfaction, implying that the multimodal interface is a measurable component of an agent's social performance.
- The counterfactual intervention procedure is a reusable training signal: the same remove-and-reinfer logic could be applied to any reasoning task where faithfulness to cited evidence matters.
Where Pith is reading between the lines
- The counterfactual faithfulness check is a general diagnostic: applying it to any chain-of-thought output could reveal unsupported claims, independent of the Werewolf setting, whenever ground-truth roles or outcomes are available.
- The paper does not isolate whether the win-rate gains come from the causal-aware training or from the extra multimodal information; an experiment that feeds the same visual description as text to the reasoner, or ablates video input, would separate these factors.
- Because the faithfulness reward depends on a single judge model, its validity hinges on that judge's stability; a cheap, direct check would compare reward labels across several judge models and prompt phrasings, or against random premise deletion as a null baseline.
- The causal premise that behavior is shaped by hidden roles is strong in Werewolf but may weaken in games with less role-behavior coupling; testing the method in games like Avalon or in negotiation tasks would reveal how far the approach generalizes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. CaM-Wolf adds a multimodal perception/performer pipeline to a Werewolf agent: video inputs are transcribed and described as text, and outputs are rendered as talking avatars. The central technical novelty is a causal-aware Reasoner trained by GRPO with a reward that removes premises from the game log and asks a separate 14B judge to re-infer the conclusion; if the conclusion changes, the premise is deemed relevant. The paper reports agent-vs-agent win rates (Table 1), role identification accuracy (Table 2), ablations (Table 3), a human study (Table 4), and user-preference rankings (Figure 6).
Significance. If the central claim holds, CaM-Wolf is a meaningful step toward natural multimodal interaction in social deduction games and a potentially useful training signal for faithful, evidence-grounded reasoning. The system is described with enough detail to reproduce the pipeline, and the code is promised. However, the novelty is currently supported by three pieces of evidence that are not yet convincing: (1) the performance differences are within/near statistical noise at the reported sample sizes; (2) the counterfactual-judge reward is not validated as a measure of causal dependence; and (3) the human study lacks enough detail and power to support the interaction-quality claims. These are fixable with additional experiments and analysis, so the manuscript has merit but requires revision.
major comments (4)
- [§4.2, Table 1; §4.3, Table 3; §4.4, Table 4] All headline results are reported as point estimates without confidence intervals, standard errors, or significance tests. With 50 games per setting, the standard error of a win rate is about 7 percentage points, making a 10-point gap (e.g., 60% vs 50% against GPT-4o) statistically non-significant; even the 70% vs 50% gap against Qwen2.5-72B is only borderline after multiple comparisons. The human study has 16 participants playing 5 games each and no per-agent exposure counts. Please report intervals/tests, or run additional games, before claiming 'superior' and 'clear advantage.'
- [§3.3.2–3.3.3, Eq. (7)] The faithfulness reward is computed by removing a premise and asking a separate judge (Qwen2.5-14B-Instruct) to re-infer the conclusion, but the judge is never run on the un-intervened log, and no controls are reported for sampling stochasticity or judge/policy disagreement. A changed conclusion may therefore reflect the judge's own bias or variance, not genuine dependence on the removed premise. This is load-bearing: the ablation in Table 3 removes the reward as a whole, so it does not validate the proxy. Please provide (a) judge consistency on un-intervened logs, (b) agreement between the judge and the policy's own self-intervention, (c) a human spot-check of premise removals, and (d) at least one robustness check with a different judge family. Without this, Eq. (7) is an unvalidated noisy reward.
- [§3.3 and Conclusion] The text repeatedly describes the intervention procedure as 'causal discovery' and says the Reasoner learns 'causal relations between observable behaviors and hidden roles.' What is actually measured is whether a separate judge's conclusion changes when a premise is removed; no causal graph, structural equation, or exogenous variation is identified. The claim is stronger than the operationalization. Please either temper the causal language or provide a formal argument that the intervention-based δ is a valid causal effect estimate for the policy's own inference. This is a correctness-risk issue because 'causal-aware' is the paper's main novelty.
- [§4.4, Table 4] The human-AI mixed-game results are not adequately powered or described. It is unclear how many games each agent appeared in, how the random selection of four agents was balanced, whether participants were blinded to agent identity, and what the variance was. The claim that CaM-Wolf receives the 'fewest votes' (0.21 vs. 0.46 for LSPO) may be driven by a few games. Please provide per-agent game counts, distributions, and statistical tests (e.g., mixed-effects model with participant as random effect).
minor comments (5)
- [Table 4] The baseline is called 'LSA' in Table 4 but 'SLA' in the body and Table 1. Please unify the name.
- [§4.1.1] 'across 16 H20 GPUs' should be 'on 16 H20 GPUs'; also specify the VRAM/parallelism details if the training time is claimed to be 6 hours.
- [§4.2] The observation that Team Werewolf consistently wins more is explained by 'high-entropy environments'; a citation or additional analysis of role balance would help, since it is a notable asymmetry.
- [Figure 3] The diagram mixes reward values and examples ('+1', '−1/2') with the training loop; consider separating the reward-scheme schematic from the example log for readability.
- [Abstract/Introduction] The phrase 'the first SDG agent that integrates multimodal perception and generation' should be softened or supported with a more explicit comparison to the closest prior systems (§2.2), since authors later acknowledge that some prior work uses multimodal analysis.
Circularity Check
No significant circularity: central claims rest on external benchmarks and an empirical RL reward, not on a self-referential derivation.
full rationale
The paper's main claims—multimodal gameplay performance and human-AI interaction quality—are evaluated against external baselines (Tables 1, 2, and 4) using win rates and role-identification accuracy, which are independent of any fitted parameter or prior self-cited result. The causal-aware Reasoner is trained with a reward (Eq. 7) whose faithfulness component depends on a separate lightweight judge model re-inferring after premise removal. This is an empirical operationalization of 'causal evidence' rather than a definitional reduction: the reward does not by construction equal the reported performance, and the judge is not the policy itself. Any concern that the judge may measure disagreement or prompt sensitivity instead of genuine causal dependence is a validity threat to the training signal, not a circularity in the derivation chain. The paper's self-citations appear only in related-work context (e.g., refs. [15], [52], [54]) and are not load-bearing for the proposed method or its evaluation. No uniqueness theorem, ansatz-via-citation, or fitted-input-called-prediction pattern is present. Accordingly, no circular step meets the evidentiary standard required by the rubric.
Axiom & Free-Parameter Ledger
free parameters (3)
- Reward coefficients =
-1 (correct), -1/|P_i| and -0.5 (faithfulness), -1 (format), -1 (repeat)
- Training data size =
500 self-play games, 3000 speaking turns
- GRPO hyperparameters =
G=8, ε=0.2, β=0.04, batch 128, LoRA rank 16, lr=1e-6, 2 epochs
axioms (4)
- domain assumption Players' observable behaviors (speech, gestures, facial expressions) are causally shaped by their hidden roles.
- domain assumption Counterfactual premise intervention with Qwen2.5-14B-Instruct reliably determines whether a conclusion depends on a premise.
- standard math GRPO update rule from DeepSeek-Math is a valid policy optimization for LLMs.
- domain assumption The Werewolf game rules (7-player and 5-player variants) are a valid testbed for social deduction.
read the original abstract
Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents. These games require complex social skills such as reasoning, deception, and collaboration. While recent advances in large language models (LLMs) have driven significant progress in SDG agents, current approaches are predominantly text-based, overlooking the multimodal nature that is fundamental to human social interaction. To bridge this gap, we introduce CaM-Wolf, the first SDG agent that integrates multimodal perception and generation. CaM-Wolf processes video inputs from other players, employs a causal-aware Reasoner trained via reinforcement learning to establish logical chains between observable behaviors and hidden roles, and presents itself through an animated avatar. Our experiments and user study show that CaM-Wolf achieves superior agent gameplay performance and enhances the quality of human-AI interaction. This work represents a significant advancement towards creating more human-like AI agents capable of participating in nuanced social dynamics. Our code is available at https://3dagentworld.github.io/avatar_wolf.
Figures
Reference graph
Works this paper leans on
-
[1]
Qi Chai, Zhang Zheng, Junlong Ren, Deheng Ye, Zichuan Lin, and Hao Wang. 2025. CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative Tasks. InFindings of the Association for Computational Linguistics: EMNLP 2025, Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (Eds.). Association for Computational Linguistics...
-
[2]
David Maxwell Chickering and Christopher Meek. 2015. Selective Greedy Equiv- alence Search: finding optimal Bayesian networks using a polynomial number of score evaluations. InProceedings of the Thirty-First Conference on Uncertainty in Artificial Intelligence. 211–219
2015
-
[3]
Kai-Hendrik Cohrs, Gherardo Varando, Emiliano Diaz, Vasileios Sitokonstantinou, and Gustau Camps-Valls. 2024. Large Language Models for Constrained-Based Causal Discovery.arXiv preprint arXiv:2406.07378(2024)
Pith/arXiv arXiv 2024
-
[4]
Yang Dai, Oubo Ma, Longfei Zhang, Xingxing Liang, Shengchao Hu, Mengzhu Wang, Shouling Ji, Jincai Huang, and Li Shen. 2024. Is mamba compatible with trajectory optimization in offline reinforcement learning?. InProceedings of the 38th International Conference on Neural Information Processing Systems(Vancou- ver, BC, Canada)(NIPS ’24). Curran Associates In...
2024
-
[5]
Honghao Fu, Junlong Ren, Qi Chai, Deheng Ye, Yujun Cai, and Hao Wang. 2025. VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft. arXiv:2508.18722 [cs.AI] https://arxiv.org/abs/2508.18722
arXiv 2025
-
[6]
Qijun Gan, Ruizi Yang, Jianke Zhu, Shaofei Xue, and Steven Hoi. 2025. Omni- Avatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation. arXiv:2506.18866 [cs.CV] https://arxiv.org/abs/2506.18866
Pith/arXiv arXiv 2025
-
[7]
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrish- nan, Fiona Ryan, Jayant Sharma, Michael Wray, Mengmeng Xu, Eric Zhongcong Xu, Chen Zhao, Siddhant Bansal, Dhruv Batra, Vincent Car...
2022
-
[8]
Longxiang He, Deheng Ye, Junbo Tan, Xueqian Wang, and Li Shen. 2025. Robust Policy Expansion for Offline-to-Online RL under Diverse Data Corruption. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=F7y7JMaTvj
2025
-
[9]
Yuya Hirata, Michimasa Inaba, Kenichi Takahashi, Fujio Toriumi, Hirotaka Osawa, Daisuke Katagami, and Kousuke Shinoda. 2016. Werewolf Game Modeling Using Action Probabilities Based on Play Log Analysis. InComputers and Games. https://api.semanticscholar.org/CorpusID:37838481
2016
-
[10]
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. InInternational Conference on Learning Representations. https: //openreview.net/forum?id=nZeVKeeFYf9
2022
-
[11]
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2025. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.ACM Trans. Inf. Syst.43, 2, Article 42 (Jan. 2025), 55 pages. doi:10.1145/3703155
doi:10.1145/3703155 2025
-
[12]
Xuanfa Jin, Ziyan Wang, Yali Du, Meng Fang, Haifeng Zhang, and Jun Wang. 2024. Learning to Discuss Strategically: A Case Study on One Night Ultimate Werewolf. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=1f82rnwCbl
2024
-
[13]
Emre Kıcıman, Robert Ness, Amit Sharma, and Chenhao Tan. 2023. Causal reasoning and large language models: Opening a new frontier for causality.arXiv preprint arXiv:2305.00050(2023)
Pith/arXiv arXiv 2023
-
[14]
Bolin Lai, Hongxin Zhang, Miao Liu, Aryan Pariani, Fiona Ryan, Wenqi Jia, Shirley Anugrah Hayati, James Rehg, and Diyi Yang. 2023. Werewolf Among Us: Multimodal Resources for Modeling Persuasion Behaviors in Social Deduction Games. InFindings of the Association for Computational Linguistics: ACL 2023, Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (E...
doi:10.18653/v1/2023 2023
-
[15]
Yihuai Lan, Zhiqiang Hu, Lei Wang, Yang Wang, Deheng Ye, Peilin Zhao, Ee-Peng Lim, Hui Xiong, and Hao Wang. 2024. LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Asso...
-
[16]
Sangmin Lee, Bolin Lai, Fiona Ryan, Bikram Boote, and James M. Rehg. 2024. Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations. In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 14585–14595. doi:10.1109/CVPR52733. 2024.01382
arXiv 2024
-
[17]
Toups Dugas, Gillian Smith, and Rose Bohrer
Shano Liang, Max Chen, Phoebe O. Toups Dugas, Gillian Smith, and Rose Bohrer
-
[18]
Jonathan Light, Min Cai, Sheng Shen, and Ziniu Hu. 2023. From Text to Tactic: Evaluating LLMs Playing the Game of Avalon. InNeurIPS 2023 Foundation Models for Decision Making Workshop. https://openreview.net/forum?id=ltUrSryS0K
2023
-
[19]
Stephanie Long, Alexandre Piché, Valentina Zantedeschi, Tibor Schuster, and Alexandre Drouin. 2023. Causal Discovery with Language Models as Imperfect Experts. InICML 2023 Workshop on Structured Probabilistic Inference & Generative Modeling. https://openreview.net/forum?id=RXlvYZAE49
2023
-
[20]
Haohao Luo, Jiayi Kuang, Wei Liu, Ying Shen, Jian Luan, and Yang Deng. 2025. Browsing Like Human: A Multimodal Web Agent with Experiential Fast-and- Slow Thinking. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (E...
doi:10.18653/v1/ 2025
-
[21]
Weiyu Ma, Qirui Mi, Yongcheng Zeng, Xue Yan, Runji Lin, Yuqiao Wu, Jun Wang, and Haifeng Zhang. 2024. Large Language Models Play StarCraft II:Benchmarks and A Chain of Summarization Approach. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id= kEPpD7yETM
2024
-
[22]
Noritsugu Nakamura, Michimasa Inaba, Kenichi Takahashi, Fujio Toriumi, Hi- rotaka Osawa, Daisuke Katagami, and Kousuke Shinoda. 2016. Constructing a Human-like agent for the Werewolf Game using a psychological model based multiple perspectives.2016 IEEE Symposium Series on Computational Intelligence (SSCI)(2016), 1–8. https://api.semanticscholar.org/Corpu...
2016
-
[23]
Rodriguez, Montek Kalsi, Nicolas Chapados, M
Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Juan A. Rodriguez, Montek Kalsi, Nicolas Chapados, M. Tamer Özsu, Aishwarya Agrawal, David Vazquez, Christopher Pal, Perouz Taslakian, Spandana Gella, and Sai Rajeswar. 2025. UI- Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction. InForty-second International Conference on Machine...
2025
-
[24]
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li,...
Pith/arXiv arXiv 2025
-
[25]
Joseph D Ramsey. 2015. Scaling up greedy causal search for continuous variables. arXiv preprint arXiv:1507.07749(2015)
Pith/arXiv arXiv 2015
-
[26]
Parkes, and Joshua B
Jack Serrino, Max Kleiman-Weiner, David C. Parkes, and Joshua B. Tenenbaum. 2019.Finding friend and foe in multi-agent games. Curran Associates Inc., Red Hook, NY, USA
2019
-
[27]
Xiao Shao, Weifu Jiang, Fei Zuo, and Mengqing Liu. 2024. SwarmBrain: Em- bodied agent for real-time strategy game StarCraft II via large language models. arXiv:2401.17749 [cs.AI] https://arxiv.org/abs/2401.17749
Pith/arXiv arXiv 2024
-
[28]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024. DeepSeek- Math: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300 [cs.CL] https://arxiv.org/abs/2402.03300
Pith/arXiv arXiv 2024
-
[29]
Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. 2025. HybridFlow: A Flexible and Efficient RLHF Framework. InProceedings of the Twentieth European Conference on Computer Systems(Rotterdam, Netherlands)(EuroSys ’25). Association for Computing Machinery, New York, NY, USA, 1279–1297. doi:10....
doi:10.1145/3689031 2025
-
[30]
Zijing Shi, Meng Fang, Shunfeng Zheng, Shilong Deng, Ling Chen, and Yali Du
-
[31]
2001.Causation, prediction, and search
Peter Spirtes, Clark Glymour, and Richard Scheines. 2001.Causation, prediction, and search. MIT press
2001
-
[32]
Peter L Spirtes, Christopher Meek, and Thomas S Richardson. 2013. Causal inference in the presence of latent variables and selection bias.arXiv preprint MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Zheng Zhang et al. arXiv:1302.4983(2013)
Pith/arXiv arXiv 2013
-
[33]
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2024. Voyager: An Open-Ended Embodied Agent with Large Language Models.Transactions on Machine Learning Research (2024). https://openreview.net/forum?id=ehfRiF0R3a
2024
-
[34]
Shenzhi Wang, Chang Liu, Zilong Zheng, Siyuan Qi, Shuo Chen, Qisen Yang, Andrew Zhao, Chaofei Wang, Shiji Song, and Gao Huang. 2023. Avalon’s Game of Thoughts: Battle Against Deception through Recursive Contempla- tion. arXiv:2310.01320 [cs.AI] https://arxiv.org/abs/2310.01320
Pith/arXiv arXiv 2023
-
[35]
Shenzhi Wang, Chang Liu, Zilong Zheng, Siyuan Qi, Shuo Chen, Qisen Yang, Andrew Zhao, Chaofei Wang, Shiji Song, and Gao Huang. 2024. Boosting LLM Agents with Recursive Contemplation for Effective Deception Handling. InFind- ings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association fo...
-
[36]
Tianhe Wang and Tomoyuki Kaneko. 2018. Application of Deep Reinforcement Learning in Werewolf Game Agents.2018 Conference on Technologies and Appli- cations of Artificial Intelligence (TAAI)(2018), 28–33. https://api.semanticscholar. org/CorpusID:57191228
2018
-
[37]
Xiaoqiang Wang and Bang Liu. 2025. OSCAR: Operating System Control via State- Aware Reasoning and Re-Planning. InThe Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=VuTrZzrPfn
2025
-
[38]
Yiqin Wang, Haoji Zhang, Jingqi Tian, and Yansong Tang. 2025. Ponder & Press: Advancing Visual GUI Agent towards General Computer Control. InFindings of the Association for Computational Linguistics: ACL 2025, Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 146...
doi:10.18653/v1/2025 2025
-
[39]
Dekun Wu, Haochen Shi, Zhiyuan Sun, and Bang Liu. 2024. Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mys- tery Games. InFindings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computa- tional Linguistics, Bangkok, Thailand, 822...
-
[40]
Junda Wu, Tong Yu, Xiang Chen, Haoliang Wang, Ryan Rossi, Sungchul Kim, Anup Rao, and Julian McAuley. 2024. DeCoT: Debiasing Chain-of-Thought for Knowledge-Intensive Tasks in Large Language Models via Causal Intervention. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Lun-Wei Ku, Andre Ma...
-
[41]
Shuang Wu, Liwen Zhu, Tao Yang, Shiwei Xu, Qiang Fu, Yang Wei, and Haobo Fu. 2024. Enhance Reasoning for Large Language Models in the Game Werewolf. arXiv:2402.02330 [cs.AI] https://arxiv.org/abs/2402.02330
Pith/arXiv arXiv 2024
-
[42]
Jing Xiang and Seyoung Kim. 2013. A* Lasso for learning a sparse Bayesian network structure for continuous variables.Advances in neural information processing systems26 (2013)
2013
-
[43]
Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, Bin Zhang, Xiong Wang, Yunfei Chu, and Junyang Lin. 2025. Qwen2.5-Omni Technical Report. arXiv:2503.20215 [cs.CL] https://arxiv.org/abs/2503.20215
Pith/arXiv arXiv 2025
-
[44]
Yuzhuang Xu, Shuo Wang, Peng Li, Fuwen Luo, Xiaolong Wang, Weidong Liu, and Yang Liu. 2023. Exploring large language models for communication games: An empirical study on werewolf.arXiv preprint arXiv:2309.04658(2023)
Pith/arXiv arXiv 2023
-
[45]
Zelai Xu, Wanjun Gu, Chao Yu, Yi Wu, and Yu Wang. 2025. Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization. InForty-second International Conference on Machine Learning. https: //openreview.net/forum?id=N2mOBiSqhc
2025
-
[46]
Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli. 2025. Hallucination is Inevitable: An Innate Limitation of Large Language Models. arXiv:2401.11817 [cs.CL] https: //arxiv.org/abs/2401.11817
Pith/arXiv arXiv 2025
-
[47]
Zelai Xu, Chao Yu, Fei Fang, Yu Wang, and Yi Wu. 2024. Language agents with reinforcement learning for strategic play in the Werewolf game. InProceedings of the 41st International Conference on Machine Learning(Vienna, Austria)(ICML’24). JMLR.org, Article 2285, 31 pages
2024
-
[48]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. InInternational Conference on Learning Representations (ICLR)
2023
-
[49]
Junkun Yuan, Xu Ma, Ruoxuan Xiong, Mingming Gong, Xiangyu Liu, Fei Wu, Lanfen Lin, and Kun Kuang. 2023. Instrumental Variable-Driven Domain Gener- alization with Unobserved Confounders.ACM Trans. Knowl. Discov. Data17, 8, Article 118 (June 2023), 21 pages. doi:10.1145/3595380
-
[50]
Matej Zečević, Moritz Willig, Devendra Singh Dhami, and Kristian Kersting. 2023. Causal Parrots: Large Language Models May Talk Causality But Are Not Causal. Transactions on Machine Learning Research(2023)
2023
-
[51]
Yuzhe Zhang, Yipeng Zhang, Yidong Gan, Lina Yao, and Chen Wang. 2024. Causal graph discovery with retrieval-augmented generation based large language mod- els.arXiv preprint arXiv:2402.15301(2024)
Pith/arXiv arXiv 2024
-
[52]
Zheng Zhang, Yihuai Lan, Yangsen Chen, Lei Wang, Xiang Wang, and Hao Wang
-
[53]
Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. 2024. GPT-4V(ision) is a generalist web agent, if grounded. InProceedings of the 41st International Conference on Machine Learning(Vienna, Austria)(ICML’24). JMLR.org, Article 2538, 37 pages
2024
-
[54]
Zhang Zheng, Deheng Ye, Peilin Zhao, and Hao Wang. 2026. The Stackelberg Speaker: Optimizing Persuasive Communication in Social Deduction Games. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Maria Liakata, Viviane P. Moreira, Jiajun Zhang, and David Jurgens (Eds.). Association for Comput...
-
[55]
DVM: Towards Controllable LLM Agents in Social Deduction Games. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1–5. doi:10.1109/ICASSP49660.2025.10888525
arXiv 2025
-
[58]
Qiyang Zhou, Xu Ruihang, Peng Wang, WenJie Lu, Xiaochun Cao, Naiqiang Tan, and Li Shen. 2026. HTAC: Hierarchical Task-Aware Composition for Contin- ual Offline Reinforcement Learning. InForty-third International Conference on Machine Learning. https://openreview.net/forum?id=akfJfpUEBj
2026
-
[2023]
arXiv:2312.17515 [cs.CL] https://arxiv.org/abs/2312.17515
Cooperation on the Fly: Exploring Language Agents for Ad Hoc Teamwork in the Avalon Game. arXiv:2312.17515 [cs.CL] https://arxiv.org/abs/2312.17515
-
[2025]
The Collaborative Sensemaking Play of Jubensha Games: A Deconstruction, Taxonomy, and Analysis.ACM Games3, 1, Article 6 (March 2025), 34 pages. doi:10.1145/3721121
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.