REVIEW 4 major objections 5 minor 1 cited by
MafiaScope claims it can tell whether an LLM agent in a social deduction game lost by misreading the situation or by wasting a correct read, and its case study finds that most losses are locked in by wrong beliefs.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 07:09 UTC pith:DBYV2PCG
load-bearing objection Genuinely reusable testbed; headline causal claim rests on probe self-reports whose validation is weak on the pinned corpus — referee it, but push for stronger belief-faithfulness evidence. the 4 major comments →
MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
MafiaScope's central claim is that an LLM agent's private beliefs in a social deduction game can be measured non-invasively and used to explain the agent's outcomes. The method pauses the game after every public utterance and night action, puts each agent's message history on a throwaway copy, and asks structured questions—what roles the agent believes, who it suspects, what it plans to do, and who it thinks suspects it—then logs the answers without letting them affect the game. Those answers are scored against engine ground truth, or against the target agent's own same-step reports for second-order beliefs. The paper then uses saved snapshots to resample each agent's utterance 500 times fro
What carries the argument
The central instrument is the probe battery plus the snapshot-and-replay fork engine. The probe engine pauses after each public event and privately asks each agent structured questions on a throwaway copy of its message list—never written back—so the probed game is distributionally identical to an unprobed one (an A/B run finds Mafia win rates 70.0% vs. 74.2%, Fisher p=0.57). The default battery includes role beliefs, role assessment with confidence, suspicion ranking, planned action, and a second-order 'social map' of how each player feels about the agent; the social map is scored against the target player's own same-step report. The replay engine saves every agent's full context at every s
Load-bearing premise
The load-bearing premise is that the private answers agents give when probed truthfully reflect what they believe; the paper itself notes in its limitations that these are self-reports that may be unfaithful, and the main validation—that probes predict votes—is not clearly better than simply reading the transcript in the primary corpus.
What would settle it
Take a game where the agent is secretly shown who the Mafia is. If the probe does not report that identity with high confidence, or if the agent's subsequent vote does not follow that probed belief, then the probe is not measuring the agent's decision-relevant beliefs. A cheaper check: re-probe the same frozen context several times and see whether the agent's enforced vote changes as often as the probe does; the paper reports a 34.8% top-suspect flip on re-probing, so a divergence between probe instability and vote stability would expose unfaithful self-reports.
If this is right
- Outcome-only evaluation is incomplete: an agent that loses can be either wrong about the world or right-but-ineffective, and MafiaScope claims to separate these with causal replay.
- The counterfactual corpus can serve as a debugging resource: each forked step records the belief state, the resampled utterance, and the resulting outcome, so a practitioner can localize a failure to a specific belief or action.
- The probe's non-invasiveness, supported by the A/B comparison showing no significant change in Mafia win rate, means the belief measurements are not an artefact of the elicitation changing play—a precondition for using the tool as a measurement device.
- The calibration and suspicion-overprediction findings, if they hold across models, imply that LLM verbalized confidence should be recalibrated before being interpreted as probability, and that second-order 'they suspect me' judgments carry a systematic egocentric bias.
Where Pith is reading between the lines
- Editorial inference: if the 34.8% test-retest flip rate on frozen contexts generalizes, then a large fraction of the apparent 'changing one's mind' in LLM social reasoning is sampling noise; belief-dynamics studies should always report a test-retest floor.
- Editorial inference: the weak separation from a transcript heuristic on the pinned corpus suggests the probe's added value is not vote prediction but the structured, per-timestep attribution channel; users should validate the belief readout against independent behavioral commitments before relying on it.
- Editorial inference: the result that wrong-belief votes are locked in and resampling cannot rescue them implies, if general, that improving LLM agents' social reasoning may require interventions on belief acquisition and updating rather than on utterance selection or action policy.
- Editorial inference: the same probe-and-fork pattern could be transferred to other hidden-role games (Werewolf, Resistance-style missions) to test whether the 'loss decided before the utterance' pattern is specific to Mafia or recurrent across social-deduction settings; the paper notes the engine already supports these reskins.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. MafiaScope is an open-source testbed for non-invasive belief probing of LLM agents in the social deduction game Mafia. After every public utterance, each agent privately answers probe questions that never re-enter the game; answers are scored against engine ground truth and the target agents' own same-step reports. The system includes a visualizer, counterfactual fork-and-replay, and a released corpus of probed games and forks. A case study across two model families reports four findings: (F1) belief trajectories converge toward ground truth on DeepSeek but not gpt-4o-mini; (F2) stated confidence is poorly calibrated (ECE 0.168); (F3) agents over-predict being suspected by a factor of 1.53; (F4) probed beliefs track votes with 64.9% top-1 accuracy on the pinned corpus, not significantly above a transcript heuristic; and (F5) probe token budgets affect truncation. A counterfactual experiment on 16 votes (8 policy-gap, 8 perception-gap) claims that outcome flips concentrate at votes made under a correct assessment, while wrong-assessment votes are behaviorally locked in.
Significance. The paper makes a commendable systems contribution: a reusable, well-engineered, open-source pipeline for recording private beliefs, visualizing them, and running counterfactual forks, with a substantial released dataset and replication corpora. The case study is honest about its exploratory nature and reveals interesting phenomena (overconfidence, a spotlight effect, probe-budget effects). The central claim—that the tool can causally distinguish policy failures from perception failures—would be a meaningful advance for LLM social-reasoning evaluation if the evidence supports it. The release of code, data, and 512 forks is a concrete strength. The main weakness is that the causal distinction rests on probe-derived labels whose faithfulness is not convincingly validated on the corpus used for the causal experiment.
major comments (4)
- [§6, §7 F4, §9] The central policy/perception distinction in §6 is built on quadrant labels from the same probe instrument whose faithfulness is the paper's own stated limitation ('Probe answers are self-reports and may be unfaithful', §9). On the pinned 32-game corpus that produces the §6 labels, the behavioral validation is weak: pre-vote probe top-1 vote prediction (64.9%) is not significantly above a transcript heuristic (58.8%, McNemar p=0.66), and post-vote probes match the vote 95.1%, consistent with re-describing the action rather than revealing a stable belief. If probes are artifacts of the elicitation, the headline 'flips concentrate where the agent had read the game right' is measuring the probe, not the agent. Please provide an independent validation of probe faithfulness on the pinned corpus (e.g., transcript-only belief estimators, pre-registered vote predictions) or restrict the causal c
- [§6, statistical analysis] The Fisher exact test on pooled forks (7/46 vs 2/114, p=0.0025) treats forks as independent, but each fork is nested within one of 16 parent votes (8 per quadrant), and variants from the same parent share an utterance distribution. The paper itself states 'the unit of evidence (not any single fork)' yet pools fork-level counts. This can lead to anti-conservative p-values. Report per-parent flip rates with confidence intervals, or use a cluster bootstrap / mixed-effects model. Also, with only one replay per fork, outcome variance includes stochastic game continuation; at least report the variance across a few repeated replays.
- [§6, vote selection] The selection of the 16 'selected' innocent day votes (8 policy-gap, 8 perception-gap) is not described. Were they randomly sampled from all votes, or chosen to maximize separation? The control (8 randomly picked votes) is described in one sentence without per-vote outcomes or a matching protocol. This selection ambiguity undermines the generalizability of the quadrant contrast and could introduce bias that is not canceled by 'same procedure' reasoning. Please specify the sampling protocol, the matching variables (round, target proximity, confidence), and report the control's results in the same format as the main experiment.
- [Abstract, §1, §7] The abstract and §1 state that MafiaScope 'can tell' whether an agent lost by misreading the game or by wasting a correct read, and §6 uses the word 'causally,' while §7 explicitly disclaims that the analyses 'illustrate how MafiaScope supports model inspection rather than establish new behavioural claims.' The causal claim is load-bearing and currently outruns the evidence base (16 hand-selected votes, one replay per fork, unvalidated probe labels on the pinned corpus). Either provide the additional validation requested above, or soften the abstract/introduction to describe the tool's capability as a demonstrative framework rather than an established causal attribution.
minor comments (5)
- [§4] Typo: 'Y AML' should be 'YAML' (two occurrences in the probe DSL paragraph).
- [§4, A/B test] The non-invasiveness A/B test reports 'Mafia wins at the same rate in both arms (70.0% vs. 74.2%)'; clarify which arm is probed and which is unprobed, and whether the 4.2-point difference direction is consistent with a probe effect.
- [§6, Figure 5] The caption says 'here 4 of 20 flip'; clarify whether these are forks with a single replay per fork or from the pooled counts, to avoid confusion with the 15.2% claim.
- [Appendix B] The test-retest floor is reported as 34.8% (CI [23.9, 47.1]); the CI seems wide for 40 points re-probed 5 times; consider reporting the number of comparisons per point and the reweighting procedure.
- [References] Reference list formatting is inconsistent (some entries contain 'ArXiv' vs 'arXiv', and a few have extra whitespace); please normalize.
Circularity Check
No significant circularity: measurements anchor to engine ground truth or behavior; self-report caveats are explicit, not hidden reductions.
full rationale
The central measurements are anchored outside the probe instrument. F1 accuracy is scored against engine-known roles; F2 calibration compares stated confidence against ground-truth accuracy; F4 compares pre-vote probe rankings to actual votes (external behavior). The only self-referential quantity is F3, but the paper explicitly defines it as "a consistency between two self-reports, not ToM against engine ground truth" (Appendix C), so the 1.53 over-prediction ratio is an operational comparison of two probe answers, not a hidden fit or a forced equation. The §6 counterfactual uses engine-simulated replay outcomes; quadrant labels come from the probe instrument, but the paper flags "quadrant labels from the same probe instrument" and the flip rates are generated by 500 resampled utterances, not derived from the labels by construction. There are no load-bearing self-citations and no imported uniqueness theorems. The limitations the paper concedes (probe unfaithfulness, non-significant F4 gain on the pinned corpus, post-vote report matching action) are validity and robustness concerns, not circularity. The derivation chain is therefore self-contained in the sense required by this review.
Axiom & Free-Parameter Ledger
free parameters (3)
- replay temperature =
1.2
- variant selection size =
20 (19 + factual), farthest-point cosine over e5 embeddings
- probe token budgets =
960 (case study), 400 (replication batch)
axioms (3)
- domain assumption Private probe self-reports are usable measures of agent belief
- domain assumption Probing does not change the game distribution
- domain assumption Frozen-context resampling at temperature 1.2 is a representative counterfactual distribution
read the original abstract
An LLM agent's public behaviour reveals little about its social reasoning: an agent that votes correctly may be guessing, and an agent that lies well leaves no trace of what it actually believes. We present MafiaScope, an open testbed that turns the social deduction game Mafia into a measurement instrument for machine Theory of Mind. It distinguishes whether an agent lost because it misread the game or because it failed to act on a correct assessment, a distinction that is invisible from outcomes and dialogue transcripts alone. After every public utterance, each agent privately answers structured probe questions whose responses never re-enter the game and are scored against the ground truth known to the engine. An interactive visualizer replays games from the perspective of an individual agent's beliefs, displays timeline-aligned accuracy and calibration, and supports counterfactual replay from any recorded step. In a case study across two model families comprising tens of thousands of parsed probe responses, we find that stated confidence is poorly calibrated, agents overestimate how often they are suspected by a factor of 1.5, and single-vote counterfactual replays rarely change game outcomes: outcome flips occur primarily when the agent had already formed a correct belief state, whereas decisions made under an incorrect model of the world remain largely unchanged under resampling. The engine, visualizer, recorded games, and counterfactual replay corpus are released under an open-source licence. Code: https://github.com/karpovilia/mafiascope. Live demo: https://karpovilia.github.io/mafiascope/. Screencast: https://vimeo.com/1208920221.
Figures
Forward citations
Cited by 1 Pith paper
-
Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
Frontier LLMs win Secret Hitler matches and can deceive, but most fail to keep a consistent false persona as evidence accumulates, with DRR often falling below 50%.
Reference graph
Works this paper leans on
-
[1]
2023 , eprint =
Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf , author =. 2023 , eprint =
2023
-
[2]
Proceedings of the 41st International Conference on Machine Learning , series =
Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game , author =. Proceedings of the 41st International Conference on Machine Learning , series =. 2024 , url =
2024
-
[3]
2024 , eprint =
Enhance Reasoning for Large Language Models in the Game Werewolf , author =. 2024 , eprint =
2024
-
[4]
Werewolf Arena: A Case Study in
Bailis, Suma and Friedhoff, Jane and Chen, Feiyang , year =. Werewolf Arena: A Case Study in. 2407.13943 , archivePrefix =
-
[5]
Xia, Xinyuan and Song, Yuanyi and Ma, Haomin and Cai, Jinyu , year =. 2506.12841 , archivePrefix =
-
[6]
Light, Jonathan and Cai, Min and Shen, Sheng and Hu, Ziniu , year =. 2310.05036 , archivePrefix =
-
[7]
2023 , eprint =
Avalon's Game of Thoughts: Battle Against Deception through Recursive Contemplation , author =. 2023 , eprint =
2023
-
[8]
Chi, Yizhou and Mao, Lingjun and Tang, Zineng , year =. 2407.16521 , archivePrefix =
-
[9]
2025 , eprint =
Among Us: A Sandbox for Measuring and Detecting Agentic Deception , author =. 2025 , eprint =
2025
-
[10]
2023 , eprint =
Hoodwinked: Deception and Cooperation in a Text-Based Game for Language Models , author =. 2023 , eprint =
2023
-
[11]
Putting the Con in Context: Identifying Deceptive Actors in the Game of Mafia
Ibraheem, Samee and Zhou, Gaoyue and DeNero, John. Putting the Con in Context: Identifying Deceptive Actors in the Game of Mafia. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. doi:10.18653/v1/2022.naacl-main.11
-
[12]
2025 , eprint =
Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia , author =. 2025 , eprint =
2025
-
[13]
Scientific Reports , volume =
Finding deceivers in social context with large language models and how to find them: the case of the Mafia game , author =. Scientific Reports , volume =. 2024 , doi =
2024
-
[14]
Science , volume =
Human-level play in the game of Diplomacy by combining language models with strategic reasoning , author =. Science , volume =. 2022 , doi =
2022
-
[15]
Advances in Neural Information Processing Systems 32 (NeurIPS 2019) , pages =
Finding Friend and Foe in Multi-Agent Games , author =. Advances in Neural Information Processing Systems 32 (NeurIPS 2019) , pages =. 2019 , url =
2019
-
[16]
2025 , doi =
Zhang, Zheng and Xiao, Nuoqian and Chai, Qi and Ye, Deheng and Wang, Hao , booktitle =. 2025 , doi =
2025
-
[17]
Suspicion-Agent: Playing Imperfect Information Games with Theory of Mind Aware
Guo, Jiaxian and Yang, Bo and Yoo, Paul and Lin, Bill Yuchen and Iwasawa, Yusuke and Matsuo, Yutaka , year =. Suspicion-Agent: Playing Imperfect Information Games with Theory of Mind Aware. 2309.17277 , archivePrefix =
-
[18]
Revisiting the Evaluation of Theory of Mind through Question Answering
Le, Matthew and Boureau, Y-Lan and Nickel, Maximilian. Revisiting the Evaluation of Theory of Mind through Question Answering. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. doi:10.18653/v1/D19-1598
-
[19]
Wu, Yufan and He, Yinghui and Jia, Yilin and Mihalcea, Rada and Chen, Yulong and Deng, Naihao. Hi- T o M : A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models. Findings of the Association for Computational Linguistics: EMNLP 2023. 2023. doi:10.18653/v1/2023.findings-emnlp.717
-
[20]
FANT o M : A Benchmark for Stress-testing Machine Theory of Mind in Interactions
Kim, Hyunwoo and Sclar, Melanie and Zhou, Xuhui and Le Bras, Ronan and Kim, Gunhee and Choi, Yejin and Sap, Maarten. FANT o M : A Benchmark for Stress-testing Machine Theory of Mind in Interactions. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.18653/v1/2023.emnlp-main.890
-
[21]
Xu, Hainiu and Zhao, Runcong and Zhu, Lixing and Du, Jinhua and He, Yulan. O pen T o M : A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. doi:10.18653/v1/2024.acl-long.466
-
[22]
Advances in Neural Information Processing Systems 36 (NeurIPS 2023), Datasets and Benchmarks Track , year =
Understanding Social Reasoning in Language Models with Language Models , author =. Advances in Neural Information Processing Systems 36 (NeurIPS 2023), Datasets and Benchmarks Track , year =
2023
-
[23]
MMT o M - QA : Multimodal Theory of Mind Question Answering
Jin, Chuanyang and Wu, Yutong and Cao, Jing and Xiang, Jiannan and Kuo, Yen-Ling and Hu, Zhiting and Ullman, Tomer and Torralba, Antonio and Tenenbaum, Joshua and Shu, Tianmin. MMT o M - QA : Multimodal Theory of Mind Question Answering. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. ...
-
[24]
arXiv preprint arXiv:2302.08399 , year =
Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks , author =. arXiv preprint arXiv:2302.08399 , year =. doi:10.48550/arXiv.2302.08399 , url =
-
[25]
Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LM s
Sap, Maarten and Le Bras, Ronan and Fried, Daniel and Choi, Yejin. Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LM s. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. doi:10.18653/v1/2022.emnlp-main.248
-
[26]
Proceedings of the National Academy of Sciences , volume =
Evaluating large language models in theory of mind tasks , author =. Proceedings of the National Academy of Sciences , volume =. 2024 , publisher =. doi:10.1073/pnas.2405460121 , url =
-
[27]
Theory of Mind in Large Language Models: Examining Performance of 11 State-of-the-Art models vs
van Duijn, Max and van Dijk, Bram and Kouwenhoven, Tom and de Valk, Werner and Spruit, Marco and van der Putten, Peter. Theory of Mind in Large Language Models: Examining Performance of 11 State-of-the-Art models vs. Children Aged 7-10 on Advanced Tests. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL). 2023. doi:10.18...
-
[28]
arXiv preprint arXiv:2207.05221 , year =
Language Models (Mostly) Know What They Know , author =. arXiv preprint arXiv:2207.05221 , year =
-
[29]
Advances in Neural Information Processing Systems , volume =
Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting , author =. Advances in Neural Information Processing Systems , volume =. 2023 , url =
2023
-
[30]
arXiv preprint arXiv:2410.13787 , year =
Looking Inward: Language Models Can Learn About Themselves by Introspection , author =. arXiv preprint arXiv:2410.13787 , year =
-
[31]
Transformer Circuits Thread , year =
Emergent Introspective Awareness in Large Language Models , author =. Transformer Circuits Thread , year =
-
[32]
and Levinstein, Benjamin A
Herrmann, Daniel A. and Levinstein, Benjamin A. , journal =. Standards for Belief Representations in. 2025 , doi =
2025
-
[33]
Do Language Models Have Beliefs?
Hase, Peter and Diab, Mona and Celikyilmaz, Asli and Li, Xian and Kozareva, Zornitsa and Stoyanov, Veselin and Bansal, Mohit and Iyer, Srinivasan , journal =. Do Language Models Have Beliefs?. 2021 , url =
2021
-
[34]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , year =
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback , author =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , year =. doi:10.18653/v1/2023.emnlp-main.330 , url =
-
[35]
2024 , address =
Liu, Ziyi and Anand, Abhishek and Zhou, Pei and Huang, Jen-tse and Zhao, Jieyu , booktitle =. 2024 , address =
2024
-
[36]
2024 , url =
Zhou, Xuhui and Zhu, Hao and Mathur, Leena and Zhang, Ruohong and Yu, Haofei and Qi, Zhengyang and Morency, Louis-Philippe and Bisk, Yonatan and Fried, Daniel and Neubig, Graham and Sap, Maarten , booktitle =. 2024 , url =
2024
-
[37]
Generative Agents: Interactive Simulacra of Human Behavior , author =. Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST '23) , year =. doi:10.1145/3586183.3606763 , url =
-
[38]
2024 , url =
Chen, Weize and Su, Yusheng and Zuo, Jingwei and Yang, Cheng and Yuan, Chenfei and Chan, Chi-Min and Yu, Heyang and Lu, Yaxi and Hung, Yi-Hsin and Qian, Chen and Qin, Yujia and Cong, Xin and Xie, Ruobing and Liu, Zhiyuan and Sun, Maosong and Zhou, Jie , booktitle =. 2024 , url =
2024
-
[39]
2023 , howpublished =
Wu, Yuxiang and Jiang, Zhengyao and Khan, Akbir and Fu, Yao and Ruis, Laura and Grefenstette, Edward and Rockt. 2023 , howpublished =
2023
-
[40]
and Burger, Doug and Wang, Chi , booktitle =
Wu, Qingyun and Bansal, Gagan and Zhang, Jieyu and Wu, Yiran and Li, Beibin and Zhu, Erkang and Jiang, Li and Zhang, Xiaoyun and Zhang, Shaokun and Liu, Jiale and Awadallah, Ahmed Hassan and White, Ryen W. and Burger, Doug and Wang, Chi , booktitle =. 2024 , url =
2024
-
[41]
Interactive Debugging and Steering of Multi-Agent
Epperson, Will and Bansal, Gagan and Dibia, Victor and Fourney, Adam and Gerrits, Jack and Zhu, Erkang and Amershi, Saleema , booktitle =. Interactive Debugging and Steering of Multi-Agent. 2025 , publisher =. doi:10.1145/3706598.3713581 , url =
arXiv 2025
-
[42]
Language Models with Rationality , author =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , month = dec, year =. doi:10.18653/v1/2023.emnlp-main.877 , url =
-
[43]
Coscia, Adam and Guo, Shunan and Koh, Eunyee and Endert, Alex , booktitle =. 2025 , publisher =. doi:10.1145/3746059.3747746 , url =
arXiv 2025
-
[44]
Tan, Chao-Hong and Gu, Jia-Chen and Ling, Zhen-Hua , booktitle =. Is. 2023 , address =. doi:10.18653/v1/2023.findings-emnlp.326 , url =
-
[45]
arXiv preprint arXiv:2304.13835 , year =
Multi-Party Chat: Conversational Agents in Group Settings with Humans and Models , author =. arXiv preprint arXiv:2304.13835 , year =
-
[46]
Inoue, Koji and Lala, Divesh and Elmers, Mikey and Ochi, Keiko and Kawahara, Tatsuya , booktitle =. An. 2025 , address =
2025
-
[47]
2024 , url =
Kuratov, Yuri and Bulatov, Aydar and Anokhin, Petr and Rodkin, Ivan and Sorokin, Dmitry and Sorokin, Artyom and Burtsev, Mikhail , booktitle =. 2024 , url =
2024
-
[48]
2024 , address =
Jiang, Hang and Zhang, Xiajie and Cao, Xubo and Breazeal, Cynthia and Roy, Deb and Kabbara, Jad , booktitle =. 2024 , address =
2024
-
[49]
arXiv preprint arXiv:2307.00184 , year =
Personality Traits in Large Language Models , author =. arXiv preprint arXiv:2307.00184 , year =
-
[50]
2024 , address =
Wang, Xintao and Xiao, Yunze and Huang, Jen-tse and Yuan, Siyu and Xu, Rui and Guo, Haoran and Tu, Quan and Fei, Yaying and Leng, Ziang and Wang, Wei and Chen, Jiangjie and Li, Cheng and Xiao, Yanghua , booktitle =. 2024 , address =
2024
-
[51]
Journal of Personality and Social Psychology , volume =
The Spotlight Effect in Social Judgment: An Egocentric Bias in Estimates of the Salience of One's Own Actions and Appearance , author =. Journal of Personality and Social Psychology , volume =. 2000 , doi =
2000
-
[52]
Psychological Review , volume =
The Trouble with Overconfidence , author =. Psychological Review , volume =. 2008 , doi =
2008
-
[53]
Proceedings of the 34th International Conference on Machine Learning , series =
On Calibration of Modern Neural Networks , author =. Proceedings of the 34th International Conference on Machine Learning , series =. 2017 , url =
2017
-
[54]
2025 , note =
Kazutoshi Shinoda and Nobukatsu Hojo and Kyosuke Nishida and Saki Mizuno and Keita Suzuki and Ryo Masumura and Hiroaki Sugiyama and Kuniko Saito , booktitle =. 2025 , note =
2025
-
[55]
Proceedings of the 24th International Conference on Autonomous Agents and Multiagent Systems (AAMAS) , year =
Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning , author =. Proceedings of the 24th International Conference on Autonomous Agents and Multiagent Systems (AAMAS) , year =
-
[56]
Proceedings of the 41st International Conference on Machine Learning (ICML) , series =
Language Models Represent Beliefs of Self and Others , author =. Proceedings of the 41st International Conference on Machine Learning (ICML) , series =. 2024 , note =
2024
-
[57]
Findings of the Association for Computational Linguistics: EMNLP 2023 , pages =
Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models , author =. Findings of the Association for Computational Linguistics: EMNLP 2023 , pages =. 2023 , address =
2023
-
[58]
Proceedings of the 42nd International Conference on Machine Learning (ICML), Position Paper Track , year =
Position: Theory of Mind Benchmarks are Broken for Large Language Models , author =. Proceedings of the 42nd International Conference on Machine Learning (ICML), Position Paper Track , year =
-
[59]
Mrinal Agarwal and Saad Rana and Theo Sundoro and Hermela Berhe and Spencer Kim and Vasu Sharma and Sean O'Brien and Kevin Zhu , year =
-
[60]
Yuan, Ye and Song, Rui and Li, Weien and Li, Zeyu and Liu, Haochen and Kong, Xiangyu and Han, Changjiang and Yang, Yonghan and Zhao, Zichen and Dong, Zixuan and Lyu, Fuyuan and He, Bowei and Wu, Haolun and Kang, Jikun and Liu, Xue , year =
-
[61]
2026 , note =
Bayesian Social Deduction with Graph-Informed Language Models , author =. 2026 , note =
2026
-
[62]
2020 , note =
Temporal Graph Networks for Deep Learning on Dynamic Graphs , author =. 2020 , note =
2020
-
[63]
International Conference on Learning Representations (ICLR) , year =
Inductive Representation Learning on Temporal Graphs , author =. International Conference on Learning Representations (ICLR) , year =
-
[64]
Journal of Machine Learning Research , volume =
Representation Learning for Dynamic Graphs: A Survey , author =. Journal of Machine Learning Research , volume =
-
[65]
2024 , note =
Learning to Discuss Strategically: A Case Study on One Night Ultimate Werewolf , author =. 2024 , note =
2024
-
[66]
2024 , publisher =
Chen, Zhuang and Wu, Jincenzi and Zhou, Jinfeng and Wen, Bosi and Bi, Guanqun and Jiang, Gongyao and Cao, Yaru and Hu, Mengting and Lai, Yunghwei and Xiong, Zexuan and Huang, Minlie , booktitle =. 2024 , publisher =
2024
-
[67]
Findings of the Association for Computational Linguistics: ACL 2023 , year =
Werewolf Among Us: Multimodal Resources for Modeling Persuasion Behaviors in Social Deduction Games , author =. Findings of the Association for Computational Linguistics: ACL 2023 , year =
2023
-
[68]
2024 , publisher =
Lan, Yihuai and Hu, Zhiqiang and Wang, Lei and Wang, Yang and Ye, Deheng and Zhao, Peilin and Lim, Ee-Peng and Xiong, Hui and Wang, Hao , booktitle =. 2024 , publisher =
2024
-
[69]
2026 , eprint =
Wang, Kevin and Th. 2026 , eprint =
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.