Pith. sign in

Paper Citation Record · LEDGER

Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2407.13943.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.13943 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:20:49.596088Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:18.497365Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation adb3bcfd-249e-4689-9a0c-d2cee680b55d · inbound

A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios cites this paper.

A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:58:45.706246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:58:45.706246Z digest=sha256:cdd69cfc36f6fa5dc0cc958412a977a072d234f417bc8ddbe3f0de4c6e508969

Observation 1642e4c6-d0a5-47e0-9312-1f1565f81288 · inbound

Who Speaks Next? Multi-party AI Discussion Leveraging the Systematics of Turn-taking in Murder Mystery Games cites this paper.

Who Speaks Next? Multi-party AI Discussion Leveraging the Systematics of Turn-taking in Murder Mystery Games Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:10:37.865517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:10:37.865517Z digest=sha256:9e574ac2be5d5876405066f4bdac0981f17c96f98f500cc5efc14785018e1615

Observation a0595463-d9c1-42c6-8e6d-42b7648a6a6f · inbound

Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization cites this paper.

Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T21:58:39.391646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:58:39.391646Z digest=sha256:feec61366ca8ebc51ce3b6d755cce0ce0c4e261bfaeefe405acf50201079f586

Observation b354bf91-0a7a-4a52-beff-d7b843a97177 · inbound

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers cites this paper.

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T22:50:32.949249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:50:32.949249Z digest=sha256:34fcdfce02e476e0be4d9758efb0ee2bbb377a1f36b7a574ba3550c89fca97a5

Observation cc9f7743-f87c-4dd4-92b1-1cfc14760251 · inbound

Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets cites this paper.

Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:22.412122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:22.412122Z digest=sha256:7d4e8164da9f734638872a365abc71033923563751349a83ee1bc345539bde51

Observation 6df2bdba-089d-4aee-bafa-a3f3e4da6171 · inbound

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models cites this paper.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.845155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.845155Z digest=sha256:6d87e06752d7ecd3e4704c4a6c7246ce6f27a32b392b13813e432552c52a23c4

Observation 724b08bb-f7b8-44c3-a214-03db4ff78204 · inbound

TextAtari: 100K Frames Game Playing with Language Agents cites this paper.

TextAtari: 100K Frames Game Playing with Language Agents Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:57.923142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:57.923142Z digest=sha256:c2981a6dd84c871077809168ce034fbb89c10a94ec94687b9803d4b541895d3c

Observation 3be38aa0-8869-492d-812c-2af4af6b75b7 · inbound

Strategy Adaptation in Large Language Model Werewolf Agents cites this paper.

Strategy Adaptation in Large Language Model Werewolf Agents Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:43.264423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:42:43.264423Z digest=sha256:ac694a39372b1e8cdad158173d1b818483b1dcf17a17f4ef1ce816edcd528983

Observation 7eeaeb2a-9457-43d6-85e2-6642c708532a · inbound

Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia cites this paper.

Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:26:24.820227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-18T13:25:17.313704Z digest=sha256:ff539632aa5f84ff45c828106a47b3ee78b7220d2f97be2c5714b03213da35b8

Observation 14bb5245-6fcd-4294-b7f8-0e5807abde3d · inbound

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum cites this paper.

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:43:49.911579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T05:08:54.560648Z digest=sha256:d5be1ffe7fe4d6208d95382b7a6cad693d4a62f6d0059866b47a95e9f0b2bcd3

Observation 41bb56e2-2d5c-40f1-b85f-5f907eb5614c · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.123141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T03:25:24.844859Z digest=sha256:92683ffd63747f7329808890e39bd8433689da31d03be4e037e598bb3ae91610

Observation d4372387-bfcd-4f7f-a16a-7ad8176b37e7 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:26.904795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T06:44:28.552513Z digest=sha256:072cdb632a7c227f473d58aac2c2a6e8cdcbc3fab9ad2abc0d1eb5b85b202823

Observation 9d9b41be-bcdb-4364-a6c3-024f314de6cf · inbound

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs cites this paper.

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:13.222505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T07:15:27.939886Z digest=sha256:ff8721f727d85a8d5d41e865b1084a045af7d5b98c0e610b987d4da0a98df3ba

Observation 45e7a251-a33f-4ef9-82d8-9dec8bf25823 · inbound

RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue cites this paper.

RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:33.728065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T06:34:39.457798Z digest=sha256:aeef204736f4da5f895c1068103478f70a920efc14717aa03af5ea4f0f85c425

Observation 62b9d302-6e6f-4bc0-a0b9-8a6ae95a92a3 · inbound

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play cites this paper.

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:18.499911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T20:57:49.840546Z digest=sha256:5c2e0508d4c17eb50f4fbdd2020073562a0c53657b934777d3438be408333cf7

Observation 4c68c23d-7cae-4249-944a-47e68d77ebdc · inbound

Thinking Out Loud: Real-Time Deception Monitoring in Asymmetric LLM Negotiations cites this paper.

Thinking Out Loud: Real-Time Deception Monitoring in Asymmetric LLM Negotiations Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:35.459536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T07:11:25.568064Z digest=sha256:c52e8f6160d45cd6565a637f76c5f7240c3f43611902ef5e8944579a1f3654e2

Observation 5b6be818-7b78-48d2-aa8f-de906eedda25 · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.661711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:839f193eddd0c577e3a225bd2dc9aa41c870a3f91a25059a028c5303a604c0d2

Observation fd43ebee-2b79-4cf7-a707-d86854e5e49e · inbound

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games cites this paper.

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T10:15:59.479435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:15:59.479435Z digest=sha256:1b5c4553a51af969e03f9e27a1c242ef22a8492f34efe3243c243830b2d02b5b

Observation 3b95b283-e15e-4d31-aedd-e5b1279b29c4 · inbound

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games cites this paper.

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T07:14:59.839689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:14:59.839689Z digest=sha256:057feccc55ed8a169297deaad53c2062efd9dcc584a69d8292da6981c9343591

Observation d3fc3623-1d5b-4819-9f6d-f4875f209f35 · inbound

Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games cites this paper.

Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T09:04:16.361753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:04:16.361753Z digest=sha256:09ad1d8aba5ae0b85663d5cdee659ef1041f2ee5c8ae87fc45b889940726e0aa

Observation bcfcad8a-1b0e-47b7-9701-120e26b7325d · inbound

Cumulative suspicion and absorption dynamics in an agent-based Mafia game cites this paper.

Cumulative suspicion and absorption dynamics in an agent-based Mafia game Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T08:47:10.407663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:47:10.407663Z digest=sha256:90d210342f326c6cad59cfc2f01348cff9366d25340c5cca47662b1b4ff80294

Observation 4d19343b-bada-4ef0-933b-0773c7bfddbb · inbound

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems cites this paper.

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T00:51:15.781283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:51:15.781283Z digest=sha256:806161901be468f38952700a8d06efec44e0c370693cc502d565b049444fdecf

Observation 885495d6-0512-401a-8651-ec9a1ee43306 · inbound

Persuasive and Compliant Tendencies Predict Group Decision-Making in Humans and Language Models cites this paper.

Persuasive and Compliant Tendencies Predict Group Decision-Making in Humans and Language Models Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:49.596088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:49.596088Z digest=sha256:e28d4507b2ab204a3b3f8efbca46d9a124bc807f5184e0c199530c1b17d52abe