Pith. sign in

REVIEW 2 cited by

PLAYER*: Enhancing LLM-based Multi-Agent Communication and Interaction in Murder Mystery Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.17662 v5 pith:BQ7HSZVU submitted 2024-04-26 cs.CL

classification cs.CL
keywords reasoningplayergamesmmgsagentschallengescomplexinteraction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce WellPlay, a reasoning dataset for multi-agent conversational inference in Murder Mystery Games (MMGs). WellPlay comprises 1,482 inferential questions across 12 games, spanning objectives, reasoning, and relationship understanding, and establishes a systematic benchmark for evaluating agent reasoning abilities in complex social settings. Building on this foundation, we present PLAYER*, a novel framework for Large Language Model (LLM)-based agents in MMGs. MMGs pose unique challenges, including undefined state spaces, absent intermediate rewards, and the need for strategic reasoning through natural language. PLAYER* addresses these challenges with a sensor-based state representation and an information-driven strategy that optimises questioning and suspect pruning. Experiments show that PLAYER* outperforms existing methods in reasoning accuracy, efficiency, and agent-human interaction, advancing reasoning agents for complex social scenarios.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles

    cs.AI 2025-08 unverdicted novelty 6.0 of 10

    TurtleSoup-Bench is a new interactive benchmark showing that LLMs struggle with imaginative reasoning compared to humans.

  2. Sparse Activation Editing for Reliable Instruction Following in Narratives

    cs.CL 2025-05 conditional novelty 6.0 of 10

    An unsupervised SAE-based method that localizes and adjusts instruction-relevant neurons improves instruction adherence and reduces refusals on a new 1,212-example narrative benchmark.

Pith tools