Pith. sign in

REVIEW 1 cited by

Digital Player: Evaluating Large Language Models based Human-like Agent in Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.20807 v1 pith:Z2ZGQUOJ submitted 2025-02-28 cs.LG

classification cs.LG
keywords digitalplayersagentshuman-likegamelanguagelargellm-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the rapid advancement of Large Language Models (LLMs), LLM-based autonomous agents have shown the potential to function as digital employees, such as digital analysts, teachers, and programmers. In this paper, we develop an application-level testbed based on the open-source strategy game "Unciv", which has millions of active players, to enable researchers to build a "data flywheel" for studying human-like agents in the "digital players" task. This "Civilization"-like game features expansive decision-making spaces along with rich linguistic interactions such as diplomatic negotiations and acts of deception, posing significant challenges for LLM-based agents in terms of numerical reasoning and long-term planning. Another challenge for "digital players" is to generate human-like responses for social interaction, collaboration, and negotiation with human players. The open-source project can be found at https:/github.com/fuxiAIlab/CivAgent.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. State-Inference-Based Prompting for Natural Language Trading with Game NPCs

    cs.AI 2025-07 conditional novelty 5.0 of 10

    A prompt framework that makes LLM game merchants follow a six-state trading flow achieves over 97% state compliance, over 95% item accuracy, and 99.7% price accuracy in simulated dialogues.

Pith tools