Pith. sign in

REVIEW 2 cited by

How Different AI Chatbots Behave? Benchmarking Large Language Models in Behavioral Economics Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.12362 v1 pith:H5VN6ASX submitted 2024-12-16 cs.AI cs.CL

classification cs.AIcs.CL
keywords behavioralbenchmarkingchatbotsdecision-makingdeploymenteconomicsgameslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The deployment of large language models (LLMs) in diverse applications requires a thorough understanding of their decision-making strategies and behavioral patterns. As a supplement to a recent study on the behavioral Turing test, this paper presents a comprehensive analysis of five leading LLM-based chatbot families as they navigate a series of behavioral economics games. By benchmarking these AI chatbots, we aim to uncover and document both common and distinct behavioral patterns across a range of scenarios. The findings provide valuable insights into the strategic preferences of each LLM, highlighting potential implications for their deployment in critical decision-making roles.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Be.FM: Open Foundation Models for Human Behavior

    cs.AI 2025-05 reject novelty 5.0 of 10

    Be.FM fine-tunes Llama models on behavioral data and claims improved behavior prediction, but its headline evaluation is compromised by testing on the same data it trained on.

  2. AI Agent Behavioral Science

    q-bio.NC 2025-06 conditional novelty 4.0 of 10

    AI agents should be studied as behavioral entities shaped by context and interaction, not only as trained models.

Pith tools