Pith. sign in

REVIEW 4 cited by

Decision-Making Behavior Evaluation Framework for LLMs under Uncertain Context

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.05972 v2 pith:ER7HQYFE submitted 2024-06-10 cs.AI cs.CYcs.HCcs.LGecon.TH

classification cs.AIcs.CYcs.HCcs.LGecon.TH
keywords llmsdecision-makingaversionbehaviorriskethicallosswhen
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

When making decisions under uncertainty, individuals often deviate from rational behavior, which can be evaluated across three dimensions: risk preference, probability weighting, and loss aversion. Given the widespread use of large language models (LLMs) in decision-making processes, it is crucial to assess whether their behavior aligns with human norms and ethical expectations or exhibits potential biases. Several empirical studies have investigated the rationality and social behavior performance of LLMs, yet their internal decision-making tendencies and capabilities remain inadequately understood. This paper proposes a framework, grounded in behavioral economics, to evaluate the decision-making behaviors of LLMs. Through a multiple-choice-list experiment, we estimate the degree of risk preference, probability weighting, and loss aversion in a context-free setting for three commercial LLMs: ChatGPT-4.0-Turbo, Claude-3-Opus, and Gemini-1.0-pro. Our results reveal that LLMs generally exhibit patterns similar to humans, such as risk aversion and loss aversion, with a tendency to overweight small probabilities. However, there are significant variations in the degree to which these behaviors are expressed across different LLMs. We also explore their behavior when embedded with socio-demographic features, uncovering significant disparities. For instance, when modeled with attributes of sexual minority groups or physical disabilities, Claude-3-Opus displays increased risk aversion, leading to more conservative choices. These findings underscore the need for careful consideration of the ethical implications and potential biases in deploying LLMs in decision-making scenarios. Therefore, this study advocates for developing standards and guidelines to ensure that LLMs operate within ethical boundaries while enhancing their utility in complex decision-making environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Recovering Event Probabilities from Large Language Model Embeddings via Axiomatic Constraints

    cs.CL 2025-05 conditional novelty 7.0 of 10

    A sign-flip constraint on one latent dimension of a VAE trained on LLM embeddings yields complementary event probabilities that sum to near one and track true probabilities on held-out dice events.

  2. Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty

    cs.AI 2025-08 unverdicted novelty 6.0 of 10

    Prospect Theory parameters estimated for LLMs become unstable when epistemic markers are injected into prompts, indicating the model is not robust for decision-making under linguistic uncertainty.

  3. Adaptive Preference Optimization with Uncertainty-aware Utility Anchor

    cs.LG 2025-09 conditional novelty 5.0 of 10

    UAPO decomposes pairwise preference loss into two anchor-based terms, enabling offline LLM alignment with unpaired feedback at competitive benchmark scores.

  4. A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios

    cs.CL 2024-12 conditional novelty 3.0 of 10

    LLM-based game-playing agents are surveyed across choice-focused and communication-focused games, with a comparative performance table and future directions.

Pith tools