Pith. sign in

REVIEW 4 cited by

STEER: Assessing the Economic Rationality of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.09552 v2 pith:MWNBY2HP submitted 2024-02-14 cs.CL econ.GNq-fin.EC

classification cs.CLecon.GNq-fin.EC
keywords shouldagenteconomicllmsassessingdifferentelementsexhibit
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

There is increasing interest in using LLMs as decision-making "agents." Doing so includes many degrees of freedom: which model should be used; how should it be prompted; should it be asked to introspect, conduct chain-of-thought reasoning, etc? Settling these questions -- and more broadly, determining whether an LLM agent is reliable enough to be trusted -- requires a methodology for assessing such an agent's economic rationality. In this paper, we provide one. We begin by surveying the economic literature on rational decision making, taxonomizing a large set of fine-grained "elements" that an agent should exhibit, along with dependencies between them. We then propose a benchmark distribution that quantitatively scores an LLMs performance on these elements and, combined with a user-provided rubric, produces a "STEER report card." Finally, we describe the results of a large-scale empirical experiment with 14 different LLMs, characterizing the both current state of the art and the impact of different model sizes on models' ability to exhibit rational behavior.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Position Auctions in AI-Generated Content

    cs.GT 2025-06 conditional novelty 6.0 of 10

    New mechanism-design results for position auctions with context-dependent click-through rates under multinomial logit and cascade user models, with exact optimality in the MNL case and an O(log m) approximation in the...

  2. AI Agent Behavioral Science

    q-bio.NC 2025-06 conditional novelty 4.0 of 10

    AI agents should be studied as behavioral entities shaped by context and interaction, not only as trained models.

  3. An Economy of AI Agents

    econ.GN 2025-09 accept novelty 2.0 of 10

    A survey chapter that maps open economic questions about AI agents in markets, organizations, and institutions, arguing that current theories may need extension.

  4. Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs

    cs.LG 2025-02

Pith tools