Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Should I State or Should I Show? Aligning AI with Human Preferences

T0 review · 3 major / 6 minor · reviewed 2026-07-13 · grok-4.5

Pith's one-line read Revealed lottery choices align AI agents with human risk preferences better than written prompts, because people struggle to state what their choices already show.

desk verdict Clean incentivized head-to-head: revealed lottery choices beat free-text prompts for frontier LLM alignment on risk, with real mechanism checks and a model-dependent conflict wrinkle. read the letter →

arxiv 2603.29317 v2 pith:73NXTGGJ submitted 2026-03-31 econ.GN q-fin.EC

classification econ.GNq-fin.EC
keywords AIalignmentrevealedpreferencestatedchoiceunderriskpromptengineeringhuman-AIdelegationlargelanguagemodelsprincipal-agentproblem
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

As AI agents take on more autonomous decisions, a core problem is that people often cannot fully write down the preferences those agents should follow. This paper tests that problem in choice under risk. Subjects first reveal preferences by choosing between binary lotteries, then write free-text instructions for an AI agent, then face a new set of similar lotteries that serve as ground truth. On average, an AI agent given only the past choices predicts the later choices more accurately than an AI agent given only the written prompt. The gap is largest for people whose choices show Allais-type inconsistencies, and evidence points to writing difficulty rather than incomplete preference data: AI-generated prompts based on the same information close most of the gap, and subjects' own prompts only partly predict even the choices they made before writing. People also misjudge which source works better and often delegate to the weaker agent. When both sources are given together, some models overweight the less accurate prompt when the two conflict. The paper therefore treats revealed preference as a strong communication channel for AI alignment, but one that needs careful design around human self-knowledge and how models resolve conflicts.

What carries the argument

Out-of-sample match rate: the share of Part II lottery choices correctly predicted by Data-AI (trained on Part I choices) versus Prompt-AI (given the free-text instruction), with Both-AI used to study how models resolve conflicts between the two sources.

What would settle it

Replicate the design with the same Part I/Part II lottery structure and frontier models; if Prompt-AI's match rate equals or exceeds Data-AI's once subjects receive prompt coaching or structured templates, or if Both-AI no longer underperforms Data-AI when models are forced to weight data over text on conflicts, the central ranking fails.

Watch

Extended reading notes

Core claim

On average, an AI agent given a subject's revealed-preference lottery choices predicts that subject's later choices under risk more accurately than an AI agent given the same subject's written prompt. The performance gap is driven mainly by subjects' difficulty translating their own preferences into clear written instructions, not by a lack of information in the choice data itself.

Load-bearing premise

That subjects' later incentivized lottery choices are the right ground truth for the preferences the AI should implement, and that a short set of binary lotteries plus a free-text prompt stand in for how people will actually hand preferences to agentic systems.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper studies how well an AI agent can implement a human principal's preferences over binary lotteries when given either revealed-preference choice data (Data-AI) or free-text stated-preference prompts (Prompt-AI). In an incentivized online experiment (N=290 after exclusions), subjects complete Part I lottery choices, write a prompt, and choose which information source to delegate; Part II choices (structurally similar lotteries) serve as held-out ground truth for out-of-sample match rates. On average Data-AI predicts Part II choices more accurately than Prompt-AI. The gap is attributed to subjects' difficulty articulating preferences: prompts only partially recover Part I choices; AI-generated prompts close the gap; and semantic similarity of human prompts to AI prompts shrinks the performance difference. Subjects often misperceive relative agent quality and frequently fail to select the better source. When both sources are provided (Both-AI), Claude tends to follow the less accurate prompt on conflicts, while GPT follows data more often; main Data vs Prompt ranking is robust to GPT-5.4. The paper concludes that revealed preference is a powerful communication channel for AI alignment under risk, but implementation (delegation, conflict resolution) matters.

Significance. The paper addresses a first-order principal-agent problem for agentic AI: misalignment from incomplete preference specification rather than conflicting incentives. The comparative design is careful and well-suited to the claim—incentivized choices and prompts, Part II ground truth withheld as a benchmark at decision time, random lottery payment, comprehension checks, and multi-session Prolific recruitment. Mechanism evidence (within-sample prompt predictiveness, AI-generated prompts, semantic similarity, Allais-type heterogeneity) and model robustness (Claude vs GPT conflict resolution; GPT-5.4 replication of main ranking) strengthen the contribution beyond a pure horse race. If the result holds, it gives a concrete, implementable message for AI-alignment and human-AI interaction literatures: revealed-preference data can outperform free-text instructions for choice under risk, but humans may not opt into it and multi-source agents may overweight stated goals depending on post-training. Strengths include transparent experimental incentives, clear agent definitions, and explicit acknowledgment of model-dependent conflict resolution.

major comments (3)
  1. External validity of the evaluation target and communication channels is the main load-bearing limit on how far the central claim can travel. Part II binary lottery match rate is a clean within-design benchmark (Section 3), but the paper's broader framing (agentic AI, irreversible decisions, high-dimensional tasks) is not tested. The 13-item Part I set plus free-text prompt may understate both the value of rich interaction and the difficulty of scaling revealed-preference elicitation. The manuscript should either (i) more tightly bound claims to choice under risk with fixed binary menus, or (ii) add discussion/evidence on how the ranking would change with multi-attribute or sequential agentic tasks, interactive prompting, or larger choice histories.
  2. Both-AI results (Section 4.4 / Result 4) are important for the 'careful implementation' conclusion but rest on a single conflict-resolution pattern that is model-specific. Claude follows Prompt-AI on 66% of conflicts despite lower accuracy; GPT follows Data-AI more often and restores Both-AI performance. The paper correctly notes constitutional vs RLHF training differences, but the policy implication—that 'more information does not necessarily improve performance'—depends on which frontier model is used. Strengthen by reporting full Both-AI match rates under GPT for the entire sample (not only the 891 conflicting pairs), and by clarifying whether the recommended design is 'data only,' 'data with model-aware conflict rules,' or something else.
  3. Delegation and belief results (subjects overestimate absolute gaps; 36% fail to choose the better agent) are central to welfare implications but need tighter linkage to actual welfare under the paper's payment rule. Match-rate differences are reported, yet the mapping from wrong delegation to expected payoff loss under the random-lottery incentive is not fully quantified. A short calculation of expected payoff loss from suboptimal delegation (and from Both-AI's conflict bias under Claude) would make the 'large portion fail to select the more informative source' claim more precise and would discipline how large the practical cost is.
minor comments (6)
  1. Several passages in the provided manuscript text contain garbled or placeholder characters (e.g., '�������� ����������approach', Result 4 block). Clean all OCR/encoding artifacts before publication.
  2. Abstract and introduction slightly overstate generality ('aligning AI with human preferences') relative to the lottery domain; align wording with the actual experimental domain earlier.
  3. Appendix EU comparison (≈75% match, closer to Data-AI) is useful; report confidence intervals or subject-level distributions for Data-AI vs Prompt-AI vs EU side by side in the main text or a single figure for readability.
  4. Clarify exclusion criteria for the six 'completely uninformative' prompts and report sensitivity of main gaps to including them.
  5. State whether prompts, choice data, and analysis code will be publicly released; reproducibility is especially valuable for LLM-agent experiments.
  6. Define the easy/hard/behavioral lottery taxonomy more explicitly in the main text (or a short table) so readers need not reconstruct categories from the appendix alone.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: out-of-sample match rates against held-out Part II choices are external benchmarks, not quantities defined by the inputs being compared.

full rationale

The paper is an incentivized online experiment comparing two information sources (Part I binary lottery choices vs free-text prompts) as inputs to frontier LLMs, scored by out-of-sample prediction accuracy on subjects' Part II choices. Data-AI and Prompt-AI are instantiated from Part I only; Part II is the evaluation target and was not used to fit either agent. The performance gap, mechanism evidence (prompts only partially recover Part I choices; AI-generated prompts close the gap; semantic similarity closes the gap; Allais heterogeneity), delegation misperception, and Both-AI conflict-resolution patterns are all empirical comparisons against that external benchmark. The EU appendix estimates risk parameters from Part I and predicts Part II in the same out-of-sample manner. There is no self-definitional loop, no parameter fitted to the outcome and renamed a prediction, no uniqueness theorem imported from the authors, and no load-bearing self-citation chain. Related-work citations are to independent literature. The design is self-contained against its own held-out human choices; circularity score is zero.

Assumptions & free parameters 2 free parameters · 4 assumptions · 1 invented entities

Empirical experimental paper. Load-bearing content is design choices and measurement assumptions rather than free parameters or new physical entities. The central claim rests on treating Part II choices as ground truth, treating free-text prompts and 13 binary lotteries as the two communication channels, and treating match rate as the alignment metric. No fitted constants drive the main result; EU risk parameters appear only in an appendix benchmark.

free parameters (2)
  • Minimum prompt-writing time (60 seconds)
    Design parameter chosen to reduce opportunity-cost shirking; not estimated from data but affects prompt quality and thus the stated-preference arm.
  • Lottery set size and difficulty taxonomy (13 easy/hard/behavioral items per part)
    Hand-designed stimulus set that defines the preference domain; results could shift with different menus or continuous risk tasks.
assumptions (4)
  • domain assumption Subjects' Part II binary lottery choices are the correct ground-truth measure of the preferences an AI should implement on their behalf.
    Stated in Section 3 evaluation design; without it, match rate is not an alignment metric.
  • domain assumption Free-text written instructions and a short menu of binary lottery choices are representative channels for stated versus revealed preference communication to AI agents.
    Core design choice in Part I tasks; external validity to high-dimensional real-world agent tasks is assumed rather than tested.
  • domain assumption Frontier LLMs (Claude Opus 4.5, GPT-5.4) can be treated as stable preference-implementing agents when given either choice data or prompts under the paper's prompting regime.
    Implicit throughout Results; model-specific conflict resolution is acknowledged but base capability is taken as given.
  • standard math Standard expected-utility and Allais-type classifications of lottery behavior are useful for heterogeneity analysis.
    Used to stratify subjects; classic decision-theory background, not invented here.
invented entities (1)
  • Data-AI / Prompt-AI / Both-AI agent labels
    purpose: Name the three experimental information conditions under which the same base LLM is queried.
    Operational labels for experimental arms, not new theoretical objects with independent existence; independent_evidence false by construction.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Should I State or Should I Show? Aligning AI with Human Preferences." pith.science (2026). https://pith.science/paper/73NXTGGJ

@misc{pith2026260329317,
  author       = {Pith},
  title        = {Pith review of: Should I State or Should I Show? Aligning AI with Human Preferences},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/73NXTGGJ}},
  note         = {Machine review of arXiv:2603.29317}
}
read the original abstract

As AI agents become more autonomous, properly aligning their objectives with human preferences becomes increasingly important. We study how effectively an AI agent learns a human principal's preference in choice under risk via stated versus revealed preferences. We conduct an online experiment in which subjects state their preferences through written instructions ("prompts") and reveal them through choices in a series of binary lottery questions ("data"). We find that on average, an AI agent given revealed-preference data predicts subjects' choices more accurately than an AI agent given stated-preference prompts. Further analysis suggests that the gap is driven by subjects' difficulty in translating their own preferences into written instructions. When given a choice between which information source to give to an AI agent, a large portion of subjects fail to select the more informative one. Moreover, when predictions from the two sources conflict, we find that the AI agent aligns more frequently with the prompt, despite its lower accuracy. Overall, these results highlight the revealed preference approach as a powerful mechanism for communicating human preferences to AI agents, but its success depends on careful implementation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Innate Economic Preferences of Language Models

    econ.EM 2026-07 conditional novelty 6.0 of 10

    Language models' softmax token choice is exactly a random utility model, letting logits identify preferences: twelve models show risk aversion, IIA violations, and fine-tuning can set a target risk attitude.

Pith tools

Reviewed July 13, 2026 · model on record in the stance chip above.