REVIEW 3 major objections 6 minor 1 cited by
Should I State or Should I Show? Aligning AI with Human Preferences
T0 review · 3 major / 6 minor · reviewed 2026-07-13 · grok-4.5
Pith's one-line read Revealed lottery choices align AI agents with human risk preferences better than written prompts, because people struggle to state what their choices already show.
desk verdict Clean incentivized head-to-head: revealed lottery choices beat free-text prompts for frontier LLM alignment on risk, with real mechanism checks and a model-dependent conflict wrinkle. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Out-of-sample match rate: the share of Part II lottery choices correctly predicted by Data-AI (trained on Part I choices) versus Prompt-AI (given the free-text instruction), with Both-AI used to study how models resolve conflicts between the two sources.
What would settle it
Replicate the design with the same Part I/Part II lottery structure and frontier models; if Prompt-AI's match rate equals or exceeds Data-AI's once subjects receive prompt coaching or structured templates, or if Both-AI no longer underperforms Data-AI when models are forced to weight data over text on conflicts, the central ranking fails.
Extended reading notes
Core claim
On average, an AI agent given a subject's revealed-preference lottery choices predicts that subject's later choices under risk more accurately than an AI agent given the same subject's written prompt. The performance gap is driven mainly by subjects' difficulty translating their own preferences into clear written instructions, not by a lack of information in the choice data itself.
Load-bearing premise
That subjects' later incentivized lottery choices are the right ground truth for the preferences the AI should implement, and that a short set of binary lotteries plus a free-text prompt stand in for how people will actually hand preferences to agentic systems.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies how well an AI agent can implement a human principal's preferences over binary lotteries when given either revealed-preference choice data (Data-AI) or free-text stated-preference prompts (Prompt-AI). In an incentivized online experiment (N=290 after exclusions), subjects complete Part I lottery choices, write a prompt, and choose which information source to delegate; Part II choices (structurally similar lotteries) serve as held-out ground truth for out-of-sample match rates. On average Data-AI predicts Part II choices more accurately than Prompt-AI. The gap is attributed to subjects' difficulty articulating preferences: prompts only partially recover Part I choices; AI-generated prompts close the gap; and semantic similarity of human prompts to AI prompts shrinks the performance difference. Subjects often misperceive relative agent quality and frequently fail to select the better source. When both sources are provided (Both-AI), Claude tends to follow the less accurate prompt on conflicts, while GPT follows data more often; main Data vs Prompt ranking is robust to GPT-5.4. The paper concludes that revealed preference is a powerful communication channel for AI alignment under risk, but implementation (delegation, conflict resolution) matters.
Significance. The paper addresses a first-order principal-agent problem for agentic AI: misalignment from incomplete preference specification rather than conflicting incentives. The comparative design is careful and well-suited to the claim—incentivized choices and prompts, Part II ground truth withheld as a benchmark at decision time, random lottery payment, comprehension checks, and multi-session Prolific recruitment. Mechanism evidence (within-sample prompt predictiveness, AI-generated prompts, semantic similarity, Allais-type heterogeneity) and model robustness (Claude vs GPT conflict resolution; GPT-5.4 replication of main ranking) strengthen the contribution beyond a pure horse race. If the result holds, it gives a concrete, implementable message for AI-alignment and human-AI interaction literatures: revealed-preference data can outperform free-text instructions for choice under risk, but humans may not opt into it and multi-source agents may overweight stated goals depending on post-training. Strengths include transparent experimental incentives, clear agent definitions, and explicit acknowledgment of model-dependent conflict resolution.
major comments (3)
- External validity of the evaluation target and communication channels is the main load-bearing limit on how far the central claim can travel. Part II binary lottery match rate is a clean within-design benchmark (Section 3), but the paper's broader framing (agentic AI, irreversible decisions, high-dimensional tasks) is not tested. The 13-item Part I set plus free-text prompt may understate both the value of rich interaction and the difficulty of scaling revealed-preference elicitation. The manuscript should either (i) more tightly bound claims to choice under risk with fixed binary menus, or (ii) add discussion/evidence on how the ranking would change with multi-attribute or sequential agentic tasks, interactive prompting, or larger choice histories.
- Both-AI results (Section 4.4 / Result 4) are important for the 'careful implementation' conclusion but rest on a single conflict-resolution pattern that is model-specific. Claude follows Prompt-AI on 66% of conflicts despite lower accuracy; GPT follows Data-AI more often and restores Both-AI performance. The paper correctly notes constitutional vs RLHF training differences, but the policy implication—that 'more information does not necessarily improve performance'—depends on which frontier model is used. Strengthen by reporting full Both-AI match rates under GPT for the entire sample (not only the 891 conflicting pairs), and by clarifying whether the recommended design is 'data only,' 'data with model-aware conflict rules,' or something else.
- Delegation and belief results (subjects overestimate absolute gaps; 36% fail to choose the better agent) are central to welfare implications but need tighter linkage to actual welfare under the paper's payment rule. Match-rate differences are reported, yet the mapping from wrong delegation to expected payoff loss under the random-lottery incentive is not fully quantified. A short calculation of expected payoff loss from suboptimal delegation (and from Both-AI's conflict bias under Claude) would make the 'large portion fail to select the more informative source' claim more precise and would discipline how large the practical cost is.
minor comments (6)
- Several passages in the provided manuscript text contain garbled or placeholder characters (e.g., '�������� ����������approach', Result 4 block). Clean all OCR/encoding artifacts before publication.
- Abstract and introduction slightly overstate generality ('aligning AI with human preferences') relative to the lottery domain; align wording with the actual experimental domain earlier.
- Appendix EU comparison (≈75% match, closer to Data-AI) is useful; report confidence intervals or subject-level distributions for Data-AI vs Prompt-AI vs EU side by side in the main text or a single figure for readability.
- Clarify exclusion criteria for the six 'completely uninformative' prompts and report sensitivity of main gaps to including them.
- State whether prompts, choice data, and analysis code will be publicly released; reproducibility is especially valuable for LLM-agent experiments.
- Define the easy/hard/behavioral lottery taxonomy more explicitly in the main text (or a short table) so readers need not reconstruct categories from the appendix alone.
Circularity Check
No circularity: out-of-sample match rates against held-out Part II choices are external benchmarks, not quantities defined by the inputs being compared.
full rationale
The paper is an incentivized online experiment comparing two information sources (Part I binary lottery choices vs free-text prompts) as inputs to frontier LLMs, scored by out-of-sample prediction accuracy on subjects' Part II choices. Data-AI and Prompt-AI are instantiated from Part I only; Part II is the evaluation target and was not used to fit either agent. The performance gap, mechanism evidence (prompts only partially recover Part I choices; AI-generated prompts close the gap; semantic similarity closes the gap; Allais heterogeneity), delegation misperception, and Both-AI conflict-resolution patterns are all empirical comparisons against that external benchmark. The EU appendix estimates risk parameters from Part I and predicts Part II in the same out-of-sample manner. There is no self-definitional loop, no parameter fitted to the outcome and renamed a prediction, no uniqueness theorem imported from the authors, and no load-bearing self-citation chain. Related-work citations are to independent literature. The design is self-contained against its own held-out human choices; circularity score is zero.
Assumptions & free parameters
free parameters (2)
- Minimum prompt-writing time (60 seconds)
- Lottery set size and difficulty taxonomy (13 easy/hard/behavioral items per part)
assumptions (4)
- domain assumption Subjects' Part II binary lottery choices are the correct ground-truth measure of the preferences an AI should implement on their behalf.
- domain assumption Free-text written instructions and a short menu of binary lottery choices are representative channels for stated versus revealed preference communication to AI agents.
- domain assumption Frontier LLMs (Claude Opus 4.5, GPT-5.4) can be treated as stable preference-implementing agents when given either choice data or prompts under the paper's prompting regime.
- standard math Standard expected-utility and Allais-type classifications of lottery behavior are useful for heterogeneity analysis.
invented entities (1)
-
Data-AI / Prompt-AI / Both-AI agent labels
Cite this review
Pith. "Pith review of Should I State or Should I Show? Aligning AI with Human Preferences." pith.science (2026). https://pith.science/paper/73NXTGGJ
@misc{pith2026260329317,
author = {Pith},
title = {Pith review of: Should I State or Should I Show? Aligning AI with Human Preferences},
year = {2026},
howpublished = {\url{https://pith.science/paper/73NXTGGJ}},
note = {Machine review of arXiv:2603.29317}
}
read the original abstract
As AI agents become more autonomous, properly aligning their objectives with human preferences becomes increasingly important. We study how effectively an AI agent learns a human principal's preference in choice under risk via stated versus revealed preferences. We conduct an online experiment in which subjects state their preferences through written instructions ("prompts") and reveal them through choices in a series of binary lottery questions ("data"). We find that on average, an AI agent given revealed-preference data predicts subjects' choices more accurately than an AI agent given stated-preference prompts. Further analysis suggests that the gap is driven by subjects' difficulty in translating their own preferences into written instructions. When given a choice between which information source to give to an AI agent, a large portion of subjects fail to select the more informative one. Moreover, when predictions from the two sources conflict, we find that the AI agent aligns more frequently with the prompt, despite its lower accuracy. Overall, these results highlight the revealed preference approach as a powerful mechanism for communicating human preferences to AI agents, but its success depends on careful implementation.
Forward citations
Cited by 1 Pith paper
-
The Innate Economic Preferences of Language Models
Language models' softmax token choice is exactly a random utility model, letting logits identify preferences: twelve models show risk aversion, IIA violations, and fine-tuning can set a target risk attitude.
Reviewed July 13, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.