REVIEW 2 major objections 6 minor 22 references
Transparent AI delegation lets hybrid expert markets beat human-only markets.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
LLM experts in credence goods markets reduce efficiency and consumer surplus unless liability or transparent prosocial objectives operate, and expert delegation with transparent objectives can outperform human-only markets.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection Interesting and careful set of experiments on LLMs in credence goods markets, but the headline HAH-vs-HH claim is not statistically established because it compares across separate experiments with no joint test. the 2 major comments →
From Digital Distrust to Codified Honesty: Experimental Evidence on Generative AI in Credence Goods Markets
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
In a credence goods setting with four experts and four consumers, the paper finds that a majority of human experts delegate pricing and treatment to an LLM when allowed, and do so more often when they can choose the LLM's objective prompt. The most popular objectives are 'self-interested' and 'efficiency-loving,' meaning a slight majority of delegating experts incorporate the consumer's payoff. When consumers can see which objective an expert chose, approach rates jump, especially for efficiency-loving and inequity-averse AI experts; self-interested AI experts and non-delegating human experts lose custom. In the no-institution and verifiability conditions, disclosed chosen objectives raise r
What carries the argument
The central mechanism is the LLM's codified objective function: a short prompt that instructs the AI to maximize its own payoff, maximize total payoff ('efficiency-loving'), equalize payoffs ('inequity-averse'), or pursue no explicit objective. The paper treats this prompt as a policy-relevant institution, akin to a contract term consumers can observe. The argument runs through consumer belief formation: because consumers know the AI's objective, they can predict whether under- or overtreatment is profitable and choose accordingly; competition among experts then lets prosocial AI experts undercut dishonest ones. The design's other machinery is the standard one-shot credence goods market with
Load-bearing premise
The headline claim that transparent Human-AI-Human markets outperform Human-Human markets rests on comparing two separately run experiments with different participant pools, recruitment dates, and payment schemes; the paper reports no cross-experiment statistical test, so the comparison holds only if those pools are interchangeable.
What would settle it
Run one preregistered experiment that randomly assigns consumers to either a Human-Human market or a Human-AI-Human market with chosen, transparent objectives, using identical platform, payment, and recruitment timing, and test the efficiency difference between arms. If the transparent hybrid arm does not exceed the human-only arm—or if the advantage disappears when the disclosed objective is not enforced—the central claim is falsified.
If this is right
- Left to default self-interest, LLM experts defraud consumers and destroy trade: AI-AI markets see zero successful interactions without liability, so deploying untrained LLMs as autonomous experts can reduce market efficiency.
- Human experts are willing to delegate to LLMs, and agency over the AI's objective increases delegation, implying real-world experts may choose AI proxies if given control over their goals.
- Mandatory transparency of AI objectives can substitute partially for liability: in the no-institution condition, disclosed chosen objectives raise efficiency from about 0.5 to 0.74, while obfuscation leaves efficiency unchanged at about 0.52.
- Expert surplus rises with LLM delegation even when total efficiency falls, so experts have an individual incentive to adopt LLMs at consumers' expense—a misalignment regulators may need to address.
- The gains from prosocial AI objectives depend on the institution: the efficiency-loving prompt reduces fraud under verifiability but increases overcharging under liability, so objective guidelines should not be centrally mandated across all domains.
Where Pith is reading between the lines
- A natural extension: if transparency about AI objectives became a regulatory requirement in repair, financial, or health advice markets, the efficiency gains seen in the lab might generalize, but only when consumers believe the disclosed objective is enforceable; unregulated self-reports would be cheap talk.
- The paper's comparison of Human-Human and Human-AI-Human markets draws on separately run experiments; a single unified experiment with common recruitment and payments could test whether the headline outperformance survives random assignment.
- One could test the mechanism directly by varying whether the AI's objective is merely stated or verifiably enforced; the model predicts efficiency gains only in the enforceable-transparency condition.
- The result that efficiency-loving AI experts attract the most consumers may imply that firms can use AI objectives as a competitive signaling device, making objective disclosure a strategic variable rather than an altruistic one.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a series of one-shot credence-goods experiments comparing AI-AI, Human-Human, Human-AI, and Human-AI-Human markets under three institutions (No Institution, Verifiability, Liability). LLM behavior is elicited with objective-prompt manipulations (self-interested, inequity-averse, efficiency-loving) and, in the Human-AI-Human design, human experts may delegate to an LLM and may choose the LLM's objective, which is either transparent or obfuscated to consumers. The paper finds that AI-AI markets break down without liability; Human-Human markets achieve moderate efficiency; substituting human experts by self-interested LLMs lowers consumer surplus and raises expert surplus; and, in Human-AI-Human markets, transparency about a chosen prosocial LLM objective raises consumer approach rates and efficiency relative to other HAH treatments. The abstract's headline claim is that Human-AI-Human markets outperform Human-Human markets under transparency rules.
Significance. If the headline comparison were statistically established, this would be a meaningful contribution to the emerging literature on LLMs in markets and to the policy debate on AI transparency. The paper is careful in several respects: pre-registered sample sizes, comprehension checks with two attempts, strategy-method elicitation, a common underlying market environment, and full replication materials on OSF. The paper also reports several machine-checked LLM protocols. The main empirical contribution is the within-HAH finding that transparency about a chosen efficiency-loving objective raises efficiency relative to obfuscation and fixed self-interested objectives. However, the broader claim that HAH markets outperform HH markets depends on a comparison across experiments with different consumer pools, payments, and information conditions, and that comparison is not tested. This is the central weakness that needs to be addressed.
major comments (2)
- [Section 6.2, Tables 7/A4/A5 vs Table 3] The abstract's claim that 'Human-AI-Human markets outperform Human-Human markets under transparency rules' is not supported by any statistical test that includes the Human-Human baseline. The t-tests reported in Section 6.2 (e.g., Chosen_Transparent vs Chosen_Obfuscated, t=2.98, p=0.004) compare HAH treatments within the HAH experiment only. The HAH point estimates (0.74, 0.72, 0.94 in No Institution, Verifiability, Liability) are merely juxtaposed with the HH point estimates (0.61, 0.65, 0.84 from Table 3). The two experiments used different recruitment dates, different consumer base payments ($1.50 for HH, $2.00 for HAH), and different information sets (HH consumers saw only price pairs; HAH consumers also saw delegation and, in transparency conditions, the LLM objective). A pooled regression with an experiment indicator and interactions, or a permutation test across groups, is necessa
- [Section 3.1, Prediction 'No Institution & Transparent Efficiency-Loving Expert Preferences'] The predicted price pair P={pbar,5} is not derived from the stated efficiency-loving objective. Total payoff is invariant to price, since price is a transfer between consumer and expert. The text argues that the problem becomes 'analogous to Liability under the assumption of self-interest,' but that step imports the self-interested expert's profit-maximizing competition logic. Under the stated objective, any price pair that induces consumers to enter yields the same total payoff, so price and expert surplus are indeterminate. The prediction that experts earn 1 and consumers earn 5 is therefore an additional benchmark assumption, not an implication of efficiency-loving preferences. This matters because the theoretical framing in Section 6.2 uses this prediction to interpret the HAH efficiency gains. The empirical measurements can stand, but the theoretical prediction should be reformulate
minor comments (6)
- [Abstract] The sentence beginning 'Consequently, Human-AI-Human markets outperform...' should be qualified as 'in our point estimates' or accompanied by the cross-experiment test described in the major comments.
- [Footnote 11] Experts did not know that their chosen LLM objective would be transparent to consumers. This is a sensible design choice for eliciting preferences, but the Discussion should more explicitly state that the HAH efficiency result combines two separate decisions: experts choosing objectives non-strategically, and consumers reacting to those objectives.
- [Tables 3 and 7/A4/A5] The surplus definitions differ across tables: Table 3 refers to 'average cumulative income of all consumers in a group,' while Tables 7/A4/A5 refer to 'average income per consumer in a group.' Please align the definitions or clarify that the levels are not directly comparable across tables.
- [Figures 5, 6, and 7] Approach shares and delegation rates are reported without uncertainty intervals. Adding 95% confidence intervals or error bars would help readers assess the reported chi-square tests.
- [Section 7] Typo: 'Social planers' should be 'Social planners.' Also, Section 6.2 has 'alower share' for 'a lower share.'
- [Section 5.3] The statement that problem types are drawn randomly across 1000 simulations to compute Table 5 should clarify whether standard errors reflect this simulation randomness; the current text could be read as reporting deterministic group averages.
Circularity Check
No significant circularity: theoretical predictions are derived from standard model assumptions, and the empirical results are measurements rather than quantities forced by construction or fitted inputs.
full rationale
The paper's Section 3 predictions are derived from a standard self-interested credence goods model (Dulleck and Kerschbamer 2006) using explicit parameter values (h=0.5, V=10, c=6, c=2, outside option 1.6). These are theoretical benchmarks used to compare against experimental observations, not parameters fitted to the experimental data. The social-preference predictions (efficiency-loving, inequity-averse) follow directly from the stated objective functions and are used as comparative benchmarks; they are not estimated from the data. The central claims about Human-AI-Human markets are empirical comparisons of measured market outcomes across treatments (Tables 7, A4, A5 vs Table 3). Even though the headline claim relies on a cross-experiment comparison that is not jointly tested — a genuine statistical-validity concern — that is not a form of circularity: the compared quantities are independently measured outcomes, not definitions or fitted values. The paper's self-citations (Erlei et al. 2022; Erlei and Meub 2025) appear only in the literature review as background and are not load-bearing for any derivation. No uniqueness theorem is imported from the authors' prior work, no ansatz is smuggled in via self-citation, and no known empirical pattern is merely renamed. The observed inefficiencies and efficiency gains are behavioral findings, and the paper does not present any 'prediction' that reduces by construction to an input or fitted parameter.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Standard one-shot credence goods model: consumers know only p(h)=0.5, V=10, treatment costs 6 and 2, outside option 1.6; experts receive perfectly accurate diagnostic signals.
- domain assumption Indifferent consumers visit an expert and indifferent experts choose honest treatment.
- domain assumption Self-interested experts maximize expected monetary payoff in one-shot interactions (no reputation).
- ad hoc to paper LLM agents act according to their system and user prompts (role instructions, objective prompt, chain-of-thought), i.e., prompt adherence.
Cite this review
Pith. "Pith review of From Digital Distrust to Codified Honesty: Experimental Evidence on Generative AI in Credence Goods Markets." pith.science (2026). https://pith.science/paper/AROKPAJY
@misc{pith2026250906069,
author = {Pith},
title = {Pith review of: From Digital Distrust to Codified Honesty: Experimental Evidence on Generative AI in Credence Goods Markets},
year = {2026},
howpublished = {\url{https://pith.science/paper/AROKPAJY}},
note = {Machine review of arXiv:2509.06069}
}
read the original abstract
Generative AI is transforming the provision of expert services. This article uses a series of one-shot experiments to quantify the behavioral, welfare and distribution consequences of large language models (LLMs) on AI-AI, Human-Human, Human-AI and Human-AI-Human expert markets. Using a credence goods framework where experts have private information about the optimal service for consumers, we find that Human-Human markets generally achieve higher levels of efficiency than AI-AI and Human-AI markets through pro-social expert preferences and higher consumer trust. Notably, LLM experts still earn substantially higher surplus than human experts -- at the expense of consumer surplus - suggesting adverse incentives that may spur the harmful deployment of LLMs. Concurrently, a majority of human experts chooses to rely on LLM agents when given the opportunity in Human-AI-Human markets, especially if they have agency over the LLM's (social) objective function. Here, a large share of experts prioritizes efficiency-loving preferences over pure self-interest. Disclosing these preferences to consumers induces strong efficiency gains by marginalizing self-interested LLM experts and human experts. Consequently, Human-AI-Human markets outperform Human-Human markets under transparency rules. With obfuscation, however, efficiency gains disappear, and adverse expert incentives remain. Our results shed light on the potential opportunities and risks of disseminating LLMs in the context of expert services and raise several regulatory challenges. On the one hand, LLMs can negatively affect human trust in the presence of information asymmetries and partially crowd-out experts' other-regarding preferences through automation. On the other hand, LLMs allow experts to codify and communicate their objective function, which reduces information asymmetries and increases efficiency.
Reference graph
Works this paper leans on
-
[1]
Personal research, second opinions, and the diagnostic effort of experts
Agarwal, Ritu, Che-Wei Liu, and Kislaya Prasad. 2019 . “Personal research, second opinions, and the diagnostic effort of experts. ”Journal of Economic Behavior & Organization158: 44–61. Ahmadi, Iman
work page 2019
-
[5]
Serving consumers in an uncertain world: A credence goods experiment
“Serving consumers in an uncertain world: A credence goods experiment. ”MPI Collective Goods Discussion Paper(2023/11). Balafoutas, Loukas, and Rudolf Kerschbamer
work page 2023
-
[6]
Turning large language models into cognitive models
“Turning large language models into cognitive models. ”arXiv preprint arXiv:2306.03917. Brookins, Philip, and Jason Matthew DeBacker
-
[7]
Generative AI Triggers Welfare-Reducing Decisions in Humans
“Generative AI Triggers Welfare-Reducing Decisions in Humans. ”arXiv preprint arXiv:2401.12773. Erlei, Alexander, Richeek Das, Lukas Meub, Avishek Anand, and Ujwal Gadiraju
work page internal anchor Pith review Pith/arXiv arXiv
-
[10]
Algorithmic Collusion by Large Language Models
“ Algorithmic Collusion by Large Language Models. ”arXiv preprint arXiv:2404.00806. Goktas, Denizalp, Amy Greenwald, Takayuki Osogami, Roma Patel, Kevin Leyton-Brown, Grant Schoenebeck, Daphne Cornelisse, Constantinos Daskalakis, Ian Gemp, John Horton et al
-
[11]
Gpt in game theory experiments
“Gpt in game theory experiments. ”arXiv preprint arXiv:2305.05516. 36 Hammond, Lewis, Alan Chan, Jesse Clifton, Jason Hoelscher-Obermaier, Akbir Khan, Euan McLean, Chan- dler Smith, Wolfram Barfuss, Jakob Foerster, Tomáš Gavenčiak et al
-
[12]
Multi-agent risks from advanced ai
“Multi-agent risks from advanced ai. ”arXiv preprint arXiv:2502.14143. Hansen, Anne Lundgaard, John J Horton, Sophia Kazinnik, Daniela Puzzello, and Ali Zarifhonarvar
-
[14]
Credence goods markets, online infor- mation and repair prices: A natural field experiment
“Credence goods markets, online infor- mation and repair prices: A natural field experiment. ”Journal of Public Economics222: 104891. Kerschbamer, Rudolf, Matthias Sutter, and Uwe Dulleck. 2017 . “How social preferences shape incentives in (experimental) markets for credence goods. ”The Economic Journal127 (600): 393–416. Lai, Jinqi, Wensheng Gan, Jiayang...
work page 2017
-
[15]
I’m sorry, but I cannot fulfill this request as It goes against OpenAI Use policy
“I’m sorry, but I cannot fulfill this request as It goes against OpenAI Use policy. ”The Verge. See 37 https://www. theverge. com/2024/1/12/24036156/openai-policy-amazon-ai-listings (accessed 1 February 2024). Lorè, Nunzio, and Babak Heydari
work page 2024
-
[16]
“The Good, the Bad, and the Hulk-like GPT: Analyzing Emotional De- cisions of Large Language Models in Cooperation and Bargaining Games. ”arXiv preprint arXiv:2406.03299. Nay, John J, David Karamardian, Sarah B Lawsky, Wenting Tao, Meghana Bhat, Raghav Jain, Aaron Travis Lee, Jonathan H Choi, and Jungo Kasai
work page internal anchor Pith review Pith/arXiv arXiv
-
[17]
Investigating emergent goal-like behaviour in large language models using experimental economics
“Investigating emergent goal-like behaviour in large language models using experimental economics. ”arXiv preprint arXiv:2305.07970. Rahwan, Iyad, Manuel Cebrian, Nick Obradovich, Josh Bongard, Jean-François Bonnefon, Cynthia Breazeal, Jacob W Crandall, Nicholas A Christakis, Iain D Couzin, Matthew O Jackson et al. 2019 . “Machine be- haviour. ”Nature568 ...
Pith/arXiv arXiv 2019
-
[18]
“People over trust AI-generated medical responses and view them to be as valid as doctors, despite low accuracy. ” arXiv preprint arXiv:2408.15266. Spatharioti, Sofia Eleni, David M Rothschild, Daniel G Goldstein, and Jake M Hofman
-
[19]
Self-consistency improves chain of thought reasoning in language models
“Self-consistency improves chain of thought reasoning in language models. ”arXiv preprint arXiv:2203.11171. Wang, Zhen, Ruiqi Song, Chen Shen, Shiya Yin, Zhao Song, Balaraju Battu, Lei Shi, Danyang Jia, Talal Rahwan, and Shuyue Hu
-
[20]
“Large Language Models Overcome the Machine Penalty When Acting Fairly but Not When Acting Selfishly or Altruistically. ”arXiv preprint arXiv:2410.03724. Wei, Jason, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou et al
-
[21]
Bloomberggpt: A large language model for finance
“Bloomberggpt: A large language model for finance. ”arXiv preprint arXiv:2303.17564. Zhang, Ceyao, Kaijie Yang, Siyi Hu, Zihao Wang, Guanghe Li, Yihang Sun, Cheng Zhang, Zhaowei Zhang, Anji Liu, Song-Chun Zhu et al
-
[22]
Revolutionizing finance with llms: An overview of applications and insights
“Revolutionizing finance with llms: An overview of applications and insights. ”arXiv preprint arXiv:2401.11641. 39 Appendix TABLEA1. Expert Price Setting Across Treatments No Institution Verifiability Liability Treatment ¯ p ¯pN ¯ p ¯pN ¯ p ¯pN No Objective 4.0 5.0 160 4.0 7 .0 160 5.0 7 .0 160 Maximize Payoff 5.0 8.0 160 5.0 5.0 160 4.0 7 .0 160 Inequity...
-
[2016]
Medical insurance and free choice of physician shape patient overtreatment: A laboratory experiment
“Medical insurance and free choice of physician shape patient overtreatment: A laboratory experiment. ”Journal of Economic Behavior & Organization131: 78–105. Inderst, Roman, Kiryl Khalmetski, and Axel Ockenfels. 2019 . “Sharing guilt: How better access to information may backfire. ”Management Science65 (7): 3322–3336. Ivanov, Dima, Paul Dütting, Inbal Ta...
work page 2019
-
[2018]
How artificial intelligence will affect the practice of law
“How artificial intelligence will affect the practice of law. ”University of Toronto Law Journal68 (supplement 1): 106–124. Anagol, Santosh, Shawn Cole, and Shayak Sarkar. 2017 . “Understanding the advice of commissions-motivated agents: Evidence from the Indian life insurance market. ”Review of Economics and Statistics99 (1): 1–15. Aobdia, Daniel, Saad S...
work page 2017
-
[2022]
“For what it’s worth: Humans overwrite their economic self-interest to avoid bargaining with AI systems. ” InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems,: 1–18. Erlei, Alexander, and Lukas Meub
work page 2022
-
[2023]
Playing repeated games with large language models
“Playing repeated games with large language models. ”arXiv preprint arXiv:2305.16867. Alarie, Benjamin, Anthony Niblett, and Albert H Yoon
-
[2024]
“Who is asking ChatGPT health questions? Analysis of a nationally representative Australian community sample. ”medRxiv: 2024–10. Balafoutas, Loukas, Adrian Beck, Rudolf Kerschbamer, and Matthias Sutter
work page 2024
-
[2025]
Reputation-Driven Adoption and Avoidance of Algorithmic Decision Aids in Credence Goods Markets
“Reputation-Driven Adoption and Avoidance of Algorithmic Decision Aids in Credence Goods Markets. ”arXiv preprint arXiv:2401.17929. Fehr, Ernst, and Urs Fischbacher
work page internal anchor Pith review Pith/arXiv arXiv
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.