REVIEW 2 major objections 4 minor 12 references
Does AI and Human Advice Mitigate Punishment for Selfish Behavior? An Experiment on AI ethics From a Psychological Perspective
T0 review · 2 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that selfish behavior is punished less after selfish advice and more after prosocial advice, and that the AI-versus-human source of the advice does not change punishment.
desk verdict A clean, well-powered pre-registered experiment showing lay punishment tracks advice content and not source, under an abstract-label design that needs a robustness check with real advice texts. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine of the design is the costly third-party punishment paradigm: each evaluator receives 70 points and can spend up to 20 to deduct three times as many points from a decision-maker, making punishment behaviorally real rather than hypothetical. On top of this, the design pairs the advice-content manipulation (selfish vs prosocial) with a source manipulation (AI vs human, plus a no-advice control) and uses attribution theory to predict that external persuasion lowers perceived intentionality and therefore punishment. The central contrast that carries the argument is the two-by-two-plus-control comparison of punishment across advice content and source.
What would settle it
Show evaluators the real advice texts—including GPT-4's verbatim output and the human advisors' texts—instead of only their labels, and observe whether punishment of selfish decision-makers diverges by source; if AI-advised selfishness is punished less than human-advised selfishness when content is visible, the paper's source-null claim fails.
Extended reading notes
Core claim
The study's core discovery is a double dissociation: the content of advice moves punishment, and the source does not. Evaluators spent 5.27 points on average to punish selfish behavior, versus 1.96 for prosocial behavior. Among selfish decision-makers, those who had received selfish advice were punished less (M = 4.34) than those who received no advice (M = 5.41), while those who had received prosocial advice were punished more (M = 6.12). Comparing AI to human advice produced no significant difference in punishment for either selfish or prosocial advice. A separate exploratory measure showed evaluators attributed more responsibility to human than AI advisors when decision-makers followed the advice, yet this did not translate into more costly punishment, yielding what the paper calls a perception-behavior gap.
Load-bearing premise
Evaluators were told only the type and source of the advice, never its actual wording, so the core comparison assumes that punishment reactions to real advice can be reproduced with abstract labels like 'selfish advice from an AI'.
Editorial extensions
If this is right
- If the source-insensitivity result holds, people who follow selfish AI advice cannot expect cheaper punishment than people who follow selfish human advice; blame lands on the decision-maker in both cases.
- Selfish advice acts as a partial excuse: it reduces punishment relative to no advice, but it does not make selfishness costless; selfish behavior remains punished much more than prosocial behavior.
- Prosocial advice raises the bar: acting selfishly after being advised to act prosocially draws more punishment than acting selfishly with no advice, consistent with the idea that flouting a clear moral nudge is judged as especially intentional.
- Because responsibility and punishment diverge, legal or organizational policies that adjust culpability based on responsibility attributions may not align with the punishment that lay people actually choose to impose.
Reading between the lines
- Editorial extension: because evaluators saw only labels and not the actual advice text, the source-null result is best read as about the label 'AI' versus 'human'; real persuasive wording could interact with source, so a study that shows verbatim advice from GPT-4 and from human advisors is the natural stress test.
- Editorial extension: the result that ignoring prosocial advice is punished more than acting selfishly without advice suggests norm-enforcement systems may be designed to reward compliance signals as much as outcomes; this could be tested in workplace discipline or online moderation where both outcome and 'warning ignored' are observable.
- Editorial extension: the perception-behavior gap implies that self-reported responsibility scales and costly punishment tap different constructs; future work could vary the cost of punishment continuously to estimate the 'price' at which the AI-human gap appears.
- Editorial extension: since the experiment was run in the US, cultural variation in AI trust and norm enforcement could moderate the null source effect; a cross-country replication would establish whether the null is universal.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a pre-registered, three-stage experiment on how people punish selfish and prosocial behavior that follows AI, human, or no advice. Stage 1 collected human-written and GPT-4-generated advice texts, Stage 2 collected real investment decisions from decision-makers after they read such advice, and Stage 3 used a representative US sample of 633 evaluators who made costly punishment decisions and responsibility ratings for all combinations of advice source (AI/human/none), advice type (selfish/prosocial/none), and decision-maker behavior (selfish/prosocial). The main findings are that selfish behavior is punished much more than prosocial behavior; among selfish behavior, punishment is lower after selfish advice and higher after prosocial advice than after no advice; and punishment does not significantly differ between AI and human advice sources, although evaluators do attribute more responsibility to human than to AI advisors when advice is followed. The paper concludes that behavior and advice content shape punishment, whereas the advice source does not.
Significance. If the results hold, the study makes a valuable contribution to the psychology of AI ethics and the experimental literature on third-party punishment in hybrid human-AI settings. The design has notable strengths: it is pre-registered, uses real AI-generated and human-written advice, employs an incentivized costly-punishment measure rather than self-reports only, and draws on a large representative US sample with mixed-effects analyses. The comparison of a behavioral punishment measure with an exploratory responsibility measure is also informative. The central claim, however, rests on a design choice that is not validated, which limits the strength of the conclusions that can be drawn from the data.
major comments (2)
- [Part 3 - Evaluation stage (Methods, pp. 18-19)] The decision to withhold the actual advice texts from evaluators is load-bearing for both halves of the headline result. The manuscript states: 'We chose not to disclose the exact content of the advice to isolate the effect of the advice's source (AI or human) and type (selfish or prosocial) on punishment.' This assumes that the abstract labels 'selfish advice' and 'prosocial advice' reproduce the psychological effects of real persuasive content, and that the source label 'AI' versus 'human' with no accompanying wording captures how advice source actually affects punishment. The paper's own literature review cites evidence that wording, framing, and justifications shape ethical behavior (Pittarello et al., 2015; Leib et al., 2019; Kobis et al., 2024), which makes this abstraction assumption particularly vulnerable. No validation is provided that label-only responses match responses to real advice texts, and the follow-up survey described in the Discussion (34% of participants focusing on the decision-maker) does not address this. Consequently, the content effects (H3a/H3b) and the source null (H4) could both be artifacts of the labeling manipulation rather than results about actual advice content or actual AI/human advice. The authors should either provide evidence that the abstraction is valid (e.g., a validation study comparing label-only with full-text presentation), or substantially temper the conclusions and clearly frame the results as being about advice-type labels and source tags rather than about real advice.
- [Results, section (4), and Table 1, model 4] The conclusion that punishment 'does not vary' between AI and human advice is too strong given the precision of the study. The 95% confidence intervals for the AI interaction terms (bselfish x AI advice = -0.019, 95% CI [-0.336, 0.298]; bprosocial x AI advice = -0.175, 95% CI [-0.492, 0.141]) exclude only differences larger than roughly one-third to one-half of a punishment point on the 0-20 scale. The sample was powered to detect a one-point difference, so smaller but theoretically meaningful differences cannot be excluded. The abstract and Discussion should phrase the result as 'we found no evidence of a difference' or 'the difference, if any, is small,' rather than as a strong null claim that the advice source does not shape punishment.
minor comments (4)
- [Sample and power calculations (p. 23)] The description of the power analysis is difficult to follow: 'when comparing AI advice to human advice, the sample size allows us to achieve 91% power to simultaneously detect (i) a reduction of one punishment point (out of 20) after receiving prosocial advice and (ii) an increase of one punishment point after receiving selfish advice.' It is not clear from this sentence what the target effect size for the AI-versus-human comparison is; please reword to state explicitly that the test is powered to detect a one-point difference between the AI and human advice conditions.
- [Discussion (p. 39)] The text cites 'Fehr & Gächter, 2004' for costly punishment, but the reference list contains only Fehr & Fischbacher (2004). Either add the Fehr and Gächter reference or correct the in-text citation.
- [References (p. 56)] The Roe and Just (2009) reference contains a stray 'r' at the end of the page range ('1266-1271.r'); please correct this typographical error.
- [Abstract and Figure 1] The abstract states that evaluators 'could punish real decision-makers who received AI, human, or no advice,' but it does not mention that evaluators never saw the actual advice content, only its type and source. Adding this qualification would improve transparency and align the abstract with the methods as described.
Circularity Check
No significant circularity: the headline effects are pre-registered empirical comparisons, and self-citations are background rather than load-bearing.
full rationale
The paper's central claims are empirical hypotheses tested with a preregistered, financially incentivized experiment, not quantities derived from fitted parameters or from the definitions of the variables. The outcome variable—costly punishment—is measured behaviorally and is not constructed from the advice labels or from the cited prior work. The advice-type and advice-source conditions are manipulated, and the reported differences (e.g., selfish behavior punished less after selfish advice and more after prosocial advice) are contingent empirical results; they are not analytically entailed by the labels 'selfish advice' or 'prosocial advice.' The paper explicitly tests the alternative compliance/defiance explanation rather than treating it as a tautology. The authors cite their own prior work (e.g., Leib et al., 2024; Köbis et al., 2024) to motivate the research question and to contextualize findings, but these citations are not the evidential basis for the new punishment results, which come from the experiment's own data. There is no fitted input renamed as a prediction, no imported uniqueness theorem, and no ansatz smuggled in via citation. The label-only presentation of advice is a potential construct-validity limitation—real advice texts were collected but not shown to evaluators—but this is a design concern about whether the manipulation reflects actual advice content, not a circular derivation: punishment decisions were not logically forced by the labels. Overall, the derivation chain is self-contained and the circularity burden is minimal.
Assumptions & free parameters
assumptions (3)
- domain assumption Attribution theory's assumption that behavior perceived as externally influenced is punished less applies to AI advisors as well as human advisors.
- domain assumption Costly punishment in the Fehr-Fischbacher paradigm is a valid behavioral proxy for real-world punishment.
- ad hoc to paper Evaluators' punishment decisions can be studied using only the labels 'selfish advice', 'prosocial advice', 'AI', and 'human', with actual advice content withheld.
Cite this review
Pith. "Pith review of Does AI and Human Advice Mitigate Punishment for Selfish Behavior? An Experiment on AI ethics From a Psychological Perspective." pith.science (2026). https://pith.science/paper/OZFK334X
@misc{pith2026250719487,
author = {Pith},
title = {Pith review of: Does AI and Human Advice Mitigate Punishment for Selfish Behavior? An Experiment on AI ethics From a Psychological Perspective},
year = {2026},
howpublished = {\url{https://pith.science/paper/OZFK334X}},
note = {Machine review of arXiv:2507.19487}
}
read the original abstract
People increasingly rely on AI-advice when making decisions. At times, such advice can promote selfish behavior. When individuals abide by selfishness-promoting AI advice, how are they perceived and punished? To study this question, we build on theories from social psychology and combine machine-behavior and behavioral economic approaches. In a pre-registered, financially-incentivized experiment, evaluators could punish real decision-makers who (i) received AI, human, or no advice. The advice (ii) encouraged selfish or prosocial behavior, and decision-makers (iii) behaved selfishly or, in a control condition, behaved prosocially. Evaluators further assigned responsibility to decision-makers and their advisors. Results revealed that (i) prosocial behavior was punished very little, whereas selfish behavior was punished much more. Focusing on selfish behavior, (ii) compared to receiving no advice, selfish behavior was penalized more harshly after prosocial advice and more leniently after selfish advice. Lastly, (iii) whereas selfish decision-makers were seen as more responsible when they followed AI compared to human advice, punishment between the two advice sources did not vary. Overall, behavior and advice content shape punishment, whereas the advice source does not.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
N., Meyers, S., Gray, K., & Bigman, Y
Arnestad, M. N., Meyers, S., Gray, K., & Bigman, Y. E. (2024). The existence of manual mode increases human blame for AI mistakes. Cognition, 252, 105931. Aschauer, F., Sohn, M., & Hirsch, B. (2023). Managerial advice‐taking—Sharing responsibility with (non)human advisors trumps decision accuracy. European Management Review, 21(1), 186-203. Awad, E., Dsou...
work page 2024
-
[10]
Hebl, M. R., King, E. B., Glick, P., Singletary, S. L., & Kazama, S. (2007). Hostile and benevolent reactions toward pregnant women: complementary interpersonal punishments and rewards that maintain traditional roles. The Journal of Applied Psychology, 92(6), 1499–1511. Heider, F. (2013). The psychology of interpersonal relations. Psychology Press. Heinri...
work page 2007
-
[35]
Cushman, F. (2008). Crime and punishment: distinguishing the roles of causal and intentional analyses in moral judgment. Cognition, 108(2), 353–380. Cushman, F. (2015). Punishment in humans: From intuitions to institutions. Philosophy Compass, 10(2), 117-133. Dong, M., Conway, J. R., Bonnefon, J., Shariff, A., & Rahwan, I. (2024). Fears about artificial i...
work page 2008
-
[72]
Jacobsen, W. C. (2020). School punishment and interpersonal exclusion: Rejection, withdrawal, and separation from friends. Criminology; an Interdisciplinary Journal, Punishment following AI advice 53 58(1), 35–69. Jobin, A., Ienca, M., & Vayena, E. (2019). The global landscape of AI ethics guidelines. Nature machine intelligence, 1(9), 389-399. Kelley, H....
work page 2020
-
[235]
Ma, D., Akram, H., & Chen, I. H. (2024). Artificial Intelligence in Higher Education: a cross-cultural examination of students’ behavioral intentions and attitudes. International Review of Research in Open and Distributed Learning, 25(3), 134–157. Malle, B. F., Scheutz, M., Cusimano, C., Voiklis, J., Komatsu, T., Thapa, S., & Aladia, S. (2025). People’s j...
work page 2024
-
[254]
Köbis, N., Bonnefon, J.-F., & Rahwan, I. (2021). Bad machines corrupt good morals. Nature Human Behaviour, 5(6), 679–685. Köbis, N., Rahwan, Z., Bersch, C., Ajaj, T., Bonnefon, J. F., & Rahwan, I. (2024). Experimental evidence that delegating to intelligent machines can increase dishonest behaviour (No. dnjgz). Center for Open Science. Krügel, S., Osterma...
work page 2021
-
[397]
Starke, C., Ventura, A., Bersch, C., Cha, M., de Vreese, C., Doebler, P., Dong, M., Krämer, N., Leib, M., Peter, J., Schäfer, L., Soraperra, I., Szczuka, J., Tuchtfeld, E., Wald, R., & Köbis, N. (2024). Risks and protective measures for synthetic relationships. Nature Human Behaviour, 8(10), 1834–1836. Shalvi, S., Gino, F., Barkan, R., & Ayal, S. (2015). ...
work page 2024
-
[439]
K., Xu, R., Rathje, S., & Van Bavel, J
Globig, L. K., Xu, R., Rathje, S., & Van Bavel, J. J. (2024). Perceived (Mis) alignment in generative Artificial Intelligence Varies Across Cultures. Preprint. DOI,
work page 2024
Show all 12 references
-
[443]
Baumert, A., Halmburger, A., & Schmitt, M. (2013). Interventions against norm violations: dispositional determinants of self-reported and real moral courage: Dispositional determinants of self-reported and real moral courage. Personality & Social Psychology Bulletin, 39(8), 10...
2013 arXiv
-
[453]
Pavey, L., & Sparks, P. (2009). Reactance, autonomy and paths to persuasion: Examining perceptions of threats to freedom and informational value. Motivation and Emotion, 33(3), 277–290. Pillutla, M. M., & Murnighan, J. K. (1996). Unfairness, anger, and spite: Emotional rejecti...
2009
-
[3432]
E., Reeder, G
Monroe, A. E., Reeder, G. D., & James, L. (2015). Perceptions of intentionality for goal- related action: behavioral description matters. PloS One, 10(3), e0119841. Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to...
2015
-
[4569]
M., & Hagens, M
Leib, M., Köbis, N., Rilke, R. M., & Hagens, M. & Irlenbusch, B. (2024). Corrupted by algorithms? How AI-generated and human-written advice shape (dis)honesty. The Punishment following AI advice 54 Economic Journal, 134(658), 766–784. Leib, M., Pittarello, A., Gordon-Hecker, T...
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.