{"id":"c457d2ee-48a4-41b2-b073-337cc7286955","arxiv_id":"2507.19487","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a pre-registered, financially incentivized study, evaluators punished selfish behavior less after selfish advice and more after prosocial advice, with no difference between AI and human advice.","lead":"This experiment measured whether people punish selfish decisions less when a human or an AI advisor encouraged them. Punishment depends on what the advice said, not on whether it came from AI or a person.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Label-only advice presentation may drive both headline results: the content effect and the AI-human null could change when real advice texts are shown to evaluators.","rationale":"The reader's weakest_assumption identifies exactly the claim I would attack, and I see it as the single most load-bearing concern. The paper's headline has two parts: advice content changes punishment, and advice source does not. Both depend on the stimulus being a faithful instantiation of 'advice.' The experiment deliberately removed the actual advice text, replacing it with a label. This is not a minor design choice: attribution theory, the paper's stated framework, is about perceived intentionality and persuasion; a label 'selfish advice' is a different psychological object from a specific argument that urges selfishness. The paper's own literature review shows that wording and framing matter for ethical behavior, so the abstraction assumption is not safe. If the label-only responses do not generalize to real texts, the central claim is an artifact of the method—an internal-validity problem, not merely a question of external generalization. Secondary concerns such as the null-result interpretation and the overstated discussion are real but less fundamental; they are already acknowledged by the reader. The abstraction issue threatens both the positive content effect and the null source effect simultaneously, because both are measured without the actual stimulus. A full-text replication is the natural, feasible test since the advice texts already exist. If that test shows different patterns, the conditional verdict should remain conditional or move toward rejection; if it replicates, the claim is substantially strengthened. Given the reader already set a CONDITIONAL verdict based on this concern, my read does not change that verdict.","tokens_in":22388,"tokens_out":5304,"duration_ms":59401,"concrete_test":"Run a pre-registered follow-up evaluation study using the actual advice texts already collected: present evaluators with the real text (randomly drawn from the five human and five AI texts) in each source×type cell, plus the no-advice control, and ask for the same punishment decisions. Pre-specify three tests: (i) an interaction between condition (label-only vs. real-text) and advice source (AI vs. human) for selfish behavior; if significant, the original source null is not robust. (ii) An interaction between condition and advice type; if significant, the content effect depends on labels. (iii) Variance components for individual texts within each source×type cell; if the intraclass correlation for text identity is substantial, specific wording matters beyond the abstract label. This directly tests whether the abstraction from real advice content changes punishment decisions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—punishment follows advice content but not advice source—rests on presenting evaluators with abstract labels ('selfish advice,' 'prosocial advice,' 'AI,' 'human') rather than the actual advice texts collected in Stage 1. The paper states: 'We chose not to disclose the exact content of the advice to isolate the effect of the advice's source (AI or human) and type (selfish or prosocial) on punishment' (Methods, Evaluation stage). This design assumes abstract labels reproduce the psychological effect of real persuasive content. That assumption is load-bearing for both halves of the headline. For the content effect (H3a/H3b), the label directly encodes normative valence, so evaluators can punish defiance/compliance without ever processing the argument that would have made the advice persuasive; real advice may be ambiguous or weakly persuasive, weakening or even reversing the effect. For the source null (H4), the source manipulation is a mere tag; no actual wording distinguishes human from AI advice, so a null difference may reflect a weak manipulation rather than true equivalence. The paper itself cites evidence that wording and framing shape ethical behavior (Pittarello et al., 2015; Leib et al., 2019; Köbis et al., 2024), making the abstraction assumption especially vulnerable. No validation is provided that label-only responses match responses to real advice; the follow-up survey mentioned in the Discussion (34% of participants focused only on the decision-maker when punishing) further suggests evaluators may not have engaged with the source cue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a pre-registered, three-stage experiment on how people punish selfish and prosocial behavior that follows AI, human, or no advice. Stage 1 collected human-written and GPT-4-generated advice texts, Stage 2 collected real investment decisions from decision-makers after they read such advice, and Stage 3 used a representative US sample of 633 evaluators who made costly punishment decisions and responsibility ratings for all combinations of advice source (AI/human/none), advice type (selfish/prosocial/none), and decision-maker behavior (selfish/prosocial). The main findings are that selfish behavior is punished much more than prosocial behavior; among selfish behavior, punishment is lower after selfish advice and higher after prosocial advice than after no advice; and punishment does not significantly differ between AI and human advice sources, although evaluators do attribute more responsibility to human than to AI advisors when advice is followed. The paper concludes that behavior and advice content shape punishment, whereas the advice source does not.","tokens_in":22643,"tokens_out":6736,"duration_ms":66227,"significance":"If the results hold, the study makes a valuable contribution to the psychology of AI ethics and the experimental literature on third-party punishment in hybrid human-AI settings. The design has notable strengths: it is pre-registered, uses real AI-generated and human-written advice, employs an incentivized costly-punishment measure rather than self-reports only, and draws on a large representative US sample with mixed-effects analyses. The comparison of a behavioral punishment measure with an exploratory responsibility measure is also informative. The central claim, however, rests on a design choice that is not validated, which limits the strength of the conclusions that can be drawn from the data.","major_comments":[{"comment":"The decision to withhold the actual advice texts from evaluators is load-bearing for both halves of the headline result. The manuscript states: 'We chose not to disclose the exact content of the advice to isolate the effect of the advice's source (AI or human) and type (selfish or prosocial) on punishment.' This assumes that the abstract labels 'selfish advice' and 'prosocial advice' reproduce the psychological effects of real persuasive content, and that the source label 'AI' versus 'human' with no accompanying wording captures how advice source actually affects punishment. The paper's own literature review cites evidence that wording, framing, and justifications shape ethical behavior (Pittarello et al., 2015; Leib et al., 2019; Kobis et al., 2024), which makes this abstraction assumption particularly vulnerable. No validation is provided that label-only responses match responses to real advice texts, and the follow-up survey described in the Discussion (34% of participants focusing on the decision-maker) does not address this. Consequently, the content effects (H3a/H3b) and the source null (H4) could both be artifacts of the labeling manipulation rather than results about actual advice content or actual AI/human advice. The authors should either provide evidence that the abstraction is valid (e.g., a validation study comparing label-only with full-text presentation), or substantially temper the conclusions and clearly frame the results as being about advice-type labels and source tags rather than about real advice.","section":"Part 3 - Evaluation stage (Methods, pp. 18-19)"},{"comment":"The conclusion that punishment 'does not vary' between AI and human advice is too strong given the precision of the study. The 95% confidence intervals for the AI interaction terms (bselfish x AI advice = -0.019, 95% CI [-0.336, 0.298]; bprosocial x AI advice = -0.175, 95% CI [-0.492, 0.141]) exclude only differences larger than roughly one-third to one-half of a punishment point on the 0-20 scale. The sample was powered to detect a one-point difference, so smaller but theoretically meaningful differences cannot be excluded. The abstract and Discussion should phrase the result as 'we found no evidence of a difference' or 'the difference, if any, is small,' rather than as a strong null claim that the advice source does not shape punishment.","section":"Results, section (4), and Table 1, model 4"}],"minor_comments":[{"comment":"The description of the power analysis is difficult to follow: 'when comparing AI advice to human advice, the sample size allows us to achieve 91% power to simultaneously detect (i) a reduction of one punishment point (out of 20) after receiving prosocial advice and (ii) an increase of one punishment point after receiving selfish advice.' It is not clear from this sentence what the target effect size for the AI-versus-human comparison is; please reword to state explicitly that the test is powered to detect a one-point difference between the AI and human advice conditions.","section":"Sample and power calculations (p. 23)"},{"comment":"The text cites 'Fehr & Gächter, 2004' for costly punishment, but the reference list contains only Fehr & Fischbacher (2004). Either add the Fehr and Gächter reference or correct the in-text citation.","section":"Discussion (p. 39)"},{"comment":"The Roe and Just (2009) reference contains a stray 'r' at the end of the page range ('1266-1271.r'); please correct this typographical error.","section":"References (p. 56)"},{"comment":"The abstract states that evaluators 'could punish real decision-makers who received AI, human, or no advice,' but it does not mention that evaluators never saw the actual advice content, only its type and source. Adding this qualification would improve transparency and align the abstract with the methods as described.","section":"Abstract and Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The label-only advice presentation is the core validity concern. The authors make a deliberate methodological choice, but it is not validated, and it directly affects the manuscript's central conclusions. I think the correct response is major revision: the authors should either provide evidence that their abstraction is valid or reframe their claims as being about the effects of advice-type and source labels. The paper is otherwise well executed and would be a solid contribution once this issue is addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a well-executed, pre-registered experiment with a genuine new result—costly, incentivized punishment is sensitive to whether advice was selfish or prosocial, but not to whether an AI or a human gave it. The paper deserves a serious referee; the main caveat is the label-only advice manipulation.\n\nWhat's new: previous work on blame and responsibility for AI used hypothetical scenarios or self-reports. Here evaluators spend real points to penalize decision-makers, and the decision-makers' behavior and advice are real (collected in earlier stages). The design is clean: representative US sample, pre-registered hypotheses, transparent tables, OSF materials. The headline finding, that punishment follows behavior and advice content but not source, is consistent with the reported analysis. The responsibility measure shows a separate, interesting pattern—human advisors are seen as more responsible than AI advisors when advice is followed—which suggests the source cue did register, making the punishment null more credible.\n\nSoft spots: the decision to withhold the actual advice text and present only labels ('selfish advice', 'human advice') is load-bearing. The paper says this isolates source and type, but it assumes abstract labels reproduce the effect of real persuasive content. If actual advice wording is ambiguous or weak, the content effect could shrink, and the source null could be a weak-manipulation artifact rather than true equivalence. The paper does report a follow-up survey but doesn't validate label-only responses against real-advice responses. That's a genuine limitation, not a fatal flaw. The null for AI vs human is also only powered to detect one point out of twenty, so small differences can't be ruled out. And the discussion overreaches a bit when it talks about attributing moral agency to AI; the data show people punish similarly, not that they see AI as a moral agent.\n\nVerdict: this is a solid empirical contribution for behavioral ethics and human-AI interaction. It deserves peer review. I'd cite it and bring it to reading group. For a revision or follow-up, I'd want a robustness condition where evaluators read the actual advice texts.","headline":"A clean, well-powered pre-registered experiment showing lay punishment tracks advice content and not source, under an abstract-label design that needs a robustness check with real advice texts.","tokens_in":23146,"tokens_out":1865,"would_cite":true,"duration_ms":19356,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that selfish behavior is punished less after selfish advice and more after prosocial advice, and that the AI-versus-human source of the advice does not change punishment.","keywords":["artificial intelligence","advice","costly punishment","selfish behavior","prosocial behavior","responsibility attribution","machine behavior","third-party punishment"],"falsifier":"Show evaluators the real advice texts—including GPT-4's verbatim output and the human advisors' texts—instead of only their labels, and observe whether punishment of selfish decision-makers diverges by source; if AI-advised selfishness is punished less than human-advised selfishness when content is visible, the paper's source-null claim fails.","tokens_in":22171,"feed_emoji":"⚖️","tokens_out":7022,"duration_ms":70994,"temperature":0.7,"pith_summary":"People punish selfish behavior less when the decision-maker had received advice encouraging selfishness, and punish it more when the decision-maker ignored advice encouraging prosociality. The paper's central claim is that advice content shapes costly punishment in hybrid human-AI settings, whereas the advice source—human or AI—does not. In a pre-registered, financially incentivized experiment with 633 evaluators, spending to punish selfish decision-makers was statistically identical whether the advisor was a person or an AI. The finding matters because legal and organizational responses to AI-influenced misconduct often assume special leniency for 'the machine made me do it'; this experiment suggests lay punishers do not grant that leniency with their own money.","feed_headline":"Advice content—not AI vs human—drives punishment for selfish acts","feed_subtitle":"Experiment: selfish advice softens punishment, prosocial advice hardens it, and the AI-versus-human label makes no difference.","key_machinery":"The engine of the design is the costly third-party punishment paradigm: each evaluator receives 70 points and can spend up to 20 to deduct three times as many points from a decision-maker, making punishment behaviorally real rather than hypothetical. On top of this, the design pairs the advice-content manipulation (selfish vs prosocial) with a source manipulation (AI vs human, plus a no-advice control) and uses attribution theory to predict that external persuasion lowers perceived intentionality and therefore punishment. The central contrast that carries the argument is the two-by-two-plus-control comparison of punishment across advice content and source.","core_discovery":"The study's core discovery is a double dissociation: the content of advice moves punishment, and the source does not. Evaluators spent 5.27 points on average to punish selfish behavior, versus 1.96 for prosocial behavior. Among selfish decision-makers, those who had received selfish advice were punished less (M = 4.34) than those who received no advice (M = 5.41), while those who had received prosocial advice were punished more (M = 6.12). Comparing AI to human advice produced no significant difference in punishment for either selfish or prosocial advice. A separate exploratory measure showed evaluators attributed more responsibility to human than AI advisors when decision-makers followed the advice, yet this did not translate into more costly punishment, yielding what the paper calls a perception-behavior gap.","pith_inferences":["Editorial extension: because evaluators saw only labels and not the actual advice text, the source-null result is best read as about the label 'AI' versus 'human'; real persuasive wording could interact with source, so a study that shows verbatim advice from GPT-4 and from human advisors is the natural stress test.","Editorial extension: the result that ignoring prosocial advice is punished more than acting selfishly without advice suggests norm-enforcement systems may be designed to reward compliance signals as much as outcomes; this could be tested in workplace discipline or online moderation where both outcome and 'warning ignored' are observable.","Editorial extension: the perception-behavior gap implies that self-reported responsibility scales and costly punishment tap different constructs; future work could vary the cost of punishment continuously to estimate the 'price' at which the AI-human gap appears.","Editorial extension: since the experiment was run in the US, cultural variation in AI trust and norm enforcement could moderate the null source effect; a cross-country replication would establish whether the null is universal."],"forward_implications":["If the source-insensitivity result holds, people who follow selfish AI advice cannot expect cheaper punishment than people who follow selfish human advice; blame lands on the decision-maker in both cases.","Selfish advice acts as a partial excuse: it reduces punishment relative to no advice, but it does not make selfishness costless; selfish behavior remains punished much more than prosocial behavior.","Prosocial advice raises the bar: acting selfishly after being advised to act prosocially draws more punishment than acting selfishly with no advice, consistent with the idea that flouting a clear moral nudge is judged as especially intentional.","Because responsibility and punishment diverge, legal or organizational policies that adjust culpability based on responsibility attributions may not align with the punishment that lay people actually choose to impose."],"supporting_citations":[{"why":"Supplies the costly third-party punishment task that measures punishment behaviorally.","marker":"Fehr & Fischbacher, 2004"},{"why":"Supplies the attribution-theory prediction that externally influenced behavior is judged less intentional.","marker":"Heider, 2013"},{"why":"Supplies the causal-attribution logic that external pressure shifts responsibility away from the decision-maker.","marker":"Kelley, 1973"},{"why":"Supplies the machine-behavior approach of studying AI systems as observable actors alongside humans.","marker":"Rahwan et al., 2019"},{"why":"Supplies prior evidence that AI-generated advice can shape dishonest behavior, the behavioral baseline this design builds on.","marker":"Leib et al., 2024"},{"why":"Supplies evidence that delegating to AI increases dishonesty, framing why selfish AI advice is a live ethical concern.","marker":"Köbis et al., 2024"},{"why":"Supplies the competing claim that humans are blamed more than machines for mistakes, which the null source effect in costly punishment qualifies.","marker":"Awad et al., 2020"},{"why":"Supplies the contrasting finding that human wrongdoing is judged more immoral than machine wrongdoing, against which the behavioral null is interpreted.","marker":"Wilson et al., 2022"}],"fun_headline_variants":["Punishment hinges on advice content, not AI label","Selfish advice shields, prosocial advice blames","AI vs human? Punishment ignores source","Advice source irrelevant, content rules punishment"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Evaluators were told only the type and source of the advice, never its actual wording, so the core comparison assumes that punishment reactions to real advice can be reproduced with abstract labels like 'selfish advice from an AI'.","fun_headline_variants_meta":{"raw":{"variants":["Punishment hinges on advice content, not AI label","Selfish advice shields, prosocial advice blames","AI vs human? Punishment ignores source","Advice source irrelevant, content rules punishment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1238,"prompt_tokens":939,"completion_tokens":299,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":239}},"tokens_in":555,"tokens_out":299,"duration_ms":3299,"temperature":1.0,"reasoning_tokens":239,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:54:47.929507+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Show evaluators the real advice texts—including GPT-4's verbatim output and the human advisors' texts—instead of only their labels, and observe whether punishment of selfish decision-makers diverges by source; if AI-advised selfishness is punished less than human-advised selfishness when content is visible, the paper's source-null claim fails.","supporting_citations":[{"cited_title":"M., & Hagens, M","cited_arxiv_id":null,"evidence_quote":"Supplies prior evidence that AI-generated advice can shape dishonest behavior, the behavioral baseline this design builds on."}],"review_version":1}