{"id":"a110f95e-d486-469c-a48b-d334559f45d7","arxiv_id":"2601.20487","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In anonymous group feedback, an AI-labelled team member changed neither cooperation nor norm perceptions compared with a human label—group actions drove behaviour.","lead":"In a four-player public goods game where one member was secretly a bot, people contributed the same whether the bot was labelled human or AI; group behavior, not identity, predicted cooperation. The result suggests cooperative norms can extend to mixed human-AI groups, but only when feedback hides individual contributions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Aggregate feedback hides the bot's individual behavior, so the null label effect may reflect attributional ignorance rather than normative equivalence; the paper's 'behaviour, not identity' conclusion is not actually tested.","rationale":"The paper is a transparent, preregistered study with a plausible bounded finding: under anonymous aggregate feedback, telling participants that one of four players is an AI does not change cooperation. That is a legitimate empirical result. The problem is the central theoretical interpretation. The authors claim that 'behaviour, not identity' drives cooperation, but the design intentionally prevents participants from observing the bot's individual behaviour. Therefore the null label effect may simply reflect that participants had no information with which to attribute any specific action to the AI. This is precisely the reader's weakest assumption: the conclusion is bounded by the design's opacity. The paper's own limitation section concedes that aggregate feedback dampened bot-strategy effects, which is essentially an admission that the key behavioural cue was invisible. I agree with the reader that this is the weakest point. The reader's verdict of CONDITIONAL is appropriate: the bounded claim is fine, but the stronger theoretical conclusion should be conditioned on the information structure. My stress-test does not move the verdict. I considered whether the abstract's TOST contradiction would be more load-bearing, but that is a reporting inconsistency; even if resolved, the attributional-opacity problem remains. A replication with individual-level feedback would settle whether the equivalence is real or an artifact of ignorance.","tokens_in":13730,"tokens_out":4866,"duration_ms":60916,"concrete_test":"Run a close replication of Experiment 1 with one change: after each PGG round, the results page lists each group member's contribution individually (e.g., 'You: X, Player 2: Y, Player 3: Z, Player 4: W'), while keeping the human/AI label condition identical. If the null label effect and null label-by-strategy interaction persist when the bot's individual contributions are visible, normative equivalence survives the stronger test. If label-contingent differences appear (e.g., participants punish an AI free-rider less harshly, or reward an AI cooperator less), the original finding is attributable to information opacity rather than normative equivalence. Use the same preregistered TOST bound (±5 tokens) and group-level clustered standard errors.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim — \"behaviour, not identity, drives cooperation\" — is not supported by this design. In §2.2, participants \"did not see the individual contributions of specific group members\"; they saw only the group total and their own payoff. The bot's strategy (unconditional cooperation, conditional cooperation, free-riding) is therefore never directly observable. Under this information structure, the AI label is a semantic cue with no behavioral referent. The null label effect could mean either (a) normative equivalence — people treat AI the same even when its behavior is visible — or (b) attributional ignorance — participants could not connect any observed behavior to the bot, so the label was inert. The paper repeatedly draws conclusion (a), e.g., §4.1: \"participants rely on observable behaviour as the primary normative cue rather than the agent's identity,\" yet the agent's behavior is never observable. The authors partially acknowledge this in §4.3: \"the impact of the bot's strategy was likely dampened by the aggregate feedback mechanism.\" That admission undercuts the theoretical interpretation. \"Bounded normative equivalence\" is a valid empirical description of this anonymous aggregate setting, but the title and abstract overclaim by implying identity cues are generally subordinated to behavior.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an online experiment (N=236) using a four-player repeated Public Goods Game with three human participants and one bot framed as either human or AI, with three bot strategies (unconditional cooperator, conditional cooperator, free-rider). Participants received only aggregate group feedback, not individual contributions. The authors find no significant effect of the AI label on PGG contributions, PD cooperation, or norm perceptions; cooperation is driven primarily by the group's lagged contribution and the participant's own lagged contribution. They interpret this as 'normative equivalence' and conclude that behaviour, not identity, drives cooperation in mixed human–AI groups.","tokens_in":1594,"tokens_out":1449,"duration_ms":57839,"significance":"If the empirical null effect is taken at face value, the study provides a useful descriptive baseline: under anonymous aggregate feedback and minimal social presence, an AI label does not change cooperative behaviour in a small-group PGG. The paper is genuinely preregistered (AsPredicted #234846), reports deviations honestly, and includes robustness checks excluding suspicious participants. However, the central theoretical claim—that behavioural signals override identity cues—is not actually tested by the design, because the bot's individual behaviour was never observable to participants. The statistical analysis also ignores group-level nesting. The contribution is better positioned as a bounded empirical finding about an information structure than as evidence about the relative weight of behaviour vs. identity.","major_comments":[{"comment":"The central claim that 'behaviour, not identity' drives cooperation (title; §4.1) is not testable with this information structure. §2.2 states participants 'did not see the individual contributions of specific group members,' so they could never observe the bot's strategy. §4.3 admits 'the impact of the bot's strategy was likely dampened by the aggregate feedback mechanism.' The observed null label effect is equally compatible with attributional ignorance: the label was inert because it had no behavioural referent. To support the claim, the authors would need a condition with individual-level feedback, or evidence that participants could infer the bot's behaviour from aggregate totals. At minimum, the interpretation must be reframed as 'no label effect under anonymous aggregate feedback' rather than 'behaviour overrides identity.'","section":"§2.2, §4.1, §4.3"},{"comment":"The abstract states that a 'formal equivalence test (TOST)' indicated a label effect smaller than ±5 tokens, but §3.2.1 reports a 90% CI from estimated marginal means and explicitly states 'we cannot claim formal statistical equivalence.' No TOST procedure is described anywhere in the paper. The equivalence bound of 5 tokens appears post hoc and was not preregistered. Please either conduct and report a pre-specified TOST (with the bound justified and the test statistics) or correct the abstract and remove the term 'formal equivalence.' This discrepancy between abstract and full text is load-bearing for the paper's main conclusion.","section":"Abstract; §3.2.1"},{"comment":"The mixed-effects model accounts for participant-level random effects but not group-level clustering. Each group consists of three human participants interacting with the same bot and with each other for ten rounds, so contributions within a group are likely correlated. Ignoring group as a random effect (or using cluster-robust standard errors) can understate standard errors and inflate the apparent precision of the null label effect. The reported 90% CI for the label contrast [–3.92, 2.94] should be re-estimated with group-level clustering before claiming that effects larger than ±4 tokens can be ruled out. This also affects the p-values reported for the strategy comparisons.","section":"§3.2.1, Table 1"},{"comment":"The bot strategies were not disclosed and were not individually observable; the lack of strategy effects therefore does not support the interpretation that 'two other human moderators buffered the group against the extreme behaviours' exhibited by the bot. Because participants only saw the group total, the difference among the three bot strategies (always 100, always 0, or the previous group average) was diluted by the two human contributions and by the group average. The strategy manipulation may have failed to manipulate participants' perceptions of the bot's behaviour. This makes the null strategy effects difficult to interpret and undermines the 'human buffer' interpretation offered in §3.2.1.","section":"§3.2.1, §2.2"}],"minor_comments":[{"comment":"The sentence 'neither the human-AI label nor the specific bot strategy produced caused differences' contains a typo ('produced caused differences' should be 'produced differences').","section":"§3.1"},{"comment":"The Figure 3 caption refers to 'the first letter' and 'the second letter' but does not explain the letter codes in the figure legend. Please clarify what C and D denote in both the figure and the caption.","section":"Figure 3"},{"comment":"The preregistration is referenced as AsPredicted #234846 at one point and as 'https://aspredicted.org/wd7j-jyg5.pdf' elsewhere. Please confirm these refer to the same document and make the citation consistent.","section":"§2.4"},{"comment":"The wording 'not bound to specific to human or AI agents contributions' is grammatically unclear. Presumably 'not bound to specific human or AI agents' contributions' was intended. Please revise.","section":"§2.5"},{"comment":"The conclusion states 'When behaviour is transparent, individuals appear to rely on shared group signals rather than categorical distinctions.' This conflicts with the design, where behaviour was not transparent. This sentence should be corrected to refer to the aggregate feedback setting actually studied.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"The empirical finding—no label effect under anonymous aggregate feedback—is potentially publishable if reframed as a bounded, descriptive result. The current title and abstract overstate the theoretical reach of the design, and the statistical analysis should account for group-level clustering. The conceptual limitation is not fixable without new data, but a careful rewrite that acknowledges attributional ignorance as an alternative explanation would bring the claims in line with the evidence. I would support revision rather than rejection, but the manuscript needs substantial reworking rather than minor edits."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good paper to know about, but read the full text before quoting the abstract. The headline result — no detectable label effect in a repeated PGG when a bot is labelled human vs AI — is credible and preregistered. Yet the paper's own interpretation goes beyond what the design can support, and the abstract overstates the statistical claim.\n\nWhat's new: three humans plus one bot in a repeated four-player PGG, with three bot strategies, a PD follow-up, and norm elicitation. That fills a real gap; the prior literature is mostly dyadic or all-machine. The preregistration, the suspicion robustness check, and the descriptive reporting are careful. The null result is a useful baseline for mixed human-majority groups.\n\nThe soft spots: the title and §4.1 claim that 'behaviour, not identity' drives cooperation. But participants never saw individual contributions — only the group total. So they couldn't tell what the bot did. The null label effect is exactly what you'd expect if the label had no behavioural referent. The paper half-acknowledges this in §4.3 but doesn't follow through. 'Bounded normative equivalence' is a fair description of the setting; 'behaviour not identity' is not. Also, the abstract's TOST claim — that equivalence was formally established — is contradicted by the full text, which correctly says no formal equivalence can be claimed. That needs fixing. Two smaller points: the mixed models include participant random effects but no group-level random effects, which could understate uncertainty; and data/code are only available on request, with an inconsistent preregistration link, making verification harder.\n\nWho it's for: experimentalists working on human-AI cooperation, especially the machine-penalty literature. It will be cited as a group-level null result. It deserves a serious referee, but I'd expect major revision to align the claims with the information structure. Send it out, but push the authors to narrow the interpretation.","headline":"A credible null result with an overambitious interpretation — the design hides the bot's behavior, so 'behaviour, not identity' isn't actually tested.","tokens_in":14468,"tokens_out":3131,"would_cite":true,"duration_ms":32984,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In a four-player public goods game, labelling one teammate as AI leaves cooperation, norm persistence, and norm perceptions unchanged when participants only see aggregate group feedback; cooperative behaviour tracks the group's previous con","keywords":["human-AI cooperation","public goods game","social norms","normative equivalence","conditional cooperation","algorithm aversion","aggregate feedback","group behaviour"],"falsifier":"Run the same experiment but show participants an itemized breakdown of each group member's contribution after every round; if cooperation levels or norm ratings then diverge between the human-labelled and AI-labelled conditions, the claimed normative equivalence is an artefact of aggregated feedback.","tokens_in":13669,"feed_emoji":"🤖","tokens_out":2678,"duration_ms":35257,"temperature":0.7,"pith_summary":"This paper tries to establish that cooperative social norms extend to artificial agents in small groups when individual actions are not observable. In a repeated public-goods experiment with three humans and one scripted bot, framing the bot as human or AI made no detectable difference in how much people contributed, how they responded to others' contributions, or how norms carried over to a later prisoner's dilemma. The authors argue this 'normative equivalence' is bounded by the informational structure: aggregate feedback prevents participants from attributing the bot's behaviour, so group behaviour, not partner labels, drives cooperation. If right, this challenges dyadic algorithm-aversion findings and suggests that in opaque collective settings, AI agents can be absorbed into human cooperative norms without special treatment.","feed_headline":"AI label changes nothing when group feedback stays anonymous","feed_subtitle":"Public goods experiment: cooperation tracks group behaviour, not whether a teammate is human or AI.","key_machinery":"The load-bearing design element is anonymous aggregate feedback: after each round participants saw only the total group contribution and their own payoff, never who contributed what. This mimics the opacity of large-scale collective action and prevents participants from attributing the bot's individual strategy or identity from its actions. Across conditions, the bot followed one of three predefined strategies—unconditional cooperation, conditional cooperation, or free-riding—and was labelled either human or AI. The statistical core is a mixed-effects regression showing that lagged group contribution and own lagged contribution predict current contribution, with a precision analysis (90% con","core_discovery":"The central finding is a null result with bounded precision: across 236 participants, contributions in a ten-round public goods game differed by only about one token between human-labelled and AI-labelled conditions (b = 1.09, p = .738), with a 90% confidence interval ruling out label effects larger than roughly ±4 tokens. Cooperation was instead driven by the group's previous average contribution and by one's own previous contribution, and these mechanisms were statistically indistinguishable across labels. A follow-up one-shot prisoner's dilemma showed no label-based difference in norm persistence, and post-game norm ratings (social appropriateness, empirical expectations, injunctive expec","pith_inferences":["A direct testable extension the authors did not run: if the same design is repeated with itemized feedback showing each member's contribution each round, label effects should reappear—this would confirm that aggregate feedback, not normative flexibility, drives the null result.","The equivalence may have accountability consequences: if people treat AI teammates as norm-following group members, responsibility for collective outcomes could diffuse as readily to AI agents as to humans, raising questions about blame and credit.","The absence of bot-strategy effects hints that a single extreme actor is diluted by two human peers; varying the proportion of AI agents could reveal a threshold where the machine penalty emerges, consistent with prior all-machine findings.","Because the setting is anonymous, short-term, and low-stakes, the equivalence may not hold in repeated, identifiable, or high-stakes collaborations; status-based differentiation could resurface when reputation is at stake."],"forward_implications":["If the central claim is correct, algorithm aversion and machine-penalty effects documented in dyadic interactions do not automatically scale to mixed human–AI groups when individual contributions are hidden behind aggregate feedback.","Cooperative norms can absorb artificial agents without requiring anthropomorphic design or human-like labelling, as long as the agent's behaviour is not individually identifiable.","In groups where humans remain the majority, a single AI member's extreme strategy (always cooperate, conditionally cooperate, or free-ride) has little measurable impact on overall cooperation, suggesting human peers buffer against the bot's behaviour.","Norm persistence from a mixed group to a subsequent one-on-one interaction is no weaker than from an all-human group, implying that norms formed in hybrid groups carry over similarly.","The results provide a baseline: introducing communication, adaptive agents, or transparent individual feedback would be the natural next step to test when this equivalence breaks."],"fun_headline_variants":["Anonymous group feedback makes AI label irrelevant","Cooperation follows group, not AI label, in anonymous games","No AI effect when group feedback is anonymous","What drives cooperation? Group's past moves, not teammate's label"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The conclusion depends on participants being unable to infer the bot's individual contributions or strategy from the aggregate feedback they received; if they could, the observed null effect would reflect ignorance rather than genuinely equivalent normative logic.","fun_headline_variants_meta":{"raw":{"variants":["Anonymous group feedback makes AI label irrelevant","Cooperation follows group, not AI label, in anonymous games","No AI effect when group feedback is anonymous","What drives cooperation? Group's past moves, not teammate's label"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1163,"prompt_tokens":799,"completion_tokens":364,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":313}},"tokens_in":543,"tokens_out":364,"duration_ms":4821,"temperature":1.0,"reasoning_tokens":313,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T06:13:33.643720+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same experiment but show participants an itemized breakdown of each group member's contribution after every round; if cooperation levels or norm ratings then diverge between the human-labelled and AI-labelled conditions, the claimed normative equivalence is an artefact of aggregated feedback.","supporting_citations":[],"review_version":1}