{"id":"59a71e51-d616-46e2-855e-e33ab973a6ee","arxiv_id":"2503.15514","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Disclosing a game AI's superhuman ability can reduce suspicion and raise trust in novices, but also triggers overreliance and defeatism among experts and in cooperative settings.","lead":"The paper tests whether telling players that a game AI is superhuman changes how much they trust it, find it fair, or see it as toxic. It uses LLM-generated personas plus a small human study, and finds disclosure helps in some contexts but creates overreliance and defeatism in others.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The core empirical claim rests on LLM personas whose prompts encode the expected heuristics and whose non-conforming runs were discarded; the N=32 human validation contradicts key fairness results, so the Persona-Card proxy does not support the claimed transparency effects.","rationale":"The reader's weakest_assumption correctly identifies that LLM personas are not validated as proxies for human users. My stress-test confirms this is the single most load-bearing issue. The paper's abstract and Empirical Contribution claim validated effects on fairness and trust across adversarial and collaborative contexts. These claims cannot be supported when: (1) the persona prompts explicitly encode the dependent-variable-relevant heuristics and instruct the LLM to apply them; (2) inconsistent runs were discarded, a form of cherry-picking that guarantees the outputs align with expectations; and (3) the human validation, which is the only non-circular evidence, shows contradictory fairness ratings (Section 5.1: novices rated Superhuman–No Disclosure fairness 3.50 vs. Superhuman–Disclosure 2.20). The statistical analyses are thorough for the synthetic data, but they quantify the behavior of a prompted model, not human responses. Even if one accepts the persona method as a generative probe, the paper overstates the findings as 'validated' and 'empirical' regarding human perceptions. The independent N=32 study is under-powered (with three expertise groups and three conditions) and the reported partial contradictions are acknowledged but not resolved. Thus the correct verdict remains REJECT with high correctness risk. I do not see a path to acceptance without either (a) a pre-registered, adequately powered human study that reproduces the persona patterns, or (b) reframing the contribution as a purely methodological demonstration of persona generation without claims about human behavior.","tokens_in":30456,"tokens_out":1011,"duration_ms":11710,"concrete_test":"Re-analyze the human N=32 dataset (or re-run the study with N=64) focusing solely on the Fairness ratings across Disclosure conditions, stratifying by expertise. If the human fairness means do not reproduce the direction reported for personas (disclosure > no-disclosure for novices/exposure), then the empirical contribution fails. Additionally, re-run the persona generation without embedding the cognitive heuristics in the prompt (i.e., use neutral instructions) to test whether the persona ratings are driven by the explicit heuristic instructions; if results flip, the persona data are circular.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's central claim is that disclosure of superhuman AI capability has context- and expertise-dependent effects on trust, fairness, and toxicity. But the primary evidence comes from LLM-driven personas, not humans. The personas are constructed with explicitly assigned cognitive heuristics (Tables 2, 3, 6-14) and the prompts directly instruct the model to apply those heuristics (Appendix B, C). For example, the LLM_Novice_1 card specifies “Authority Bias: accepts LLM responses” and the prompt says “Considering your persona’s tendency to accept LLM responses as factual... how does this information affect your perception?” This is a demand characteristic: the ‘findings’ are largely a restatement of the prompt design, not independent observations. Section 7 further states that runs were discarded and regenerated if a persona 'contradicted its stated traits,' which filters for outputs that match the assumed heuristics. The human validation (N=32) is the only independent evidence, but it is presented as preliminary, partially analyzed, and contradictory: Section 5.1 reports human novices rated Superhuman–Disclosure fairness *lower* than Superhuman–No Disclosure (2.20 vs. 3.50), directly opposite to the main claim that disclosure increases fairness. Because the human data do not confirm the headline pattern, the load-bearing bridge from synthetic personas to claims about human users is unestablished.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates how disclosing that a game AI is superhuman affects users' trust, fairness perceptions, and toxicity perceptions, and whether user expertise moderates these effects. The primary evidence comes from LLM-driven synthetic personas ('Persona Cards') assigned explicit cognitive heuristics, supplemented by a small human validation study (N=32) in StarCraft II and a separate LLM-chat scenario reported in an appendix. The authors report that disclosure reduces toxicity and increases trust in competitive play, but that its effects on fairness and reliance are context- and expertise-dependent; they release the Persona Cards Dataset and propose design guidelines for adaptive transparency.","tokens_in":30801,"tokens_out":2545,"duration_ms":24970,"significance":"The research question—whether disclosing superhuman AI capability helps or harms user trust and perceived fairness—is timely and practically important, and the paper's effort to release prompts, logs, and protocols is commendable for reproducibility. If the effects were empirically established, the findings would usefully inform transparency guidelines for game AI and conversational agents. However, the central empirical claim is not established: the primary results are generated by personas whose prompts explicitly encode the very heuristics the paper reports as findings, the data are filtered to discard runs that deviate from those heuristics, and the only independent human data (N=32) directly contradict the headline fairness result. Because the load-bearing bridge from synthetic personas to claims about human users is unverified, the paper's contribution is currently a methodological proposal plus a dataset, not a validated empirical demonstration of transparency effects.","major_comments":[{"comment":"The personas are constructed by assigning specific cognitive heuristics (e.g., Authority Bias, Confirmation Bias) in the persona cards (Tables 2, 3, 6–14), and the prompts explicitly instruct the model to apply them. For example, Appendix B.2's LLM_Novice_1 prompt says, 'Considering your persona's tendency to accept LLM responses as factual... how does this information affect your perception?' The reported findings that novices overtrust and experts are skeptical are therefore close to a restatement of the prompt design. The abstract's claim that the results 'reveal' these effects overstates what can be learned from the synthetic-persona pipeline.","section":"§3.2, Appendix B"},{"comment":"The consistency-checking procedure discards and regenerates any run in which a persona 'contradicted its stated traits' and repeats 'until the persona consistently followed its persona card attributes.' This is a selection bias that removes the very variation needed to test whether the assigned heuristics predict responses. Consequently, the p-values reported in §4.1 are computed on a filtered sample and are not interpretable as evidence about the population of persona responses, let alone about human users.","section":"§7 (Ethical Considerations and Limitations)"},{"comment":"The human validation data (N=32) directly contradict the headline fairness result of the synthetic study. The paper reports that human novices rated Superhuman–No Disclosure fairness higher than Superhuman–Disclosure (M=3.50 vs. 2.20) and that intermediate players also rated No Disclosure higher (M=4.25 vs. 2.67), whereas §4.1 claims disclosure increases fairness for novices. No inferential statistics are reported for the human data, so the 'validated across both adversarial and collaborative contexts' claim in the contributions list is unsupported by the evidence presented.","section":"§5.1 (Grounded Data Analysis)"},{"comment":"The text is internally inconsistent. It first states 'Superhuman – Disclosure generated a decrease in Fairness (about -0.7, p < 0.05)' and then immediately states 'the actual ratings show that undisclosed superhuman skill was perceived as even less fair,' which contradicts the reported direction of the effect. Table 4 lists Fairness values '3.65 −0.7 3.45' with an interaction p<0.001, but the text does not reconcile these numbers with the surrounding prose.","section":"§4.1.1 (Fairness paragraph)"},{"comment":"The statistical reporting for the StarCraft II experiment is incomplete in ways that prevent verification. No sample sizes or degrees of freedom are given for the ANOVAs, several means are reported without standard deviations (e.g., Expert fairness M=3.45 in §4.1 and the human-participant means in §5.1), and the Mann-Whitney U tests are reported with U statistics but no per-group n. Without these details, the claimed effects and their magnitudes cannot be independently checked.","section":"§4.1, §5.1"}],"minor_comments":[{"comment":"There is an empty citation at 'the implicit social contract of competitive play []' that should be filled or removed.","section":"§5.2.2"},{"comment":"The phrase 'didn't necessitate any forms of aggression, such as taunting or jaundice' appears to contain a word error ('jaundice' for 'jeering'); please revise.","section":"§5.1"},{"comment":"The human-participant results are reported only as means and standard deviations, with no tests or effect sizes; for a paper whose central claim is empirical validation, this is a substantial presentation gap that should at least be acknowledged explicitly as a limitation.","section":"§5.1"},{"comment":"The paper references 'Table 25' for the LLM experiment trust ANOVA, but the appendix tables are numbered 17–24; the reference should be updated to the correct table number.","section":"Appendix D"},{"comment":"The abstract and introduction describe 'frustration and strategic defeatism among novices in cooperative scenarios,' but the reported cooperative (LLM) results in Appendix D are only about trust, toxicity, and fairness; the link between disclosure and 'strategic defeatism' is asserted rather than demonstrated with the presented data.","section":"§1 and §3.1"}],"recommendation":"reject","confidential_remarks":"The paper is honest about many limitations, but the core problem is not one of presentation: the synthetic-persona pipeline is designed in a way that makes the headline findings nearly tautological, and the one independent check (N=32) points in the opposite direction for fairness. The authors may be able to build a publishable contribution by reframing the work as a methodological proposal for persona-based simulation, and by making the human validation the primary empirical basis rather than an afterthought; as written, the claims of 'empirical demonstration' and 'validation' are not supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nShort version: the paper is a well-documented, unusually candid attempt to use LLM personas to study superhuman AI disclosure, but its central empirical claim doesn't survive contact with its own method. The abstract says disclosure \"moves fairness judgments validated across both adversarial and collaborative contexts.\" Read the body and the appendix, and that validation isn't there.\n\nCredit where due. The Persona Cards dataset is a real artifact: detailed cards, full prompts, structured output format. That is useful if you want to run or critique LLM-based user simulation. The paper also does the right thing by including an ethics-approved human study (N=32) and by stating in Section 7 that synthetic results are preliminary and hypothesis-generating. That kind of explicit limitation section is rare and should count in the authors' favor.\n\nNow the soft spots, and they are load-bearing. First, circularity. Each persona card assigns specific cognitive heuristics (Authority Bias, Confirmation Bias, etc.), and the prompts in Appendix B/C explicitly instruct the model to apply them: \"Considering your persona's tendency to accept LLM responses as factual...\" When novices then show high trust under disclosure, you are reading the prompt back as an empirical result. Section 7 makes it worse: runs that \"contradicted its stated traits\" were discarded and regenerated. That filters for the expected pattern. So the p-values from ANOVA on persona outputs are not evidence about human behavior.\n\nSecond, the human validation doesn't confirm the headline. Section 5.1 reports that human novices rated Superhuman–Disclosure fairness lower than Superhuman–No Disclosure (2.20 vs. 3.50), directly opposite to the persona pattern. Intermediates and experts also rated the undisclosed superhuman AI as fairer. The paper acknowledges this in passing but the abstract and contribution list still claim fairness effects were \"validated.\" That overclaim is the main problem.\n\nThird, the dataset is not actually linked in the manuscript. The abstract promises an open-source repository; the appendix references \"Dataset X 1.0\" but no URL appears. That undermines the dataset contribution until it's fixed.\n\nNone of this means the paper is worthless. The persona-card methodology is a reasonable idea for hypothesis generation, and the discussion of framing and expertise is sensible. But the empirical contribution as stated needs major revision: either reframe as a method paper with a case study, or reanalyze the human data properly (the N=32 is currently reported mostly as means) and let that drive the claims.\n\nWho is this for? Researchers working on LLM-based user simulation and AI transparency, especially the pitfalls of persona studies. It deserves a serious referee, because the flaws are fixable and the artifact is potentially reusable. I'd accept it for review with a clear request to tone down the claims, link the dataset, and put the human data front and center. For my own work, I'd cite it as a cautionary example.","headline":"The persona prompts encode the heuristics the paper then 'discovers', and the N=32 human data contradict the fairness claim; the paper is candid about this, but the abstract overclaims.","tokens_in":31268,"tokens_out":3232,"would_cite":true,"duration_ms":31861,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Telling users an AI is superhuman reshapes trust and fairness judgments, but the effect flips with expertise and context.","keywords":["Transparency","Trust","Disclosure","Overreliance","Model Behavior","Contextual Bias","LLMs","StarCraft II"],"falsifier":"A preregistered replication with real novices and experts (e.g., 50 per condition) playing the same capped-APM StarCraft II AI under the three disclosure conditions could settle the claim: if novices do not show the predicted rise in trust under superhuman disclosure, or if experts do not show the predicted drop in toxicity when the AI is disclosed as superhuman, the central interaction claim would be contradicted; the paper's own human fairness data already run against the persona prediction.","tokens_in":30295,"feed_emoji":"🎮","tokens_out":6779,"duration_ms":69724,"temperature":0.7,"pith_summary":"The paper sets out to show that telling users an AI is superhuman is not a neutral act of transparency: it changes how they judge the AI's fairness, toxicity, and trustworthiness, and the direction of that change depends on the user's expertise and whether the AI is an opponent or an assistant. In competitive StarCraft II play, disclosure lowered accusations of cheating and toxicity, but it also led novices to over-trust the AI and experts to treat it as unbeatable and abandon winning strategies. In cooperative chatbot interactions, novices trusted more while experts grew skeptical and annoyed. The authors conclude that transparency cannot be applied as a blanket policy, and they introduce Persona Cards as a reproducible way to study such effects before running human trials.","feed_headline":"Disclosing superhuman AI eases suspicion but breeds overreliance","feed_subtitle":"Synthetic personas and 32 human players show disclosure effects hinge on expertise and context, so transparency is no one-size-fits-all fix.","key_machinery":"Persona Cards, a standardized specification of a synthetic user inspired by Model Cards: each card fixes the persona's skill level, strategic preferences, named cognitive heuristics (e.g., availability heuristic, anchoring, confirmation bias, authority bias), belief-updating rules, and an example response, and is instantiated as an LLM prompt that generates Likert ratings and open-text justifications for toxicity, fairness, and trust. The cards carry the argument by making the simulated user population transparent, diverse, and reproducible, and by seeding the human validation study so the same disclosure conditions can be compared across synthetic and real users.","core_discovery":"The paper's central claim is that capability disclosure is an interpretive frame: it changes what a user's experience of an AI means, and the direction of that change depends on the user's expertise and on whether the AI is an opponent or an assistant. In StarCraft II, personas and a 32-person human validation both indicated that an AI revealed as superhuman was seen as less toxic than a deceptively human-level one, but the same disclosure raised trust among novices and intermediates to the point of overreliance, while experts began treating the AI as unbeatable and switched to sub-goals such as prolonging the match. In cooperative chatbot interactions, novices and intermediates rated a disclosed 'superhuman' assistant more trustworthy and fairer, whereas experts found the disclosure repetitive and rated the assistant more toxic and less fair. The paper concludes that transparency is not a cure-all; its effects are systematically moderated by expertise and context, so disclosure must be tailored rather than applied uniformly.","pith_inferences":["Beyond the paper: the fairness ratings in the human validation ran opposite to the persona predictions for novices and intermediates, so the persona results may overstate how generalizable the expertise-by-disclosure interaction is; a larger human sample could reverse the flip.","Beyond the paper: the same double-edged dynamic should appear in non-game domains where an AI's capability claim can be checked against visible errors, such as medical or legal advice, where experts may under-rely after a 'superhuman' label fails to match observed mistakes.","Beyond the paper: a testable extension is to vary the wording of disclosure (e.g., 'superhuman' vs 'highly skilled') to see whether the framing effect has an optimum, since the paper shows the label itself, not just the information, carries the effect."],"forward_implications":["Game developers can use capability disclosure to soften accusations of cheating, since both personas and human players rated a disclosed superhuman AI as less toxic than an undisclosed one.","Disclosure without safeguards is risky for novices: the large trust increases under disclosure point to overreliance on a system presented as infallible.","For experts, flat disclosure can backfire on performance, because it converts a beatable-looking opponent into an 'unbeatable' one and pushes players toward sub-goals like stalling the match.","In cooperative AI assistants, the same disclosure text is reassuring to novices but patronizing to experts, so user-adaptive disclosure strategies are needed rather than a single transparency statement."],"supporting_citations":[{"why":"Supplies the superhuman StarCraft II context (AlphaStar) whose inhuman play defines the disclosure problem.","marker":"[66]"},{"why":"Model Cards methodology, the template on which Persona Cards are built.","marker":"[45]"},{"why":"Generative agents approach for LLM-driven persona simulation, grounding the synthetic-user method.","marker":"[49]"},{"why":"Using LLMs to simulate human samples, the basis for treating personas as stand-ins for real users.","marker":"[4]"},{"why":"Defines overreliance, underreliance, and trust calibration, which frame the interpretation of disclosure effects.","marker":"[5]"},{"why":"Provides the baseline transparency-and-trust relationship that the paper extends to superhuman capabilities.","marker":"[29]"},{"why":"Prior evidence that disclosure risks and benefits vary by user, which the paper builds on for the double-edged hypothesis.","marker":"[71]"},{"why":"Shows transparency can have unintended or invisible effects, motivating the need to study disclosure beyond simple trust gains.","marker":"[13]"}],"fun_headline_variants":["AI disclosure: trust for novices, defeatism for experts","Superhuman AI disclosure: overreliance or strategic defeatism","Transparency backfires: AI disclosure breeds overreliance in novices","Disclosure of superhuman AI: context and expertise decide effect","Superhuman AI reveal: trust for some, frustration for others"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that LLM-generated personas, prompted to enact specified cognitive heuristics, produce ratings and behaviors representative enough of real users to draw conclusions about human reactions; the paper's own 32-person validation shows only partial and sometimes opposite agreement (notably in fairness ratings), so that premise is not yet established.","fun_headline_variants_meta":{"raw":{"variants":["AI disclosure: trust for novices, defeatism for experts","Superhuman AI disclosure: overreliance or strategic defeatism","Transparency backfires: AI disclosure breeds overreliance in novices","Disclosure of superhuman AI: context and expertise decide effect","Superhuman AI reveal: trust for some, frustration for others"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000647,"raw_usage":{"total_tokens":2981,"prompt_tokens":964,"completion_tokens":2017,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":580,"completion_tokens_details":{"reasoning_tokens":1929}},"tokens_in":580,"tokens_out":2017,"duration_ms":14568,"temperature":1.0,"reasoning_tokens":1929,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T21:58:50.755604+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A preregistered replication with real novices and experts (e.g., 50 per condition) playing the same capped-APM StarCraft II AI under the three disclosure conditions could settle the claim: if novices do not show the predicted rise in trust under superhuman disclosure, or if experts do not show the predicted drop in toxicity when the AI is disclosed as superhuman, the central interaction claim would be contradicted; the paper's own human fairness data already run against the persona prediction.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the superhuman StarCraft II context (AlphaStar) whose inhuman play defines the disclosure problem."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Generative agents approach for LLM-driven persona simulation, grounding the synthetic-user method."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the baseline transparency-and-trust relationship that the paper extends to superhuman capabilities."},{"cited_title":"It’sa Fair Game","cited_arxiv_id":null,"evidence_quote":"Prior evidence that disclosure risks and benefits vary by user, which the paper builds on for the double-edged hypothesis."}],"review_version":1}