REVIEW 4 major objections 6 minor 45 references
A mass multiplayer game where students built LLM bots to sway a fake election did not boost their bot-spotting confidence and blunted their discomfort with spreading misinformation.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-07-31 23:18 UTC pith:KYSVQ3SK
load-bearing objection A genuinely impressive large-scale deployment and honest experience report, but the headline claims about inoculation theory and desensitization rest on survey measures that do not support them. the 4 major comments →
On Capturing the Narrative: Social Media Manipulation Wargaming for Cyberliteracy
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's core discovery is empirical and partly counterintuitive: a large, carefully built role-reversal game, with real incentives and a realistic multi-agent platform, failed to make participants better at spotting bots and instead made them emotionally colder toward misinformation. On self-report measures, 256 students before and 83 after the game showed no significant change in concern about bots, perceived bot influence, or confidence in detecting bots (all p > .09). In contrast, all emotion-related measures—regret, guilt, shame, anxiety about consequences, and motivation to correct falsehoods—moved in the direction of reduced sensitivity, with p-values from 1.3e-5 down to 9.2e-7. Th
What carries the argument
The central machinery is the Capture the Narrative environment itself: a custom social-media platform (Legit Social) populated by 4,000 LLM-backed NPC citizens, each with a 40-dimensional profile and a probabilistic opinion-update rule, plus seven 'special' NPCs—journalists, columnists, candidates, and an outgoing president—who write news stories based on platform activity. Player teams get a public API and can run up to 40 bot accounts each; only NPCs can vote. The scoring system does the explanatory work: story score credits teams whenever NPC attitudes shift, weighted toward election day, while engagement score rewards likes, reposts, and trending placement. The paper shows how those two
Load-bearing premise
The load-bearing premise is that self-reported confidence in detecting bots and self-rated emotional reactions, gathered from the 83 students who completed the post-survey out of 256, accurately capture the inoculation and desensitization effects the game is claimed to have—despite 67.5 percent attrition and no direct measure of actual detection ability.
What would settle it
Run the same competition with a behavioral outcome: before and after, have participants classify a held-out set of posts as human, NPC, or player-bot, and measure both accuracy and willingness to share each post, with a control group that does not play. If post-game detection accuracy improves markedly while confidence is unchanged, the paper's 'no inoculation' claim would be overturned; if emotional discomfort does not decline when debriefing is added, the desensitization claim would be narrowed.
If this is right
- Active inoculation through attacker-side play cannot be assumed to work at scale; the paper's own results argue for adding a defensive 'blue-team' phase.
- Digital literacy interventions need behavioral outcome measures, such as real bot-detection accuracy, because self-reported confidence did not move even when the game was massive and immersive.
- The engagement-score effect reproduces real-world platform dynamics inside the classroom, meaning the simulation is a credible model of how GenAI lowers the barrier to influence operations.
- Educators who assign misinformation-creation exercises should expect an emotional desensitization side effect and plan structured debriefs to counter it.
- Future iterations should redesign scoring to reward sustained, quality influence with diminishing returns on volume, as the paper recommends.
Where Pith is reading between the lines
- A direct test of the paper's interpretation would pair the same game with a pre/post behavioral task—having participants classify held-out posts as NPC-generated, player-bot-generated, or human-authored—and a no-game control group; if detection accuracy improves while confidence stays flat, the 'no inoculation' conclusion would need revision.
- The observed 'spam pivot' suggests that manipulation strategy is shaped more by the reward structure than by players' ethical priors; if true, the same logic implies real-world platform incentive redesign could do more to curb spam misinformation than audience-side literacy training.
- The emotional desensitization result, if it generalizes, implies that short, high-intensity exercises in creating misinformation may be net harmful for media literacy, and that ethical reflection has to be built into the exercise itself, not just surveyed afterward.
- The paper's framing of 'inoculation failure' may actually point to a boundary condition: inoculation games work when tactics are static and finite, but may fail when the adversary is a generative model that invents new tactics continuously.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Capture the Narrative (CTN), a four-week multi-university competition in which 108 student teams built LLM-powered bots to influence a simulated election on a custom social-media platform with 4,000 NPC citizens. The authors report platform-scale descriptive metrics (7,068,206 player-bot posts, ~60% of platform content), pre/post survey data from 256 participants before and 83 after the competition, and thematic analyses of participant strategies. The central empirical claims are that (i) participants did not become more confident at spotting bots, contrary to inoculation theory, and (ii) participants' self-reported emotional discomfort with spreading misinformation attenuated after the competition. The paper concludes with design lessons and future improvements such as a blue-team phase.
Significance. The engineering and deployment achievement is substantial: a large-scale, multi-agent LLM-driven competition at 18 universities is a novel contribution to misinformation literacy education. The descriptive operational data (team counts, posting volumes, strategic pivots) and the candid reflection on incentive misalignment are useful for educators. However, the empirical claims about inoculation and emotional desensitization rest on unvalidated self-report instruments and a severely attritioned sample with no comparison group. The paper is best read as an experience report, but the abstract and conclusion frame it as an evaluative study; that framing requires major revision. If the authors reframe the claims as exploratory and provide effect sizes plus transparent limitations, the paper could make a valuable contribution to the CSCI/CY community.
major comments (4)
- [Section 5.4 / Section 7] The central conclusion that participants' 'emotional sensitivity to spreading misinformation appeared to decline' is supported only by self-report Likert items about guilt, regret, shame, anxiety, and motivation to correct misinformation. There is no behavioral measure (e.g., actual posting decisions, deletion rates, response to detection, or subsequent sharing behavior in a transfer task). Attenuation of reported emotion could reflect response-shift bias, social desirability at pretest, or regression to the mean. Without a control group, the observed pre/post change cannot be attributed to CTN. Please present these results as exploratory descriptive findings and temper the causal language in the abstract and Section 7.
- [Section 5.1 / Section 5.3 / Section 7] The conclusion that CTN 'did not deliver the bot-detection inoculation' conflates confidence in detecting bots with the construct of inoculation, which predicts resistance to persuasion. The survey item in Section 5.1 measures self-perceived confidence, not actual detection accuracy or resistance to misleading content. The only quasi-behavioral measure, Section 5.3's attribution of bot posts, is post-only (n=69), does not compare accuracy against ground truth, and is not linked to the pre/post confidence measure. To support the inoculation-theory claim, the study would need a behavioral resistance measure (e.g., willingness to share/engage with misinformation post-intervention, or a validated detection task). As written, the claim is not testable with the reported data.
- [Sections 5.1, 5.4, 6.1] The inferential statistics are underpowered and potentially biased. Attrition is 67.5% (83 of 256). The Mann-Whitney U test in Section 6.1 compares completers versus non-completers only on technical skills; it does not rule out selection on attitudinal or engagement variables. Additionally, multiple Wilcoxon signed-rank tests are run without any multiple-comparison correction; the p-values in Section 5.4 (e.g., 9.2e-7) are extreme for n=83 and the paper does not report effect sizes or the distribution of difference scores. Given the Likert scale's ordinal nature and the matched-pair design, the authors should report median/quartile changes, effect sizes, and corrected p-values, and explicitly acknowledge that the absence of a control group precludes causal attribution.
- [Section 5.5 / Section 6] The claim that 'most teams prioritised high-volume posting over nuanced influence' is supported only by a thematic coding of open-ended responses from 44 post-competition respondents (of whom 7 explicitly described a pivot to spamming). This is an anecdotal subset, not a systematic content analysis. If platform logs contain posting volumes and score breakdowns, the paper should present those data to substantiate the volume-over-quality claim; otherwise, this assertion should be softened to 'some teams reported'.
minor comments (6)
- [Section 5.1] The reported medians (e.g., 'influence Mdn=1, concern and frequency Mdn=2; detection Mdn=3') are confusing without a clear statement that lower values indicate stronger agreement for items in the Ethics and Emotion block, while other items may use a different direction. Please add a note on scale direction for all Likert items.
- [General / Author Template] The ACM reference format line and page footer say '2018' and 'June 03–05, 2018'—likely a leftover template artifact; please update to the actual submission year/venue.
- [Reference [12]] Reference [12] is a news article reporting the 'one in five' bot statistic; consider citing the underlying peer-reviewed study or providing a more robust source for the claim.
- [Section 5.2] The sentence '89% technical, with 11% having no prior Python experience' is redundant; also clarify what 'technical' means (field of study vs. programming experience).
- [Tables 2 and 3] Table 3 reports pre-competition n=42, but Section 5.5 says 59 respondents reported a predefined plan; clarify the relationship between these numbers (e.g., 42 provided a written strategy).
- [Section 6.1] The limitation paragraph is candid and appreciated, but it is placed late; the abstract and introduction should already signal the exploratory nature of the survey findings.
Circularity Check
No significant circularity: the paper's claims are empirical survey results, not derivations from fitted inputs or self-citations.
full rationale
This is an empirical study rather than a formal derivation. The central findings—that students did not become more confident at spotting bots and that reported emotional sensitivity to spreading misinformation declined—come from pre/post self-report surveys analyzed with Wilcoxon signed-rank tests (Sections 5.1 and 5.4). These outcomes are not constructed from the game's scoring rules or NPC parameters; the story-score attribution rule and engagement-score incentives influence observed team behavior but are not used to compute the survey outcomes. The discussion of attrition, technical cohort skew, and the lack of a defensive phase is presented as limitation, not as a circular justification. The few self-citations in the reference list (e.g., Akhtar et al.) support background claims about bot detection and manipulation and are not load-bearing for the paper's empirical conclusions. No equation is shown to equal its own input, and no fitted parameter is renamed as a prediction. The skeptical concerns about construct validity—confidence in bot detection as a proxy for inoculation, and self-reported emotion as evidence of desensitization—are legitimate evaluation concerns but are not circularity in the sense of a claim reducing to its inputs by definition or by self-citation.
Axiom & Free-Parameter Ledger
free parameters (3)
- NPC update-function parameters =
not specified
- Story vs engagement score weighting =
not specified
- Attribution credit window for story score =
not specified
axioms (4)
- domain assumption NPC attitude changes are attributable to posts they consumed (causal attribution).
- domain assumption Self-reported confidence in detecting bots is a meaningful outcome for evaluating inoculation.
- domain assumption LLM-backed NPCs produce diverse, context-sensitive behavior approximating real social media users.
- standard math Wilcoxon signed-rank test assumptions hold and no multiple-comparison correction is applied.
invented entities (2)
-
4,000 AI-driven NPC citizens
no independent evidence
-
Special NPCs (journalists, columnists, candidates, outgoing president)
no independent evidence
read the original abstract
Misinformation is deeply embedded in online discourse, with nearly one in five posts during global events generated by bots that amplify false content. In recent years, the use of Generative AI has further lowered the barrier to producing convincing misinformation, yet most digital literacy education still relies on static checklists and single-player inoculation games built for an earlier media landscape. This paper describes how we addressed this educational gap through Capture the Narrative, a four-week multi-university competition in which student teams build LLM-powered bots to influence a simulated election. We report on our custom social-media platform, the competition environment and design of its 4,000 AI-driven Non-Player Character (NPC) citizens, and what running Capture the Narrative at scale actually involved. In our first iteration, 108 teams from 18 Australian universities produced 7,068,206 player-bot posts, approximately 60% of all platform content. We surveyed 256 students before and 83 after the competition to understand their perceptions of misinformation and the game itself and found that students did not become more confident at spotting bots, contrary to what inoculation theory predicts. Because engagement was rewarded, most teams prioritised high-volume posting over nuanced influence, mirroring real-world platform dynamics. We close with recommendations for educators considering similar interventions, and propose future improvements, such as including a blue-team defensive phase.
Figures
Reference graph
Works this paper leans on
-
[1]
Mohammad Majid Akhtar, Navid Shadman Bhuiyan, Rahat Masood, Muhammad Ikram, and Salil S. Kanhere. 2025. BotSSCL: Social Bot Detection with Self- Supervised Contrastive Learning.Online Social Networks and Media48 (2025), 100318. doi:10.1016/j.osnem.2025.100318
arXiv 2025
-
[2]
Mohammad Majid Akhtar, Rahat Masood, Muhammad Ikram, and Salil S Kanhere
-
[3]
Mohammad Majid Akhtar, Rahat Masood, Muhammad Ikram, and Salil S Kan- here. 2026. TBTrackerX: Fantastic Trigger Bots and Where to Find Malicious Campaigns on X.. InNDSS
2026
-
[4]
Mayte Santos Albardía, Simón Peña-Fernández, and Irati Agirreazkuenaga. 2025. Technology, education and critical media literacy: potential, challenges, and opportunities.Frontiers in Human Dynamics7 (2025). doi:10.3389/fhumd.2025. 1608911
-
[5]
Khudejah Ali, Cong Li, Khawaja Zain-ul abdin, and Syed Ali Muqtadir. 2022. The effects of emotions, individual attitudes towards vaccination, and social endorse- ments on perceived fake news credibility and sharing motivations.Computers in Human Behavior134 (9 2022). doi:10.1016/j.chb.2022.107307
arXiv 2022
-
[6]
Lu Cheng, Ruocheng Guo, Kai Shu, and Huan Liu. 2021. Causal understanding of fake news dissemination on social media. InProceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining. 148–157
2021
-
[7]
John Cook, Ullrich K.H. Ecker, Melanie Trecek-King, Gunnar Schade, Karen Jeffers-Tracy, Jasper Fessmann, Sojung Claire Kim, David Kinkead, Margaret Orr, Emily Vraga, Kurt Roberts, and Jay McDowell. 2023. The cranky uncle game—combining humor and gamification to build student resilience against climate misinformation.Environmental Education Research29, 4 (...
arXiv 2023
-
[8]
Juan Echeverr a, Emiliano De Cristofaro, Nicolas Kourtellis, Ilias Leontiadis, Gianluca Stringhini, and Shi Zhou. 2018. LOBO: Evaluation of Generalization Deficiencies in Twitter Bot Classifiers. InAnnual Computer Security Applications Conference(San Juan, PR, USA)(ACSAC ’18). ACM, 137–146. doi:10.1145/3274694. 3274738
-
[9]
L. K. Fazio, N. M. Brashier, B. K. Payne, and E. J. Marsh. 2015. Knowledge Does Not Protect Against Illusory Truth.Journal of Experimental Psychology: General 144, 5 (2015), 993–1002. doi:10.1037/xge0000098.supp
-
[10]
Emilio Ferrara et al. 2016. The rise of social bots.Commun. ACM59, 7 (2016), 96–104
2016
-
[11]
Ira Bruce Gaultney, Todd Sherron, and Carrie Boden. 2022. Political polarization, misinformation, and media literacy.Journal of Media Literacy Education14, 1 (2022), 59–81. doi:10.23860/JMLE-2022-14-1-5
-
[12]
Dominic Giannini. 2025. Bots influencing election discussion on social media. The Canberra Times(21 April 2025). https://www.canberratimes.com.au/story/ 8946759/bots-influencing-election-discussion-on-social-media/
2025
- [13]
-
[14]
Josh A Goldstein. 2023. Generative Language Models and Automated Influ- ence Operations: Emerging Threats and Potential Mitigations.arXiv preprint arXiv:2301.04246(2023)
Pith/arXiv arXiv 2023
-
[15]
Lindsay Grace and Bob Hone. 2019. Factitious: Large scale computer game to fight fake news and improve news literacy. InConference on Human Factors in Computing Systems - Proceedings. Association for Computing Machinery. doi:10. 1145/3290607.3299046
arXiv 2019
-
[16]
Bin Guo, Yasan Ding, Lina Yao, Yunji Liang, and Zhiwen Yu. 2020. The future of false information detection on social media: New perspectives and trends.ACM Computing Surveys (CSUR)53, 4 (2020), 1–36
2020
-
[17]
Bing He, Yibo Hu, Yeon-Chang Lee, Soyoung Oh, Gaurav Verma, and Srijan Kumar. 2025. A survey on the role of crowds in combating online misinformation: Annotators, evaluators, and creators.ACM Transactions on Knowledge Discovery from Data19, 1 (2025), 1–30
2025
-
[18]
Buyun He, Yingguang Yang, Qi Wu, Hao Liu, Renyu Yang, Hao Peng, Xiang Wang, Yong Liao, and Pengyuan Zhou. 2024. Dynamicity-aware social bot detection with dynamic graph transformers. InInternational Joint Conference on Artificial Intelligence(Jeju, Korea)(IJCAI ’24). Article 646, 9 pages. doi:10.24963/ijcai.2024/ 646
-
[19]
Ali Jarrahi and Leila Safari. 2023. Evaluating the effectiveness of publishers’ features in fake news detection on social media.Multimedia Tools and Applications 82, 2 (2023), 2913–2939
2023
-
[20]
Soveatin Kuntur, Anna Wróblewska, Marcin Paprzycki, and Maria Ganzha. 2024. Under the influence: A survey of large language models in fake news detection. IEEE Transactions on Artificial Intelligence6, 2 (2024), 458–476
2024
-
[21]
Chei Sian Lee and Long Ma. 2012. News sharing in social media: The effect of gratifications and prior experience.Computers in Human Behavior28, 2 (3 2012), 331–339. doi:10.1016/j.chb.2011.10.002
-
[22]
Mingxiang Liao, Qixiang Ye, Wangmeng Zuo, Fang Wan, Tianyu Wang, et al
-
[23]
Long Ma, Chei Sian Lee, and Dion H. Goh. 2013. Understanding news sharing in social media from the diffusion of innovations perspective. InProceedings - 2013 IEEE International Conference on Green Computing and Communications and IEEE Internet of Things and IEEE Cyber, Physical and Social Computing, GreenCom- iThings-CPSCom 2013. 1013–1020. doi:10.1109/Gr...
-
[24]
Advances in Neural Information Processing Systems37 (2024), 109790–109816
Evaluation of text-to-video generation models: A dynamics perspective. Advances in Neural Information Processing Systems37 (2024), 109790–109816
2024
-
[25]
Sarah McGrew. 2024. Teaching lateral reading: Interventions to help people read like fact checkers. doi:10.1016/j.copsyc.2023.101737
arXiv 2024
-
[26]
Sarah McGrew. 2020. Learning to evaluate: An intervention in civic online reasoning.Computers and Education145 (2 2020). doi:10.1016/j.compedu.2019. 103711
-
[27]
Mohamed Mostafa, Ahmad S Almogren, Muhammad Al-Qurishi, and Majed Alrubaian. 2024. Modality deep-learning frameworks for fake news detection on social networks: A systematic literature review.Comput. Surveys57, 3 (2024), 1–50
2024
-
[28]
Nicholas Micallef, Mihai Avram, Filippo Menczer, and Sameer Patil. 2021. Fakey: A Game Intervention to Improve News Literacy on Social Media.Proceedings of the ACM on Human-Computer Interaction5, CSCW1 (4 2021). doi:10.1145/3449080
-
[29]
Jonathan Osborne and Daniel Pimentel. 2023. Science education in an age of misinformation.Science Education107, 3 (5 2023), 553–571. doi:10.1002/sce.21790
-
[30]
Tham Thi Nguyen, Duy Cao Nguyen, Hien Thu Nguyen, Hoa Thi Do, Toan Ngo, Anh Bao Gia Pham, Trang Quynh Tran, Linh Phuong Hoang, Hoa Dang, Laurent Boyer, Guillaume Fond, Pascal Auquier, Carl A. Latkin, Roger C.M. Ho, Cyrus S.H. Ho, and Melvyn W.B. Zhang. 2025. Exposure to fake news on social media, coping mechanisms, and mental health impact among Vietnames...
-
[31]
Jon Roozenbeek and Sander van der Linden. 2019. Fake news game confers psy- chological resistance against online misinformation.Palgrave Communications5, 1 (12 2019). doi:10.1057/s41599-019-0279-9
-
[32]
Traberg, and Sander Van Der Linden
Jon Roozenbeek, Cecilie S. Traberg, and Sander Van Der Linden. 2022. Technique- based inoculation against real-world misinformation.Royal Society Open Science 9, 5 (2022). doi:10.1098/rsos.211719
-
[33]
Mohammed Saeed, Nicolas Traub, Maelle Nicolas, Gianluca Demartini, and Paolo Papotti. 2022. Crowdsourced Fact-Checking at Twitter: How Does the Crowd Compare With Experts?. InProceedings of the 31st ACM International Conference on Information & Knowledge Management. 1736–1746
2022
-
[34]
Jon Roozenbeek and Sander van der Linden. 2020. Breaking Harmony Square: A game that “inoculates” against political misinformation.Harvard Kennedy School Misinformation Review1, 8 (2020). doi:10.37016/mr-2020-47
-
[35]
Erin M. Steffes and Lawrence E. Burgee. 2009. Social ties and online word of mouth.Internet Research19, 1 (2009), 42–59. doi:10.1108/10662240910927812
-
[36]
Kate Starbird, Ahmer Arif, and Tom Wilson. 2019. Disinformation as collaborative work: Surfacing the participatory nature of strategic information operations.ACM on HCI3, CSCW (2019), 1–26
2019
-
[37]
Debbe Thompson, Tom Baranowski, Richard Buday, Janice Baranowski, Victo- ria Thompson, Russell Jago, and Melissa Juliano Griffith. 2010. Serious video games for health: How behavioral science guided the development of a seri- ous video game.Simulation and Gaming41, 4 (2010), 587–606. doi:10.1177/ 1046878108328087
2010
-
[38]
2021.Content Matters, Fake Or Not: Media Content Influence On Perceived Intergroup Threat
Nili Steinfeld and Sabina Lissitsa. 2021.Content Matters, Fake Or Not: Media Content Influence On Perceived Intergroup Threat. Paper presented at AoIR 2021: The 22nd Annual Conference of the Association of Internet Researchers. Technical Report. http://spir.aoir.org
2021
-
[39]
Soroush Vosoughi, Deb Roy, and Sinan Aral. 2018. The spread of true and false news online.Science359, 6380 (2018), 1146–
2018
-
[40]
Sander van der Linden. 2021. Some recommendations for doing high-impact research in social psychological science.Asian Journal of Social Psychology24, 1 (3 2021), 37–41. doi:10.1111/ajsp.12463
-
[41]
Kuuku Nyameye Wilson, Benjamin Ghansah, Patricia Ananga, Stephen Opoku Oppong, Winston Kwamina Essibu, and Einstein Kow Essibu. 2025. Exploring the efficacy of computer games as a pedagogical tool for teaching and learning programming: A systematic review.Education and Information Technologies30, 4 (3 2025), 4157–4184. doi:10.1007/s10639-024-13005-2
-
[42]
Soeun Yang, Jae Woo Lee, Hyoung Jee Kim, Minji Kang, Eun Ryung Chong, and Eun mee Kim. 2021. Can an online educational game contribute to developing information literate citizens?Computers and Education161 (2 2021). doi:10.1016/ j.compedu.2020.104057
arXiv 2021
-
[43]
Weiguang Wang, Tianning Zang, and Xiaoyu Zhang. 2025. Temporal-Aware Social Bot Detection with Graph Contrastive Learning. InInternational Conference on Computational Science. Springer, 3–18
2025
-
[1151]
arXiv:https://www.science.org/doi/pdf/10.1126/science.aap9559 doi:10.1126/science.aap9559
-
[2024]
InProceedings of the 19th ACM Asia Conference on Computer and Communications Security
Sok: False information, bots and malicious campaigns: Demystifying elements of social media manipulations. InProceedings of the 19th ACM Asia Conference on Computer and Communications Security. 1784–1800. On Capturing the Narrative: Social Media Manipulation Wargaming for Cyberliteracy Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
2018
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.