Pith. sign in

REVIEW 4 major objections 5 minor 88 references

"GenAI Defaults to Bias!" Gamify AI Literacy Through Reflections on Prompts

T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read A party game can teach adults to recognize GenAI's hidden defaults and biases.

desk verdict A genuinely useful game design wrapped in an overclaimed quantitative evaluation. read the letter →

arxiv 2509.13679 v2 pith:N2XZBTMU submitted 2025-09-17 cs.HC

classification cs.HC
keywords AIliteracygenerativereflectiveplaygame-basedlearningpromptengineeringdemographicbiastext-to-imagegenerationhuman-AIinteraction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that a short social game—where players compete to recreate an image with the fewest prompt words—can make adults more accurately predict how generative image models behave, including their biases and quirks. It claims that this 'reflective play' shifts players' mental models: 73.8% of measured understanding outcomes ended in a calibrated state, and 94% of outcomes stayed the same or improved after one session. The authors contend this matters because most existing AI literacy efforts target children or classrooms, while adults who use GenAI get little structured practice with its failure modes. If the finding holds, a lightweight party game could be a scalable way to improve critical GenAI literacy.

What carries the argument

The load-bearing mechanism is the 'shortest prompt' constraint inside the game's reverse-engineering loop. Players see a target image and must reproduce it with as few words as possible, which pushes them to omit details and thereby expose their assumptions about what the model will fill in. The loop is completed by a voting turn, a reveal-and-discuss turn, and a quick-draw chatbot that lets players test alternative prompts. This combination of hypothesis anchoring, structured contrast, social comparison, and repeated exposure is what the paper claims drives reflective play and learning.

What would settle it

Have two independent coders apply the paper's codebook to the same pre- and post-survey responses; if inter-rater agreement is low (e.g., Cohen's kappa below 0.6), the reported 73.8% calibrated estimate is not stable. A randomized controlled trial with a no-game control group would also settle whether the improvement is caused by the game or by taking the survey twice.

Watch

Extended reading notes

Core claim

The paper's central discovery is that constraining prompts—forcing players to write the shortest possible description of an image—turns the model's hidden defaults into visible disruptions. When a player writes 'a man' and the generator returns a white man with a beard, or writes 'CEO' and gets a middle-aged white man, the mismatch between player assumptions and model output becomes a moment of reflection. After six rounds of this loop, participants gave more specific, conditional accounts of GenAI behavior, distinguishing persistent biases (demographic defaults) from improved capacities (text rendering). The authors interpret this as evidence that structured play, not instruction, can calib

Load-bearing premise

The main quantitative result assumes that one researcher's post-hoc coding of open-ended survey responses reliably distinguishes calibrated from flawed understanding.

Editorial extensions

If this is right

  • Adults can update their mental models of GenAI in a single 60-minute session without explicit instruction.
  • Players come away with practical prompting strategies: specify critical details, gamble on defaults, iterate, and contrast prompts to isolate what matters.
  • The game's design targets underspecification—a stable human trait—rather than specific model bugs, so it may stay relevant as models improve.
  • Reflection quality varies with group composition; diverse groups surface more bias-related hypotheses, while homogeneous groups may miss seeded defaults.
  • A scaled version could generate large datasets of natural-language prompts, outputs, and human similarity judgments for AI research.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the effect is real, similar 'shortest prompt' mechanics could be adapted to text-based LLMs, exposing defaults in writing style, tone, or content rather than images.
  • The 73.8% figure rests on a single coder's post-hoc classification; a replication with two coders and a pre-registered codebook would establish whether the effect is an artifact of coding.
  • The paper's group-composition observation suggests a testable intervention: deliberately mixing novices and experts, or people of different backgrounds, might amplify the learning effect beyond what the paper measured.
  • The game's logged prompts and votes could serve as a human-centered benchmark for perceptual similarity, complementing automated metrics.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents ImaginAItion, a multiplayer web-based party game for adult GenAI literacy. Players compete to reconstruct a target image using the shortest possible prompt, then vote, compare outputs, and use a Quick Draw chatbot to test alternative prompts. The design is grounded in the reflective play framework and targets three reflection goals: underspecification/model defaults, human–AI misalignment, and model unpredictability. The authors report a playtest with 30 adults in 10 trios, analyzing pre–post surveys, gameplay logs, and interviews through a 3×3 codebook of understanding-state transitions. The abstract claims the game "significantly improved players' understanding of GenAI behaviors by 35% in accuracy," while the body reports descriptive percentages, notably that 73.8% of understanding outcomes landed in calibrated states (Section 5.1.1, Figure 2). Qualitative analyses examine category-specific successes and failures, group dynamics, expertise effects, and prompting strategies. The paper also contributes design insights from iterative prototyping and open-sourced game code and study materials.

Significance. If the central quantitative claim were supported, this would be a valuable contribution to scalable adult GenAI literacy, addressing a clear gap in game-based interventions. The paper's strengths include a detailed design rationale tied to the reflective play framework, a reproducible artifact with code and study materials, rich qualitative excerpts showing concrete reflection episodes, and an honest discussion of category-level failures and group-composition effects. However, the quantitative headline is not currently supported: there are no inferential statistics, no control group, no inter-rater reliability check, and the abstract is stronger than the body. The qualitative and design contributions are useful as an exploratory study, but the evidence base does not justify "significantly improved...by 35% in accuracy."

major comments (4)
  1. [Abstract vs. §5.1.1/Figure 2] The abstract's claim of "significantly improved players' understanding of GenAI behaviors by 35% in accuracy" is not supported by the body. No significance test is reported, and the number 35% does not appear in the results section. The transition matrix (Figure 2a) shows 19.6% Aligned + 13.6% Enlightened = 33.2%, which is not the claimed 35%, and these are descriptive percentages from a single coding pass. Please correct the abstract to match the exploratory, descriptive framing used in §6.2, or add the missing inferential analysis.
  2. [§4.3.2 / Figure 2] The main quantitative outcome—73.8% calibrated states and 94% stable-or-better—is produced by one author applying a codebook that was "iteratively and reflexively developed" from the same data it labels. There is no inter-rater reliability check, no blind coding, no pre-registration, and no report of how many responses were discarded as "completely empty or unrelated." Because the categories require subjective judgment (e.g., "calibrated" vs. "flawed," "overly generic," "nuanced"), the proportions are not independently verifiable. Add a second coder and report reliability (e.g., Cohen's κ) on a subset, or explicitly designate these as author-coded illustrative categories and remove the quantitative headline.
  3. [§5.1.1 vs. §6.2] Even with reliable coding, the pre–post design without a control group cannot support a claim of "significant improvement." The self-rated confidence increase (4.2→5.1) and perceived ease of prompting (4.0→4.9) are reported without significance tests. The paper itself acknowledges in §6.2 that it "did not attempt to use pre–post survey design to measure learning gains in a traditional sense." The abstract and Section 5 claims need to be reconciled with this stated limitation.
  4. [§4.2 / Appendix A.2 / Table 2] The pre–post survey questions are one-to-one with the game's prompt categories (demographic, cultural, number/spatial, text, body parts, realism, co-occurrence), and the codebook was developed from the same data. Thus the "calibrated" transitions may partly reflect rehearsal of game-specific examples rather than generalized GenAI literacy. This near-transfer issue is not addressed in the paper's framing; please discuss it explicitly and avoid wording that implies generalized understanding beyond the game context.
minor comments (5)
  1. [§5.1.1 / Figure 2] The percentages 94% and 73.8% are given without denominators and without clarifying whether responses or participants are the unit of analysis. Since each participant answered multiple questions, per-participant aggregation should be reported.
  2. [Figure 4] Per-category percentages are based on very small counts. Please include raw counts or annotate the figure with n per category so the reader can judge the stability of the differences.
  3. [§5.2.1] Participant IDs G0-P3 and G0-P2 appear in the text, but Table 3 only lists groups G1–G10. Check the labeling and correct the references.
  4. [Table 2] The "Ex. image" column is empty in the manuscript text; either remove the column or provide the intended images in the final version.
  5. [General] The manuscript uses the Woodstock ’18 ACM template with placeholder conference metadata and DOI. These must be updated for the intended venue.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the evaluation is an empirical, self-contained pre-post coding analysis; self-citations are motivational and not load-bearing.

full rationale

Walking the claimed derivation chain, the paper does not derive a formal result from an input. ImaginAItion is designed from reflective-play literature and evaluated through pre/post surveys and gameplay logs. The central quantitative claims — 73.8% of understanding outcomes landing in calibrated states, 94% stable-or-better, and the self-rated 5.7/7 helpfulness — are descriptive statistics over a coded transition matrix (Section 5.1.1, Figure 2), not quantities that are equal by construction to the game's prompts or to any fitted parameter. The codebook in Table 4 was developed from the same pre/post data and applied by one author (Section 4.3.2), which is a genuine validity threat; and the abstract's 'significantly improved' and '35% in accuracy' wording outruns the body's cautious framing (Section 6.2: 'we did not attempt to use pre–post survey design to measure learning gains in a traditional sense, or detect significant effects'). However, these are evidentiary weaknesses, not circularity: the outcome numbers depend on participants' actual open-ended responses, and no prediction is statistically forced by a fitted input. The self-citations ([9], [48], [49]) support the motivational framing around underspecification and prompting, but they are not the source of the evaluation results nor do they forbid alternative interpretations. No equation or definition reduces the reported outcomes to the paper's own inputs, so no circular step can be exhibited.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No numeric parameters are fitted to data in the paper. Design choices such as the one-point longest-prompt penalty, six rounds, and category pools are hand-chosen game parameters, not fitted to optimize the measured outcome. The Quick Draw chatbot and scoring system are interface features, not new scientific entities.

assumptions (4)
  • domain assumption The Reflective Play framework (Miller et al., 2024) provides a valid set of mechanisms for evoking transformative reflection in this game context.
    Invoked in Section 3.2 when mapping game mechanics to Disruptions, Slowdowns, Questioning, Revisiting, and Enhancers.
  • domain assumption gpt-image-1's outputs for the curated prompts exhibit the persistent biases and limitations listed in Table 2 consistently enough for players to observe them during a given session.
    Section 3.4 notes the model improved significantly even within the study period; the game's disruptive moments depend on the model behaving as typified.
  • domain assumption Open-ended self-report survey responses before and after play validly reflect participants' understanding of GenAI behaviors.
    Section 4.1 uses pre/post surveys as the primary instrument; response effort and social desirability can bias self-report.
  • ad hoc to paper The 3x3 codebook of reflection outcomes is exhaustive and codable without inter-rater reliability.
    Section 4.3.2 describes a single coder developing the codebook from the data; no independent validation is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of "GenAI Defaults to Bias!" Gamify AI Literacy Through Reflections on Prompts." pith.science (2026). https://pith.science/paper/N2XZBTMU

@misc{pith2026250913679,
  author       = {Pith},
  title        = {Pith review of: "GenAI Defaults to Bias!" Gamify AI Literacy Through Reflections on Prompts},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/N2XZBTMU}},
  note         = {Machine review of arXiv:2509.13679}
}
read the original abstract

As Generative AI (GenAI) becomes widespread, it is increasingly important for the public to understand the model's behaviors and biases. However, existing AI literacy efforts miss opportunities to engage the general public to reflect on enduring GenAI bias and behaviors (e.g., how GenAI defaults to its internal bias in response to ambiguous or challenging prompts). In this work, we introduce ImaginAItion, a multiplayer game to help adults better reflect on GenAI bias and understand GenAI behaviors. ImaginAItion is grounded in reflective play to surface GenAI limitations by encouraging players to manipulate prompt specificity (e.g., an underspecified prompt "CEO" defaults to a white man). From ten sessions (n=30), we find that the game significantly improved players' understanding of GenAI behaviors by 35% in accuracy. Qualitative analysis showed how game mechanisms supported player reflections, including on prompting strategies to mitigate GenAI bias. Our work demonstrates a viable pathway to scale GenAI literacy through playful, social interventions resilient to rapidly evolving technologies.

Figures

Figures reproduced from arXiv: 2509.13679 by the authors.

Figure 1
Figure 1. An example sequence from ImaginAItion gameplay. The real game could be played at: ImaginAItion game online. A video of a game round is available here. behaviors [2, 29, 51, 81]. As GenAI evolves rapidly, such interven￾tions risk becoming outdated, offering insights that no longer hold. Designing games that expose both persistent and evolving model behaviors remains an open challenge — it demands systems that foster … view at source ↗
Figure 2
Figure 2. Pre- to post-game reflection outcomes in participants’ flawed/no/calibrated understandings of GenAI behaviors, [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. G7-P2’s quick-draw experiment to reduce the prompt length using just “mother”. The reproduced image is considered more similar to the original image than her prompt “portrait of mother holding her baby” [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Reflection outcomes distribution across prompt [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]
Figure 5
Figure 5. Figure 5: Reflection outcomes distribution by expertise. [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]
Figure 6
Figure 6. Figure 6: How to Play Description. Players can view the instructions on how to play the game with an overview of the prompting, [PITH_FULL_IMAGE:figures/full_fig_p020_6.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

88 extracted references · 6 canonical work pages

  1. [1]

    Jihyun Janice Ahn and Wenpeng Yin. 2025. Prompt-reverse inconsistency: Llm self-inconsistency beyond generative randomness and prompt paraphrasing. arXiv preprint arXiv:2504.01282(2025)

  2. [2]

    Safinah Ali, Vishesh Kumar, and Cynthia Breazeal. 2023. AI audit: a card game to reflect on everyday AI systems. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 15981–15989

  3. [3]

    Haesol Bae and Aras Bozkurt. 2024. The Untold Story of Training Students with Generative AI: Are We Preparing Students for True Learning or Just Personaliza- tion?Online Learning28, 3 (2024)

  4. [4]

    Maalvika Bhat and Duri Long. 2024. Designing Interactive Explainable AI Tools for Algorithmic Literacy and Transparency. InProceedings of the 2024 ACM Designing Interactive Systems Conference(Copenhagen, Denmark)(DIS ’24). Association for Computing Machinery, New York, NY, USA, 939–957. https://doi.org/10.1145/3643834.3660722

  5. [5]

    Ali Borji. 2023. Qualitative failures of image generation models and their applica- tion in detecting deepfakes.arXiv [cs.CV](March 2023). arXiv:2304.06470 [cs.CV] http://arxiv.org/abs/2304.06470

  6. [6]

    Aras Bozkurt. 2024. Why generative AI literacy, why now and why it matters in the educational landscape?: Kings, queens and GenAI dragons. , 283–290 pages

  7. [7]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology3, 2 (2006), 77–101

  8. [8]

    Lorena Casal-Otero, Alejandro Catala, Carmen Fernández-Morante, Maria Taboada, Beatriz Cebreiro, and Senén Barro. 2023. AI literacy in K-12: a system- atic literature review.International Journal of STEM Education10, 1 (April 2023), 1–17. https://doi.org/10.1186/s40594-023-00418-7

Show all 88 references
  1. [9]

    Yang Chenyang, Shi Yike, Ma Qianou, Michael Xieyang Liu, Kästner Chris- tian, and Wu Tongshuang. 2025. What prompts don’t say: Understanding and managing underspecification in LLM prompts.arXiv [cs.CL](May 2025). arXiv:2505.13360 [cs.CL] http://arxiv.org/abs/2505.13360

  2. [10]

    Judeth Oden Choi, Jodi Forlizzi, Michael Christel, Rachel Moeller, Mackenzie Bates, and Jessica Hammer. 2016. Playtesting with a Purpose. InProceedings of the 2016 Annual Symposium on Computer-Human Interaction in Play. ACM, New York, NY, USA. https://doi.org/10.1145/2967934.2968103

  3. [11]

    Wikipedia contributors. [n. d.]. Pictionary. https://en.wikipedia.org/wiki/ Pictionary. Accessed: December 8, 2024

  4. [12]

    Andrew Cox. 2024. Algorithmic literacy, AI literacy and responsible generative AI literacy.Journal of Web Librarianship(2024), 1–18

  5. [13]

    2018.A process tool for the development of Transformational games

    Sabrina Culyba. 2018.A process tool for the development of Transformational games. CMU ETS Press. http://pstorage-cmu-348901238291901.s3.amazonaws. com/13117568/TransformationalFramework_2018.pdf

  6. [14]

    Simone Daniotti, Johannes Wachs, Xiangnan Feng, and Frank Neffke. 2025. Who is using AI to code? Global diffusion and impact of generative AI.arXiv preprint arXiv:2506.08945(2025)

  7. [15]

    Wesley Hanwen Deng, Wang Claire, Howard Ziyu Han, Jason I Hong, Ken- neth Holstein, and Motahhare Eslami. 2025. WeAudit: Scaffolding user audi- tors and AI practitioners in auditing Generative AI.arXiv [cs.HC](Jan. 2025). arXiv:2501.01397 [cs.HC] http://arxiv.org/abs/2501.01397

  8. [16]

    Chiara Di Lodovico, Federico Torrielli, Luigi Di Caro, and Amon Rapp. 2025. How do people develop folk theories of generative AI text-to-image models? A qualitative study on how people strive to explain and make sense of GenAI.Int. J. Hum. Comput. Interact.(April 2025), 1–25. ...

  9. [17]

    Leyla Dogruel, Philipp Masur, and Sven Joeckel. 2022. Development and valida- tion of an algorithm literacy scale for internet users.Communication Methods and Measures16, 2 (2022), 115–133

  10. [18]

    Stefania Druga, Nancy Otero, and Amy J. Ko. 2022. The Landscape of Teaching Resources for AI Education. InProceedings of the 27th ACM Conference on on Innovation and Technology in Computer Science Education Vol. 1(Dublin, Ireland) (ITiCSE ’22). Association for Computing Machin...

  11. [19]

    Xiaoxue Du, Zhichun Liu, Xi Wang, and Qiping Tang. 2024. Fostering AI Lit- eracy through Interactive Game-Based Learning: A Case Study on Enhancing Algorithmic Thinking Skills. InSociety for Information Technology & Teacher Edu- cation International Conference. Association for...

  12. [20]

    Jonathan St BT Evans and Keith E Stanovich. 2013. Dual-process theories of higher cognition: Advancing the debate.Perspectives on psychological science8, 3 (2013), 223–241

  13. [21]

    Jennifer Fereday and Eimear Muir-Cochrane. 2006. Demonstrating rigor using thematic analysis: A hybrid approach of inductive and deductive coding and theme development.International journal of qualitative methods5, 1 (2006), 80–92

  14. [22]

    Jackbox Games. [n. d.]. Drawful. https://steamdb.info/app/442070/charts/. https: //www.jackboxgames.com/games/drawful-2 Accessed: December 8, 2024

  15. [23]

    Raghu Garud. 1997. Know-how, know-why, and know-what.Advances in strategic management14 (1997), 81–101

  16. [24]

    embedded

    Mary Flanagan Geoff Kaufman. 2015. A psychologically “embedded” approach to designing games for prosocial causes.Cyberpsychology: Journal of Psychosocial Research on Cyberspace(2015)

  17. [25]

    Ronald N Giere and Barton Moffatt. 2003. Distributed Cognition:: Where the Cognitive and the Social Merge.Soc. Stud. Sci.33, 2 (April 2003), 301–310. https: //doi.org/10.1177/03063127030332017

  18. [26]

    Google. [n. d.]. Quick, Draw! https://quickdraw.withgoogle.com/. Accessed: December 8, 2024

  19. [27]

    Anuj Gupta, Yasser Atef, Anna Mills, and Maha Bali. 2024. Assistant, parrot, or colonizing loudspeaker? ChatGPT metaphors for developing critical AI Literacies. Open Praxis16, 1 (2024), 37–53

  20. [28]

    Daniel Hershcovich, Stella Frank, Heather Lent, Miryam De Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, et al. 2022. Challenges and strategies in cross-cultural NLP.arXiv preprint arXiv:2203.10020(2022)

  21. [29]

    Olli Hilke, Nicolas Pope, Juho Kahila, Henriikka Vartiainen, Teemu Roos, Tuomo Parkki, and Matti Tedre. 2025. Breakable Machine: A K-12 Classroom Game for Transformative AI Literacy Through Spoofing and eXplainable AI (XAI).arXiv preprint arXiv:2508.14201(2025)

  22. [30]

    Ting-Chia Hsu and Tai-Ping Hsu. 2025. Teaching AI with games: the impact of generative AI drawing on computational thinking skills.Education and Informa- tion Technologies(2025), 1–20

  23. [31]

    Viegas, and Martin Wattenberg

    Minsuk Kahng, Nikhil Thorat, Duen Horng Polo Chau, Fernanda B. Viegas, and Martin Wattenberg. 2019. GAN Lab: Understanding Complex Deep Gen- erative Models using Interactive Visual Experimentation.IEEE Transactions on Visualization and Computer Graphics25, 1 (Jan. 2019), 310–3...

  24. [32]

    Dongyeop Kang and Eduard Hovy. 2021. Style is NOT a single variable: Case stud- ies for cross-stylistic language understanding. InProceedings of the 59th annual meeting of the association for computational linguistics and the 11th international joint conference on natural lang...

  25. [33]

    Maria Kasinidou. 2024. Development of Personalised Educational Tools for AI Literacy Using Participatory Design. InAdjunct Proceedings of the 32nd ACM Conference on User Modeling, Adaptation and Personalization(Cagliari, Italy) (UMAP Adjunct ’24). Association for Computing Mac...

  26. [34]

    Embedded

    Geoffrey Kaufman and Mary Flanagan. 2015. A Psychologically “Embedded” Approach to Designing Games for Prosocial Causes.Cyberpsychology, Behavior, and Social Networking18, 4 (2015), 241–248. https://doi.org/10.1089/cyber.2014. 0513

  27. [35]

    Geoff Kaufman, Mary Flanagan, and Gili Freedman. 2019. Not just for girls: Encouraging cross-gender role play and reducing gender stereotypes with a strategy game. InProceedings of the Annual Symposium on Computer-Human Interaction in Play. ACM, New York, NY, USA. https://doi....

  28. [36]

    Newman, and Allison Woodruff

    Patrick Gage Kelley, Yongwei Yang, Courtney Heldreth, Christopher Moessner, Aaron Sedley, Andreas Kramm, David T. Newman, and Allison Woodruff. 2021. Exciting, Useful, Worrying, Futuristic: Public Perception of Artificial Intelligence in 8 Countries. InProceedings of the 2021 ...

  29. [37]

    Kenneth R Koedinger, Paulo F Carvalho, Ran Liu, and Elizabeth A McLaughlin

  30. [38]

    Valentin Kuleto, Milena Ilić, Mihail Dumangiu, Marko Ranković, Oliva MD Mar- tins, Dan Păun, and Larisa Mihoreanu. 2021. Exploring opportunities and chal- lenges of artificial intelligence and machine learning in higher education institu- tions.Sustainability13, 18 (2021), 10424

  31. [39]

    Matthias Carl Laupichler, Alexandra Aster, Jana Schirch, and Tobias Raupach

  32. [40]

    Mosh Levy, Zohar Elyoseph, and Yoav Goldberg. 2025. Humans Perceive Wrong Narratives from AI Reasoning Texts.arXiv preprint arXiv:2508.16599(2025)

  33. [41]

    From Unseen Needs to Classroom Solutions

    Hanqi Li, Ruiwei Xiao, Hsuan Nieu, Ying-Jui Tseng, and Guanze Liao. 2025. “From Unseen Needs to Classroom Solutions”: Exploring AI Literacy Challenges & Opportunities with Project-Based Learning Toolkit in K-12 Education. In Proceedings of the AAAI Conference on Artificial Int...

  34. [42]

    Xiaomin Li, Zhou Yu, Zhiwei Zhang, Xupeng Chen, Ziji Zhang, Yingying Zhuang, Narayanan Sadagopan, and Anurag Beniwal. 2025. When thinking fails: The pit- falls of reasoning for instruction-following in llms.arXiv preprint arXiv:2505.11423 Conference’17, July 2017, Washington, ...

  35. [43]

    Andreas Lieberoth. 2015. Shallow gamification: Testing psychological effects of framing an activity as a game.Games and Culture10, 3 (2015), 229–248

  36. [44]

    Bill Yuchen Lin, Frank F Xu, Kenny Zhu, and Seung-won Hwang. 2018. Mining cross-cultural differences and similarities in social media. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 709–719

  37. [45]

    Michael Xieyang Liu, Tongshuang Wu, Tianying Chen, Franklin Mingzhe Li, Aniket Kittur, and Brad A Myers. 2024. Selenite: Scaffolding Online Sensemak- ing with Comprehensive Overviews Elicited from Large Language Models. In Proceedings of the 2024 CHI Conference on Human Factor...

  38. [46]

    Vivian Liu and Lydia B Chilton. 2022. Design guidelines for prompt engi- neering text-to-image generative models. InCHI Conference on Human Fac- tors in Computing Systems (CHI ’22). ACM, New York, NY, USA, 1–23. https: //doi.org/10.1145/3491102.3501825

  39. [47]

    Duri Long and Brian Magerko. 2020. What is AI literacy? Competencies and design considerations. InProceedings of the 2020 CHI conference on human factors in computing systems. 1–16

  40. [48]

    Qianou Ma, Anika Jain, Jini Kim, Megan Chai, and Geoff Kaufman. 2025. Imag- inAItion: Promoting generative AI literacy through game-based learning. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–9...

  41. [49]

    Qianou Ma, Weirui Peng, Chenyang Yang, Hua Shen, Ken Koedinger, and Tong- shuang Wu. 2025. What should we engineer in prompts? Training humans in requirement-driven LLM use.ACM Trans. Comput. Hum. Interact.32, 4 (Aug. 2025), 1–27. https://doi.org/10.1145/3731756

  42. [50]

    Qianou Ma, Hua Shen, Kenneth Koedinger, and Sherry Tongshuang Wu. 2024. How to teach programming in the AI era? Using LLMs as a teachable agent for debugging. InProceedings of International Conference on Artificial Intelligence in Education. Springer Nature Switzerland, Cham, ...

  43. [51]

    Pragati Maheshwary, Aditi Haiman, Emma Brown, Aarnav Sangekar, and Canwen Wang. 2025. Case 429: A Murder Mystery Game for Teaching Bias in Generative AI. (2025)

  44. [53]

    A I Meta. 2024. Create, edit, and animate images with Meta AI. https://www. meta.ai/ai-image-generator. Accessed: 2024-11-8

  45. [54]

    Josh Aaron Miller, Kutub Gandhi, Matthew Alexander Whitby, Mehmet Kosa, Seth Cooper, Elisa D Mekler, and Ioanna Iacovides. 2024. A design framework for reflective play. InProceedings of the CHI Conference on Human Factors in Computing Systems, Vol. 91. ACM, New York, NY, USA, ...

  46. [55]

    Katelyn Morrison, Mayank Jain, Jessica Hammer, and Adam Perer. 2023. Eye into AI: Evaluating the Interpretability of Explainable AI Techniques through a Game with a Purpose.Proc. ACM Hum.-Comput. Interact.7, CSCW2, Article 273 (Oct. 2023), 22 pages. https://doi.org/10.1145/3610064

  47. [56]

    Davy Tsz Kit Ng, Jac Ka Lok Leung, Samuel Kai Wah Chu, and Maggie Shen Qiao. 2021. Conceptualizing AI literacy: An exploratory review.Computers and Education: Artificial Intelligence2, 100041 (Jan. 2021), 100041. https://doi.org/10. 1016/j.caeai.2021.100041

  48. [57]

    Davy Tsz Kit Ng, Chen Xinyu, Jac Ka Lok Leung, and Samuel Kai Wah Chu

  49. [58]

    The Op. [n. d.]. Telestrations. https://theop.games/collections/telestrations. Ac- cessed: December 8, 2024

  50. [59]

    OpenAI. 2024. DALL-E 3 Image generation. https://platform.openai.com/docs/ guides/image-generation?image-generation-model=dall-e-3. Accessed: 2024-12- 8

  51. [60]

    OpenAI. 2025. Introducing 4o Image Generation. https://openai.com/index/ introducing-4o-image-generation/. Accessed: 2025-9-8

  52. [61]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback.Advances in neural information processing systems35 (...

  53. [62]

    Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2019. Language models as knowledge bases?arXiv preprint arXiv:1909.01066(2019)

  54. [63]

    Yiwei Qin, Kaiqiang Song, Yebowen Hu, Wenlin Yao, Sangwoo Cho, Xiaoyang Wang, Xuansheng Wu, Fei Liu, Pengfei Liu, and Dong Yu. 2024. InFoBench: Evaluating Instruction Following Ability in Large Language Models. InFindings of the Association for Computational Linguistics: ACL 2...

  55. [64]

    Amon Rapp, Chiara Di Lodovico, and Luigi Di Caro. 2025. How do people react to ChatGPT’s unpredictable behavior? Anthropomorphism, uncanniness, and fear of AI: A qualitative study on individuals’ perceptions and understandings of LLMs’ nonsensical hallucinations.Int. J. Hum. C...

  56. [65]

    Paulius Rauba, Qiyao Wei, and Mihaela van der Schaar. 2024. Quantifying perturbation impacts for large language models.arXiv preprint arXiv:2412.00868 (2024)

  57. [66]

    Danielle Reynolds and Scott Brady. [n. d.]. CautionSigns. https://cautionsigns. app/. Accessed: December 8, 2024

  58. [67]

    Benjamin Sanchez-Lengeling, Emily Reif, Adam Pearce, and Alexander B Wiltschko. 2021. A gentle introduction to graph neural networks.Distill6, 9 (2021), e33

  59. [68]

    Camilo Chacón Sartori. 2025. Architectures of Error: A Philosophical Inquiry into AI and Human Code Generation.arXiv preprint arXiv:2505.19353(2025)

  60. [69]

    Shreya Shankar, JD Zamfirescu-Pereira, Björn Hartmann, Aditya Parameswaran, and Ian Arawjo. 2024. Who validates the validators? aligning llm-assisted evalu- ation of llm outputs with human preferences. InProceedings of the 37th Annual ACM Symposium on User Interface Software a...

  61. [70]

    Yijia Shao, Humishka Zope, Yucheng Jiang, Jiaxin Pei, David Nguyen, Erik Bryn- jolfsson, and Diyi Yang. 2025. Future of Work with AI Agents: Auditing Au- tomation and Augmentation Potential across the US Workforce.arXiv preprint arXiv:2506.06576(2025)

  62. [71]

    Mustafa Shukor, Enrico Fini, Victor Guilherme Turrisi da Costa, Matthieu Cord, Joshua Susskind, and Alaaeldin El-Nouby. 2025. Scaling laws for native multi- modal models.arXiv preprint arXiv:2504.07951(2025)

  63. [72]

    Hannu Simonen, Atte Kiviniemi, and Jonas Oppenlaender. 2025. An explo- ration of default images in text-to-image generation.arXiv [cs.HC](July 2025). arXiv:2505.09166 [cs.HC] http://arxiv.org/abs/2505.09166

  64. [73]

    Anselm Strauss and Juliet Corbin. 1998. Basics of qualitative research techniques. (1998)

  65. [74]

    Piiastiina Tikka, Miia Laitinen, Iikka Manninen, and Harri Oinas-Kukkonen

  66. [75]

    Khoi Trinh, Scott Seidenberger, Raveen Wijewickrama, Murtuza Jadliwala, and Anindya Maiti. 2025. A picture is worth a thousand prompts? Efficacy of iterative human-driven prompt refinement in image regeneration tasks. InProceedings of the Thirty-Third International Joint Confe...

  67. [76]

    Jessica Vandenberg, Wookhee Min, Veronica Cateté, Danielle Boulden, and Brad- ford Mott. 2022. Promoting AI education for rural middle grades students with digital game design. InProceedings of the 54th ACM Technical Symposium on Computer Science Education V. 2. 1388–1388

  68. [77]

    Jennifer Villareale, Gabriele Cimolino, and Daniel Gomme. 2023. Playing with dezgo: Adapting human-AI interaction to the context of play. InProceedings of the 18th International Conference on the Foundations of Digital Games. ACM, New York, NY, USA. https://doi.org/10.1145/358...

  69. [78]

    I want to see how smart this AI really is

    Jennifer Villareale, Casper Harteveld, and Jichen Zhu. 2022. “I want to see how smart this AI really is”: Player mental model development of an adversarial AI player.Proc. ACM Hum. Comput. Interact.6, CHI PLAY (Oct. 2022), 1–26. https://doi.org/10.1145/3549482

  70. [79]

    Luis Von Ahn and Laura Dabbish. 2008. Designing games with a purpose.Com- mun. ACM51, 8 (2008), 58–67

  71. [80]

    Junling Wang, Anna Rutkiewicz, April Yi Wang, and Mrinmaya Sachan. 2025. Generating pedagogically meaningful visuals for math word problems: A new benchmark and analysis of text-to-Image models.arXiv [cs.CL](June 2025). arXiv:2506.03735 [cs.CL] http://arxiv.org/abs/2506.03735

  72. [81]

    Ning Wang, Eric Greenwald, Ryan Montgomery, and Maxyn Leitner. 2022. ARIN- 561: An educational game for learning artificial intelligence for high-school stu- dents. InInternational Conference on Artificial Intelligence in Education. Springer, 528–531

  73. [82]

    Zijie J Wang, Robert Turko, Omar Shaikh, Haekyu Park, Nilaksh Das, Fred Hohman, Minsuk Kahng, and Duen Horng Polo Chau. 2020. CNN explainer: learning convolutional neural networks with interactive visualization.IEEE Transactions on Visualization and Computer Graphics27, 2 (202...

  74. [83]

    Jiancong Xiao, Ziniu Li, Xingyu Xie, Emily Getzen, Cong Fang, Qi Long, and Weijie J Su. 2024. On the algorithmic bias of aligning large language models with RLHF: Preference collapse and matching regularization.arXiv preprint arXiv:2405.16455(2024)

  75. [84]

    Michelle Yavorskiy and Lydia Kim. 2024. MeadowMinds: An AI Literacy Game. (2024)

  76. [85]

    I don’t know

    Marvin Zammit, Iro Voulgari, Antonios Liapis, and Georgios N Yannakakis. 2021. The road to AI literacy education: from pedagogical needs to tangible game design. Academic Conferences International. ImaginAItion: GenAI Literacy Game Conference’17, July 2017, Washington, DC, USA...

  77. [2018]

    InPersuasive Technology

    Reflection through gaming: Reinforcing health message response through gamified rehearsal. InPersuasive Technology. Springer International Publishing, Cham, 200–212. https://doi.org/10.1007/978-3-319-78978-1_17

  78. [2022]

    2022), 100101

    Artificial intelligence literacy in higher and adult education: A scoping literature review.Computers and Education: Artificial Intelligence3, 100101 (Jan. 2022), 100101. https://doi.org/10.1016/j.caeai.2022.100101

  79. [2023]

    An astonishing regularity in student learning rate.Proc. Natl. Acad. Sci. U. S. A.120, 13 (March 2023), e2221311120. https://doi.org/10.1073/pnas.2221311120

  80. [2024]

    Fostering students’ AI literacy development through educational games: AI knowledge, affective and cognitive engagement.Journal of Computer Assisted Learning(2024)

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.