Pith. sign in

REVIEW 3 major objections 5 minor 101 references

"AI just keeps guessing": Using ARC Puzzles to Help Children Identify Reasoning Errors in Generative AI

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A game that puts a child's puzzle solution next to an AI's grid lets six-year-olds see where generative AI's reasoning goes wrong.

desk verdict Genuinely new ARC-based AI-literacy tool with solid descriptive findings; the unvalidated parser pipeline and over-scoped RQ2 deserve referee attention. read the letter →

arxiv 2505.16034 v1 pith:76MKGKCX submitted 2025-05-21 cs.HC cs.CY

classification cs.HCcs.CY
keywords AIliteracyGenerativeChildrenAbstractionandReasoningCorpusARCpuzzlesParticipatorydesignMultimedialearningErrordetection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Generative AI's fluent, confident text can make its mistakes hard to see, and children are especially prone to trusting it. This paper argues that a better starting point is a visual task where the correct answer is easy for a child to verify and hard for the AI to produce. The authors built AI Puzzlers, a browser game using Abstraction and Reasoning Corpus (ARC) grid puzzles in which a child solves a puzzle, asks GPT-4o to solve the same puzzle, and compares the AI's grid with their own, optionally reading the AI's step-by-step explanation. Across two participatory design sessions with 21 children aged 6 to 11, they report that children quickly spotted incorrect AI solutions by eye, cross-examined the AI's explanations against its visual output, caught contradictions in its reasoning, and, in a hint-giving mode, iteratively debugged the AI's behavior. Their conclusion is that visual-verbal comparison reduces cognitive load and lets children who are not yet fluent readers recognize that generative AI "just keeps guessing."

What carries the argument

The central object is AI Puzzlers, a browser game built on the Abstraction and Reasoning Corpus (ARC), a set of visual grid puzzles in which solvers infer a transformation rule from example input-output pairs and apply it to a test grid. The game's mechanism is the side-by-side comparison: the child solves a puzzle, clicks "Ask AI to Solve" to have GPT-4o produce a grid, and sees the AI's attempt next to the correct solution, with an optional "Ask AI to Explain" text. The design rationale is the dual-channel assumption of the Cognitive Theory of Multimedia Learning, which holds that distributing information across visual and verbal channels prevents cognitive overload. Under the hood, a parser converts the puzzle into a textual grid for the model and another parser renders the model's textual output back into a visual grid, so every repeated request also shows the AI's output varying from attempt to attempt.

What would settle it

A control study in which children get the same ARC tasks and the same AI outputs presented only as text — grid coordinates or verbal descriptions, with no visual grid — would settle the mechanism: if error detection and critique are just as fast and frequent, the visual comparison is not the active ingredient. A transfer probe would add a second check: compare whether children who played AI Puzzlers question a ChatGPT answer on an unrelated topic more than a matched control group would.

Watch

Extended reading notes

Core claim

The paper's central claim is that presenting generative AI's reasoning as a visually comparable artifact — the AI's grid solution placed next to the child's own correct solution and the AI's textual explanation — turns error detection into a perceptual task instead of a domain-knowledge task. Children aged 6 to 11, including those not yet fluent readers, quickly noticed when the AI's grid was wrong, often reacting with surprise and amusement when puzzles they considered easy stumped the AI. They went beyond noticing: they found contradictions between the AI's explanation and its output, described the AI as "scientific" but opaque, and concluded from its changing answers that it was guessing rather than reasoning. In the hint-giving mode, they refined their instructions step by step, demonstrating an emergent understanding that the AI needs explicit, unambiguous guidance. The authors take these observations as evidence that a dual visual-verbal presentation, grounded in the Cognitive Theory of Multimedia Learning, reduces cognitive overload and scaffolds critical evaluation of AI outputs.

Load-bearing premise

The load-bearing premise is that children's observed error detection and critical questioning came from the visual-verbal design of AI Puzzlers, not from the puzzles being easy, the AI's failures being unusually obvious, or the children's prior experience in participatory design making them unusually comfortable speaking up.

Editorial extensions

If this is right

  • Children who start with high expectations that AI will solve easy puzzles revise that belief when they see the AI's incorrect grid, so comparison-based exercises can address overtrust directly.
  • AI literacy tools for children should show outputs in an inspectable visual form paired with a brief verbal explanation, rather than relying on polished text alone.
  • Giving children a hint-typing "assist" mode turns them from passive consumers into debuggers who refine instructions and test hypotheses about how the AI processes information.
  • Watching the AI produce different answers on repeated attempts gives children concrete evidence that the AI is guessing rather than reasoning, which supports their understanding of its limitations.
  • The design's success depends on tasks where the correct answer is easy for children to verify visually; extending the approach means finding new tasks that preserve that asymmetry as AI improves.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An extension the paper leaves implicit is that the same side-by-side comparison could transfer to any output a child can judge faster than the AI can produce, such as diagrams, maps, code, or math solutions; the mechanism would fail wherever the child lacks the knowledge to act as referee.
  • A direct test of the design's mechanism would vary only the output modality: one group sees the AI's answer as a colored grid, another sees the same answer as text coordinates; equal detection rates would mean the visual channel is not the active ingredient.
  • The study is observational, so a transfer experiment with a pre/post measure of whether children question AI answers in an unrelated free-form chat would convert the claimed AI-literacy benefit into a measurable outcome.
  • Because newer AI models are already improving on ARC, the specific corpus may eventually stop producing the competence gap the game depends on; the durable contribution would then be the general pattern of letting the learner judge tasks the AI still fails, which would need refreshed benchmarks.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents AI Puzzlers, a web-based educational game built on Abstraction and Reasoning Corpus (ARC) puzzles, in which children solve visual puzzles, ask GPT-4o for solutions and explanations, and compare the AI's visual outputs with their own. The authors report two participatory design sessions with 21 children aged 6–11, using Cooperative Inquiry. The qualitative findings describe how children reacted with surprise to AI errors, iteratively debugged AI by refining textual hints, identified inconsistencies between AI explanations and AI-generated grids, and articulated differences between human and AI problem-solving. The paper claims that the visual, side-by-side design helped children, including younger non-fluent readers, detect and analyze errors in generative AI reasoning, and it frames the design through Mayer and Moreno's Cognitive Theory of Multimedia Learning.

Significance. If the findings hold, the paper makes a useful design contribution to child-facing AI literacy: it offers a low-barrier, game-based method for making generative AI failures visible to children and provides rich vignettes of how children articulate AI's limitations. The open-source codebase and the transparent description of the participatory design process are strengths, as is the authors' willingness to acknowledge limits on transferability and generalizability. However, the central contribution is descriptive rather than causal: the design does not isolate the visual modality, does not compare against text-only AI output, and does not validate that the displayed AI grids faithfully represent GPT-4o's raw output. These gaps matter for the paper's stronger interpretive claims, even though the basic observation that children noticed visual mismatches is credible and well illustrated.

major comments (3)
  1. [Section 3.2.2 and Findings (Section 5)] The paper does not validate the grid_parser()/response_parser() round-trip that converts ARC puzzles into text for GPT-4o and renders the model's textual output back into visual grids. If a parser default, color-token mismatch, or ambiguous serialization produced the child-visible outputs, then the 'AI errors' children detected may be system artifacts rather than GPT-4o reasoning failures. Because the central claim is that children identify and analyze errors in genAI reasoning, the authors should provide evidence that the displayed grids correspond to the model's raw output—for example, by reporting parser audits, showing raw model responses alongside rendered grids for the 12 puzzles, or documenting failure cases and how they were handled.
  2. [Section 1 (RQ2) and Sections 4.2/6.1] RQ2 asks how presenting information across visual and textual modalities influences children's ability to critically assess AI outputs, and the Discussion interprets the findings through CTML, but the study has no manipulation or comparison condition isolates the visual-verbal presentation. There is no text-only condition, no condition without the side-by-side comparison, and no baseline task controlling for puzzle ease or the obviousness of GPT-4o's errors. The observed detection and reflection are therefore equally consistent with the puzzles being easy for children (the playtest in Section 3.2 reports M = 2.38) and with the AI's errors being unusually salient. The authors should either add a comparison condition in future work or re-frame the contribution as a design exploration whose modality-related claims are hypotheses, not findings.
  3. [Abstract and Section 1] The abstract and introduction state that 'even younger children, who were not yet fluent readers, quickly detected inconsistencies in AI-generated solutions,' but no measure of reading fluency is reported anywhere in the paper. The vignettes do not identify which children were non-fluent readers, and the data include no reading-assessment or age-disaggregated analysis. This claim currently exceeds the evidence; it should either be removed, softened to 'including the youngest participants,' or supported with explicit evidence about the children's reading abilities and their error-detection behavior.
minor comments (5)
  1. [Section 3.2] The one-sample t-test is reported with t(103) = −6.48 for N = 106 playtest children, but a one-sample test should have 105 degrees of freedom; please verify the reported statistic and clarify whether N refers to children or completed puzzle ratings.
  2. [Sections 3.2.3, 4.2.2, and 5.1.2] The interaction mode where children provide hints is called 'Assist Mode' in the system description and findings, but 'Human-AI mode' in Session 2; please use consistent terminology throughout.
  3. [Figure 15] The narrative in Section 5.2.2 describes three AI attempts, but the figure caption lists four attempts; the mismatch between text and figure should be corrected.
  4. [Section 3.2.2] The authors do not specify the GPT-4o API parameters (e.g., temperature, max tokens) or the exact date of data collection; given the paper's emphasis on variability and future replication, adding this information would be helpful.
  5. [Section 7] The limitations section is candid about the single-site, co-design-experienced sample and the lack of transfer measures, but it does not acknowledge the absence of a modality comparison or the parser-validation issue; adding these to the limitations would make the scope of the claims clearer.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation found: the paper's claims are empirical observations from co-design sessions, not predictions derived from fitted inputs or self-citation.

full rationale

The paper makes no formal derivation; its central claims are qualitative empirical findings from two participatory design sessions with 21 children. The design choices invoke external theories (Mayer and Moreno's CTML, Wittrock's generative learning, Cooperative Inquiry) and an external benchmark (ARC), and the AI behavior is generated live by GPT-4o and observed rather than fitted. No equation or quantitative model links a fitted parameter to a reported 'prediction,' and no result is shown to be equivalent to its input by construction. The self-citations (e.g., references [19], [20], and [92]) provide prior context, methodological framing, or recruitment details and do not carry the main claim that children detected inconsistencies; that claim rests on video-recorded observations, analytic memos, and direct quotes. The paper explicitly limits its conclusions, stating that findings should be understood as 'formative theoretical generalizations rather than statistical generalization,' and it acknowledges that transfer to other genAI interactions was not assessed. The unvalidated grid_parser and response_parser pipeline is a possible validity threat to whether displayed outputs faithfully reflect GPT-4o reasoning, but it is an empirical instrumentation concern, not circularity: the children's judgments are not defined in terms of the parser outputs, and no central claim reduces to an input-output identity. Therefore, no significant circularity is present.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper makes no fitted numerical claims. Its load-bearing assumptions are domain and methodological: the ARC puzzles are a suitable proxy for reasoning tasks, the CTML dual-channel account applies to children in this setting, the participatory design sessions provide valid evidence, and GPT-4o is a representative generative AI system. These are stated or implied in Sections 3 and 4.

assumptions (5)
  • domain assumption ARC puzzles require no prior knowledge and are a valid benchmark for AI reasoning, while AI systems struggle and humans excel.
    Invoked in Section 3 to justify puzzle choice; no performance data for the 12 selected puzzles is reported.
  • domain assumption Information is processed through separate visual and verbal channels, and distributing content across them reduces cognitive load.
    CTML from Mayer and Moreno is used as the theoretical basis for the design in Sections 3.1 and 6.2, but the study does not measure cognitive load.
  • domain assumption Cooperative Inquiry with children yields authentic and valid design insights.
    Methodological premise for the participatory design sessions in Section 4; the children had prior experience with this method.
  • domain assumption GPT-4o's behavior during the sessions is representative of generative AI reasoning errors relevant to children.
    The system uses GPT-4o via API, and the paper generalizes from its failures to genAI limitations in Sections 3.2 and 5.
  • domain assumption Children's verbal reports and physical reactions during video analysis indicate genuine reasoning rather than facilitator cuing.
    Data coding in Section 4.3 treats observed dialogue as evidence of learning and critical evaluation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of "AI just keeps guessing": Using ARC Puzzles to Help Children Identify Reasoning Errors in Generative AI." pith.science (2026). https://pith.science/paper/76MKGKCX

@misc{pith2026250516034,
  author       = {Pith},
  title        = {Pith review of: "AI just keeps guessing": Using ARC Puzzles to Help Children Identify Reasoning Errors in Generative AI},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/76MKGKCX}},
  note         = {Machine review of arXiv:2505.16034}
}
read the original abstract

The integration of generative Artificial Intelligence (genAI) into everyday life raises questions about the competencies required to critically engage with these technologies. Unlike visual errors in genAI, textual mistakes are often harder to detect and require specific domain knowledge. Furthermore, AI's authoritative tone and structured responses can create an illusion of correctness, leading to overtrust, especially among children. To address this, we developed AI Puzzlers, an interactive system based on the Abstraction and Reasoning Corpus (ARC), to help children identify and analyze errors in genAI. Drawing on Mayer & Moreno's Cognitive Theory of Multimedia Learning, AI Puzzlers uses visual and verbal elements to reduce cognitive overload and support error detection. Based on two participatory design sessions with 21 children (ages 6 - 11), our findings provide both design insights and an empirical understanding of how children identify errors in genAI reasoning, develop strategies for navigating these errors, and evaluate AI outputs.

Figures

Figures reproduced from arXiv: 2505.16034 by the authors.

Figure 1
Figure 1. Overview of AI Puzzlers: (A) children first solve an ARC puzzle independently, then (B) test whether genAI can solve [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Comparison of the correct vs AI-generated solutions. The visual nature of AI Puzzlers makes AI errors easy to spot. [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. An example of an ARC puzzle with instructions for solving it. The correct answer shows the shortest bar is colored [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (13 more)
Figure 4
Figure 4. Figure 4: Annotated screenshot of Manual Mode in AI Puzzlers. The interface consists of several key components: (A) Puzzle [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Screenshot of Assist Mode in AI Puzzlers. Certain interface elements are enlarged to highlight key interactive features, [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Children engaging with AI Puzzlers alongside adult facilitators. [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Panel (1) shows children collaboratively solving a puzzle, while panel (2) presents the AI’s attempt at the same puzzle. [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: Panel (1) shows children collaboratively solving a puzzle, while panel (2) presents the AI’s attempt at the same puzzle. [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]
Figure 9
Figure 9. Figure 9: Panel (1) shows children collaboratively solving a puzzle, while panel (2) presents the AI’s attempt at the same puzzle. [PITH_FULL_IMAGE:figures/full_fig_p011_9.png]
Figure 10
Figure 10. Figure 10: Children iteratively refined their instructions to guide the AI towards solving the puzzle. The sequence showcases [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 11
Figure 11. Figure 11: Children’s Interaction with Assist Mode in AI Puzzlers. (A) Displays examples of example patterns to infer the [PITH_FULL_IMAGE:figures/full_fig_p013_11.png]
Figure 12
Figure 12. Figure 12: Panel (1) shows children collaboratively solving a puzzle, while panel (2) presents the AI’s attempt at the same puzzle. [PITH_FULL_IMAGE:figures/full_fig_p015_12.png]
Figure 13
Figure 13. Figure 13: Panel (1) shows children collaboratively solving a puzzle, while panel (2) presents the AI’s attempt at the same puzzle. [PITH_FULL_IMAGE:figures/full_fig_p015_13.png]
Figure 14
Figure 14. Figure 14: Panel (1) shows children collaboratively solving a puzzle, while panel (2) presents the AI’s attempt at the same puzzle. [PITH_FULL_IMAGE:figures/full_fig_p016_14.png]
Figure 15
Figure 15. Figure 15: The correct solution (left) is compared with AI-generated attempts (right). [PITH_FULL_IMAGE:figures/full_fig_p016_15.png]
Figure 16
Figure 16. Figure 16: The correct solution (left) is compared with AI-generated attempts (right). [PITH_FULL_IMAGE:figures/full_fig_p017_16.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

101 extracted references · 42 canonical work pages

  1. [1]

    Luca Ambrosio, Jordy Schol, Vincenzo Amedeo La Pietra, Fabrizio Russo, Gianluca Vadalà, and Daisuke Sakai. 2024. Threats and opportunities of using ChatGPT in scientific writing—The risk of getting spineless. JOR Spine 7, 1 (2024), e1296–n/a

  2. [2]

    Valentina Andries and Judy Robertson. 2023. Alexa doesn’t have that many feelings: Children’s understanding of AI through interactions with smart speakers in their homes. Computers and Education: Artificial Intelligence 5 (2023), 100176. doi:10.1016/j.caeai.2023.100176

  3. [3]

    Zied Bahroun, Chiraz Anane, Vian Ahmed, and Andrew Zacca. 2023. Transform- ing education: A comprehensive review of generative artificial intelligence in educational settings through bibliometric and content analysis. Sustainability 15, 17 (2023), 12983

  4. [4]

    Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big?. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency. 610–623

  5. [5]

    Gernot Beutel, Eline Geerits, and Jan T Kielstein. 2023. Artificial hallucination: GPT on LSD? Critical Care 27, 1 (2023), 148

  6. [6]

    Melanie Birks, Ysanne Chapman, and Karen Francis. 2008. Memoing in qualitative research: Probing data and processes. Journal of research in nursing 13, 1 (2008), 68–75

  7. [7]

    Georges-Pierre Bonneau, Hans-Christian Hege, Chris R Johnson, Manuel M Oliveira, Kristin Potter, Penny Rheingans, and Thomas Schultz. 2014. Overview and state-of-the-art of uncertainty visualization. Scientific visualization: Uncer- tainty, multifield, biomedical, and scalable visualization (2014), 3–27

  8. [8]

    Ali Borji. 2023. Qualitative failures of image generation models and their appli- cation in detecting deepfakes. Image and Vision Computing 137 (2023), 104771

Show all 101 references
  1. [9]

    Acey Boyce and Tiffany Barnes. 2010. BeadLoom Game: using game elements to increase motivation and learning. In Proceedings of the fifth international conference on the foundations of digital games . 25–31

  2. [10]

    Michelle Carney, Barron Webster, Irene Alvarado, Kyle Phillips, Noura Howell, Jordan Griffith, Jonas Jongejan, Amit Pitaru, and Alexander Chen. 2020. Teach- able machine: Approachable Web-based tool for exploring machine learning classification. In Extended abstracts of the 20...

  3. [11]

    Pew Research Center. 2025. About a Quarter of US Teens Have Used ChatGPT for Schoolwork, Double the Share in 2023 . https://www.pewresearch.org/short- reads/2025/01/15/about-a-quarter-of-us-teens-have-used-chatgpt-for- schoolwork-double-the-share-in-2023/

  4. [12]

    Amanda Chaffin and Tiffany Barnes. 2010. Lessons from a course on serious games research and prototyping. In Proceedings of the Fifth International Confer- ence on the Foundations of Digital Games . 32–39

  5. [13]

    Liuqing Chen, Shuhong Xiao, Yunnong Chen, Yaxuan Song, Ruoyu Wu, and Lingyun Sun. 2024. ChatScratch: an AI-Augmented System Toward Autonomous Visual Programming Learning for Children Aged 6-12. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Ho...

  6. [14]

    François Chollet. 2019. On the measure of intelligence. arXiv preprint arXiv:1911.01547 (2019)

  7. [15]

    Rudrajit Choudhuri, Ambareesh Ramakrishnan, Amreeta Chatterjee, Bianca Trinkenreich, Igor Steinmacher, Marco Gerosa, and Anita Sarma. 2024. Insights from the Frontline: GenAI Utilization Among Software Engineering Students. (2024)

  8. [16]

    Clark, Emily E

    Douglas B. Clark, Emily E. Tanner-Smith, and Stephen S. Killingsworth. 2016. Digital Games, Design, and Learning: A Systematic Review and Meta-Analysis. Review of Educational Research86, 1 (2016), 79–122. doi:10.3102/0034654315582065 arXiv:https://doi.org/10.3102/0034654315582...

  9. [17]

    Daniel C Cliburn. 2006. The effectiveness of games as assignments in an intro- ductory programming course. In Proceedings. Frontiers in Education. 36th Annual Conference. IEEE, 6–10

  10. [18]

    I Want to Think Like an SLP

    Aayushi Dangol, Aaleyah Lewis, Hyewon Suh, Xuesi Hong, Hedda Meadan, James Fogarty, and Julie A Kientz. 2025. “I Want to Think Like an SLP”: A Design Exploration of AI-Supported Home Practice in Speech Therapy. In Proceedings of the 2025 CHI Conference on Human Factors in Comp...

  11. [19]

    Kientz, Jason Yip, and Caroline Pitt

    Aayushi Dangol, Michele Newman, Robert Wolfe, Jin Ha Lee, Julie A. Kientz, Jason Yip, and Caroline Pitt. 2024. Mediating Culture: Cultivating Socio-cultural Understanding of AI in Children through Participatory Design. In Proceedings of the 2024 ACM Designing Interactive Syste...

  12. [20]

    Aayushi Dangol, Robert Wolfe, Runhua Zhao, JaeWon Kim, Trushaa Ramanan, Katie Davis, and Julie A. Kientz. 2025. Children’s Mental Models of AI Reasoning: Implications for AI Literacy Education. InProceedings of the Interaction Design and Children Conference (IDC ’25) . Associa...

  13. [21]

    Riddhi Divanji, Aayushi Dangol, Ella J Lombard, Katharine Chen, and Jennifer D Rubin. 2024. Togethertales RPG: Prosocial skill development through digitally mediated collaborative role-playing. In Proceedings of the 23rd Annual ACM Interaction Design and Children Conference . ...

  14. [22]

    Stefania Druga. 2018. Growing up with AI: Cognimates: from coding to teaching machines. Ph. D. Dissertation. Massachusetts Institute of Technology

  15. [23]

    Stefania Druga and Amy J Ko. 2021. How do children’s perceptions of machine intelligence change when training and coding smart programs?. In Proceedings of the 20th annual ACM interaction design and children conference . 49–61

  16. [24]

    Allison Druin. 1999. Cooperative inquiry: developing new technologies for children with children. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Pittsburgh, Pennsylvania, USA) (CHI ’99). Association for Computing Machinery, New York, NY, USA, 59...

  17. [25]

    Allison DRUIN. 2002. The role of children in the design of new technology. Behaviour & information technology 21, 1 (2002), 1–25

  18. [26]

    Utkarsh Dwivedi, Jaina Gandhi, Raj Parikh, Merijke Coenraad, Elizabeth Bon- signore, and Hernisa Kacorri. 2021. Exploring Machine Teaching with Children. In 2021 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC), Vol. 2021. IEEE, United States, 1–11

  19. [27]

    Jennifer Fereday and Eimear Muir-Cochrane. 2006. Demonstrating rigor using thematic analysis: A hybrid approach of inductive and deductive coding and theme development. International journal of qualitative methods 5, 1 (2006), 80–92

  20. [28]

    Douglas A Gentile. 2009. Video games affect the brain—for better and worse. Cerebrum: The DANA Foundation (2009)

  21. [29]

    MARK GRIFFITHS. 1997. Computer Game Playing in Early Adolescence. Youth & society 29, 2 (1997), 223–237

  22. [30]

    Jessica Grose. 2024. What Teachers Told Me About A.I. in School. The New York Times (14 Aug. 2024). https://www.nytimes.com/2024/08/14/opinion/ai-schools- teachers-students.html Opinion

  23. [31]

    Shuchi Grover. 2024. Teaching AI to K-12 Learners: Lessons, Issues, and Guid- ance. In Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1 (Portland, OR, USA) (SIGCSE 2024). Association for Computing Machinery, New York, NY, USA, 422–428. doi:10....

  24. [32]

    Ariel Han and Zhenyao Cai. 2023. Design implications of generative AI systems for visual storytelling for young learners. In Proceedings of the 22nd Annual ACM Interaction Design and Children Conference . 470–474

  25. [33]

    Sungsoo Ray Hong, Jessica Hullman, and Enrico Bertini. 2020. Human factors in model interpretability: Industry practices, challenges, and needs. Proceedings of the ACM on Human-Computer Interaction 4, CSCW1 (2020), 1–26

  26. [34]

    Nicola Jones. 2025. How should we test AI for human-level intelligence? OpenAI’s o3 electrifies quest. Nature 637, 8047 (2025), 774–775

  27. [35]

    Brian Jordan, Nisha Devasia, Jenna Hong, Randi Williams, and Cynthia Breazeal

  28. [36]

    Ken Kahn and Niall Winters. 2021. Constructionism and AI: A history and possible futures. British Journal of Educational Technology 52, 3 (2021), 1130–1142

  29. [37]

    Jeongah Kim and Jaekwoun Shim. 2022. Development of an AR-based AI educa- tion app for non-majors. IEEE Access 10 (2022), 14149–14156

  30. [38]

    Kirschner, John Sweller, and Richard E

    Paul A. Kirschner, John Sweller, and Richard E. Clark. 2006. Why Minimal Guidance During Instruction Does Not Work: An Analysis of the Failure of Con- structivist, Discovery, Problem-Based, Experiential, and Inquiry-Based Teaching. Educational psychologist 41, 2 (2006), 75–86

  31. [39]

    Marta Laupa. 1991. Children’s Reasoning About Three Authority Attributes: Adult Status, Knowledge, and Social Position. Developmental psychology 27, 2 (1991), 321–329

  32. [40]

    Michael J Lee. 2014. Gidget: An online debugging game for learning and engage- ment in computing education. In 2014 ieee symposium on visual languages and human-centric computing (vl/hcc). IEEE, 193–194

  33. [41]

    Duri Long and Brian Magerko. 2020. What is AI Literacy? Competencies and Design Considerations. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–16. doi:10.1...

  34. [42]

    Duri Long, Sophie Rollins, Jasmin Ali-Diaz, Katherine Hancock, Samnang Nuon- sinoeun, Jessica Roberts, and Brian Magerko. 2023. Fostering AI Literacy with Embodiment & Creativity: From Activity Boxes to Museum Exhibits. In Proceed- ings of the 22nd Annual ACM Interaction Desig...

  35. [43]

    Rebecca Marrone, Victoria Taddeo, and Gillian Hill. 2022. Creativity and artificial intelligence—A student perspective. Journal of Intelligence 10, 3 (2022), 65

  36. [44]

    RE Mayer. 2005. Cognitive theory of multimedia learning . The Cambridge Hand- book of Visuospatial Thinking/Cambridge University Press

  37. [45]

    Richard E. Mayer. 2009. Multimedia learning

  38. [46]

    Richard E Mayer and Roxana Moreno. 2003. Nine ways to reduce cognitive load in multimedia learning. Educational psychologist 38, 1 (2003), 43–52

  39. [47]

    Nora McDonald, Sarita Schoenebeck, and Andrea Forte. 2019. Reliability and Inter-rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice. Proc. ACM Hum.-Comput. Interact. 3, CSCW, Article 72 (Nov. 2019), 23 pages. doi:10.1145/3359174

  40. [48]

    Common Sense Media. 2025. New Report Shows Students Are Em- bracing Artificial Intelligence Despite Lack of Parent A wareness . https: //www.commonsensemedia.org/press-releases/new-report-shows-students- are-embracing-artificial-intelligence-despite-lack-of-parent-awareness-an...

  41. [49]

    Pekka Mertala, Janne Fagerlund, and Oscar Calderon. 2022. Finnish 5th and 6th Grade Students’ Pre-Instructional Conceptions of Artificial Intelligence (AI) and Their Implications for AI Literacy Education. 3 (2022), 100095. doi:10.1016/j. caeai.2022.100095

  42. [50]

    Tilman Michaeli, Stefan Seegerer, Lennard Kerber, and Ralf Romeike. 2023. Data, Trees, and Forests–Decision Tree Learning in K-12 Education. arXiv preprint arXiv:2305.06442 (2023)

  43. [51]

    Geddes, Peter C

    Scott Monteith, Tasha Glenn, John R. Geddes, Peter C. Whybrow, Eric Achtyes, and Michael Bauer. 2024. Artificial intelligence and increasing misinformation. British journal of psychiatry 224, 2 (2024), 33–35

  44. [52]

    It’s smart and it’s stupid:

    Luis Morales-Navarro, Phillip Gao, Eric Yang, and Yasmin B Kafai. 2024. " It’s smart and it’s stupid:" Youth’s conflicting perspectives on LLMs’ language com- prehension and ethics. In Proceedings of the 19th WiPSCE Conference on Primary and Secondary Computing Education Resea...

  45. [53]

    Terran Mott, Alexandra Bejarano, and Tom Williams. 2022. Robot co-design can help us engage child stakeholders in ethical reflection. In 2022 17th ACM/IEEE International Conference on Human-Robot Interaction (HRI) . IEEE, 14–23

  46. [54]

    I want it to talk like Darth Vader

    Michele Newman, Kaiwen Sun, Ilena B Dalla Gasperina, Grace Y Shin, Matthew Kyle Pedraja, Ritesh Kanchi, Maia B Song, Rannie Li, Jin Ha Lee, and Jason Yip. 2024. " I want it to talk like Darth Vader": Helping Children Construct Creative Self-Efficacy with Generative AI. In Proc...

  47. [55]

    Davy Tsz Kit Ng, Chen Xinyu, Jac Ka Lok Leung, and Samuel Kai Wah Chu

  48. [56]

    Mohammad Obaid, Gökçe Elif Baykal, Güncel Kırlangıc, Tilbe Göksun, and Asım Evren Yantaç. 2024. Collective co-design activities with children for design- ing classroom robots. In Proceedings of the 4th African Human Computer Interac- tion Conference (East London, South Africa)...

  49. [57]

    OpenAI. 2024. Early access for safety testing. OpenAI Blog , (Dec 2024),

  50. [58]

    Allan Paivio and Kalman Csapo. 1973. Picture superiority in free recall: Imagery or dual coding? Cognitive psychology 5, 2 (1973), 176–206

  51. [59]

    Ben- jamin Shapiro, Danielle Albers Szafir, Edd V

    William Christopher Payne, Yoav Bergner, Mary Etta West, Carlie Charp, R. Ben- jamin Shapiro, Danielle Albers Szafir, Edd V. Taylor, and Kayla DesPortes. 2021. danceON: Culturally Responsive Creative Computing. In Proceedings of the 2021 CHI Conference on Human Factors in Comp...

  52. [60]

    Kristin Potter, Paul Rosen, and Chris R Johnson. 2012. From quantification to visualization: A taxonomy of uncertainty visualization approaches. InUncertainty Quantification in Scientific Computing: 10th IFIP WG 2.5 Working Conference, WoCoUQ 2011, Boulder, CO, USA, August 1-4...

  53. [61]

    Junaid Qadir. 2023. Engineering education in the era of ChatGPT: Promise and pitfalls of generative AI for education. In 2023 IEEE Global Engineering Education Conference (EDUCON). IEEE, 1–9

  54. [62]

    Are you smart?

    Kai Quander, Tanzila Roushan Milky, Natalie Aponte, Natalia Caceres Carrascal, and Julia Woodward. 2024. “Are you smart?”: Children’s Understanding of “Smart” Technologies. In Proceedings of the 23rd Annual ACM Interaction Design and Children Conference. 625–638

  55. [63]

    Muhammad Raees, Inge Meijerink, Ioanna Lykourentzou, Vassilis-Javed Khan, and Konstantinos Papangelis. 2024. From explainable to interactive AI: A litera- ture review on current trends in human-AI interaction. International Journal of Human-Computer Studies (2024), 103301

  56. [64]

    Mitchel Resnick. 2024. Generative AI and creative learning: Concerns, opportu- nities, and choices. (2024)

  57. [65]

    Mitchel Resnick and Brian Silverman. 2005. Some reflections on designing construction kits for kids. In Proceedings of the 2005 Conference on Interaction Design and Children (Boulder, Colorado) (IDC ’05). Association for Computing Machinery, New York, NY, USA, 117–122. doi:10....

  58. [66]

    AI just keeps guessing

    Richard Rogers. 2018. Coding and writing analytic memos on qualitative data: A review of Johnny Saldaña’s the coding manual for qualitative researchers. The “AI just keeps guessing”: Using ARC Puzzles to Help Children Identify Reasoning Errors in Generative AI IDC ’25, June 23...

  59. [67]

    Donald A. Schön. 1990. Educating the reflective practitioner: toward a new design for teaching and learning in the professions . Jossey-Bass, San Francisco

  60. [68]

    Marita Skjuve, Asbjørn Følstad, and Petter Bae Brandtzaeg. 2023. The user experience of ChatGPT: findings from a questionnaire study of early users. In Proceedings of the 5th international conference on conversational user interfaces . 1–10

  61. [69]

    Jaemarie Solyst, Amy Ogan, and Jessica Hammer. 2023. Intergenerational Games to Learn About AI and Ethics. In Proceedings of the 54th ACM Tech- nical Symposium on Computer Science Education V. 2 (Toronto ON, Canada) (SIGCSE 2023). Association for Computing Machinery, New York,...

  62. [70]

    Jiahong Su and Weipeng Yang. 2023. Unlocking the power of ChatGPT: A framework for applying generative AI in education. ECNU Review of Education 6, 3 (2023), 355–366

  63. [71]

    Hyewon Suh, Aayushi Dangol, Hedda Meadan, Carol A Miller, and Julie A Kientz

  64. [72]

    Yujie Sun, Dongfang Sheng, Zihan Zhou, and Yifei Wu. 2024. AI hallucination: towards a comprehensive classification of distorted information in artificial intelligence-generated content. Humanities and Social Sciences Communications 11, 1 (2024), 1–14

  65. [73]

    John Sweller. 2010. Element Interactivity and Intrinsic, Extraneous, and Germane Cognitive Load. Educational psychology review 22, 2 (2010), 123–138

  66. [74]

    In Proceedings of the 3rd Annual Meeting of the Symposium on Human-Computer Interaction for Work

    Opportunities and challenges for AI-based support for speech-language pathologists. In Proceedings of the 3rd Annual Meeting of the Symposium on Human-Computer Interaction for Work. 1–14

  67. [75]

    Anastasios Theodoropoulos. 2022. Participatory design and participatory de- bugging: Listening to students to improve computational thinking by creating games. International Journal of Child-Computer Interaction 34 (2022), 100525

  68. [76]

    David Touretzky, Christina Gardner-McCune, Fred Martin, and Deborah See- horn. 2019. Envisioning AI for K-12: What Should Every Child Know about AI? Proceedings of the AAAI Conference on Artificial Intelligence 33, 01 (Jul. 2019), 9795–9799. doi:10.1609/aaai.v33i01.33019795

  69. [77]

    Lev Tankelevitch, Viktor Kewenig, Auste Simkute, Ava Elizabeth Scott, Advait Sarkar, Abigail Sellen, and Sean Rintel. 2024. The metacognitive demands and opportunities of generative AI. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems . 1–24

  70. [78]

    Iro Voulgari, Marvin Zammit, Elias Stouraitis, Antonios Liapis, and Georgios Yannakakis. 2021. Learn to Machine Learn: Designing a Game Based Approach for Teaching Machine Learning to Primary and Secondary Education Students. In Proceedings of the 20th Annual ACM Interaction D...

  71. [79]

    Lev S Vygotsky. 1978. Mind in society (M. Cole, V. John-Steiner, S. Scribner, & E. Souberman, Eds.)

  72. [80]

    Alexa, Can I Program You?

    Jessica Van Brummelen, Viktoriya Tabunshchyk, and Tommy Heng. 2021. “Alexa, Can I Program You?”: Student Perceptions of Conversational Artificial Intel- ligence Before and After Programming Alexa. In Proceedings of the 20th An- nual ACM Interaction Design and Children Conferen...

  73. [81]

    Dakuo Wang, Elizabeth Churchill, Pattie Maes, Xiangmin Fan, Ben Shneiderman, Yuanchun Shi, and Qianying Wang. 2020. From human-human collaboration to Human-AI collaboration: Designing AI systems that can work together with people. In Extended abstracts of the 2020 CHI conferen...

  74. [82]

    Heidi C Webb and Mary Beth Rosson. 2011. Exploring careers while learning Alice 3D: a summer camp for middle school girls. In Proceedings of the 42nd ACM technical symposium on Computer science education . 377–382

  75. [83]

    Greg Walsh, Elizabeth Foss, Jason Yip, and Allison Druin. 2013. FACIT PD: a framework for analysis and creation of intergenerational techniques for par- ticipatory design. In proceedings of the SIGCHI Conference on Human Factors in Computing Systems. 2893–2902

  76. [84]

    Robert Wolfe, Aayushi Dangol, Alexis Hiniker, and Bill Howe. 2024. Dataset Scale and Societal Consistency Mediate Facial Impression Bias in Vision-Language AI. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , Vol. 7. 1635–1647

  77. [85]

    Robert Wolfe, Aayushi Dangol, Bill Howe, and Alexis Hiniker. 2024. Representa- tion Bias of Adolescents in AI: A Bilingual, Bicultural Study. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , Vol. 7. 1621–1634

  78. [86]

    Wittrock

    Merlin C. Wittrock. 1989. Generative Processes of Comprehension. Educational psychologist 24, 4 (1989), 345–376

  79. [87]

    It Would Be Cool to Get Stampeded by Dinosaurs

    Julia Woodward, Feben Alemu, Natalia E. López Adames, Lisa Anthony, Jason C. Yip, and Jaime Ruiz. 2022. “It Would Be Cool to Get Stampeded by Dinosaurs”: Analyzing Children’s Conceptual Model of AR Headsets Through Co-Design. In Proceedings of the 2022 CHI Conference on Human ...

  80. [88]

    Julia Woodward, Zari McFadden, Nicole Shiver, Amir Ben-Hayon, Jason C Yip, and Lisa Anthony. 2018. Using co-design to examine how children conceptualize intelligent interfaces. In Proceedings of the 2018 CHI conference on human factors in computing systems. 1–14

  81. [89]

    Robert Wolfe and Tanushree Mitra. 2024. The Impact and Opportunities of Gener- ative AI in Fact-Checking. InThe 2024 ACM Conference on Fairness, Accountability, and Transparency. 1531–1543

  82. [90]

    Lixiang Yan, Samuel Greiff, Ziwen Teuber, and Dragan Gašević. 2024. Promises and challenges of generative artificial intelligence for human learning. Nature Human Behaviour 8, 10 (2024), 1839–1850

  83. [91]

    Robert K Yin. 2013. Validity and generalization in future case study evaluations. Evaluation 19, 3 (2013), 321–332

  84. [92]

    Yi Wu. 2023. Integrating generative AI in education: how ChatGPT brings challenges for future learning and teaching. Journal of Advanced Research in Education 2, 4 (2023), 6–10

  85. [93]

    Yip, Kiley Sobel, Caroline Pitt, Kung Jin Lee, Sijin Chen, Kari Nasu, and Laura R

    Jason C. Yip, Kiley Sobel, Caroline Pitt, Kung Jin Lee, Sijin Chen, Kari Nasu, and Laura R. Pina. 2017. Examining Adult-Child Interactions in Intergenerational Participatory Design. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems (Denver, Colora...

  86. [94]

    Eunice Yiu, Eliza Kosoy, and Alison Gopnik. 2024. Transmission Versus Truth, Imitation Versus Innovation: What Children Can Do That Large Lan- guage and Language-and-Vision Models Cannot (Yet). Perspectives on Psy- chological Science 19, 5 (2024), 874–883. doi:10.1177/17456916...

  87. [95]

    Money shouldn’t be money!

    Jason C Yip, Frances Marie Tabio Ello, Fumi Tsukiyama, Atharv Wairagade, and June Ahn. 2023. " Money shouldn’t be money!": An Examination of Financial Literacy and Technology for Children Through Co-Design. In Proceedings of the 22nd Annual ACM Interaction Design and Children ...

  88. [96]

    Chao Zhang, Cheng Yao, Jianhui Liu, Zili Zhou, Weilin Zhang, Lijuan Liu, Fang- tian Ying, Yijun Zhao, and Guanyun Wang. 2021. StoryDrawer: A Co-Creative Agent Supporting Children’s Storytelling through Collaborative Drawing. In Extended Abstracts of the 2021 CHI Conference on ...

  89. [97]

    Jiahua Zhao, Gwo-Jen Hwang, Shao-Chen Chang, Qi-fan Yang, and Artorn Nokkaew. 2021. Effects of gamified interactive e-books on students’ flipped learn- ing performance, motivation, and meta-cognition tendency in a mathematics course. Educational Technology Research and Develop...

  90. [98]

    Yen Na Yum, Neil Cohn, and Way Kwok-Wai Lau. 2021. Effects of picture-word integration on reading visual narratives in L1 and L2. Learning and Instruction 71 (2021), 101397

  91. [101]

    Christopher Zorn, Chadwick A Wingrave, Emiko Charbonneau, and Joseph J LaViola Jr. 2013. Exploring Minecraft as a conduit for increasing interest in programming.. In FDG. 352–359

  92. [2021]

    In Proceedings of the AAAI Conference on Artificial Intelligence , Vol

    PoseBlocks: A toolkit for creating (and dancing) with AI. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 15551–15559

  93. [2024]

    Journal of computer assisted learning 40, 5 (2024), 2049–2064

    Fostering students’ AI literacy development through educational games: AI knowledge, affective and cognitive engagement. Journal of computer assisted learning 40, 5 (2024), 2049–2064

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.