Pith. sign in

REVIEW 4 major objections 5 minor 105 references

Children's Mental Models of AI Reasoning: Implications for AI Literacy Education

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Children's views of AI reasoning sort into three mental models, and the belief that AI reasons because it is inherently smart largely disappears by seventh grade.

desk verdict A genuinely useful taxonomy of children's AI mental models built on ARC puzzles, but the grade-level trend is not statistically established and needs an exact test or softer claims. read the letter →

arxiv 2505.16031 v1 pith:EFIGL5WV submitted 2025-05-21 cs.AI cs.HC

classification cs.AIcs.HC
keywords AIliteracymentalmodelsreasoningchildrenARCpuzzlesinductivedeductivegrade-leveldifferences
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks how children in grades 3-8 think an AI reasons, using grid-based ARC puzzles as a concrete focus. It claims that children's mental models fall into three types: Inherent (AI reasons because intelligence is built into it), Deductive (AI applies rules or knowledge it was given), and Inductive (AI learns patterns from data). The central finding is an age-linked shift: younger children mostly attribute reasoning to AI's intrinsic smartness, while by grade 7 most children describe AI as a pattern recognizer trained on data. If this developmental picture is correct, it matters for AI literacy education because curricula and explainable-AI tools can be pitched to the mental model children actually hold, and can target the misconceptions that persist at every grade.

What carries the argument

The central mechanism is a two-phase design: an initial co-design session with 8 children generated the three-part taxonomy, and a field study with 106 children then tested it. ARC puzzles are the scaffolding object—they are visual grid puzzles in which a solver must infer a transformation from example input-output pairs and apply it to a new grid, so they provide a content-free, developmentally appropriate way for children to talk about reasoning without relying on verbal or domain-specific knowledge. The argument is carried by coding children's open-ended written reflections into the Inherent, Deductive, and Inductive models, then chi-square testing those codes against grade level and prior AI use.

What would settle it

A replication with roughly equal numbers of children from each grade (for example, 40 or more per grade) recruited from schools rather than a voluntary STEM outreach event, using the same ARC-puzzle prompt, would settle the claim. If seventh- and eighth-graders in that sample still give intrinsic-intelligence explanations in any number, or if the grade-by-model association is no longer significant, then the paper's developmental story does not hold as stated.

Watch

Extended reading notes

Core claim

On the paper's own terms, children's reasoning-about-AI falls into a stable three-code taxonomy, and the distribution of those codes changes with grade. In a field study of 106 children, 35.9% gave Inductive explanations, 32.0% gave Deductive explanations, and 32.0% gave Inherent explanations; a chi-square test of grade by model was significant, with $\chi^2(10)=32.00$, $p<.001$, and Cramer's $V=0.39$. The proportion coded as Inherent falls across grades and disappears from the sample after grade 6, while Inductive becomes the predominant model by grade 7. The paper also identifies four domains where children expect AI reasoning to fail—emotional and social reasoning, conceptual and categorical reasoning, non-literal reasoning, and reasoning with unfamiliar representations—and three educational tensions around the overlap and gaps between data, computational, and AI literacies, generalizing AI reasoning across contexts, and keeping AI literacy curricula current.

Load-bearing premise

The paper's central age trend rests on the assumption that the sample's grade mix represents children generally; with only two third-graders and eight eighth-graders, all recruited at a voluntary STEM outreach event, the decline in the Inherent model could be sampling noise rather than a developmental shift.

Editorial extensions

If this is right

  • AI literacy lessons aimed at pattern-finding and training data will likely land best with children in grades 6-8, while younger grades need scaffolds that directly challenge the 'AI just knows' view.
  • Explainable AI tools for children should foreground how a model does and does not reason, because children at every grade show gaps among data literacy, programming literacy, and AI literacy.
  • The ARC puzzle format itself shows promise as a low-cost elicitation and teaching tool for AI reasoning, since children engaged with it readily and could articulate mental models through it.
  • The disappearance of the Inherent model after grade 6 points to a late-elementary and middle-school window in which children's view of AI becomes more mechanistic, making it a natural time to introduce training data and its role in shaping why AI answers as it does.
  • Educators should expect that children's Inductive and Deductive models are still imperfect—for example, some children attribute AI reasoning to internet retrieval or human rule-writing—so curricula should correct those specific misconceptions even as they reinforce the data-driven view.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension is to run the same ARC-puzzle prompt with a much larger, balanced sample, because the paper's own grade 3 cell has only 2 children and grade 8 has only 8; the headline developmental trajectory could be over-stated even if the qualitative taxonomy is sound.
  • If the Inherent-to-Inductive shift is driven more by exposure to chatbots and generative AI than by age itself, then grade-based curricula would need to adapt faster than grade-level bands, directly extending the paper's pace-of-change tension.
  • The four perceived-limitation domains children named could be turned into an item bank: asking children to design 'hard for AI' puzzles could work as a formative assessment of how their mental model handles non-literal, social, conceptual, and unfamiliar-representation reasoning.
  • Only three children in the field study mentioned robots, which hints that the classic 'robot as AI' trope may be fading; an editorial extension would be to test whether abstract puzzle tasks or recent generative-AI exposure is what makes children drop the robotic framing.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper investigates children's mental models of AI reasoning through a two-phase study. In a preliminary co-design study, eight children (grades 3–8) solved ARC puzzles and designed their own puzzles, leading the authors to identify three mental models: Inductive (AI generalizes from data), Deductive (AI applies programmed rules), and Inherent (AI can reason because of its technological nature). In a field study at a STEM outreach event, 106 children (grades 3–8) solved ARC puzzles and wrote reflective responses, which were coded using the three-model framework. The authors report that the Inherent model declines with grade level and the Inductive model increases, becoming predominant by grade 7, based on a chi-square test (χ²(10) = 32.00, p < .001, Cramer's V = 0.39). They also report three perceived limitations of AI reasoning (social/emotional, non-literal, unfamiliar representations, conceptual) and three tensions for AI literacy education.

Significance. The paper addresses a timely and underexplored question—how children conceptualize AI reasoning specifically, rather than AI in general—and the ARC-puzzle scaffolding is a creative, concrete elicitation method. The qualitative taxonomy is clearly presented with illustrative quotes, and the two-phase design from co-design to field study adds ecological validity. If the developmental gradient were robustly established, the work would inform AI literacy curricula and explainable-AI design for children. However, the headline quantitative claim currently rests on a statistically flawed chi-square test, and the reliability measure used for the qualitative coding is inappropriate; these issues need correction before the findings can be fully trusted.

major comments (4)
  1. [§6.2.1, Table 2] The chi-square test of independence is not valid for the 6×3 grade-by-model table. With only 2 participants in grade 3 and 8 in grade 8, and with the Inherent category overall at 34/106, several expected cell counts are far below 5 (e.g., approximately 0.64 and 2.57 for the Inherent model in grades 3 and 8). The reported p < .001 and Cramer's V = 0.39 are therefore unreliable because the chi-square approximation fails under these sparse-cell conditions. Please report an exact test such as Fisher-Freeman-Halton, collapse sparse grade categories, or use an ordinal logistic regression that treats grade as a continuous predictor. The claim that the Inherent model "vanishes entirely" after grade 6 is not statistically supported because it rests on zero observed counts among just 13 seventh-graders and 8 eighth-graders.
  2. [§5.5] Cronbach's alpha is not an appropriate measure of inter-rater reliability for nominal, mutually exclusive coding categories. Cronbach's alpha measures internal consistency of multi-item scales and is unsuitable for agreement between coders on categorical labels. Please report an appropriate chance-corrected agreement index such as Cohen's kappa or Krippendorff's alpha for the full coded dataset, and re-state the reliability of the coding accordingly.
  3. [§4.3 and §6.1] The three mental-model categories were derived from the co-design data and then used to deductively code the field-study reflections. The paper notes that emergent coding was allowed and produced no new categories, which mitigates the concern, but confirmatory bias cannot be ruled out when the same researchers define the categories in one dataset and then look for them in another. Please discuss this methodological dependency explicitly and consider an independent blind-coding pass or a pre-registered codebook as a safeguard in future work.
  4. [§5.1, §8] The field study sample is a convenience sample from a self-selected STEM outreach event, and the grade-level claim is based on very small cell counts at the extremes (2 third-graders, 8 eighth-graders). The Limitations section acknowledges geographic scope and exclusions, but it does not address how the self-selected, STEM-oriented sample may bias the observed developmental trend. This is load-bearing for the central claim because the grade gradient may be an artifact of who attends such events rather than a general developmental pattern.
minor comments (5)
  1. [§6.2.2] The degrees of freedom reported for the chatbot and videogame AI chi-square tests appear incorrect: for a 2×3 table the degrees of freedom should be 2, not 10. With df = 2, χ² = 5.16 and 5.29 yield p-values near .076 and .071 as reported, but with df = 10 the p-values would be much larger. Please correct the degrees of freedom and re-check the reported p-values.
  2. [§4.4.3] P7 is identified as "girl, age 9, grade 5" in this section, but Table 1 lists P7 as a 9-year-old girl in grade 3. Please correct the inconsistency.
  3. [§6.1.3] The quote attributed to P72 in this section says "grade 6," but P72 is described as a girl in grade 5 in §6.1.2 and Table 3. Please verify the participant ID.
  4. [§7.1] The sentence "they often they often struggle to integrate these literacies" contains a duplicated phrase; please remove the repetition.
  5. [Figure 7] The figure's note acknowledges only two grade-three participants, but the plotted proportions for grade 3 are based on n = 2; consider displaying raw counts or confidence intervals in addition to proportions so readers can assess the sparsity of the data.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the coding framework is transparently confirmatory and the grade-level association is an empirical result, not a by-construction fit.

full rationale

The paper contains no mathematical derivation, so circularity must be assessed against the coding and statistical chain. The three mental-model categories were developed in the co-design study (Sections 3.6 and 4.3) and then applied deductively to the field study (Section 5.5). This is a confirmatory design rather than an independence-generating one, but it is not circular: the field responses were not generated by the codebook, the authors explicitly allowed emergent coding and reported that no additional categories were consistently observed, and the central quantitative claim (grade-level association, Section 6.2.1) is an empirical distribution that the codebook does not fix. The self-citations present in the paper (e.g., references [17] and [18]) are used for cultural context and for describing the ARC puzzle interface; they are not load-bearing evidence for the three models or the grade gradient. The chi-square concerns about small expected cell counts, raised in the skeptical review, are validity and correctness risks rather than circularity, and they do not affect this score.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper does not introduce physical or mathematical entities; its load-bearing premises are methodological and statistical. The coding categories are a qualitative taxonomy, not invented entities with independent falsifiable handles. The main assumptions concern the validity of ARC puzzles as a probe, the representativeness of one written response, and the applicability of chi-square to small cells.

assumptions (4)
  • domain assumption ARC puzzles are a developmentally appropriate scaffold for probing children's understanding of reasoning.
    Section 3.3-3.4 justifies this via prior work [2, 71]; the entire method depends on the puzzles being a valid framing tool.
  • domain assumption Children's written reflections after solving ARC puzzles are a valid window into their mental models of AI reasoning.
    Section 5.4 uses a single open-ended prompt after puzzle solving; the paper does not triangulate with interviews or richer elicitation.
  • domain assumption Johnson-Laird's theory of mental models applies to children's conceptualizations of AI.
    Section 2 adopts Johnson-Laird's framing [45] without testing alternative conceptualizations.
  • domain assumption Chi-square test assumptions are met for the grade-level analysis.
    Section 6.2.1 applies chi-square to a 6x3 table where grade 3 (n=2) and grade 8 (n=8) are too small for reliable expected counts; the paper does not report expected counts or corrections.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Children's Mental Models of AI Reasoning: Implications for AI Literacy Education." pith.science (2026). https://pith.science/paper/EFIGL5WV

@misc{pith2026250516031,
  author       = {Pith},
  title        = {Pith review of: Children's Mental Models of AI Reasoning: Implications for AI Literacy Education},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EFIGL5WV}},
  note         = {Machine review of arXiv:2505.16031}
}
read the original abstract

As artificial intelligence (AI) advances in reasoning capabilities, most recently with the emergence of Large Reasoning Models (LRMs), understanding how children conceptualize AI's reasoning processes becomes critical for fostering AI literacy. While one of the "Five Big Ideas" in AI education highlights reasoning algorithms as central to AI decision-making, less is known about children's mental models in this area. Through a two-phase approach, consisting of a co-design session with 8 children followed by a field study with 106 children (grades 3-8), we identified three models of AI reasoning: Deductive, Inductive, and Inherent. Our findings reveal that younger children (grades 3-5) often attribute AI's reasoning to inherent intelligence, while older children (grades 6-8) recognize AI as a pattern recognizer. We highlight three tensions that surfaced in children's understanding of AI reasoning and conclude with implications for scaffolding AI curricula and designing explainable AI tools.

Figures

Figures reproduced from arXiv: 2505.16031 by the authors.

Figure 1
Figure 1. An ARC puzzle, with two example transformations and a final grid to which the user must apply the rule. In this case, [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Participants engaging with the web-based ARC puzzle interface. Screenshots show participants solving (1) a Level 1 [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Puzzle designed by P3 (boy, grade 3) that challenges [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Puzzle designed by P2 (girl, grade 5) that requires [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: A puzzle designed by P7 (girl, grade 3) that requires social and emotional reasoning. The puzzle presents the question, [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: A “spot the difference” puzzle designed by P8 (girl, grade 8) that requires categorical and conceptual reasoning. The [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: We found that the proportion of children whose responses indicated an Inherent reasoning mental model declined as [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: This figure includes three observations about AI reasoning that would be supported by a background in computational [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

105 extracted references · 55 canonical work pages

  1. [1]

    Eman A Alasadi and Carlos R Baiz. 2023. Generative AI in education and research: Opportunities, concerns, and solutions. Journal of Chemical Education 100, 8 (2023), 2965–2971

  2. [2]

    Jennifer Amsterlaw. 2006. Children’s beliefs about everyday reasoning. Child Development 77, 2 (2006), 443–464

  3. [3]

    Valentina Andries and Judy Robertson. 2023. Alexa doesn’t have that many feel- ings: Children’s understanding of AI through interactions with smart speakers in their homes. Computers and Education: Artificial Intelligence 5 (2023), 100176

  4. [4]

    Maciej Besta, Julia Barth, Eric Schreiber, Ales Kubicek, Afonso Catarino, Robert Gerstenberger, Piotr Nyczyk, Patrick Iff, Yueling Li, Sam Houliston, et al. 2025. Reasoning Language Models: A Blueprint. arXiv preprint arXiv:2501.11223 (2025)

  5. [5]

    Rahul Bhargava and Catherine D’Ignazio. 2015. Designing tools and activities for data literacy learners. In Workshop on data literacy, Webscience

  6. [6]

    Melanie Birks, Ysanne Chapman, and Karen Francis. 2008. Memoing in qualita- tive research: Probing data and processes. Journal of research in nursing 13, 1 (2008), 68–75

  7. [7]

    Virginia Braun and Victoria Clarke. 2021. Thematic analysis: a practical guide. (2021). IDC ’25, June 23–26, 2025, Reykjavik, Iceland Dangol et al

  8. [8]

    Karen Brennan and Mitchel Resnick. 2012. New frameworks for studying and assessing the development of computational thinking. In Proceedings of the 2012 annual meeting of the American educational research association, Vancouver, Canada, Vol. 1. 25

Show all 105 references
  1. [9]

    Michelle Carney, Barron Webster, Irene Alvarado, Kyle Phillips, Noura Howell, Jordan Griffith, Jonas Jongejan, Amit Pitaru, and Alexander Chen. 2020. Teach- able machine: Approachable Web-based tool for exploring machine learning classification. In Extended abstracts of the 20...

  2. [10]

    Ismail Celik. 2023. Exploring the determinants of artificial intelligence (Ai) literacy: Digital divide, computational thinking, cognitive absorption.Telematics and Informatics 83 (2023), 102026

  3. [11]

    Ruijia Cheng, Aayushi Dangol, Frances Marie Tabio Ello, Lingyu Wang, and Sayamindu Dasgupta. 2023. Concepts, practices, and perspectives for devel- oping computational data literacy: Insights from workshops with a new data programming system. In Proceedings of the 22nd Annual ...

  4. [12]

    Joseph Chipps, Aayushi Dangol, Brittany Terese Fasy, Stacey Hancock, Mengy- ing Jiang, Aubrey Rogowski, KA Searle, and Colby Tofel-Grehl. 2022. Culturally responsive storytelling across content areas using American Indian ledger art and physical computing.. In American Society...

  5. [13]

    François Chollet. 2019. On the measure of intelligence. arXiv preprint arXiv:1911.01547 (2019)

  6. [14]

    Yun Dai. 2024. Integrating unplugged and plugged activities for holistic AI education: An embodied constructionist pedagogical approach. Education and Information Technologies (2024), 1–24

  7. [15]

    Aayushi Dangol and Sayamindu Dasgupta. 2023. Constructionist approaches to critical data literacy: A review. InProceedings of the 22nd Annual ACM Interaction Design and Children Conference . 112–123

  8. [16]

    I Want to Think Like an SLP

    Aayushi Dangol, Aaleyah Lewis, Hyewon Suh, Xuesi Hong, Hedda Meadan, James Fogarty, and Julie A Kientz. 2025. “I Want to Think Like an SLP”: A Design Exploration of AI-Supported Home Practice in Speech Therapy. In Proceedings of the 2025 CHI Conference on Human Factors in Comp...

  9. [17]

    Aayushi Dangol, Michele Newman, Robert Wolfe, Jin Ha Lee, Julie A Kientz, Jason Yip, and Caroline Pitt. 2024. Mediating Culture: Cultivating Socio-cultural Understanding of AI in Children through Participatory Design. In Proceedings of the 2024 ACM Designing Interactive System...

  10. [18]

    AI just keeps guessing

    Aayushi Dangol, Runhua Zhao, Robert Wolfe, Trushaa Ramanan, Julie A. Kientz, and Jason Yip. 2025. “AI just keeps guessing”: Using ARC Puzzles to Help Children Identify Reasoning Errors in Generative AI. In Proceedings of the Interaction Design and Children Conference (IDC ’25)...

  11. [19]

    DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei F...

  12. [20]

    Griffin Dietz, Joseph Outa, Lauren Lowe, James A Landay, and Hyowon Gweon

  13. [21]

    Dominic DiFranzo, Yoon Hyung Choi, Amanda Purington, Jessie G Taft, Janis Whitlock, and Natalya N Bazarova. 2019. Social media testdrive: Real-world social media education for the next generation. In Proceedings of the 2019 CHI conference on human factors in computing systems . 1–11

  14. [22]

    the big picture

    Andrea A DiSessa. 2018. Computational literacy and “the big picture” concerning computers in mathematics education. Mathematical thinking and learning 20, 1 (2018), 3–31

  15. [23]

    Riddhi Divanji, Aayushi Dangol, Ella J Lombard, Katharine Chen, and Jennifer D Rubin. 2024. Togethertales RPG: Prosocial skill development through digitally mediated collaborative role-playing. In Proceedings of the 23rd Annual ACM Interaction Design and Children Conference . ...

  16. [24]

    Stefania Druga. 2018. Growing up with AI: Cognimates: from coding to teaching machines. Ph. D. Dissertation. Massachusetts Institute of Technology

  17. [25]

    Stefania Druga and Amy J Ko. 2021. How do children’s perceptions of machine intelligence change when training and coding smart programs?. In Proceedings of the 20th annual ACM interaction design and children conference . 49–61

  18. [26]

    Stefania Druga, Sarah T Vu, Eesh Likhith, and Tammy Qiu. 2019. Inclusive AI literacy for kids around the world. In Proceedings of FabLearn 2019 . 104–111

  19. [27]

    Hey Google is it ok if I eat you?

    Stefania Druga, Randi Williams, Cynthia Breazeal, and Mitchel Resnick. 2017. " Hey Google is it ok if I eat you?" Initial explorations in child-agent interaction. In Proceedings of the 2017 conference on interaction design and children . 595–600

  20. [28]

    Stefania Druga, Randi Williams, Hae Won Park, and Cynthia Breazeal. 2018. How smart are the smart toys? Children and parents’ agent interaction and intelligence attribution. In Proceedings of the 17th ACM conference on interaction design and children. 231–240

  21. [29]

    Allison Druin. 1999. Cooperative inquiry: developing new technologies for children with children. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems (Pittsburgh, Pennsylvania, USA) (CHI ’99). Association for Computing Machinery, New York, NY, USA, 592...

  22. [30]

    Allison Druin. 2002. The role of children in the design of new technology. Behaviour and Information Technology 21, 1 (2002), 1–25. https://doi.org/10. 1080/01449290210147484

  23. [31]

    Bederson, Juan Pablo Hourcade, Lisa Sherman, Glenda Revelle, Michele Platner, and Stacy Weng

    Allison Druin, Benjamin B. Bederson, Juan Pablo Hourcade, Lisa Sherman, Glenda Revelle, Michele Platner, and Stacy Weng. 2001. Designing a digi- tal library for young children. In Proceedings of the 1st ACM/IEEE-CS Joint Conference on Digital Libraries (Roanoke, Virginia, USA)...

  24. [32]

    Julian Estevez, Gorka Garate, and Manuel Graña. 2019. Gentle introduction to artificial intelligence for high-school students using scratch. IEEE access 7 (2019), 179027–179036

  25. [33]

    Teresa Margaret Flanagan. 2023. GROWING UP IN THE DIGITAL AGE: INVES- TIGATING CHILDREN’S USE, JUDGMENT, AND ENGAGEMENT WITH INTER- ACTIVE TECHNOLOGIES. Ph. D. Dissertation. Cornell University

  26. [34]

    Jodi Forlizzi and Carl DiSalvo. 2006. Service robots in the domestic environment: a study of the roomba vacuum in the home. In Proceedings of the 1st ACM SIGCHI/SIGART conference on Human-robot interaction . 258–265

  27. [35]

    Dedre Gentner and Albert L Stevens. 2014. Mental models. Psychology Press

  28. [36]

    Elliot Glazer, Ege Erdil, Tamay Besiroglu, Diego Chicharro, Evan Chen, Alex Gunning, Caroline Falkman Olsson, Jean-Stanislas Denain, Anson Ho, Emily de Oliveira Santos, et al . 2024. Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai. arXiv preprint...

  29. [37]

    Eric Greenwald, Ari Krakowski, Timothy Hurt, Kelly Grindstaff, and Ning Wang

  30. [38]

    Mona Leigh Guha, Allison Druin, and Jerry Alan Fails. 2013. Cooperative Inquiry Revisited: Reflections of the Past and Guidelines for the Future of Intergenerational Co-Design. International Journal of Child-Computer Interaction 1, 1 (2013), 14–23. https://doi.org/10.1016/j.ij...

  31. [39]

    Dagmar Mercedes Heeg and Lucy Avraamidou. 2024. Young children’s under- standing of AI. Education and Information Technologies (2024), 1–24

  32. [40]

    Ji-Yeon Hong and Yungsik Kim. 2022. Development of Digital and AI teaching- learning strategies based on computational thinking for Enhancing Digital Literacy and AI Literacy of Elementary School Student. Journal of the Korean Association of Information Education 26, 5 (2022), 341–352

  33. [41]

    Sungsoo Ray Hong, Jessica Hullman, and Enrico Bertini. 2020. Human factors in model interpretability: Industry practices, challenges, and needs. Proceedings of the ACM on Human-Computer Interaction 4, CSCW1 (2020), 1–26

  34. [42]

    Jie Huang and Kevin Chen-Chuan Chang. 2023. Towards Reasoning in Large Language Models: A Survey. In The 61st Annual Meeting Of The Association For Computational Linguistics. Children’s Mental Models of AI Reasoning: Implications for AI Literacy Education IDC ’25, June 23–26, ...

  35. [43]

    Zhen Huang, Zengzhi Wang, Shijie Xia, Xuefeng Li, Haoyang Zou, Ruijie Xu, Run-Ze Fan, Lyumanshan Ye, Ethan Chern, Yixin Ye, et al. 2024. Olympicarena: Benchmarking multi-discipline cognitive reasoning for superintelligent ai. Ad- vances in Neural Information Processing Systems...

  36. [44]

    AI4K12 Initiative. 2020. Big Ideas Poster. https://ai4k12.org/resources/big- ideas-poster/

  37. [45]

    Philip Nicholas Johnson-Laird. 1983. Mental models: Towards a cognitive science of language, inference, and consciousness . Number 6. Harvard University Press

  38. [46]

    Nicola Jones. 2025. How should we test AI for human-level intelligence? Ope- nAI’s o3 electrifies quest. Nature 637, 8047 (2025), 774–775

  39. [47]

    Ken Kahn and Niall Winters. 2021. Constructionism and AI: A history and possible futures. British Journal of Educational Technology 52, 3 (2021), 1130– 1142

  40. [48]

    Eliza Kosoy, Soojin Jeong, Anoop Sinha, Alison Gopnik, and Tanya Kraljic. 2024. Children’s Mental Models of Generative Visual and Text Based AI Models.arXiv preprint arXiv:2405.13081 (2024)

  41. [49]

    Moritz Kreinsen and Sandra Schulz. 2021. Students’ conceptions of artificial intelligence. In Proceedings of the 16th Workshop in Primary and Secondary Computing Education. 1–2

  42. [50]

    Lukas Lehner and Martina Landman. 2024. Unplugged Decision Tree Learning– A Learning Activity for Machine Learning Education in K-12. In International Conference on Creative Mathematical Sciences Communication . Springer, 50–65

  43. [51]

    Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al . 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437 (2024)

  44. [52]

    Duri Long and Brian Magerko. 2020. What is AI Literacy? Competencies and Design Considerations. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–16. https://...

  45. [53]

    Scott M Lundberg, Gabriel Erion, Hugh Chen, Alex DeGrave, Jordan M Prutkin, Bala Nair, Ronit Katz, Jonathan Himmelfarb, Nisha Bansal, and Su-In Lee. 2020. From local explanations to global understanding with explainable AI for trees. Nature machine intelligence 2, 1 (2020), 56–67

  46. [54]

    Ruizhe Ma, Ismaila Temitayo Sanusi, Vaishali Mahipal, Joseph E Gonzales, and Fred G Martin. 2023. Developing machine learning algorithm literacy with novel plugged and unplugged approaches. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1 . 298–304

  47. [55]

    Nora McDonald, Sarita Schoenebeck, and Andrea Forte. 2019. Reliability and Inter-rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice. Proc. ACM Hum.-Comput. Interact. 3, CSCW, Article 72 (Nov. 2019), 23 pages. https://doi.org/10.1145/3359174

  48. [56]

    Gaspar Isaac Melsión, Ilaria Torre, Eva Vidal, and Iolanda Leite. 2021. Using explainability to help children understandgender bias in ai. In Proceedings of the 20th Annual ACM Interaction Design and Children Conference . 87–99

  49. [57]

    Pekka Mertala, Janne Fagerlund, and Oscar Calderon. 2022. Finnish 5th and 6th Grade Students’ Pre-Instructional Conceptions of Artificial Intelligence (AI) and Their Implications for AI Literacy Education. 3 (2022), 100095. https: //doi.org/10.1016/j.caeai.2022.100095

  50. [58]

    Tilman Michaeli, Stefan Seegerer, Lennard Kerber, and Ralf Romeike. 2023. Data, Trees, and Forests–Decision Tree Learning in K-12 Education. arXiv preprint arXiv:2305.06442 (2023)

  51. [59]

    Melanie Mitchell. 2025. Artificial intelligence learns to reason. , eadw5211 pages

  52. [60]

    Arseny Moskvichev, Victor Vikram Odouard, and Melanie Mitchell. 2023. The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain. Transactions on Machine Learning Research (2023)

  53. [61]

    Terran Mott, Alexandra Bejarano, and Tom Williams. 2022. Robot co-design can help us engage child stakeholders in ethical reflection. In 2022 17th ACM/IEEE International Conference on Human-Robot Interaction (HRI) . IEEE, 14–23

  54. [62]

    Davy Tsz Kit Ng, Jac Ka Lok Leung, Samuel Kai Wah Chu, and Maggie Shen Qiao. 2021. Conceptualizing AI literacy: An exploratory review. Computers and Education: Artificial Intelligence 2 (2021), 100041

  55. [63]

    Davy Tsz Kit Ng, Jiahong Su, Jac Ka Lok Leung, and Samuel Kai Wah Chu. 2023. Artificial intelligence (AI) literacy education in secondary schools: a review. Interactive Learning Environments (2023), 1–21

  56. [64]

    Donald A Norman. 2014. Some observations on mental models. In Mental models. Psychology Press, 15–22

  57. [65]

    OpenAI. 2022. Introducing ChatGPT. OpenAI Blog , (Nov 2022),

  58. [66]

    OpenAI. 2024. Learning to reason with LLMs. OpenAI Blog , (Sep 2024),

  59. [67]

    Anne Ottenbreit-Leftwich, Krista Glazewski, Minji Jeon, Katie Jantaraweragul, Cindy E Hmelo-Silver, Adam Scribner, Seung Lee, Bradford Mott, and James Lester. 2023. Lessons learned for AI education with elementary students and teachers. International Journal of Artificial Inte...

  60. [68]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback.Advances in neural information processing systems 35 ...

  61. [69]

    Blakeley H Payne. 2019. An ethics of artificial intelligence curriculum for middle school students. MIT Media Lab Personal Robots Group. Retrieved Oct 10 (2019), 2019

  62. [70]

    Stephen J Payne. 2007. Mental models in human-computer interaction. The human-computer interaction handbook (2007), 89–102

  63. [71]

    Jean Piaget. 1964. Cognitive development in children. Journal of research in science teaching 2, 2 (1964), 176–186

  64. [72]

    Are you smart?

    Kai Quander, Tanzila Roushan Milky, Natalie Aponte, Natalia Caceres Carrascal, and Julia Woodward. 2024. “Are you smart?”: Children’s Understanding of “Smart” Technologies. In Proceedings of the 23rd Annual ACM Interaction Design and Children Conference. 625–638

  65. [73]

    Alec Radford and Karthik Narasimhan. 2018. Improving Language Understand- ing by Generative Pre-Training. https://api.semanticscholar.org/CorpusID: 49313245

  66. [74]

    Muhammad Raees, Inge Meijerink, Ioanna Lykourentzou, Vassilis-Javed Khan, and Konstantinos Papangelis. 2024. From explainable to interactive AI: A literature review on current trends in human-AI interaction. International Journal of Human-Computer Studies (2024), 103301

  67. [75]

    David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. 2023. Gpqa: A graduate-level google-proof q&a benchmark. arXiv preprint arXiv:2311.12022 (2023)

  68. [76]

    Richard Rogers. 2018. Coding and writing analytic memos on qualitative data: A review of Johnny Saldaña’s the coding manual for qualitative researchers. The Qualitative Report 23, 4 (2018), 889–892

  69. [77]

    Michael T Rücker and Niels Pinkwart. 2016. Review and discussion of children’s conceptions of computers. Journal of Science Education and Technology 25 (2016), 274–283

  70. [78]

    Malik Sallam, Kholoud Al-Mahzoum, Mohammed Sallam, and Maad M Mijwil

  71. [79]

    Paula J Schwanenflugel, Robbie L Henderson, and William V Fabricius. 1998. De- veloping organization of mental verbs and theory of mind in middle childhood: evidence from extensions. Developmental Psychology 34, 3 (1998), 512

  72. [80]

    Beate Sodian and Heinz Wimmer. 1987. Children’s understanding of inference as a source of knowledge. Child development (1987), 424–433

  73. [81]

    Yukyeong Song, Xiaoyi Tian, Nandika Regatti, Gloria Ashiya Katuka, Kristy Eliz- abeth Boyer, and Maya Israel. 2024. Artificial Intelligence Unplugged: Designing Unplugged Activities for a Conversational AI Summer Camp. In Proceedings of the 55th ACM Technical Symposium on Comp...

  74. [82]

    Jiahong Su and Weipeng Yang. 2023. Unlocking the power of ChatGPT: A framework for applying generative AI in education. ECNU Review of Education 6, 3 (2023), 355–366

  75. [83]

    Jiahong Su, Yuchun Zhong, and Davy Tsz Kit Ng. 2022. A meta-review of literature on educational approaches for teaching AI at the K-12 levels in the Asia-Pacific region. Computers and Education: Artificial Intelligence 3 (2022), 100065

  76. [84]

    David Touretzky, Christina Gardner-McCune, Fred Martin, and Deborah See- horn. 2019. Envisioning AI for K-12: What should every child know about AI?. In Proceedings of the AAAI conference on artificial intelligence , Vol. 33. 9795–9799

  77. [85]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 , (2023),

  78. [86]

    Jessica Van Brummelen, Judy Hanwen Shen, and Evan W Patton. 2019. The pop- star, the poet, and the grinch: Relating artificial intelligence to the computational thinking framework with block-based coding. In Proceedings of International Conference on Computational Thinking Edu...

  79. [87]

    Alexa, Can I Program You?

    Jessica Van Brummelen, Viktoriya Tabunshchyk, and Tommy Heng. 2021. “Alexa, Can I Program You?”: Student Perceptions of Conversational Artifi- cial Intelligence Before and After Programming Alexa. In Proceedings of the 20th Annual ACM Interaction Design and Children Conference...

  80. [88]

    counter- factuals

    Mike Van Duuren, Barbara Dossett, and Dawn Robinson. 1998. Gauging chil- dren’s understanding of artificially intelligent objects: a presentation of “counter- factuals”. International Journal of Behavioral Development 22, 4 (1998), 871–889

  81. [89]

    Stella Vosniadou and William F Brewer. 1992. Mental models of the earth: A study of conceptual change in childhood. Cognitive psychology 24, 4 (1992), 535–585

  82. [90]

    Greg Walsh, Elizabeth Foss, Jason Yip, and Allison Druin. 2013. FACIT PD: a framework for analysis and creation of intergenerational techniques for par- ticipatory design. In proceedings of the SIGCHI Conference on Human Factors in Computing Systems. 2893–2902

  83. [91]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reason- ing in large language models. Advances in neural information processing systems IDC ’25, June 23–26, 2025, Reykjavik, Iceland...

  84. [92]

    Robert Wolfe, Aayushi Dangol, Alexis Hiniker, and Bill Howe. 2024. Dataset Scale and Societal Consistency Mediate Facial Impression Bias in Vision- Language AI. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, Vol. 7. 1635–1647

  85. [93]

    Robert Wolfe, Aayushi Dangol, Bill Howe, and Alexis Hiniker. 2024. Represen- tation Bias of Adolescents in AI: A Bilingual, Bicultural Study. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , Vol. 7. 1621–1634

  86. [94]

    It Would Be Cool to Get Stampeded by Dinosaurs

    Julia Woodward, Feben Alemu, Natalia E. López Adames, Lisa Anthony, Jason C. Yip, and Jaime Ruiz. 2022. “It Would Be Cool to Get Stampeded by Dinosaurs”: Analyzing Children’s Conceptual Model of AR Headsets Through Co-Design. In Proceedings of the 2022 CHI Conference on Human ...

  87. [95]

    Julia Woodward, Zari McFadden, Nicole Shiver, Amir Ben-Hayon, Jason C Yip, and Lisa Anthony. 2018. Using co-design to examine how children conceptualize intelligent interfaces. In Proceedings of the 2018 CHI conference on human factors in computing systems. 1–14

  88. [96]

    Yi Wu. 2023. Integrating generative AI in education: how ChatGPT brings challenges for future learning and teaching. Journal of Advanced Research in Education 2, 4 (2023), 6–10

  89. [97]

    Fengli Xu, Qianyue Hao, Zefang Zong, Jingwei Wang, Yunke Zhang, Jingyi Wang, Xiaochong Lan, Jiahui Gong, Tianjian Ouyang, Fanjin Meng, et al. 2025. Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models. arXiv preprint arXiv:2501.09686 (2025)

  90. [98]

    Weipeng Yang. 2022. Artificial Intelligence education for young children: Why, what, and how in curriculum design and implementation. Computers and Education: Artificial Intelligence 3 (2022), 100061

  91. [99]

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024. Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems 36 (2024)

  92. [100]

    Svetlana Yarosh, Stryker Thompson, Kathleen Watson, Alice Chase, Ashwin Senthilkumar, Ye Yuan, and AJ Bernheim Brush. 2018. Children asking questions: speech interface reformulations and personification preferences. In Proceedings of the 17th ACM conference on interaction desi...

  93. [101]

    Jason C Yip, Kung Jin Lee, and Jin Ha Lee. 2020. Design partnerships for participatory librarianship: A conceptual model for understanding librarians co designing with digital youth. Journal of the Association for Information Science and Technology 71, 10 (2020), 1242–1256. ht...

  94. [102]

    Jason C Yip, Kiley Sobel, Caroline Pitt, Kung Jin Lee, Sijin Chen, Kari Nasu, and Laura R Pina. 2017. Examining adult-child interactions in intergenerational participatory design. In Proceedings of the 2017 CHI conference on human factors in computing systems. Association for ...

  95. [2023]

    mental states

    Theory of AI Mind: How adults and children reason about the“mental states”of conversational AI. InProceedings of the Annual Meeting of the Cognitive Science Society, Vol. 45

  96. [2024]

    In Proceedings of the 23rd Annual ACM Interaction Design and Children Conference

    It’s like I’m the AI: Youth Sensemaking About AI through Metacognitive Embodiment. In Proceedings of the 23rd Annual ACM Interaction Design and Children Conference. 789–793

  97. [2025]

    DeepSeek: Is it the End of Generative AI Monopoly or the Mark of the Impending Doomsday? Mesopotamian Journal of Big Data 2025 (2025), 26–34

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.