Pith. sign in

REVIEW 4 major objections 3 minor 80 references

MindScratch: A Visual Programming Support Tool for Classroom Learning Based on Multimodal Generative AI

T0 review · 4 major / 3 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read MindScratch, a generative-AI mind-map tool, helps fifth graders finish Scratch projects aligned with teacher-set learning objectives and raises their computational-thinking and creativity scores versus plain Scratch in a 24-student…

desk verdict A genuine systems contribution with a sensible design and large headline effects, but the CT and creativity claims rest on outcome measures that cannot separate student work from AI output. read the letter →

arxiv 2412.09001 v1 pith:7I77OWML submitted 2024-12-12 cs.HC

classification cs.HC
keywords computationalthinkinggenerativeAImindmapvisualprogrammingScratchclassroomlearningmultimodalgenerationscaffolding
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

MindScratch is a visual programming support tool that wraps Scratch classroom projects in an interactive mind map generated and guided by multimodal generative AI. The paper argues, on the strength of a within-subject study with 24 fifth graders, that children using MindScratch are more likely to finish teacher-set creative programming tasks (all 24 did, versus 13 of 24 with plain Scratch) and that their projects score higher on expert-rated consistency, originality, and creativity as well as on the Dr.Scratch computational-thinking rubric. The intended upshot is that a classroom tool can keep generative AI's outputs steered by teacher-defined learning objectives while still letting students explore freely, and that this combination improves both objective alignment and computational-thinking development. If true, it offers a concrete answer to a known problem: how to use large language models in K-12 programming classes without letting the AI's unconstrained answers derail the lesson.

What carries the argument

The load-bearing mechanism is an interactive mind map whose nodes are characters, logic, and code, with teacher-set learning objectives baked into the system prompt of an LLM-driven conversational agent. The agent walks students through three stages—project planning, material creation, and code implementation—and every suggestion it makes is constrained by the retained objectives; relevant nodes are highlighted in the map, and relationships between nodes are annotated to keep the generated content interpretable. Code help is scaffolded rather than wholesale: a fine-tuned LLM produces pseudo-code as an abstract syntax tree, which is converted into individual Scratch block images by edit-distance matching, so students receive logic suggestions and key blocks instead of a complete solution. Multimodal assets (images refined from children's doodles via Stable Diffusion with ControlNet, and text-to-audio sounds) give students personal materials without leaving the classroom workflow. The mind map doubles as a shared memory and progress display, which is what the authors say reduces cognitive load and lets teachers see where each student is.

What would settle it

Hand expert raters a set of projects produced by MindScratch and an equal set generated by the AI alone from the same teacher objectives, with student identity and condition removed; if the AI-only projects score as high on originality, creativity, and Dr.Scratch as the students' projects, then the tool's measured benefits are not evidence of student learning.

Watch

Extended reading notes

Core claim

The paper's central claim is that using MindScratch, rather than Scratch alone, lets fifth-grade students produce creative programming projects that meet the teacher's explicit learning objectives, and that the tool measurably raises the computational-thinking and creativity profile of those projects. This claim rests on the comparison in Section 6: 24 of 24 MindScratch students fulfilled the predefined tasks, while only 13 of 24 Scratch students did; expert ratings favored MindScratch strongly on originality (4.38 vs 2.48), creativity (4.13 vs 2.91), and consistency (4.25 vs 3.0), with matched quality, and the Dr.Scratch total rose from 9.96 ('basic') to 14.17 ('master'). The study also reports higher mind-map node counts (52.45 vs a 35-node baseline), higher Creativity Support Index scores on five of six dimensions, and pre/post gains on a computational-thinking survey for the three students followed over four weeks. The authors' interpretation is that the tool's stepwise mind-map scaffolding, not just the AI's raw output, is what keeps students aligned with learning goals while expanding their creative range.

Load-bearing premise

The load-bearing premise is that the outcome measures—expert ratings of originality, creativity, and consistency, plus the Dr.Scratch code-quality score—capture what the students themselves learned and can do, rather than largely crediting the AI-generated logic, images, and audio that MindScratch supplies and the students assemble.

Editorial extensions

If this is right

  • If the result holds, K-12 programming tools can embed teacher-set objectives as persistent LLM prompts, so AI guidance stays on-lesson while still allowing open-ended student exploration.
  • Teachers can shift from repeatedly answering routine 'what now?' questions to giving constructive feedback, because the system's staged dialogue and mind-map visualization carry the process forward.
  • Students can generate their own images, sounds, and code scaffolds quickly, removing the asset-hunting and blank-page barriers that often stall creative Scratch projects.
  • The reported Dr.Scratch gains imply that scaffolded logic-to-code suggestions can move novice projects from 'basic' to 'master' level on computational-thinking dimensions within a single class session.
  • The mind-map trace, with node counts and highlighted objective-relevant blocks, could itself become a formative assessment artifact for teachers.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A decisive next experiment would add a third condition in which students receive the same AI-generated logic, code blocks, and assets but without the mind-map stage structure; if those students align with objectives just as well, the mind map, not the AI, is not the active ingredient.
  • The 52.45 node count may partly reflect AI-suggested nodes that students accepted with one click; a fairer creativity measure would distinguish student-initiated nodes from accepted suggestions, ideally from interaction logs.
  • Long-term transfer is untested: the paper's three-student pre/post survey shows gains while using the tool; the strong test is whether students reproduce equivalent logic in plain Scratch after the tool is withdrawn.
  • The expert raters' ICC of 0.78 establishes inter-rater agreement, not construct validity; rating projects blind to condition and comparing against an AI-only generated project would separate student skill from AI output quality.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. MindScratch is a visual programming support tool for K-12 classrooms that combines teacher-defined learning objectives with an AI-generated interactive mind map, a staged conversational agent, and multimodal asset generation (images, audio, and code/logic suggestions). The paper reports a formative interview study with six educators, a within-subject experiment with 24 fifth-grade students comparing MindScratch to plain Scratch, expert ratings of final projects, Dr.Scratch code-quality scores, a Creativity Support Index questionnaire, and a small long-term observation of three students with a CT skills survey. The central claims are that MindScratch helps students complete projects aligned with learning objectives and that it enhances students' computational thinking skills and creative thinking.

Significance. The system addresses a genuine need in K-12 programming education: balancing structured classroom objectives with creative, project-based exploration. The design process is grounded in an educator formative study, and the system offers concrete mechanisms (teacher-set prompts, scaffolded code suggestions, multimodal asset generation) that are clearly described and open-sourced. The within-subject, counterbalanced study with 24 participants is a reasonable first evaluation, and the large observed differences in expert ratings and Dr.Scratch scores are consistent with the tool being effective at producing higher-scoring final projects. However, the paper's stronger claim that using MindScratch enhances students' computational thinking skills and creative thinking is not established by the reported evidence, because the outcome instruments score final artifacts whose code and assets are substantially AI-generated. The study's positive features include the use of validated instruments, reported inter-rater agreement (ICC = 0.78), and a reproducible GitHub repository, but these do not resolve the attribution problem.

major comments (4)
  1. [Section 5.3.2 / Section 6.1] The Dr.Scratch rubric is used to infer that students 'more extensively utilized and mastered CT skills' (Section 5.3.1), yet MindScratch supplies logic nodes, code suggestions, and even specific Scratch blocks during the coding phase (Section 4.2.4), and students used image polish 6.27 times and audio generation 2.34 times on average (Section 6.3). Dr.Scratch awards credit for constructs such as loops, conditionals, parallelism, and data representation regardless of whether the student or the AI proposed them. The significant improvement in Dr.Scratch total score (14.17 vs 9.96, t=6.44) consequently cannot separate 'the tool helped the student produce code containing CT constructs' from 'the student learned and mastered CT skills.' This directly affects the abstract's claim of enhanced computational thinking. The paper should report per-student attribution data (e.g., which blocks were generated by the AI and which were modified or added by the student), or use a transfer/no-scaffold post-test. Without such evidence, the conclusion should be restricted to project code quality, not student skill acquisition.
  2. [Section 5.3.2 / Section 6.1 and Section 6.3] The expert ratings of originality and creativity are similarly confounded. Originality is defined as 'the extent to which the project reflects the student's creation rather than being derived from existing materials,' but the projects include AI-generated images and audio that the students selected and iteratively refined; the paper reports an average of 5.32 character generations, 6.27 image polishes, and 2.34 audio generations per student. This blurs the boundary between student creation and AI derivation. The reported ICC of 0.78 establishes rater agreement, not construct validity, and the paper does not state whether the three raters were blinded to condition. If raters could identify which projects were made with MindScratch, expectation bias could inflate the consistency, originality, and creativity ratings. The authors should report blinding procedures or analyze the sub-scores for items that are least affected by AI asset quality.
  3. [Tables 4-6] Several reported p-values are not consistent with the reported t statistics and df=23 under a two-tailed paired t-test. For example, Table 5 Flow Control reports t=2.07, p=0.399, but a two-tailed test gives p≈0.05 (≈0.35 after a Bonferroni correction for seven dimensions); Table 6 Immersion reports t=2.92, p=0.005, but the two-tailed p is ≈0.008 (≈0.048 after correction for six dimensions). The authors need to state explicitly whether one- or two-tailed tests were used, report exact uncorrected and corrected p-values, and correct the inconsistencies. The qualitative conclusions are unlikely to change for the primary large effects, but the reporting must be reproducible.
  4. [Section 6.2 and Section 7.2] The long-term CT-skills claim rests primarily on a pre/post survey of only three students (P3, P4, P21), with no inferential statistics, no control condition, and no adjustment for maturation. The paper reports 'each student's CT skills improved' but the differences are small (e.g., P4: 3.9 to 4.05) and the design is exploratory. This is acknowledged in part in Section 7.2, but the abstract and conclusion still assert that MindScratch 'enhances students' computational thinking skills.' The long-term data should be explicitly labeled as a qualitative/exploratory observation, and the primary CT claim should be tied to evidence that separates AI contributions from student learning.
minor comments (3)
  1. [Section 6.3] In the mind-map node-count analysis, the paper says the average count in MindScratch was 52.45 (SD=6.93) while the baseline average was 35, but it does not clarify whether the 35 baseline nodes are the teachers' initial nodes included in both conditions. If the baseline nodes are included in the MindScratch count, the student-added increment is only ~17.45 nodes, and the comparison should be reported as a difference from the teacher-provided baseline. This also lacks a statistical test.
  2. [Section 4.2.4] The fine-tuned LLM is evaluated with BLEU and F1 scores on 30 test samples, but these are generation-similarity metrics and do not directly measure pedagogical quality or whether the suggested code was appropriate for a fifth-grade student. The authors could briefly justify the metric choice or report a small qualitative evaluation.
  3. [Section 5.4 and Section 6.2] There are minor typographical and formatting issues, such as '14.17 (SD = 4.14))' with a double parenthesis, and the repeated phrase 'The participants were 24 fifth-grade students' in Section 5.1. These should be cleaned up before publication.

Circularity Check

0 steps flagged · score 2.0 of 10

No circularity in the central comparison; one minor non-load-bearing self-citation and an in-distribution LLM benchmark keep the score low.

full rationale

The paper's central claims are tested against an external baseline (plain Scratch) with independent instruments: Dr.Scratch rubric scores, expert ratings on a 5-point scale, the Creativity Support Index, and a pre/post CT skills survey. The consistency metric is an expert judgment of alignment with teacher-set objectives; although MindScratch is prompted with those objectives and highlights relevant nodes, the metric is not computed from the system's own relevance detector, so the comparison is empirical rather than self-definitional. The fine-tuned LLM is evaluated on 30 test samples drawn from the same Scratch-card source as its 98 training samples; this is an in-distribution benchmark and a generalization weakness, but it is a component check and does not reduce the main study's claims to their inputs. The only self-citation, ChatScratch [5], is used as a related-work comparison and in the formative study, not as load-bearing evidence for MindScratch's effectiveness. No step in the derivation chain equates a prediction to a fitted input or imports a uniqueness claim from the authors' prior work.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

This is an empirical systems paper, so the ledger records measurement assumptions and hand-chosen design values rather than fitted constants. The central quantitative claims rest on the validity of the outcome instruments (Dr.Scratch, expert ratings, CSI, CT survey), the correctness of the statistical analysis, and the comparability of the baseline. No new physical or conceptual entities are postulated; MindScratch is a software artifact composed of existing generative models and prompt pipelines described in Section 4.

free parameters (3)
  • Saliency highlight threshold (high/low) = not reported
    Section 4.2.2: GPT-4 is instructed to designate code blocks as high (H) or low (L) relevance, and only high-relevance blocks are highlighted. The threshold choice affects which blocks students see and therefore influences the consistency ratings.
  • Fine-tuning dataset size and test set = 98 training samples; 30 test samples
    Section 4.2.4: the fine-tuned LLM's BLEU/F1 evaluation uses 30 test samples "based on Scratch cards", the same source as the 98 training samples, so the evaluation is partly in-distribution.
  • Teacher baseline mind map size = 35 nodes
    Section 6.3: the 52.45 mean node count is compared to a teacher-created baseline of 35 nodes; this is not a Scratch-condition comparison and no test is reported.
assumptions (4)
  • domain assumption Dr.Scratch rubric scores are a valid proxy for students' computational thinking mastery
    Section 5.3.1 states "higher code quality scores indicate that students have more extensively utilized and mastered CT skills", but the projects' code is partly suggested by the LLM, so scores may reflect AI output rather than student mastery.
  • domain assumption Expert ratings on a 5-point Likert scale validly measure consistency, quality, originality, and creativity
    Section 5.3.2: three Scratch experts rated projects with ICC=0.78 (p<0.05), but the paper does not state whether raters were blinded to condition.
  • domain assumption The Chinese-adapted CT skills survey [29] is valid for fifth graders
    Section 5.3.5: the pre/post CT survey is applied to three students; no validity or reliability statistics for this sample are reported.
  • standard math Paired t-tests and the Bonferroni correction are appropriate and correctly applied
    Section 5.4 states paired t-tests with Bonferroni correction; the reported t/p pairs are internally inconsistent and several marked-significant p-values exceed the Bonferroni threshold.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MindScratch: A Visual Programming Support Tool for Classroom Learning Based on Multimodal Generative AI." pith.science (2026). https://pith.science/paper/7I77OWML

@misc{pith2026241209001,
  author       = {Pith},
  title        = {Pith review of: MindScratch: A Visual Programming Support Tool for Classroom Learning Based on Multimodal Generative AI},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7I77OWML}},
  note         = {Machine review of arXiv:2412.09001}
}
read the original abstract

Programming has become an essential component of K-12 education and serves as a pathway for developing computational thinking skills. Given the complexity of programming and the advanced skills it requires, previous research has introduced user-friendly tools to support young learners. However, our interviews with six programming educators revealed that current tools often fail to reflect classroom learning objectives, offer flexible, high-quality guidance, and foster student creativity. This highlights the need for more adaptive and reflective tools. Therefore, we introduced MindScratch, a multimodal generative AI (GAI) powered visual programming support tool. MindScratch aims to balance structured classroom activities with free programming creation, supporting students in completing creative programming projects based on teacher-set learning objectives while also providing programming scaffolding. Our user study results indicate that, compared to the baseline, MindScratch more effectively helps students achieve high-quality projects aligned with learning objectives. It also enhances students' computational thinking skills and creative thinking. Overall, we believe that GAI-driven educational tools like MindScratch offer students a focused and engaging learning experience.

Figures

Figures reproduced from arXiv: 2412.09001 by the authors.

Figure 1
Figure 1. MindScratch’s User Interface. In the mind map (a) with block palette (a4), each node represents the collaborative [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. MindScratch’s drawing board. It enables students [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. MindScratch user interaction process: Teachers set learning objectives and programming themes to prompt the LLM [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Overview of the logic-code assistance pipeline for MindScratch. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: The score distribution between MindScratch and [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: The bar plots illustrate the pre-test and post-test results of three students across five sub-dimensions of the CT skills [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: The score distribution between MindScratch and [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

80 extracted references · 59 canonical work pages

  1. [1]

    Teresa M Amabile. 1982. Social psychology of creativity: A consensual assessment technique. Journal of personality and social psychology 43, 5 (1982), 997

  2. [2]

    Brett A Becker, Paul Denny, James Finnie-Ansley, Andrew Luxton-Reilly, James Prather, and Eddie Antonio Santos. 2023. Programming is hard-or at least it used to be: Educational opportunities and challenges of ai code generation. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1. 500–506

  3. [3]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901

  4. [4]

    Xiaolin Chai, Yan Sun, and Yan Gao. 2023. Towards Data-Driving Multi-View Evaluation Framework for Scratch. Tsinghua Science and Technology 29, 2 (2023), 517–528

  5. [5]

    Liuqing Chen, Shuhong Xiao, Yunnong Chen, Ruoyu Wu, Yaxuan Song, and Lingyun Sun. 2024. ChatScratch: An AI-Augmented System Toward Autonomous Visual Programming Learning for Children Aged 6-12. In In: Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1–19. https: //doi.org/10.1145/3613904.3642229

  6. [6]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)

  7. [7]

    Erin Cherry and Celine Latulipe. 2014. Quantifying the creativity support of digital tools through the creativity support index.ACM Transactions on Computer- Human Interaction (TOCHI) 21, 4 (2014), 1–25

  8. [8]

    code.org

    code.org 2023. code.org. https://code.org/ [Accessed on Dec. 27, 2023]

Show all 80 references
  1. [9]

    Alina A Davier, Paul W Holland, and Dorothy T Thayer. 2004. The kernel method of test equating. Springer

  2. [10]

    Paul Denny, James Prather, Brett A Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N Reeves, Eddie Antonio Santos, and Sami Sarsa. 2024. Computing education in the era of generative AI.Commun. ACM 67, 2 (2024), 56–67

  3. [11]

    Griffin Dietz, Jimmy K Le, Nadin Tamer, Jenny Han, Hyowon Gweon, Eliza- beth L Murnane, and James A Landay. 2021. Storycoder: Teaching computational thinking concepts through storytelling in a voice-guided app for children. In Proceedings of the 2021 CHI Conference on Human Fa...

  4. [12]

    Griffin Dietz, Nadin Tamer, Carina Ly, Jimmy K Le, and James A Landay. 2023. Visual StoryCoder: A Multimodal Programming Environment for Children’s Creation of Stories. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–16

  5. [13]

    James Finnie-Ansley, Paul Denny, Brett A Becker, Andrew Luxton-Reilly, and James Prather. 2022. The robots are coming: Exploring the implications of openai codex on introductory programming. In Proceedings of the 24th Australasian Computing Education Conference. 10–19

  6. [14]

    James Finnie-Ansley, Paul Denny, Andrew Luxton-Reilly, Eddie Antonio Santos, James Prather, and Brett A Becker. 2023. My ai wants to know if this will be on the exam: Testing openai’s codex on cs2 programming exercises. In Proceedings of the 25th Australasian Computing Educati...

  7. [15]

    Emily Arteaga Garcia, João Felipe Pimentel, Zixuan Feng, Marco Gerosa, Igor Steinmacher, and Anita Sarma. 2023. How to Support ML End-User Programmers through a Conversational Agent. In 2024 IEEE/ACM 46th International Conference on Software Engineering (ICSE) . IEEE Computer ...

  8. [16]

    Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, and Soujanya Poria

  9. [17]

    Elena L Glassman, Lyla Fischer, Jeremy Scott, and Robert C Miller. 2015. Foobaz: Variable name feedback for student code at scale. InProceedings of the 28th Annual ACM Symposium on User Interface Software & Technology . 609–617

  10. [18]

    Michal Gordon, Eileen Rivera, Edith Ackermann, and Cynthia Breazeal. 2015. Designing a relational social robot toolkit for preschool children to explore computational concepts. In Proceedings of the 14th international conference on interaction design and children . 355–358

  11. [19]

    Hour of Code

    hourofcode 2023. Hour of Code. https://hourofcode.com/ [Accessed on Dec. 23, 2023]

  12. [20]

    Yu-Chang Hsu, Natalie Roote Irie, and Yu-Hui Ching. 2019. Computational thinking educational policy initiatives (CTEPI) across the globe. TechTrends 63 (2019), 260–270

  13. [21]

    Gwo-Jen Hwang, Shu-Yun Chien, and Wen-Shiang Li. 2021. A multidimensional repertory grid as a graphic organizer for implementing digital games to promote students’ learning performances and behaviors. British Journal of Educational Technology 52, 2 (2021), 915–933

  14. [22]

    Theresia Devi Indriasari, Andrew Luxton-Reilly, and Paul Denny. 2020. A re- view of peer code review in higher education. ACM Transactions on Computing Education (TOCE) 20, 3 (2020), 1–25

  15. [23]

    Peiling Jiang, Jude Rayan, Steven P Dow, and Haijun Xia. 2023. Graphologue: Ex- ploring large language model responses with interactive diagrams. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . 1–20. Conference’17, July 2017, Washingt...

  16. [24]

    Mukund Kapoor. 2023. Negative Prompts in Stable Diffusion: A Beginner’s Guide. https://www.greataiprompts.com/imageprompt/what-is-negative- prompt-in-stable-diffusion/ Accessed: 2023-10-25

  17. [25]

    Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, et al. 2023. ChatGPT for good? On opportunities and challenges of large language models for education. Learning an...

  18. [26]

    Majeed Kazemitabaar, Justin Chow, Carl Ka To Ma, Barbara J Ericson, David Weintrop, and Tovi Grossman. 2023. Studying the effect of ai code generators on supporting novice learners in introductory programming. In Proceedings of the 2023 CHI Conference on Human Factors in Compu...

  19. [27]

    Majeed Kazemitabaar, Xinying Hou, Austin Henley, Barbara Jane Ericson, David Weintrop, and Tovi Grossman. 2023. How novices use LLM-based code generators to solve CS1 coding tasks in a self-paced learning environment. In Proceedings of the 23rd Koli Calling International Confe...

  20. [28]

    Frank Klassner and Scott D Anderson. 2003. Lego MindStorms: Not just for K-12 anymore. IEEE robotics & automation magazine 10, 2 (2003), 12–18

  21. [29]

    Özgen Korkmaz and BAI Xuemei. 2019. Adapting computational thinking scale (CTS) for Chinese high school students and their thinking scale skills level. Participatory Educational Research 6, 1 (2019), 10–26

  22. [30]

    Anastasia Kovalkov, Benjamin Paaßen, Avi Segal, Niels Pinkwart, and Kobi Gal

  23. [31]

    Chang Liu, Loc Hoang, Andrew Stolman, and Bo Wu. 2024. HiTA: A RAG- Based Educational Platform that Centers Educators in the Instructional Loop. In International Conference on Artificial Intelligence in Education . Springer, 405–412

  24. [32]

    Kirsti Lonka, Juho Makkonen, Minna Berg, Markus Talvio, Erika Maksniemi, Milla Kruskopf, Heidi Lammassaari, Lauri Hietajärvi, and Suvi Krista Westling

  25. [33]

    Stephen MacNeil, Andrew Tran, Dan Mogil, Seth Bernstein, Erin Ross, and Ziheng Huang. 2022. Generating diverse code explanations using the gpt-3 large language model. In Proceedings of the 2022 ACM Conference on International Computing Education Research-Volume 2. 37–39

  26. [34]

    Orni Meerbaum-Salant, Michal Armoni, and Mordechai Ben-Ari. 2011. Habits of Programming in Scratch. In Proceedings of the 16th Annual Joint Conference on Innovation and Technology in Computer Science Education (Darmstadt, Germany) (ITiCSE ’11). Association for Computing Machin...

  27. [35]

    Orni Meerbaum-Salant, Michal Armoni, and Mordechai Ben-Ari. 2011. Habits of programming in scratch. In Proceedings of the 16th annual joint conference on Innovation and technology in computer science education . 168–172

  28. [36]

    Joseph Bahman Moghadam, Rohan Roy Choudhury, HeZheng Yin, and Armando Fox. 2015. AutoStyle: Toward coding style feedback at scale. In Proceedings of the Second (2015) ACM Conference on Learning@ Scale . 261–266

  29. [37]

    Jesús Moreno-León and Gregorio Robles. 2015. Dr. Scratch: A web tool to auto- matically evaluate Scratch projects. In Proceedings of the workshop in primary and secondary computing education . 132–133

  30. [38]

    Julie Mueller, Danielle Beckett, Eden Hennessey, and Hasan Shodiev. 2017. Assess- ing computational thinking across the curriculum. Emerging research, practice, and policy on computational thinking (2017), 251–267

  31. [39]

    Stable Diffusion Online. 2023. Stable Diffusion: A Latent Text-to-Image Diffusion Model. https://stablediffusionweb.com/ Accessed: 2023-09-08

  32. [40]

    Zhenhui Peng, Yuzhi Liu, Hanqi Zhou, Zuyu Xu, and Xiaojuan Ma. 2022. CReBot: Exploring interactive question prompts for critical paper reading. International Journal of Human-Computer Studies 167 (2022), 102898

  33. [41]

    Matei-Dan Popovici. 2023. ChatGPT in the classroom. Exploring its potential and limitations in a functional programming course. International Journal of Human–Computer Interaction (2023), 1–12

  34. [42]

    Dylan J Portelance and Marina Umaschi Bers. 2015. Code and Tell: Assessing young children’s learning of computational thinking using peer video interviews with ScratchJr. In Proceedings of the 14th international conference on interaction design and children. 271–274

  35. [43]

    JD Zamfirescu-Pereira Laryn Qi, Bjoern Hartmann, and John DeNero Narges Norouzi. 2023. Conversational Programming with LLM-Powered Interactive Support in an Introductory Computer Science Course. (2023)

  36. [44]

    Michael Reilly. 2013. The kindergarten coders. New scientist 2927 (2013), 21–22

  37. [45]

    Mitchel Resnick, John Maloney, Andrés Monroy-Hernández, Natalie Rusk, Evelyn Eastmond, Karen Brennan, Amon Millner, Eric Rosenbaum, Jay Silver, Brian Silverman, et al. 2009. Scratch: programming for all. Commun. ACM 52, 11 (2009), 60–67

  38. [46]

    Mitchel Resnick, John Maloney, Andrés Monroy-Hernández, Natalie Rusk, Evelyn Eastmond, Karen Brennan, Amon Millner, Eric Rosenbaum, Jay Silver, Brian Silverman, and Yasmin Kafai. 2009. Scratch: Programming for All. Commun. ACM 52, 11 (nov 2009), 60–67. https://doi.org/10.1145/...

  39. [47]

    Andoni Rivera-Pinto, Johan Kildal, and Elena Lazkano. 2023. Toward program- ming a collaborative robot by interacting with its digital twin in a mixed reality environment. International Journal of Human–Computer Interaction (2023), 1–13

  40. [48]

    Anthony Robins, Janet Rountree, and Nathan Rountree. 2003. Learning and teaching programming: A review and discussion. Computer science education 13, 2 (2003), 137–172

  41. [49]

    Margarida Romero, Therese Laferriere, and Thomas Michael Power. 2016. The move is on! From the passive multimedia learner to the engaged co-creator. ELearn 2016, 3 (2016)

  42. [50]

    Margarida Romero, Alexandre Lepage, and Benjamin Lille. 2017. Computational thinking development through creative programming in higher education. Inter- national Journal of Educational Technology in Higher Education 14 (2017), 1–15

  43. [51]

    Sherry Ruan, Liwei Jiang, Justin Xu, Bryce Joe-Kun Tham, Zhengneng Qiu, Yeshuang Zhu, Elizabeth L Murnane, Emma Brunskill, and James A Landay. 2019. Quizbot: A dialogue-based adaptive learning system for factual knowledge. In Proceedings of the 2019 CHI conference on human fac...

  44. [52]

    Sami Sarsa, Paul Denny, Arto Hellas, and Juho Leinonen. 2022. Automatic generation of programming exercises and code explanations using large language models. In Proceedings of the 2022 ACM Conference on International Computing Education Research-Volume 1. 27–43

  45. [53]

    Cynthia Selby and John Woollard. 2013. Computational thinking: the developing definition. (2013)

  46. [54]

    Karim Shabani, Mohamad Khatib, and Saman Ebadi. 2010. Vygotsky’s zone of proximal development: Instructional implications and teachers’ professional development. English language teaching 3, 4 (2010), 237–248

  47. [55]

    Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023. Role play with large language models. Nature 623, 7987 (2023), 493–498

  48. [56]

    Shute, Chen Sun, and Jodi Asbell-Clarke

    Valerie J. Shute, Chen Sun, and Jodi Asbell-Clarke. 2017. Demystifying com- putational thinking. Educational Research Review 22 (2017), 142–158. https: //doi.org/10.1016/j.edurev.2017.09.003

  49. [57]

    Simon, Judy Sheard, Michael Morgan, Andrew Petersen, Amber Settle, Jane Sinclair, Gerry Cross, and Charles Riedesel. 2016. Negotiating the maze of academic integrity in computing education. In Proceedings of the 2016 ITiCSE working group reports. 57–80

  50. [58]

    Rishabh Singh, Sumit Gulwani, and Armando Solar-Lezama. 2013. Automated feedback generation for introductory programming assignments. In Proceed- ings of the 34th ACM SIGPLAN conference on Programming language design and implementation. 15–26

  51. [59]

    Harry Stokhof, Bregje De Vries, Theo Bastiaens, and Rob Martens. 2020. Us- ing mind maps to make student questioning effective: Learning outcomes of a principle-based scenario for teacher guidance. Research in Science Education 50, 1 (2020), 203–225

  52. [60]

    Sangho Suh, Jian Zhao, and Edith Law. 2022. Codetoon: Story ideation, auto comic generation, and structure mapping for code-driven storytelling. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology . 1–16

  53. [61]

    Xiaohang Tang, Xi Chen, Sam Wong, and Yan Chen. 2023. VizPI: A Real-Time Visualization Tool for Enhancing Peer Instruction in Large-Scale Programming Lectures. In Adjunct Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–3

  54. [62]

    Christina Tikva and Efthimios Tambouris. 2021. Mapping computational think- ing through programming in K-12 education: A conceptual model based on a systematic literature Review. Computers & Education 162 (2021), 104083

  55. [63]

    Joke Voogt, Petra Fisser, Jon Good, Punya Mishra, and Aman Yadav. 2015. Com- putational thinking in compulsory education: Towards an agenda for research and practice. Education and information technologies 20 (2015), 715–728

  56. [64]

    April Yi Wang, Yan Chen, John Joon Young Chung, Christopher Brooks, and Steve Oney. 2021. Puzzleme: Leveraging peer assessment for in-class programming exercises. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–24

  57. [65]

    April Yi Wang, Andrew Head, Ashley Ge Zhang, Steve Oney, and Christopher Brooks. 2023. Colaroid: A literate programming approach for authoring ex- plorable multi-stage tutorials. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–22

  58. [66]

    be a lighting programmer

    Xinyuan Wang, Qian Xing, Qiao Jin, and Danli Wang. 2024. “be a lighting programmer”: Supporting children collaborative learning through tangible pro- gramming system. International Journal of Human–Computer Interaction 40, 10 (2024), 2622–2640

  59. [67]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35 (2022), 24824–24837

  60. [68]

    Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022. Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts. In Proceedings of the 2022 CHI conference on human factors in computing systems . 1–22

  61. [69]

    Lixiang Yan, Lele Sha, Linxuan Zhao, Yuheng Li, Roberto Martinez-Maldonado, Guanliang Chen, Xinyu Li, Yueqiao Jin, and Dragan Gašević. 2024. Practical and ethical challenges of large language models in education: A systematic scoping MindScratch: A Visual Programming Support T...

  62. [70]

    Tzu-Chi Yang and Zhi-Shen Lin. 2024. Enhancing elementary school students’ computational thinking and programming learning with graphic organizers. Computers & Education 209 (2024), 104962

  63. [71]

    Jia-Yu Yao, Kun-Peng Ning, Zhen-Hui Liu, Mu-Nan Ning, and Li Yuan. 2023. Llm lies: Hallucinations are not bugs, but features as adversarial examples. arXiv preprint arXiv:2310.01469 (2023)

  64. [72]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2022. ReAct: Synergizing Reasoning and Acting in Language Models. In The Eleventh International Conference on Learning Representations

  65. [73]

    Kangyu Yuan, Hehai Lin, Shilei Cao, Zhenhui Peng, Qingyu Guo, and Xiaojuan Ma. 2023. CriTrainer: An Adaptive Training Tool for Critical Paper Reading. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–17

  66. [74]

    Ashley Ge Zhang, Yan Chen, and Steve Oney. 2023. Vizprog: Identifying misun- derstandings by visualizing students’ coding progress. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–16

  67. [75]

    Chao Zhang, Cheng Yao, Jiayi Wu, Weijia Lin, Lijuan Liu, Ge Yan, and Fangtian Ying. 2022. StoryDrawer: A Child–AI Collaborative Drawing System to Support Children’s Creative Visual Storytelling. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems . 1–15

  68. [76]

    Lvmin Zhang and Maneesh Agrawala. 2023. Adding Conditional Control to Text-to-Image Diffusion Models. arXiv:2302.05543 [cs.CV]

  69. [77]

    Li Zhao, Xiaohong Liu, Chenhui Wang, and Yu-Sheng Su. 2022. Effect of different mind mapping approaches on primary school students’ computational thinking skills during visual programming learning. Computers & Education 181 (2022), 104445. Conference’17, July 2017, Washington,...

  70. [2018]

    Edita, Finland

    Phenomenal learning from Finland . Edita, Finland

  71. [2021]

    IEEE Transactions on Learning Technologies 14, 6 (2021), 740–753

    Automatic creativity measurement in scratch programs across modalities. IEEE Transactions on Learning Technologies 14, 6 (2021), 740–753

  72. [2023]

    arXiv preprint arXiv:2304.13731 (2023)

    Text-to-Audio Generation using Instruction Tuned LLM and Latent Diffu- sion Model. arXiv preprint arXiv:2304.13731 (2023)

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.