Pith. sign in

REVIEW 3 major objections 5 minor 99 references

Understanding Student Perceptions, Mistakes, and Debugging Approaches when Solving Natural Language Programming Tasks

T0 review · 3 major / 5 minor · reviewed 2026-07-11 · grok-4.5

Pith's one-line read Novice programmers find natural-language prompting easier than coding, but mostly fail by omitting key details and then refine by clarifying intent rather than tracing code.

desk verdict Large-N CS1 baseline on dialogue Prompt Problems: students like it, omit key details, and mostly clarify rather than read code or tests. read the letter →

arxiv 2607.05034 v1 pith:2FNKDQDZ submitted 2026-07-06 cs.CY

classification cs.CY
keywords naturallanguageprogrammingcode-generatingAIPromptProblemsstudentperceptionsmistakesdebuggingstrategiesCS1cognitiveload
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper studies how more than 900 CS1 students solve computational tasks by writing natural-language prompts for a code-generating AI instead of writing code themselves. Students generally reported that these dialogue-based Prompt Problems felt easier, more enjoyable, and better aimed at problem-solving than traditional coding exercises, because syntax load was removed. The dominant mistakes in first unsuccessful prompts were omissions of required details such as function return type, argument names, required inputs, and expected output, consistent with over-reliance on the model to fill gaps. When generated code failed, students said they mainly clarified their intent and re-read the problem depiction, far less often tracing the code or inspecting test cases. The work supplies a concrete map of what novices leave out and how they try to recover, so that instructors can decide what scaffolding is still needed.

What carries the argument

Dialogue-based Prompt Problems: students see a visual input-output specification and iteratively prompt a code model (with no direct code editing) until the generated program passes the hidden tests; unsuccessful first prompts and optional reflections are then coded for missing elements and recovery strategies.

What would settle it

Re-run the same problems with a condition that forces students to mark or correct the generated code line-by-line (or that withholds the visual problem depiction after the first failure) and measure whether omission rates and reported strategy frequencies reverse.

Watch

Extended reading notes

Core claim

In a large CS1 deployment of dialogue-based Prompt Problems, students perceived natural-language prompting as easier, more enjoyable, and more focused on problem-solving than traditional coding; their most frequent initial-prompt errors were omissions of key specification details; and their reported recovery strategies centered on clarifying intent and re-examining the problem depiction rather than tracing generated code or examining test cases.

Load-bearing premise

The claim rests on treating coded absences of required elements in first prompts, plus self-reported recovery strategies from optional reflections, as faithful measures of the pedagogically important mistakes and debugging behaviors.

Editorial extensions

If this is right

  • Instructors can expect most first-prompt failures to be missing return types, argument names, and expected outputs, so early feedback can target those omissions.
  • Curriculum designers can treat prompt-writing as a lighter-load entry to problem-solving while still planning explicit practice in code tracing and test-case reading.
  • Tool builders can add completeness checks or progressive templates before code generation to reduce over-reliance on model inference.
  • Dialogue-based prompting will not automatically teach traditional debugging habits unless the interface or pedagogy requires attention to generated code and failing tests.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same omission pattern may appear in professional ‘vibe-coding’ workflows, suggesting a shared need for lightweight completeness scaffolding.
  • If cognitive offloading is the intended benefit, later studies should measure actual germane load and transfer to unaided coding rather than only self-reported ease.
  • Pairing Prompt Problems with short code-tracing micro-tasks after each failure could convert the dominant recovery strategy into one that also builds code-comprehension skill.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports a large-scale observational study of dialogue-based Prompt Problems in a CS1 C-programming course (N>900). Students wrote natural-language prompts to GPT-4o mini for six visual-spec tasks across two lab batches, then reflected on experience (Ref1–3) and recovery strategies (Ref4). Thematic analysis of 200 reflections and coding of 1,286 initially unsuccessful first prompts show that students generally perceived prompting as easier, more enjoyable, and better for problem-solving than traditional coding (Likert M=3.80); the most frequent mistakes were omissions of return type, argument names, required inputs, expected output, and functionality details (Table 3); and reported recovery focused more on clarifying intent and re-reading the problem depiction than on tracing generated code or inspecting test cases (Figure 5). Results are interpreted through Cognitive Load Theory as evidence of beneficial offloading of syntax.

Significance. If the descriptive findings hold, the work supplies the first systematic catalog of novice prompt-level omissions and self-reported refinement strategies for dialogue-based Prompt Problems, directly informing curriculum design, scaffolding (e.g., templates, Parsons-style prompt assembly), and tool features as GenAI becomes standard in CS1. Strengths include the large authentic enrollment sample, dual-batch design with success-rate and interaction logs (Tables 1–2), inductive thematic saturation at 200 responses, and exhaustive coding of half of all incorrect first prompts. These baselines are timely and actionable even without causal claims.

major comments (3)
  1. [§3.3, Table 3] §3.3 and Table 3: The mistake taxonomy treats absence of researcher-defined elements (return type, argument names, order, expected output, etc.) as errors even when the tool does not enforce names and the model can often infer them. This coding scheme is load-bearing for the claim that “the most common mistakes are related to the omission of key details,” yet no inter-rater reliability statistic is reported (only consensus discussion) and no ablation shows that these absences, rather than other prompt properties, actually caused the incorrect code. A sensitivity analysis or explicit justification against the visual-spec interface is needed.
  2. [§4.4, Figure 5] §4.4 and Figure 5 (RQ3): Strategy themes rest entirely on optional self-reports (n=174 coded) with no triangulation against the logged messages, code executions, or reset events. The central claim that students “focused more on clarifying their intent and reflecting on the provided problem details than on tracing generated code or examining test cases” therefore risks post-hoc rationalization or social-desirability bias. At minimum, a sample of log-validated trajectories should be reported or the claim explicitly scoped to “reported strategies.”
  3. [§5.1–5.2] §5.1–5.2: Cognitive Load Theory is invoked to interpret ease and offloading, yet no direct CLT measure (subjective or dual-task) is collected. The claim that prompting frees resources for problem-solving therefore remains an untested interpretive overlay rather than an empirical result; either add a brief CLT instrument or soften the causal language linking perceptions to germane load.
minor comments (5)
  1. [Table 2] Table 2 reports means for conversations/messages only among students with an initially incorrect prompt; a parallel column for all attempters would clarify selection effects.
  2. [Figure 4] Figure 4 stacks Ref1/Ref2 counts but does not report total unique respondents per theme; adding n or percentages would aid interpretation.
  3. [§3.1] §3.1: The exact system prompt / temperature settings for GPT-4o mini are not stated; reproducibility would benefit from a short appendix note.
  4. [§2.3] Several self-citations to prior Prompt Problems work appear; a brief sentence distinguishing the present dialogue-based contribution from the earlier zero-shot studies would help readers.
  5. [§3.1] Typo/consistency: “Prompt Programming” tool name vs. “Prompt Problems” activity; standardize early.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical observational study whose findings are induced from student data, not derived by construction from inputs or self-citation.

full rationale

This paper reports a large-scale CS1 deployment of dialogue-based Prompt Problems (N>900), thematic coding of reflections (Ref1–Ref4), and coding of missing elements in unsuccessful first prompts (Tables 2–4, §3.3–§4.4). The central claims—students found prompting easier/more enjoyable, most common mistakes were omissions of return type/argument names/inputs/expected output, and recovery strategies emphasized clarifying intent and re-reading the problem depiction—are descriptive frequencies and induced themes from the collected corpus. Cognitive Load Theory is used only as an interpretive lens (§2.5, §5), not as a derivation that forces the observed frequencies. Prior Prompt Problems citations (e.g., [14], [60], [68], [70]) supply pedagogical context and tool description; they do not supply the measured student outcomes or uniqueness claims that would make the present results tautological. There are no fitted parameters re-labeled as predictions, no self-definitional equations, no uniqueness theorems imported from overlapping authors, and no ansatz smuggled via citation. The study is transparent about its limits (no A/B, self-report strategies, researcher-coded omissions; §5.4). The derivation chain is simply ‘collect data → code → report frequencies/themes,’ which is self-contained against the paper’s own corpus and does not reduce by construction to its inputs.

Assumptions & free parameters 0 free parameters · 4 assumptions · 2 invented entities

Central claims rest on standard education-research practices plus a few domain choices: that missing researcher-defined specification elements count as mistakes, that GPT-4o mini dialogue without problem-spec access is a fair proxy for authentic GenAI workflows, and that CLT is an appropriate interpretive frame. No free parameters are fitted to produce the main percentages; invented constructs are analytic codes, not physical entities.

assumptions (4)
  • ad hoc to paper A complete successful prompt for these tasks should include function name, return type, argument names/order, required inputs, functionality explanation, and (for B1) expected output; absence of any is coded as a mistake.
    §3.3 and Table 3: researchers defined required elements from problem specs; argument names coded as errors even though the autograder did not strictly require exact names.
  • domain assumption Cognitive Load Theory explains why removing syntax should free working memory for problem-solving and why over-trust can produce harmful offloading.
    §2.5 and Discussion: CLT is the interpretive lens for ease perceptions and omission patterns; not independently measured via cognitive-load instruments.
  • domain assumption GPT-4o mini responses conditioned only on student dialogue (not the hidden problem statement) adequately represent dialogue-based code-generating assistants for CS1 tasks.
    §3.1: model choice and no-spec-to-model design; results may shift with stronger models or different system prompts.
  • domain assumption Optional written reflections and coded subsets (200 reflections; ~half of incorrect first prompts) are representative enough for thematic and frequency claims.
    §3.3–4: saturation claimed by response 160; selection into optional Ref4 and incomplete B2 attempts may bias strategy themes.
invented entities (2)
  • Mistake taxonomy of missing prompt elements (return type, argument names, etc.)
    purpose: Operationalize what counts as a prompt-level error for frequency and correlation analysis.
    Inductively refined codes specific to this study’s problems; independent_evidence is false because the taxonomy is defined for this dataset, though it is falsifiable on new prompt corpora.
  • Strategy themes for prompt refinement (clarify intent, reflect on depiction, trace code, etc.)
    purpose: Summarize how students report recovering from incorrect AI code.
    Thematic codes from Ref4; useful analytic categories, not independently measured behavioral constructs outside self-report.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Understanding Student Perceptions, Mistakes, and Debugging Approaches when Solving Natural Language Programming Tasks." pith.science (2026). https://pith.science/paper/2FNKDQDZ

@misc{pith2026260705034,
  author       = {Pith},
  title        = {Pith review of: Understanding Student Perceptions, Mistakes, and Debugging Approaches when Solving Natural Language Programming Tasks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2FNKDQDZ}},
  note         = {Machine review of arXiv:2607.05034}
}
read the original abstract

Learning to communicate with code-generating AI models is an emerging skill for novice programmers. One recent pedagogical approach, Prompt Problems, has students solve computational tasks by writing natural-language prompts for code-generating AI models. However, little is known about the specific prompt-level mistakes novice programmers make, the kinds of computational details they fail to communicate, and what strategies they use to recover when generated code is incorrect. In a CS1 course, we studied attempts by more than 900 students to solve dialogue-based Prompt Problems. We analyzed student reflections, unsuccessful prompts, and reported debugging strategies. Compared to traditional coding tasks, students generally found prompting easier, more enjoyable, and better targeted at developing problem-solving skills. The most common mistakes are related to the omission of key details, suggesting both a failure to acknowledge their importance and over-reliance on AI to infer them. When prompts failed, students focused more on clarifying their intent and reflecting on the provided problem details than on tracing generated code or examining test cases.

Figures

Figures reproduced from arXiv: 2607.05034 by the authors.

Figure 1
Figure 1. Illustration of a student’s iterative refinement process while successfully solving a Prompt Problem. (a) presents the [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Illustration of a student’s iterative refinement process while successfully solving a Prompt Problem. (a) presents [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Illustration of four problems used in our study, with the remaining two shown in Figures 1a and 2a. Problems in [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Themes identified in students’ responses to Ref1 and Ref2. Bars show the number of coded occurrences for each theme [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Strategies reported to refine prompts following unsuccessful attempts on the second batch of problems (B2) ( [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

99 extracted references · 5 linked inside Pith

  1. [1]

    Cagla Acun and Ramazan Acun. 2023. GAI-Enhanced Assignment Framework: A Case Study on Generative AI Powered History Education. InNeurIPS’23 Workshop on Generative AI for Education

  2. [2]

    Umair Z Ahmed, Shubham Sahai, Ben Leong, and Amey Karkare. 2025. Feasibility Study of Augmenting Teaching Assistants with AI for CS1 Programming Feed- back. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  3. [3]

    Matin Amoozadeh, Daye Nam, Daniel Prol, Ali Alfageeh, James Prather, Michael Hilton, Sruti Srinivasa Ragavan, and Amin Alipour. 2024. Student-AI Interaction: A Case Study of CS1 Students. InProceedings of the Koli Calling International Conference on Computing Education Research (Koli Calling)

  4. [4]

    Sushmita Azad, Binglin Chen, Maxwell Fowler, Matthew West, and Craig B. Zilles

  5. [5]

    InProceedings of the International Conference on Artificial Intelligence in Education (AIED)

    Strategies for Deploying Unreliable AI Graders in High-Transparency High-Stakes Exams. InProceedings of the International Conference on Artificial Intelligence in Education (AIED)

  6. [6]

    Suma Bailis, Lara McConnaughey, Jane Friedhoff, Feiyang Chen, Chase Adams, and Jacob Moon. 2023. WordPlay: An Agent Framework for Language Learning Games. InNeurIPS’23 Workshop on Generative AI for Education

  7. [7]

    Like a Nesting Doll

    Seth Bernstein, Paul Denny, Juho Leinonen, Lauren Kan, Arto Hellas, Matt Little- field, Sami Sarsa, and Stephen MacNeil. 2024. "Like a Nesting Doll": Analyzing Recursion Analogies Generated by CS Students Using Large Language Models. In Proceedings of the Conference on Innovation and Technology in Computer Science Education (ITiCSE)

  8. [8]

    Virginia Braun and Victoria Clarke. 2022. Conceptual and Design Thinking for Thematic Analysis.Qualitative psychology9, 1 (2022)

Show all 99 references
  1. [9]

    Neil C. C. Brown, Pierre Weill-Tessier, Juho Leinonen, Paul Denny, and Michael Kölling. 2025. Howzat? Appealing to Expert Judgement for Evaluating Human and AI Next-Step Hints for Novice Programmers.ACM Trans. Comput. Educ.25 (2025)

  2. [10]

    Santos, and Matthias Hauswirth

    Luca Chiodini, Igor Moreno Santos, Andrea Gallidabino, Anya Tafliovich, André L. Santos, and Matthias Hauswirth. 2021. A Curated Inventory of Programming Language Misconceptions. InProceedings of the Conference on Innovation and Technology in Computer Science Education (ITiCSE)

  3. [11]

    Victoria Clarke and Virginia Braun. 2014. Thematic Analysis. InEncyclopedia of critical psychology

  4. [12]

    Patricia de Oliveira Santos et al. 2024. Impacts of the Usage of Generative Artificial Intelligence on Software Development Process. InProceedings of the Brazilian Symposium on Information Systems (SBSI)

  5. [13]

    Heffernan, Tanja Käser, Steven Moore, Anna N

    Paul Denny, Sumit Gulwani, Neil T. Heffernan, Tanja Käser, Steven Moore, Anna N. Rafferty, and Adish Singla. 2024. Generative AI for Education (GAIED): Understanding Students’ Experiences with Natural Language Programming Tasks ICER 2026 Vol. 1, August 11–14, 2026, Uppsala, Sw...

  6. [14]

    Paul Denny, Viraj Kumar, and Nasser Giacaman. 2023. Conversing with Copilot: Exploring Prompt Engineering for Solving CS1 Problems Using Natural Lan- guage. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  7. [15]

    Becker, and Brent N

    Paul Denny, Juho Leinonen, James Prather, Andrew Luxton-Reilly, Thezyrie Amarouche, Brett A. Becker, and Brent N. Reeves. 2024. Prompt Problems: A New Programming Exercise for the Generative AI Era. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  8. [16]

    Paul Denny, Andrew Luxton-Reilly, and Beth Simon. 2008. Evaluating a New Exam Question: Parsons Problems. InProceedings of the Conference on Interna- tional Computing Education Research (ICER)

  9. [17]

    Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N

    Paul Denny, James Prather, Brett A. Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N. Reeves, Eddie Antonio Santos, and Sami Sarsa. 2024. Computing Education in the Era of Generative AI.Commun. ACM(2024)

  10. [18]

    Smith, Max Fowler, James Prather, Brett A

    Paul Denny, David H. Smith, Max Fowler, James Prather, Brett A. Becker, and Juho Leinonen. 2024. Explaining Code with a Purpose: An Integrated Approach for Developing Code Comprehension and Prompting Skills. InProceedings of the Conference on Innovation and Technology in Compu...

  11. [19]

    Rodrigo Duran, Albina Zavgorodniaia, and Juha Sorva. 2022. Cognitive Load Theory in Computing Education Research: a Review.ACM Transactions on Computing Education22, 4 (2022)

  12. [20]

    Christof Ebert and Panos Louridas. 2023. Generative AI for Software Practitioners. IEEE Software(2023)

  13. [21]

    Ericson et al

    Barbara J. Ericson et al. 2022. Parsons Problems and Beyond: Systematic Literature Review and Empirical Study Designs. InProceedings of the Working Group Reports of the Conference on Innovation and Technology in Computer Science Education (ITiCSE)

  14. [22]

    Ericson, Lauren E

    Barbara J. Ericson, Lauren E. Margulieux, and Jochen Rick. 2017. Solving Parsons Problems versus Fixing and Writing Code. InProceedings of the Koli Calling International Conference on Computing Education Research (Koli Calling)

  15. [23]

    Wünsche, and Paul Denny

    Tony Haoran Feng, Andrew Luxton-Reilly, Burkhard C. Wünsche, and Paul Denny. 2025. From Automation to Cognition: Redefining the Roles of Educators and Generative AI in Computing Education. InProceedings of the Australasian Computing Education Conference (ACE)

  16. [24]

    Fernandez and Kimberly A

    Amanda S. Fernandez and Kimberly A. Cornell. 2024. CS1 with a Side of AI: Teaching Software Verification for Secure Code in the Era of Generative AI. In Proceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  17. [25]

    Explain in Plain English

    Max Fowler, Binglin Chen, Sushmita Azad, Matthew West, and Craig Zilles. 2021. Autograding "Explain in Plain English" Questions Using NLP. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  18. [26]

    Michael Gerlich. 2025. AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking.Societies15, 1 (2025)

  19. [27]

    Kafura, and Jeremy Ernst

    Luke Gusukuma, Austin Cory Bart, Dennis G. Kafura, and Jeremy Ernst. 2018. Misconception-Driven Feedback: Results from an Experimental Study. InPro- ceedings of the Conference on International Computing Education Research (ICER)

  20. [28]

    Mohammed Hassan, Grace Zeng, and Craig B. Zilles. 2024. Evaluating How Novices Utilize Debuggers and Code Execution to Understand Code. InProceed- ings of the Conference on International Computing Education Research (ICER)

  21. [29]

    Haynes and Barbara J

    Carl C. Haynes and Barbara J. Ericson. 2021. Problem-Solving Efficiency and Cognitive Load for Adaptive Parsons Problems vs. Writing the Equivalent Code. InProceedings of the Conference on Human Factors in Computing Systems (CHI)

  22. [30]

    Kaczmarczyk, Elizabeth R

    Lisa C. Kaczmarczyk, Elizabeth R. Petrick, J. Philip East, and Geoffrey L. Herman

  23. [31]

    InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

    Identifying Student Misconceptions of Programming. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  24. [32]

    Alan C. Kay. 1981. Generic Programming: APL and Smalltalk.SIGAPL APL Quote Quad12, 1 (1981)

  25. [33]

    Majeed Kazemitabaar, Runlong Ye, Xiaoning Wang, Austin Zachary Henley, Paul Denny, Michelle Craig, and Tovi Grossman. 2024. CodeAid: Evaluating a Classroom Deployment of an LLM-based Programming Assistant that Balances Student and Educator Needs. InProceedings of the Conferenc...

  26. [34]

    Smith, James Prather, Juho Leinonen, Andrew Luxton-Reilly, and Stephen MacNeil

    Chris Kerslake, Paul Denny, David H. Smith, James Prather, Juho Leinonen, Andrew Luxton-Reilly, and Stephen MacNeil. 2024. Integrating Natural Language Prompting Tasks in Introductory Programming Courses. InProceedings of the Virtual Global Computing Education Conference (SIGC...

  27. [35]

    Hassan Khosravi et al. 2026. Building AI Companions that Prioritise Learning over Performance.CoRRabs/2605.04816 (2026)

  28. [36]

    Unggi Lee et al. 2023. Generative Agent for Teacher Training: Designing Educa- tional Problem-Solving Simulations with Large Language Model-based Agents for Pre-Service Teachers. InNeurIPS’23 Workshop on Generative AI for Education

  29. [37]

    Juho Leinonen, Paul Denny, Stephen MacNeil, Sami Sarsa, Seth Bernstein, Joanne Kim, Andrew Tran, and Arto Hellas. 2023. Comparing Code Explanations Created by Students and Large Language Models. InProceedings of the Conference on Innovation and Technology in Computer Science E...

  30. [38]

    Reeves, Paul Denny, James Prather, and Brett A

    Juho Leinonen, Arto Hellas, Sami Sarsa, Brent N. Reeves, Paul Denny, James Prather, and Brett A. Becker. 2023. Using Large Language Models to Enhance Programming Error Messages. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  31. [39]

    Zilles, and Karrie Karahalios

    Tiffany Wenting Li, Silas Hsu, Max Fowler, Zhilin Zhang, Craig B. Zilles, and Karrie Karahalios. 2023. Am I Wrong, or Is the Autograder Wrong? Effects of AI Grading Mistakes on Learning. InProceedings of the Conference on International Computing Education Research (ICER)

  32. [40]

    Rongxin Liu, Carter Zenke, Charlie Liu, Andrew Holmes, Patrick Thornton, and David J. Malan. 2024. Teaching CS50 with AI: Leveraging Generative Artifi- cial Intelligence in Computer Science Education. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  33. [41]

    Suqing Liu, Zezhu Yu, Feiran Huang, Yousef Bulbulia, Andreas Bergen, and Michael Liut. 2024. Can Small Language Models With Retrieval-Augmented Generation Replace Large Language Models When Learning Computer Science?. InProceedings of the Conference on Innovation and Technolog...

  34. [42]

    Kuang-Chen Lu and Shriram Krishnamurthi. 2024. Identifying and Correcting Programming Language Behavior Misconceptions.Proceedings of the ACM on Programming Languages(2024)

  35. [43]

    Qianou Christina Ma, Sherry Tongshuang Wu, and Ken Koedinger. 2023. Is AI the Better Programming Partner? Human-Human Pair Programming vs. Human-AI pAIr Programming. InAIED Workshop on Empowering Education with LLMs

  36. [44]

    Stephen MacNeil, Paul Denny, Andrew Tran, Juho Leinonen, Seth Bernstein, Arto Hellas, Sami Sarsa, and Joanne Kim. 2024. Decoding Logic Errors: A Comparative Study on Bug Detection by Students and Large Language Models. InProceedings of the Australasian Computing Education Conf...

  37. [45]

    Stephen MacNeil, Zijian Ding, Kexin Quan, Thomas j Parashos, Yajie Sun, and Steven P Dow. 2021. Framing Creative Work: Helping Novices Frame Better Problems Through Interactive Scaffolding. InProceedings of the Conference on Crativity and Cognition (C&C)

  38. [46]

    Stephen MacNeil, Andrew Tran, Arto Hellas, Joanne Kim, Sami Sarsa, Paul Denny, Seth Bernstein, and Juho Leinonen. 2023. Experiences from Using Code Expla- nations Generated by Large Language Models in a Web Software Development E-Book. InProceedings of the Technical Symposium ...

  39. [47]

    Lauren Margulieux, James Prather, and Masoumeh Rahimi. 2025. The Biological Benefits of Failure on Learning and Tools to Manage the Fallout.Educational Psychology Review37, 2 (2025)

  40. [48]

    Markel, Steven G

    Julia M. Markel, Steven G. Opferman, James A. Landay, and Chris Piech. 2023. GPTeach: Interactive TA Training with GPT-based Students. InProceedings of the Conference on Learning @ Scale (L@S)

  41. [49]

    Raina Mason, Simon, Graham Cooper, and Barry Wilks. 2016. Flipping the Assessment of Cognitive Load: Why and How. InProceedings of the Conference on International Computing Education Research (ICER)

  42. [50]

    Nora McDonald, Sarita Schoenebeck, and Andrea Forte. 2019. Reliability and Inter-rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice.Proceedings of the ACM on Human-Computer Interaction(2019)

  43. [51]

    Carolina Mega, Lucia Ronconi, and Rossana De Beni. 2014. What Makes a Good Student? How Emotions, Self-Regulated Learning, and Motivation Contribute to Academic Achievement.Journal of educational psychology106, 1 (2014)

  44. [52]

    Morrison, Lauren E

    Briana B. Morrison, Lauren E. Margulieux, Barbara Ericson, and Mark Guzdial

  45. [53]

    InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

    Subgoals Help Students Solve Parsons Problems. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  46. [54]

    Laurie Murphy, Renée McCauley, and Sue Fitzgerald. 2012. ’Explain in Plain English’ Questions: Implications for Teaching. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  47. [55]

    Musser and Alexander A

    David R. Musser and Alexander A. Stepanov. 1988. Generic Programming. InPro- ceedings of the International Symposium on Symbolic and Algebraic Computation (ISSAC)

  48. [56]

    Daye Nam, Ahmed Omran, Ambar Murillo, Saksham Thakur, Abner Araujo, Marcel Blistein, Alexander Frömmgen, Vincent Hellendoorn, and Satish Chan- dra. 2025. Prompting LLMs for Code Editing: Struggles and Remedies.CoRR abs/2504.20196 (2025)

  49. [57]

    Manh Hung Nguyen, Sebastian Tschiatschek, and Adish Singla. 2024. Large Language Models for In-Context Student Modeling: Synthesizing Student’s Be- havior in Visual Programming from One-Shot Observation. InProceedings of the International Conference on Educational Data Mining (EDM)

  50. [58]

    Sydney Nguyen, Hannah McLean Babe, Yangtian Zi, Arjun Guha, Carolyn Jane Anderson, and Molly Q. Feldman. 2024. How Beginning Programmers and Code LLMs (Mis)read Each Other. InProceedings of the Conference on Human Factors in Computing Systems (CHI)

  51. [59]

    Olney, and Vasile Rus

    Priti Oli, Rabin Banjade, Andrew M. Olney, and Vasile Rus. 2024. Can LLMs Identify Gaps and Misconceptions in Students’ Code Explanations?CoRR abs/2501.10365 (2024)

  52. [60]

    OpenAI. 2024. GPT-4o mini: Advancing Cost-efficient Intelligence. https://openai. com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/

  53. [61]

    OpenAI. 2024. Introducing canvas. https://openai.com/index/introducing- canvas/. ICER 2026 Vol. 1, August 11–14, 2026, Uppsala, Sweden Victor-Alexandru Pădurean et al

  54. [62]

    Kim Ouwehand, Avalon van der Kroef, Jacqueline Wong, and Fred Paas. 2021. Measuring Cognitive Load: Are There More Valid Alternatives to Likert Rating Scales?Frontiers in Education6 (2021)

  55. [63]

    Victor-Alexandru Padurean, Paul Denny, Alkis Gotovos, and Adish Singla. 2025. Prompt Programming: A Platform for Dialogue-based Computational Problem Solving with Generative AI Models. InProceedings of the Conference on Innovation and Technology in Computer Science Education (ITiCSE)

  56. [64]

    Reinhard Pekrun. 2006. The Control-Value Theory of Achievement Emotions: Assumptions, Corollaries, and Implications for Educational Research and Practice. Educational psychology review18, 4 (2006)

  57. [65]

    Reinhard Pekrun, Stephanie Lichtenfeld, Herbert W Marsh, Kou Murayama, and Thomas Goetz. 2017. Achievement Emotions and Academic Performance: Longitudinal Models of Reciprocal Effects.Child development88, 5 (2017)

  58. [66]

    Reinhard Pekrun and Lisa Linnenbrink-Garcia. 2014. Introduction to Emotions in Education. InInternational handbook of emotions in education

  59. [67]

    Ji-Lun Peng and Su-Ling Yeh. 2025. Cognitive Offloading in Short-Term Memory Tasks: Trust Toward Tools as a Moderator.International Journal of Human– Computer Interaction(2025)

  60. [68]

    Tung Phung, José Cambronero, Sumit Gulwani, Tobias Kohn, Rupak Majumdar, Adish Singla, and Gustavo Soares. 2023. Generating High-Precision Feedback for Programming Syntax Errors using Large Language Models. InProceedings of the International Conference on Educational Data Mining (EDM)

  61. [69]

    James Prather et al. 2023. The Robots Are Here: Navigating the Generative AI Revolution in Computing Education. InProceedings of the Working Group Reports of the Conference on Innovation and Technology in Computer Science Education (ITiCSE)

  62. [70]

    James Prather et al. 2024. Beyond the Hype: A Comprehensive Review of Current Trends in Generative AI Research, Teaching Practices, and Tools. InProceedings of the Working Group Reports of the Conference on Innovation and Technology in Computer Science Education (ITiCSE)

  63. [71]

    Smith IV, Brent N

    James Prather, Paul Denny, Juho Leinonen, David H. Smith IV, Brent N. Reeves, Stephen MacNeil, Brett A. Becker, Andrew Luxton-Reilly, Thezyrie Amarouche, and Bailey Kimmel. 2024. Interactions with Prompt Problems: a New Way to Teach Programming with Large Language Models.CoRRa...

  64. [72]

    It’s Weird That it Knows What I Want

    James Prather, Brent N. Reeves, Paul Denny, Brett A. Becker, Juho Leinonen, Andrew Luxton-Reilly, Garrett B. Powell, James Finnie-Ansley, and Eddie Antonio Santos. 2023. "It’s Weird That it Knows What I Want": Usability and Interactions with Copilot for Novice Programmers.ACM ...

  65. [73]

    James Prather, Brent N. Reeves, Paul Denny, Juho Leinonen, Stephen MacNeil, Andrew Luxton-Reilly, João Orvalho, Amin Alipour, Ali Alfageeh, Thezyrie Amarouche, Bailey Kimmel, Jared Wright, Musa Blake, and Gweneth Barbre

  66. [74]

    InProceedings of the Australasian Computing Education Conference (ACE)

    Breaking the Programming Language Barrier: Multilingual Prompting to Empower Non-Native English Learners. InProceedings of the Australasian Computing Education Conference (ACE)

  67. [75]

    Reeves, Juho Leinonen, Stephen MacNeil, Arisoa S

    James Prather, Brent N. Reeves, Juho Leinonen, Stephen MacNeil, Arisoa S. Ran- drianasolo, Brett A. Becker, Bailey Kimmel, Jared Wright, and Ben Briggs. 2024. The Widening Gap: The Benefits and Harms of Generative AI for Novice Pro- grammers. InProceedings of the Conference on...

  68. [76]

    Yizhou Qian and James Lehman. 2017. Students’ Misconceptions and Other Difficulties in Introductory Programming: A Literature Review.ACM Transactions on Computing Education(2017)

  69. [77]

    Sara Rahimi et al . 2024. Saturation in Qualitative Research: an Evolutionary Concept Analysis.International Journal of Nursing Studies Advances6 (2024)

  70. [78]

    Ruchit Rawal, Victor-Alexandru Padurean, Sven Apel, Adish Singla, and Mariya Toneva. 2025. Hints Help Finding and Fixing Bugs Differently in Python and Text- based Program Representations. InProceedings of the International Conference on Software Engineering (ICSE)

  71. [79]

    Lauren L Richmond and Ryan G Taylor. 2025. The Benefits and Potential Costs of Cognitive Offloading for Retrospective Information.Nature Reviews Psychology (2025)

  72. [80]

    Kaitlin Riegel. 2021. Frustration in Mathematical Problem-Solving: a Systematic Review of Research.STEM Education1, 3 (2021)

  73. [81]

    Liam Rigby, Paul Denny, and Andrew Luxton-Reilly. 2020. A Miss is as Good as a Mile: Off-By-One Errors and Arrays in an Introductory Programming Course. InProceedings of the Australasian Computing Education Conference (ACE)

  74. [82]

    Risko and Sam J

    Evan F. Risko and Sam J. Gilbert. 2016. Cognitive Offloading.Trends in Cognitive Sciences20, 9 (2016)

  75. [83]

    Sami Sarsa, Paul Denny, Arto Hellas, and Juho Leinonen. 2022. Automatic Gen- eration of Programming Exercises and Code Explanations Using Large Language Models. InProceedings of the Conference on International Computing Education Research (ICER)

  76. [84]

    Benjamin Saunders, Julius Sim, Tom Kingstone, Shula Baker, Jackie Waterfield, Bernadette Bartlam, Heather Burroughs, and Clare Jinks. 2018. Saturation in Qualitative Research: Exploring Its Conceptualization and Operationalization. Quality & quantity52, 4 (2018)

  77. [85]

    Sauvola, Sasu Tarkoma, Mika Klemettinen, Jukka Riekki, and David S

    Jaakko J. Sauvola, Sasu Tarkoma, Mika Klemettinen, Jukka Riekki, and David S. Doermann. 2024. Future of Software Development with Generative AI.Automated Software Engineering(2024)

  78. [86]

    Mitchell

    Robin Schmucker, Meng Xia, Amos Azaria, and Tom M. Mitchell. 2024. Ruffle &Riley: Insights from Designing and Evaluating a Large Language Model-Based Conversational Tutoring System. InProceedings of the International Conference on Artificial Intelligence in Education (AIED)

  79. [87]

    Stanislaw Schukajlow, Katrin Rakoczy, and Reinhard Pekrun. 2017. Emotions and Motivation in Mathematics Education: Theoretical Considerations and Empirical Contributions.ZDM49, 3 (2017)

  80. [88]

    Smith, Paul Denny, and Max Fowler

    David H. Smith, Paul Denny, and Max Fowler. 2024. Prompting for Compre- hension: Exploring the Intersection of Explain in Plain English Questions and Prompt Writing. InProceedings of the Conference on Learning @ Scale (L@S)

  81. [89]

    Fangchen Song, Ashish Agarwal, and Wen Wen. 2024. The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot.CoRRabs/2410.02091 (2024)

  82. [90]

    John Sweller. 1988. Cognitive Load during Problem Solving: Effects on Learning. Cognitive Science12, 2 (1988)

  83. [91]

    John Sweller. 2011. Cognitive Load Theory. InPsychology of learning and motivation. Vol. 55

  84. [92]

    Ellen L Usher and Frank Pajares. 2009. Sources of Self-Efficacy in Mathematics: a Validation Study.Contemporary educational psychology34, 1 (2009)

  85. [93]

    Smith IV, Mounika Padala, Chris- tine Alvarado, Jamie Gorson Benario, and Leo Porter

    Annapurna Vadaparty, Daniel Zingaro, David H. Smith IV, Mounika Padala, Chris- tine Alvarado, Jamie Gorson Benario, and Leo Porter. 2024. CS1-LLM: Integrating LLMs into CS1 Instruction. InProceedings of the Conference on Innovation and Technology in Computer Science Education (ITiCSE)

  86. [94]

    Glassman

    Priyan Vaithilingam, Tianyi Zhang, and Elena L. Glassman. 2022. Expectation vs. Experience: Evaluating the Usability of Code Generation Tools Powered by Large Language Models. InProceedings of the Conference on Human Factors in Computing Systems (CHI)

  87. [95]

    Mitchell, and Chris Piech

    Sierra Wang, John C. Mitchell, and Chris Piech. 2024. A Large Scale RCT on Effective Error Messages in CS1. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  88. [96]

    Whalley et al

    Jacqueline L. Whalley et al. 2006. An Australasian Study of Reading and Compre- hension Skills in Novice Programmers, Using the Bloom and SOLO Taxonomies. InProceedings of the Australasian Computing Education Conference (ACE)

  89. [97]

    Stephanie Yang, Miles Baird, Eleanor O’Rourke, Karen Brennan, and Bertrand Schneider. 2024. Decoding Debugging Instruction: A Systematic Literature Re- view of Debugging Interventions.ACM Transactions on Computing Education (2024)

  90. [98]

    Jialu Zhang, José Pablo Cambronero, Sumit Gulwani, Vu Le, Ruzica Piskac, Gus- tavo Soares, and Gust Verbruggen. 2024. PyDex: Repairing Bugs in Introductory Python Assignments using LLMs.Proceedings of the ACM on Programming Languages(2024)

  91. [99]

    Rina Zviel-Girshin. 2024. The Good and Bad of AI Tools in Novice Programming Education.Education Sciences(2024)

Pith tools

Reviewed July 11, 2026 · model on record in the stance chip above.