REVIEW 3 major objections 5 minor 55 references
Is Solving Better Than Evaluating GenAI Solutions?
T0 review · 3 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Evaluating GenAI solutions produces no measurable learning gain over direct problem solving.
desk verdict A well-designed crossover study finds no broad differences between evaluating GenAI solutions and solving problems in an algorithms course, but the null is only interpretable if the two arms were informationally isolated — and that isolation is unmeasured. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a randomized A/B crossover design at the level of self-reported working groups in a live algorithms course. Each of six homework sets contained a fourth exercise deliberately hard enough that a commercial LLM could not reliably solve it; one condition solved it directly, while the other generated a solution with a standardized prompt and graded it, distinguishing minor flaws from major flaws in the algorithm or proof. A mid-semester switch gave every group both treatments. The paper's key measurement device is the structurally aligned exam item: each exam included a problem whose correct solution required the same core techniques (sorting the input, then building a
What would settle it
Re-run the crossover with working groups isolated from each other, or measure and statistically control for cross-condition contact; if a difference between conditions appears on the structurally aligned exam items under isolation, the reported null result was an artifact of contamination, and if the difference stays absent, the equivalence conclusion is strengthened.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is a bounded null result: replacing direct problem solving with structured evaluation of GenAI-generated solutions changed neither the level nor the shape of summative learning. Across the midterm, the final, overall course percentages, normalized change scores, and letter-grade distributions, the two groups were statistically indistinguishable, and this held for exam problems whose solution required the same sorting-and-table construction insight as the homework intervention. The conditions did differ on the homework instrument itself — evaluation scores ran higher — and in survey reports, students who adapted their study strategies rated the
Load-bearing premise
The study assumes students in the two experimental conditions stayed cleanly separated, so that the solving and evaluating groups did not share homework content, AI outputs, or evaluative insights; if they did, the two arms would converge and the null result would be an artifact.
Editorial extensions
If this is right
- Algorithms instructors can substitute structured GenAI-evaluation exercises for some direct problem solving without measurable harm to exam scores or course grades.
- Because the homework advantage under evaluation did not transfer to exams, the higher homework scores likely reflect task or rubric differences, not deeper learning.
- Any real learning gains from GenAI evaluation will require deliberate scaffolding — for example, requiring students to identify the first invalid step, construct counterexamples, and propose corrected invariants.
- The null results bound both extreme intuitions: evaluation tasks are neither a dangerously weak substitute for solving nor an automatically superior metacognitive exercise.
- The study does not prove the two activities are equivalent; it was not designed or powered as a non-inferiority trial, so 'no evidence of difference' is not 'evidence of no difference.'
Reading between the lines
- If the two arms were contaminated — students in the solving condition obtaining AI outputs or evaluation insights from friends in the evaluation condition — the null result would reflect treatment diffusion rather than genuine pedagogical equivalence; a replication with physically or digitally isolated working groups could reveal true differences.
- The midterm's structurally aligned item had a pronounced floor effect (median 0 in both groups), so the transfer question was effectively answered only by the final-exam item; a more sensitive midterm probe might change the conclusion about transfer.
- The survey relationship between self-reported study-habit change and perceived helpfulness hints at a metacognitive mechanism worth testing: students who actively incorporate evaluation tasks into exam preparation may be the only ones who benefit, suggesting an individual-differences dimension to the intervention.
- The homework-score advantage for evaluation (larger in the group that evaluated in the second half) is confounded by non-identical rubrics and topic difficulty, so even the study's one positive result is weak evidence that evaluation is 'easier'; a rubric-matched comparison would be needed to confirm.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a randomized A/B crossover study (N=220) in a junior-level algorithms course comparing two homework conditions: solving a challenging algorithmic problem directly versus evaluating a flawed GenAI-generated solution to the same problem. Across six homework assignments, working groups experienced both conditions in counterbalanced order. The primary outcomes—midterm scores, final exam scores, overall course grades, and exam items structurally aligned with the homework interventions—showed no statistically significant group differences, with small effect sizes (d ≤ 0.13). Homework scores were higher during evaluation periods (d = 0.39, p = .035 by Mann-Whitney), but the authors caution that rubrics differed and clustering was ignored. Survey data indicated most students reported no change in study habits; those who did change rated the evaluation tasks as more helpful. The paper concludes that GenAI-evaluation activities can be incorporated without broad performance losses, but do not automatically improve transfer.
Significance. If the null results are taken at face value, the paper provides a useful, controlled data point in a literature dominated by capability studies and perception surveys: in a post-introductory, theory-heavy course, replacing some direct problem solving with structured evaluation of flawed GenAI output neither helps nor substantially hurts measured summative outcomes. The study's strengths include a randomized crossover design at the working-group level, a baseline equivalence check on shared homework items, preplanned nonparametric analyses with effect sizes, and unusually candid limitations sections (e.g., acknowledging differing rubrics for the homework outcome, floor effects on the midterm aligned item, and the distinction between 'no evidence of difference' and 'equivalence'). These features make the study a credible empirical contribution to computing education research, provided the contamination and framing concerns below are addressed.
major comments (3)
- [§3.2, §5.1] Cross-group contamination is a load-bearing threat that is asserted but never measured. The design groups self-reported collaborators into the same condition to 'minimize interactions across conditions' (§3.2), but no evidence is presented that the arms were informationally isolated: no survey item asked about cross-condition discussion, no network analysis of the self-reported groups is reported, and the ChatGPT share links are not audited for whether evaluation-condition students received fully worked solutions from solving-condition peers. In a live 220-student course, such diffusion is plausible, and if it occurred, the null results on midterm, final, and aligned-exam items would be expected regardless of whether evaluation is inferior, equivalent, or superior. The paper's §5.1 lists clustering, mechanism, artifact variation, topic drift, and outcome breadth as threats but omits cont
- [§4.7.1] The paper states 'we proceeded with our preregistered non-parametric tests,' but no preregistration is mentioned in the Methods section (§3.6) or anywhere else in the manuscript. There is no trial registry entry, preprint with timestamp, or supplementary preregistration document. This is a reproducibility and credibility concern. If the analysis was preregistered, the registration details (URL, date, and list of planned tests) should be provided; if it was merely preplanned, the word 'preregistered' should be removed to avoid implying a public audit trail that does not exist.
- [Abstract, §5.1] The abstract's conclusion that GenAI-evaluation activities 'can be incorporated into algorithms coursework without broad performance losses' is a non-inferiority claim, but the study was designed as a superiority/equivalence test and is not powered or structured to demonstrate 'no losses.' The authors themselves acknowledge in §5.1 that 'the absence of statistically significant differences should not be interpreted as proof that the two activities are equivalent' and that establishing equivalence requires an analysis designed for that purpose. The abstract and practical implications should be reworded to 'we found no evidence of broad performance losses' rather than implying a positive non-inferiority result, unless the authors add an equivalence or non-inferiority analysis with pre-specified bounds (e.g., two one-sided tests on the course-grade difference).
minor comments (5)
- [§3.2] Please report the number of self-reported working groups and the average/range of cluster sizes. This information is directly relevant to the clustering concern raised in §5.1 and would help readers assess the effective sample size for group-level randomization.
- [§4.5.3] The large discrepancy between the Mann-Whitney p-value (.035) and the t-test p-value (.004) for the GenAI-graded homework difference suggests sensitivity to outliers or distributional features. Consider reporting trimmed means, robust effect sizes, or a permutation test that respects the group-level randomization to support this positive finding.
- [§4.6.2] The within-subject Friedman tests on comfort/confidence across topics are tangential to the research questions. They are reported without an explicit connection to the intervention. Consider moving this analysis to supplementary material or shortening it to one sentence.
- [Table 1] The table is dense and the 'GenAI Group' / 'Non-GenAI Group' labeling switches between 'A' and 'B' depending on the homework block. Adding a column that states explicitly which group was the evaluator (A or B) for each row would improve readability.
- [Abstract / §3.4] The term 'LATEX' should be 'LaTeX' for consistency in the assignment instructions. Minor typographical issues also appear in the midterm problem statement ('such that (1) you maximize...' is grammatically awkward).
Circularity Check
No circularity in the central claim; the result is an empirical comparison with independent course-outcome measures, and the only self-citations are non-load-bearing related-work mentions.
full rationale
The paper's central claim is an empirical finding from a randomized A/B crossover experiment, not a derivation: working groups were randomized, baseline equivalence was checked, and then midterm scores, final exam scores, overall course percentages, normalized change scores, and structurally aligned exam items were compared between conditions. No parameter is fitted to the target outcome and later presented as a prediction; no outcome variable is defined in terms of the explanatory variable; and no uniqueness theorem, imported ansatz, or renamed benchmark is used to force the conclusions. The structurally aligned exam questions are independent summative assessments designed to mirror the reasoning of prior homework exercises, not recomputations of homework performance. The only self-citations are references [12] and [48] in the related-work sections, used as examples of GenAI instructional frameworks and instructor-facing GenAI platforms; neither bears on the validity of the null result or rules out alternative explanations. The paper also explicitly acknowledges that the two homework conditions used different grading rubrics and cautions against interpreting the homework-score advantage as evidence of a learning effect, so that comparison is not a fitted input disguised as a prediction. Potential cross-condition contamination is a genuine internal-validity concern, but it is an empirical limitation rather than a circularity: if contamination occurred, the conclusions would be confounded, not entailed by the study design.
Assumptions & free parameters
assumptions (5)
- domain assumption Self-reported working groups accurately identify each student's real collaboration network.
- domain assumption ChatGPT-4o, prompted with the standardized problem statement, produced solutions that were consistently incomplete or incorrect enough to require substantive evaluation.
- domain assumption Summative exam scores and course grades are valid, sufficiently sensitive measures of the conceptual learning the intervention targets.
- domain assumption Student outcomes can be treated as independent in the primary analyses despite randomization at working-group level.
- standard math Standard statistical test assumptions (normality, homoscedasticity) are adequately checked and handled.
Cite this review
Pith. "Pith review of Is Solving Better Than Evaluating GenAI Solutions?." pith.science (2026). https://pith.science/paper/ULXREOBM
@misc{pith2026260727586,
author = {Pith},
title = {Pith review of: Is Solving Better Than Evaluating GenAI Solutions?},
year = {2026},
howpublished = {\url{https://pith.science/paper/ULXREOBM}},
note = {Machine review of arXiv:2607.27586}
}
read the original abstract
As Generative AI (GenAI) tools become increasingly capable of generating solutions to computing assignments, the computing education community is exploring pedagogical approaches that emphasize solution evaluation, verification, and critique alongside traditional solution generation. However, evidence regarding the impact of such evaluation-centered tasks on student learning remains limited, particularly in upper-division, theory-heavy courses. We conducted a randomized A/B crossover study (N=220) in a junior-level algorithms course to compare evaluating GenAI-generated solutions with traditional problem solving. Across six assignments, student working groups either solved challenging algorithmic problems directly or evaluated often-flawed GenAI-generated solutions, with roles reversed midway through the semester. We found no statistically significant differences between groups in midterm scores, final exam scores, overall course grades, or exam problems structurally aligned with the homework interventions. Students received significantly higher homework scores when evaluating GenAI-generated solutions, but this localized advantage did not translate into downstream summative gains. Survey data further indicated that most students reported no change in study habits in response to the intervention; however, those who reported adapting their study strategies rated the GenAI-evaluation assignments as significantly more helpful. These findings suggest that GenAI evaluation redistributes student effort from open-ended solution construction toward verification, diagnosis, and judgment, but does not automatically produce stronger conceptual transfer. We conclude that GenAI-evaluation activities can be incorporated into algorithms coursework without broad performance losses, but meaningful learning gains may require deliberate scaffolding that pushes students beyond simple error diagnosis.
Reference graph
Works this paper leans on
-
[1]
Brett A Becker, Paul Denny, James Finnie-Ansley, Andrew Luxton-Reilly, James Prather, and Eddie Antonio Santos. 2023. Programming is hard-or at least it used to be: Educational opportunities and challenges of ai code generation. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1. 500–506
2023
-
[2]
Carlo Bellettini, Michael Lodi, Violetta Lonati, Mattia Monga, and Anna Morpurgo
-
[3]
Christian Bird, Denae Ford, Thomas Zimmermann, Nicole Forsgren, Eirini Kalliamvakou, Travis Lowdermilk, and Idan Gazit. 2022. Taking Flight with Copilot: Early insights and opportunities of AI-powered pair-programming tools. Queue20, 6 (2022), 35–57
2022
-
[4]
Julie L. Booth, Karin E. Lange, Kenneth R. Koedinger, and Kristie J. Newton. 2013. Using example problems to improve student learning in algebra: Differentiating between correct and incorrect examples.Learning and Instruction25 (2013), 24–34. doi:10.1016/j.learninstruc.2012.11.002
-
[5]
Dennis J. Bouvier, Bruno Pereira Cipriano, Richard Glassey, Olga Petrovska, Emma Anderson, Anastasiia Birillo, Ryan Dougherty, Raymond Pettit, Nuno Pombo, Ebrahim Rahimi, Charanya Ramakrishnan, Alexander Steinmaurer, Shubbhi Taneja, Muhammad Usman, and Annapurna Vadaparty. 2026. The Rest of the Robots: Generative AI in Post-introductory Computing Educatio...
arXiv 2026
-
[6]
Chi, Miriam Bassok, Matthew W
Michelene T.H. Chi, Miriam Bassok, Matthew W. Lewis, Peter Reimann, and Robert Glaser. 1989. Self-explanations: How students study and use examples in learning to solve problems.Cognitive Science13, 2 (1989), 145–182. doi:10.1016/ 0364-0213(89)90002-5
1989
-
[8]
Paul Denny, Viraj Kumar, and Nasser Giacaman. 2023. Conversing with copilot: Exploring prompt engineering for solving cs1 problems using natural language. InProceedings of the 54th ACM technical symposium on computer science education V. 1. 1136–1142
2023
-
[9]
Paul Denny, Juho Leinonen, James Prather, Andrew Luxton-Reilly, Thezyrie Amarouche, Brett A Becker, and Brent N Reeves. 2024. Prompt Problems: A new programming exercise for the generative AI era. InProceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1. 296–302
2024
Show all 55 references
-
[10]
Paul Denny, James Prather, Brett A Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N Reeves, Eddie Antonio Santos, and Sami Sarsa. 2024. Computing education in the era of generative AI.Commun. ACM67, 2 (2024), 56–67
2024
-
[11]
Paul Denny, David H Smith IV, Max Fowler, James Prather, Brett A Becker, and Juho Leinonen. 2024. Explaining code with a purpose: An integrated approach for developing code comprehension and prompting skills. InProceedings of the 2024 on Innovation and Technology in Computer S...
2024
-
[12]
Ethan Dickey, Andres Bejarano, and Chirayu Garg. 2024. AI-Lab: A Framework for Introducing Generative Artificial Intelligence Tools in Computer Programming Courses.SN Computer Science5, 6 (2024), 720. doi:10.1007/s42979-024-03074-y
2024 doi
-
[13]
Nikitha Donekal Chandrashekar, Sehrish Basir Nizamani, Margaret Ellis, and Naren Ramakrishnan. 2026. Demystify, Use, Reflect: Preparing students to be informed LLM-users. InProceedings of the 57th ACM Technical Symposium on Computer Science Education V.2(USA)(SIGCSE TS 2026). ...
2026
-
[14]
Kelley Durkin and Bethany Rittle-Johnson. 2012. The effectiveness of using incorrect examples to support learning about decimal magnitude.Learning and Instruction22, 3 (2012), 206–214. doi:10.1016/j.learninstruc.2011.11.001
2012 doi
-
[15]
Becker, Andrew Luxton-Reilly, and James Prather
James Finnie-Ansley, Paul Denny, Brett A. Becker, Andrew Luxton-Reilly, and James Prather. 2022. The Robots Are Coming: Exploring the Implica- tions of OpenAI Codex on Introductory Programming. InProceedings of the 24th Australasian Computing Education Conference(Virtual Event...
2022
-
[16]
James Finnie-Ansley, Paul Denny, Andrew Luxton-Reilly, Eddie Antonio Santos, James Prather, and Brett A Becker. 2023. My ai wants to know if this will be on the exam: Testing openai’s codex on cs2 programming exercises. InProceedings of the 25th Australasian Computing Educatio...
2023
-
[17]
Christopher Hundhausen, Anukrati Agrawal, and Kyle Ryan. 2010. The design of an online environment to support pedagogical code reviews. InProceedings of the 41st ACM Technical Symposium on Computer Science Education(Milwaukee, Wisconsin, USA)(SIGCSE ’10). Association for Compu...
2010
-
[18]
Hundhausen, Anukrati Agrawal, and Pawan Agarwal
Christopher D. Hundhausen, Anukrati Agrawal, and Pawan Agarwal. 2013. Talk- ing about code: Integrating pedagogical code reviews into early computing courses.ACM Trans. Comput. Educ.13, 3, Article 14 (Aug. 2013), 28 pages. doi:10.1145/2499947.2499951
2013
-
[19]
Theresia Devi Indriasari, Andrew Luxton-Reilly, and Paul Denny. 2020. A Review of Peer Code Review in Higher Education.ACM Trans. Comput. Educ.20, 3, Article 22 (Sept. 2020), 25 pages. doi:10.1145/3403935
2020 doi
-
[20]
Breanna Jury, Angela Lorusso, Juho Leinonen, Paul Denny, and Andrew Luxton- Reilly. 2024. Evaluating LLM-generated Worked Examples in an Introductory Programming Course. InProceedings of the 26th Australasian Computing Educa- tion Conference(Sydney, NSW, Australia)(ACE ’24). A...
2024
-
[21]
Majeed Kazemitabaar, Justin Chow, Carl Ka To Ma, Barbara J Ericson, David Weintrop, and Tovi Grossman. 2023. Studying the effect of AI code generators on supporting novice learners in introductory programming. InProceedings of the 2023 CHI conference on human factors in comput...
2023
-
[22]
Juho Leinonen, Paul Denny, Stephen MacNeil, Sami Sarsa, Seth Bernstein, Joanne Kim, Andrew Tran, and Arto Hellas. 2023. Comparing code explanations created by students and large language models. InProceedings of the 2023 Conference on Innovation and Technology in Computer Scie...
2023
-
[23]
Juho Leinonen, Arto Hellas, Sami Sarsa, Brent Reeves, Paul Denny, James Prather, and Brett A Becker. 2023. Using large language models to enhance programming error messages. InProceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1. 563–569
2023
-
[24]
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al. 2022. Competition-level code generation with alphacode.Science378, 6624 (2022), 1092–1097
2022
-
[25]
Jonathan Liu, Seth Poulsen, Erica Goodwin, Hongxuan Chen, Grace Williams, Yael Gertner, and Diana Franklin. 2025. Teaching Algorithm Design: A Literature Review.ACM Trans. Comput. Educ.25, 2, Article 17 (May 2025), 20 pages. doi:10. 1145/3727987
2025
-
[26]
Becker, Michelle Craig, Paul Denny, Raymond Pettit, and James Prather
Dastyni Loksa, Lauren Margulieux, Brett A. Becker, Michelle Craig, Paul Denny, Raymond Pettit, and James Prather. 2022. Metacognition and Self-Regulation in Programming Education: Theories and Exemplars of Use.ACM Trans. Comput. Educ.22, 4, Article 39 (Sept. 2022), 31 pages. d...
2022 doi
-
[27]
Reeves, Juho Leinonen, and Rachel Louise Rossetti
Stephen MacNeil, James Prather, Rahad Arman Nabid, Sebastian Gutierrez, Silas Carvalho, Saimon Shrestha, Paul Denny, Brent N. Reeves, Juho Leinonen, and Rachel Louise Rossetti. 2025. Fostering Responsible AI Use Through Negative Expertise: A Contextualized Autocompletion Quiz....
2025
-
[28]
Stephen MacNeil, Andrew Tran, Arto Hellas, Joanne Kim, Sami Sarsa, Paul Denny, Seth Bernstein, and Juho Leinonen. 2023. Experiences from using code explanations generated by large language models in a web software development e-book. InProceedings of the 54th ACM Technical Sym...
2023
-
[29]
Margulieux, James Prather, Brent N
Lauren E. Margulieux, James Prather, Brent N. Reeves, Brett A. Becker, Gozde Cetin Uzun, Dastyni Loksa, Juho Leinonen, and Paul Denny. 2024. Self- Regulation, Self-Efficacy, and Fear of Failure Interactions with How Novices Use LLMs to Solve Programming Problems. InProceedings...
2024
-
[30]
Marcus Messer, Neil C. C. Brown, Michael Kölling, and Miaojing Shi. 2024. Au- tomated Grading and Feedback Tools for Programming Education: A System- atic Review.ACM Trans. Comput. Educ.24, 1, Article 10 (Feb. 2024), 43 pages. doi:10.1145/3636515
2024 doi
-
[31]
Kasia Muldner, Jay Jennings, and Veronica Chiarelli. 2022. A Review of Worked Examples in Programming Activities.ACM Trans. Comput. Educ.23, 1, Article 13 (Dec. 2022), 35 pages. doi:10.1145/3560266
2022 doi
-
[32]
Daye Nam, Andrew Macvean, Vincent Hellendoorn, Bogdan Vasilescu, and Brad Myers. 2024. Using an llm to help with code understanding. InProceedings of the IEEE/ACM 46th International Conference on Software Engineering. 1–13
2024
-
[33]
Nicol and Debra Macfarlane-Dick
David J. Nicol and Debra Macfarlane-Dick. 2006. Formative assess- ment and self-regulated learning: a model and seven principles of good feedback practice.Studies in Higher Education31, 2 (2006), 199–218. arXiv:https://doi.org/10.1080/03075070600572090 doi:10.1080/03075070600572090
2006 doi
-
[34]
Michael Sheinman Orenstrakh, Oscar Karnalim, Carlos Anibal Suarez, and Michael Liut. 2024. Detecting LLM-generated text in computing education: Comparative study for ChatGPT cases. In2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC). IEEE, 121–126
2024
-
[35]
Eng Lieh Ouh, Benjamin Kok Siew Gan, Kyong Jin Shim, and Swavek Wlodkowski
-
[36]
Russell A Poldrack, Thomas Lu, and Gašper Beguš. 2023. AI-assisted coding: Experiments with GPT-4.arXiv preprint arXiv:2304.13187(2023)
2023 arXiv
-
[37]
InProceedings of the 2023 Conference on Innovation and Technology in Computer Science Education V
ChatGPT, Can You Generate Solutions for my Coding Exercises? An Evaluation on its Effectiveness in an undergraduate Java Programming Course.. InProceedings of the 2023 Conference on Innovation and Technology in Computer Science Education V. 1(Turku, Finland)(ITiCSE 2023). Asso...
2023
-
[38]
Becker, Ibrahim Albluwi, Michelle Craig, Hieke Keuning, Natalie Kiesler, Tobias Kohn, Andrew Luxton- Reilly, Stephen MacNeil, Andrew Petersen, Raymond Pettit, Brent N
James Prather, Paul Denny, Juho Leinonen, Brett A. Becker, Ibrahim Albluwi, Michelle Craig, Hieke Keuning, Natalie Kiesler, Tobias Kohn, Andrew Luxton- Reilly, Stephen MacNeil, Andrew Petersen, Raymond Pettit, Brent N. Reeves, and Jaromir Savelka. 2023. The Robots Are Here: Na...
2023
-
[39]
Becker, Arto Hellas, Paul Denny, and Brent N
Seth Poulsen, Sami Sarsa, James Prather, Juho Leinonen, Brett A. Becker, Arto Hellas, Paul Denny, and Brent N. Reeves. 2024. Solving Proof Block Prob- lems Using Large Language Models. InProceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1(Portlan...
2024
-
[40]
James Prather, Raymond Pettit, Kayla McMurry, Alani Peters, John Homer, and Maxine Cohen. 2018. Metacognitive difficulties faced by novice programmers in automated assessment tools. InProceedings of the 2018 ACM Conference on International Computing Education Research. 41–50
2018
-
[41]
Reeves, Jaromir Savelka, IV Smith, David H., Sven Strickroth, and Daniel Zingaro
James Prather, Juho Leinonen, Natalie Kiesler, Jamie Gorson Benario, Sam Lau, Stephen MacNeil, Narges Norouzi, Simone Opel, Vee Pettit, Leo Porter, Brent N. Reeves, Jaromir Savelka, IV Smith, David H., Sven Strickroth, and Daniel Zingaro
-
[42]
Becker, Bailey Kimmel, Jared Wright, and Ben Briggs
James Prather, Brent N Reeves, Juho Leinonen, Stephen MacNeil, Arisoa S Randri- anasolo, Brett A. Becker, Bailey Kimmel, Jared Wright, and Ben Briggs. 2024. The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers. InProceedings of the 2024 ACM Conferen...
2024
-
[43]
Yunhan Qiao, Md Istiak Hossain Shihab, and Christopher Hundhausen. 2026. A systematic literature review of the use of GenAI assistants for code compre- hension: Implications for computing education research and practice.ACM Transactions on Computing Education26, 2 (2026), 1–33
2026
-
[44]
It’s weird that it knows what i want
James Prather, Brent N Reeves, Paul Denny, Brett A Becker, Juho Leinonen, Andrew Luxton-Reilly, Garrett Powell, James Finnie-Ansley, and Eddie Antonio Santos. 2023. “It’s weird that it knows what i want”: Usability and interactions with copilot for novice programmers.ACM trans...
2023
-
[45]
Ken Reily, Pam Ludford Finnerty, and Loren Terveen. 2009. Two peers are better than one: aggregating peer reviews for computing assignments is surprisingly accurate. InProceedings of the 2009 ACM International Conference on Supporting Group Work(Sanibel Island, Florida, USA)(G...
2009
-
[46]
Jaromir Savelka, Arav Agarwal, Marshall An, Chris Bogart, and Majd Sakr. 2023. Thrilled by Your Progress! Large Language Models (GPT-4) No Longer Struggle to Pass Assessments in Higher Education Programming Courses. InProceedings of the 2023 ACM Conference on International Com...
2023
-
[47]
Brent Reeves, Sami Sarsa, James Prather, Paul Denny, Brett A Becker, Arto Hellas, Bailey Kimmel, Garrett Powell, and Juho Leinonen. 2023. Evaluating the perfor- mance of code generation models for solving parsons problems with small prompt variations. InProceedings of the 2023...
2023
-
[48]
Anvit Sinha, Shruti Goyal, Zachary Sy, Rhianna Kuperus, Ethan Dickey, and Andres Bejarano. 2024. BoilerTAI: A Platform for Enhancing Instruction Us- ing Generative AI in Educational Forums. In2024 IEEE Frontiers in Education Conference (FIE). 1–8. doi:10.1109/FIE61694.2024.10893137
2024
-
[49]
Zamfirescu-Pereira, and Narges Norouzi
Samantha Boatright Smith, Heather Wei, Abby O’Neill, Aneesh Durai, John DeNero, J.D. Zamfirescu-Pereira, and Narges Norouzi. 2025. Spotting AI Missteps: Students Take on LLM Errors in CS1. InProceedings of the 56th ACM Technical Symposium on Computer Science Education V. 2(Pit...
2025
-
[50]
Jaromir Savelka, Arav Agarwal, Christopher Bogart, and Majd Sakr. 2023. Large language models (gpt) struggle to answer multiple-choice questions about code. arXiv preprint arXiv:2303.08033(2023)
2023 arXiv
-
[51]
Lev Tankelevitch, Viktor Kewenig, Auste Simkute, Ava Elizabeth Scott, Advait Sarkar, Abigail Sellen, and Sean Rintel. 2024. The Metacognitive Demands and Opportunities of Generative AI. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu, HI...
2024
-
[52]
Smith IV, Mounika Padala, Christine Alvarado, Jamie Gorson Benario, and Leo Porter
Annapurna Vadaparty, Daniel Zingaro, David H. Smith IV, Mounika Padala, Christine Alvarado, Jamie Gorson Benario, and Leo Porter. 2024. CS1-LLM: Integrating LLMs into CS1 Instruction. InProceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1(Mil...
2024
-
[53]
Explain-in-Plain-English
David H Smith IV and Craig Zilles. 2024. Code Generation Based Grading: Evaluating an Auto-grading Mechanism for" Explain-in-Plain-English" Questions. InProceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1. 171–177
2024
-
[56]
Priyan Vaithilingam, Tianyi Zhang, and Elena L Glassman. 2022. Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models. InChi conference on human factors in computing systems extended abstracts. 1–7. 13
2022
-
[2023]
InCSEDU 2023-15th International Conference on Computer Supported Education, Vol
DaVinci goes to Bebras: a study on the problem solving ability of GPT-3. InCSEDU 2023-15th International Conference on Computer Supported Education, Vol. 2. SCITEPRESS-Science and Technology Publications, 59–69
2023
-
[2025]
In2024 Working Group Reports on Innovation and Technology in Computer Science Education(Milan, Italy)(ITiCSE 2024)
Beyond the Hype: A Comprehensive Review of Current Trends in Genera- tive AI Research, Teaching Practices, and Tools. In2024 Working Group Reports on Innovation and Technology in Computer Science Education(Milan, Italy)(ITiCSE 2024). Association for Computing Machinery, New Yo...
2024
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.