Pith. sign in

REVIEW 4 major objections 7 minor 80 references

Howzat? Appealing to Expert Judgement for Evaluating Human and AI Next-Step Hints for Novice Programmers

T0 review · 4 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read GPT-4 tops expert educators at writing next-step code hints

desk verdict Useful, open, well-conducted study of hint quality, but the bold GPT-4-beats-humans claim needs inferential support and replication before it is taken as established. read the letter →

arxiv 2411.18151 v1 pith:MFP6TMSW submitted 2024-11-27 cs.CY cs.HC

classification cs.CYcs.HC
keywords LLMsAIJavanext-stephintscomparativejudgementprogrammingeducationGPT-4promptengineering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a large language model can generate the next-step hint that a stuck novice programmer needs, and what makes such a hint good. The authors took real snapshots of broken Java code from the Blackbox dataset, gave four LLMs and five experienced educators only that code (no statement of the assignment), and collected 25 candidate hints per snapshot. A panel of 41 experienced Java educators ranked the hints in blind pairwise comparisons. The paper's central claim is that GPT-4, used with a carefully designed multi-stage prompt, produced hints ranked above those written by the human educators. The paper also reports that educators consistently preferred hints of 80–160 words at a low reading level, valued guidance beyond a direct fix, disliked alternative solution paths, and showed no preference over sentiment.

What carries the argument

The machinery is comparative judgement, a ranking method in which judges repeatedly choose the better of two hints and an adaptive algorithm turns those pairwise choices into a full ordering, avoiding the inconsistency of absolute rating scales. The supporting machinery consists of the five carefully engineered prompts (especially prompt 3, a multi-stage prompt that first infers the student's task and then writes the hint), the four stuck-student code snapshots used as scenarios, and a random-forest analysis that identifies which hint attributes predict high rank. The random forest is what surfaces the two dominant factors, word count and reading grade level, and separates those from the weaker pedagogical and stylistic factors.

What would settle it

A larger comparative-judgement replication using more snapshots, or a randomised classroom study assigning stuck students to GPT-4 hints versus human expert hints and measuring subsequent progress and learning, would settle whether the rank advantage is real; if GPT-4's mean rank advantage disappears or student outcomes show no benefit, the paper's core conclusion collapses.

Watch

Extended reading notes

Core claim

The paper's central claim is that a state-of-the-art LLM with a well-designed prompt can outperform experienced human educators at the specific task of writing a one-shot next-step hint for a stuck novice programmer. The comparison was deliberately demanding: generators saw only the student's code, with no problem description, so the task had to be inferred. Judged by a separate panel of experienced Java educators using comparative judgement, GPT-4 obtained the highest mean rank, ahead of the five human hints, while GPT-3.5, Gemini, and Mixtral-8x7B did not beat the humans. The best prompt was a two-stage one that first asked the model to state what task the student appeared to be working on, then fed that inference back to generate the hint. The authors conclude that automatic hint generation is immediately viable, with the caveat that the right model and prompt are both required.

Load-bearing premise

The load-bearing premise is that four snapshots of stuck student code, ranked by experienced educators, accurately represent whether a hint helps a real novice learn; the paper itself warns that four data points per generator may be noise and that educator preference may not match what students find useful.

Editorial extensions

If this is right

  • Tool-builders can treat expert-validated LLM hints as immediately usable in novice programming environments, as long as the prompt is supplied by the tool rather than composed by the student.
  • Hint generators should target 80–160 words and a reading level of US grade 9 or below; hints outside that band lose several rank positions.
  • A multi-stage prompt that first infers the task from the code and then produces the hint is the most reliable way to generate good hints with current LLMs.
  • Educators and hint authors should avoid offering alternative solution approaches in a next-step hint, since judges consistently penalised them.
  • The comparative-judgement protocol itself provides a reusable benchmark for evaluating future LLM versions and prompt designs without needing absolute grading scales.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the educator rankings reflect what actually helps novices, the main cost of high-quality hinting shifts from writing hints to maintaining a validated prompt layer as models change — an inference this study does not itself test.
  • The null effects for sentiment and for explaining general principles suggest that a stuck novice's immediate need is concrete, short direction; whether that holds for older learners or for non-Java languages is an untested extension.
  • A direct classroom comparison — students randomly assigned GPT-4 hints versus human hints, with progress and learning measured — would tell whether the ranking advantage is a real pedagogical gain; the paper does not run that experiment.
  • Because the authors found large variation across prompts and models but little variation across their five human hint-writers, a natural follow-up is to treat prompt design itself as the main engineering target for hint quality.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper evaluates automatically generated next-step programming hints for novices by having experienced Java educators rank hints via adaptive comparative judgement. Hints were produced by four LLMs (Mixtral-8x7B, Gemini, GPT-3.5, GPT-4) using five prompts (two multi-stage), plus five human expert hints, for four real 'stuck' student code snapshots drawn from Blackbox. The authors report three main findings: (1) GPT-4 with a well-designed prompt outperformed the human expert hints; (2) prompt choice matters as much as model choice, with multi-stage prompt 3 best; and (3) hint quality is most strongly associated with word count (80-160 words ideal) and reading level (US grade 9 or below), while offering alternative approaches is rated negatively. The study is pre-registered, with open data and analysis scripts, and includes several checks on judging validity and inter-rater reliability.

Significance. If the central claim holds, the result is practically significant: a state-of-the-art LLM with an expert-designed prompt could produce next-step hints that experienced educators rank above hints written by experienced human educators, even without knowing the original task. The paper also offers a methodological contribution by demonstrating comparative judgement as a viable tool for computing-education research, and the open data and reproducible pipeline are exemplary. However, the headline GPT-4-versus-human comparison rests on only four rank observations per generator and lacks an inferential test, and the random-forest analysis is descriptive and sensitive to correlated predictors; these issues currently prevent the paper from fully establishing its strongest conclusion.

major comments (4)
  1. [§4.10, Fig. 8; §6.2; §9] The central claim that GPT-4 outperformed human experts is based on the raw mean RankScore difference in Figure 8, but no significance test, confidence interval, or effect size is reported for that comparison. With only four data points per generator and ranks that are zero-sum within each snapshot, a single snapshot could move the GPT-4 mean by several rank positions. The paper's own caution in Section 4.10 that per-generator results 'may be interpreting noise' applies directly to this comparison. I recommend a reanalysis: a permutation test or a mixed-effects model with Snapshot as a random effect and Generator (or Model) as a fixed effect, plus a sensitivity analysis leaving out one snapshot at a time. Unless such an analysis is provided, the conclusion in Sections 6.2 and 9 should be weakened to a descriptive statement about this sample.
  2. [§4.11, Table 5] The random-forest importance values in Table 5 are used to conclude that word count and reading level are the most important hint characteristics, and that Model is relatively unimportant (importance 3.0 vs 23.9 and 17.9). Because WordCount and Flesch-Kincaid grade level are correlated with the model (Mixtral produced longer, less readable hints), the low Model importance may simply reflect that the model effect is mediated by these text properties; it does not establish that the GPT-4 advantage is fully explained by length and readability. The analysis would be more convincing with a correlation matrix, a model omitting WordCount and FleschKincaidGradeLevel to see how Model importance changes, and cross-validated or uncertainty-aware estimates of importance rather than point values from a single forest fit on 100 hints.
  3. [§3.3 and §4.3] The human benchmark consists of hints written by the five researchers who also designed the prompts, selected the snapshots, and wrote the notes given to judges. This creates a potential conflict: the human hints may be at a disadvantage because the task is framed in the researchers' own terms, and the comparison is not against an independent sample of expert educators. The paper should explicitly acknowledge this limitation and ideally include some external human hints (or justify why the authors' hints are a fair benchmark). Without this, the claim of beating 'human experts' is narrower than stated.
  4. [§4.5] The inter-rater reliability for Snapshot 4 is reported as 0.60, below the 0.7 threshold the paper cites from the comparative-judgement literature. Since each generator contributes exactly one hint per snapshot, a low-reliability snapshot injects substantial noise into every generator's aggregate score. The paper should report the reliability per snapshot alongside the generator means, and discuss how the ranking for Snapshot 4 in particular might affect the GPT-4-versus-humans comparison.
minor comments (7)
  1. [§1] The word 'Artifical' in the Introduction should be 'Artificial'.
  2. [§6.5] The phrase 'reminisicent' should be 'reminiscent'.
  3. [§4.7, Table 2] The 'Model' attribute is a generation mechanism, not a hint characteristic; including it in the same table as WordCount and Sentiment conflates two different research questions. Consider presenting the model/prompt analysis separately from the hint-characteristic analysis, or clearly stating that Model is used only as a covariate.
  4. [§4.5] The chi-squared test for left/right bias is reported as p=0.06; since the result is borderline, it would be helpful to report the effect size or the exact test statistic with degrees of freedom so that readers can judge the strength of the check.
  5. [Table 4] The row showing a single '0' with no check marks is confusing; a footnote or reformatting would clarify that it represents hints with none of the four feedback-literacy concepts present.
  6. [§3.4 and References] The reference to 'San Verhavert and Maeyer' is incomplete in the author list; also reference [38] appears to cite Kolen and Brennan's book on test equating for the NoMoreMarking platform, which is likely an incorrect citation.
  7. [Figures 8-10] The violin plots with overlaid dots make individual data points difficult to distinguish when multiple hints share the same rank; consider adding horizontal jitter or a summary table of means and standard deviations.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central comparison rests on independent educator rankings, not on the paper's own predictions or fitted inputs.

full rationale

The paper's principal claims—that GPT-4 outperformed human experts and that hint length and reading level are the most important predictors of hint quality—derive from an external outcome variable (RankScore) produced by 44 Java educators using comparative judgement while blind to generator identity. The generators (LLM prompts and author-written human hints) are inputs to the evaluation, not outputs of it; the educators' rankings were not constructed from the paper's hypotheses or from the random forest model. The human-benchmark hints were written by the five researchers themselves, which is a potential bias or confounding factor, but it is not circularity: the comparison is still an empirical measurement, and the paper does not claim those humans were independent external experts. The random forest analysis is descriptive (importance and partial-dependence plots) and does not assume the GPT-4 conclusion; indeed Model importance is low (3.0) relative to WordCount (23.9), so the model result is not an artefact of the analysis encoding the desired outcome. The paper's own caution in Section 4.10 that per-Generator results 'may be interpreting noise' and the absence of a significance test are statistical-validity weaknesses, not instances of a claim reducing to its inputs by construction. Self-citations, such as use of the Blackbox dataset and prior work by the authors, are used as data sources or related work and are not load-bearing in the derivation of the hint-ranking findings. No equation, fitted parameter, or definition makes any reported 'prediction' equivalent to its inputs.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

This empirical study does not introduce free parameters or invented entities. The load-bearing assumptions are about the representativeness of the snapshots, the validity of educator judgement as a proxy for learning, and the suitability of the text metrics and random forest model for the data.

assumptions (4)
  • domain assumption Comparative judgement rankings by expert educators are a valid measure of hint quality.
    The entire ranking is built on this. The authors note in Section 6.3 that it remains possible educators are not good at predicting which hints help students.
  • domain assumption The four selected snapshots are representative of stuck novice programming states.
    Section 3.1 describes manual selection from Blackbox, and only 4 of 8 prepared snapshots yielded enough judges in Section 4.4.
  • domain assumption Flesch-Kincaid grade level and VADER sentiment scores are appropriate metrics for hint text.
    Section 4.7 uses these off-the-shelf tools without validating them on this text genre.
  • standard math Random forest attribute importance on 100 hints with 9 predictors is stable enough to rank factors.
    No confidence intervals or cross-validation are reported for the importance values in Table 5.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Howzat? Appealing to Expert Judgement for Evaluating Human and AI Next-Step Hints for Novice Programmers." pith.science (2026). https://pith.science/paper/MFP6TMSW

@misc{pith2026241118151,
  author       = {Pith},
  title        = {Pith review of: Howzat? Appealing to Expert Judgement for Evaluating Human and AI Next-Step Hints for Novice Programmers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MFP6TMSW}},
  note         = {Machine review of arXiv:2411.18151}
}
read the original abstract

Motivation: Students learning to program often reach states where they are stuck and can make no forward progress. An automatically generated next-step hint can help them make forward progress and support their learning. It is important to know what makes a good hint or a bad hint, and how to generate good hints automatically in novice programming tools, for example using Large Language Models (LLMs). Method and participants: We recruited 44 Java educators from around the world to participate in an online study. We used a set of real student code states as hint-generation scenarios. Participants used a technique known as comparative judgement to rank a set of candidate next-step Java hints, which were generated by Large Language Models (LLMs) and by five human experienced educators. Participants ranked the hints without being told how they were generated. Findings: We found that LLMs had considerable variation in generating high quality next-step hints for programming novices, with GPT-4 outperforming other models tested. When used with a well-designed prompt, GPT-4 outperformed human experts in generating pedagogically valuable hints. A multi-stage prompt was the most effective LLM prompt. We found that the two most important factors of a good hint were length (80--160 words being best), and reading level (US grade 9 or below being best). Offering alternative approaches to solving the problem was considered bad, and we found no effect of sentiment. Conclusions: Automatic generation of these hints is immediately viable, given that LLMs outperformed humans -- even when the students' task is unknown. The fact that only the best prompts achieve this outcome suggests that students on their own are unlikely to be able to produce the same benefit. The prompting task, therefore, should be embedded in an expert-designed tool.

Figures

Figures reproduced from arXiv: 2411.18151 by the authors.

Figure 1
Figure 1. The design of the study: We take a Snapshot of student code from Blackbox, generate 25 hints for [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. A frequency histogram based on how many participants made a given percentage of “left” decision [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. The median (line) and interquartile range (shaded area) of time taken to make each judgement, ordered [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: The frequencies of different Infit ratings, one rating per participant. [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: A plot of Infit ratings (each plotted point is a participant) against median judgement time, to see if [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]
Figure 6
Figure 6. Figure 6: The scaled scores of the hints, split by Snapshot on the X-axis. Each dot is a hint, and the violin plot [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]
Figure 7
Figure 7. Figure 7: The top shows the code from Snapshot 1 (which was reproduced exactly, including the student’s [PITH_FULL_IMAGE:figures/full_fig_p019_7.png]
Figure 8
Figure 8. Figure 8: The rank score of the hints (higher is better) split by the model that produced them, with different [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]
Figure 9
Figure 9. Figure 9: The rank score of the AI-generated hints (higher is better) split by the prompt used to generate them. [PITH_FULL_IMAGE:figures/full_fig_p020_9.png]
Figure 10
Figure 10. Figure 10: A heatmap of AI model vs prompt (plus the five humans), showing the mean rank for that Generator. [PITH_FULL_IMAGE:figures/full_fig_p021_10.png]
Figure 11
Figure 11. Figure 11: Partial dependency plots for the four most important factors in predicting which hints are best. The [PITH_FULL_IMAGE:figures/full_fig_p023_11.png]
Figure 12
Figure 12. Figure 12: A participant’s experience (a scaled score ranked by researchers using comparative judgement) against [PITH_FULL_IMAGE:figures/full_fig_p023_12.png]
Figure 13
Figure 13. Figure 13: Results of asking the participants whether the hints would better than having no hints on a 5-point [PITH_FULL_IMAGE:figures/full_fig_p025_13.png]
Figure 14
Figure 14. Figure 14: Results of asking the participants whether they felt they could make better hints than those in the [PITH_FULL_IMAGE:figures/full_fig_p026_14.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

80 extracted references · 27 canonical work pages

  1. [1]

    Alireza Ahadi, Raymond Lister, Heikki Haapala, and Arto Vihavainen. 2015. Exploring Machine Learning Methods to Automatically Identify Students in Need of Assistance. InProceedings of the Eleventh Annual International Conference on International Computing Education Research (Omaha, Nebraska, USA) (ICER ’15). Association for Computing Machinery, New York, ...

  2. [2]

    Ahmed, Nisheeth Srivastava, Renuka Sindhgatta, and Amey Karkare

    Umair Z. Ahmed, Nisheeth Srivastava, Renuka Sindhgatta, and Amey Karkare. 2020. Characterizing the Pedagogical Benefits of Adaptive Feedback for Compilation Errors by Novice Programmers. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering: Software Engineering Education and Training (Seoul, South Korea) (ICSE-SEET ’20). As...

  3. [3]

    S Bartholomew and Emily Yoshikawa-Ruesch. 2018. A systematic review of research around adaptive comparative judgement (ACJ) in K-16 education. Council on Technology an Engineering Teacher Education: Research Monograph Series 1, 1 (2018). https://doi.org/10.21061/ctete-rms.v1.c.1

  4. [4]

    Anastasiia Birillo, Elizaveta Artser, Anna Potriasaeva, Ilya Vlasov, Katsiaryna Dzialets, Yaroslav Golubev, Igor Gerasi- mov, Hieke Keuning, and Timofey Bryksin. 2024. One Step at a Time: Combining LLMs and Static Analysis to Generate Next-Step Hints for Programming Tasks. In Proceedings of the 24th Koli Calling International Conference on Computing Educa...

  5. [5]

    Neil C. C. Brown and Amjad Altadmri. 2017. Novice Java Programming Mistakes: Large-Scale Data vs. Educator Beliefs. ACM Trans. Comput. Educ. 17, 2, Article 7 (May 2017), 21 pages. https://doi.org/10.1145/2994154

  6. [6]

    Neil C. C. Brown, Jamie Ford, Pierre Weill-Tessier, and Michael Kölling. 2023. Quick Fixes for Novice Programmers: Effective but Under-Utilised. In Proceedings of the 2023 Conference on United Kingdom & Ireland Computing Education Research (UKICER ’23) . Association for Computing Machinery, New York, NY, USA, Article 3, 7 pages. https: //doi.org/10.1145/3...

  7. [7]

    Neil C. C. Brown, Michael Kölling, Davin McCall, and Ian Utting. 2014. Blackbox: A Large Scale Repository of Novice Programmers’ Activity. In Proceedings of the 45th ACM Technical Symposium on Computer Science Education (Atlanta, Georgia, USA) (SIGCSE ’14). Association for Computing Machinery, New York, NY, USA, 223–228. https: //doi.org/10.1145/2538862.2538924

  8. [8]

    Marc Brysbaert. 2019. How many words do we read per minute? A review and meta-analysis of reading rate. Journal of Memory and Language 109 (2019), 104047. https://doi.org/10.1016/j.jml.2019.104047

Show all 80 references
  1. [9]

    Francisco Enrique Vicente Castro and Kathi Fisler. 2020. Qualitative Analyses of Movements Between Task-Level and Code-Level Thinking of Novice Programmers. InProceedings of the 51st ACM Technical Symposium on Computer Science Education (Portland, OR, USA) (SIGCSE ’20). Associ...

  2. [10]

    Tyne Crow, Andrew Luxton-Reilly, and Burkhard Wuensche. 2018. Intelligent Tutoring Systems for Programming Education: A Systematic Review. In Proceedings of the 20th Australasian Computing Education Conference (Brisbane, Queensland, Australia) (ACE ’18). Association for Comput...

  3. [14]

    Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N

    Paul Denny, James Prather, Brett A. Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N. Reeves, Eddie Antonio Santos, and Sami Sarsa. 2024. Computing Education in the Era of Generative AI. Commun. ACM 67, 2 (Jan. 2024), 56–67. https://doi.or...

  4. [15]

    Becker, Catherine Mooney, John Homer, Zachary C Albrecht, and Garrett B

    Paul Denny, James Prather, Brett A. Becker, Catherine Mooney, John Homer, Zachary C Albrecht, and Garrett B. Powell

  5. [16]

    Falkner and Katrina E

    Nickolas J.G. Falkner and Katrina E. Falkner. 2012. A Fast Measure for Identifying At-Risk Students in Computer Science. In Proceedings of the Ninth Annual International Conference on International Computing Education Research (Auckland, New Zealand) (ICER ’12). Association fo...

  6. [17]

    Fiannaca, Chinmay Kulkarni, Carrie J Cai, and Michael Terry

    Alexander J. Fiannaca, Chinmay Kulkarni, Carrie J Cai, and Michael Terry. 2023. Programming without a Programming Language: Challenges and Opportunities for Designing Developer Tools for Prompt Programming. InExtended Abstracts of the 2023 CHI Conference on Human Factors in Co...

  7. [18]

    Sandy Garner, Patricia Haden, and Anthony Robins. 2005. My Program is Correct but It Doesn’t Run: A Preliminary Investigation of Novice Programmers’ Problems. In Proceedings of the 7th Australasian Conference on Computing Education - Volume 42 (Newcastle, New South Wales, Aust...

  8. [19]

    Aashish Ghimire and John Edwards. 2024. Coding with AI: How Are Tools Like ChatGPT Being Used by Students in Foundational Programming Courses. In Artificial Intelligence in Education , Andrew M. Olney, Irene-Angelica Chounta, Zitao Liu, Olga C. Santos, and Ig Ibert Bittencourt...

  9. [20]

    Glassman, Aaron Lin, Carrie J

    Elena L. Glassman, Aaron Lin, Carrie J. Cai, and Robert C. Miller. 2016. Learnersourcing Personalized Hints. In Proceedings of the 19th ACM Conference on Computer-Supported Cooperative Work & Social Computing (San Francisco, California, USA) (CSCW ’16). Association for Computi...

  10. [21]

    Guo, Julia M

    Philip J. Guo, Julia M. Markel, and Xiong Zhang. 2020. Learnersourcing at Scale to Overcome Expert Blind Spots for Introductory Programming: A Three-Year Deployment Study on the Python Tutor Website. InProceedings of the Seventh ACM Conference on Learning @ Scale (Virtual Even...

  11. [22]

    Luke Gusukuma, Dennis Kafura, and Austin Cory Bart. 2017. Authoring feedback for novice programmers in a block- based language. In 2017 IEEE Blocks and Beyond Workshop (B&B). 37–40. https://doi.org/10.1109/BLOCKS.2017.8120407

  12. [23]

    Ajanovski, Mirela Gutica, Timo Hynninen, Antti Knutas, Juho Leinonen, Chris Messom, and Soohyun Nam Liao

    Arto Hellas, Petri Ihantola, Andrew Petersen, Vangel V. Ajanovski, Mirela Gutica, Timo Hynninen, Antti Knutas, Juho Leinonen, Chris Messom, and Soohyun Nam Liao. 2018. Taxonomizing Features and Methods for Identifying At-Risk Students in Computing Courses. In Proceedings of th...

  13. [25]

    2024.Predicting Results of Social Science Experiments Using Large Language Models

    Luke Hewitt, Ashwini Ashokkumar, Isaias Ghezae, and Robb Willer. 2024.Predicting Results of Social Science Experiments Using Large Language Models . Technical Report. Working Paper. https://samim.io/dl/Predicting%20results%20of% 20social%20science%20experiments%20using%20large...

  14. [26]

    Tin Kam Ho. 1995. Random decision forests. In Proceedings of the Third International Conference on Document Analysis and Recognition (Volume 1) - Volume 1 (ICDAR ’95) . IEEE Computer Society, USA, 278

  15. [27]

    Hutto and Eric Gilbert

    C. Hutto and Eric Gilbert. 2014. VADER: A Parsimonious Rule-Based Model for Sentiment Analysis of Social Media Text. Proceedings of the International AAAI Conference on Web and Social Media 8, 1 (May 2014), 216–225. https://doi.org/10.1609/icwsm.v8i1.14550

  16. [28]

    Michelle Ichinco and Caitlin Kelleher. 2018. Semi-Automatic Suggestion Generation for Young Novice Programmers in an Open-Ended Context. In Proceedings of the 17th ACM Conference on Interaction Design and Children (Trondheim, Norway) (IDC ’18). Association for Computing Machin...

  17. [29]

    Johan Jeuring, Hieke Keuning, Samiha Marwan, Dennis Bouvier, Cruz Izu, Natalie Kiesler, Teemu Lehtinen, Dominic Lohr, Andrew Peterson, and Sami Sarsa. 2022. Towards Giving Timely Formative Feedback and Hints to Novice Programmers. In Proceedings of the 2022 Working Group Repor...

  18. [30]

    Ian Jones and Ben Davies. 2024. Comparative judgement in education research. International Journal of Research & Method in Education 47, 2 (2024), 170–181. https://doi.org/10.1080/1743727X.2023.2242273 arXiv:https://doi.org/10.1080/1743727X.2023.2242273

  19. [31]

    With Great Power Comes Great Respon- sibility!

    Ishika Joshi, Ritvik Budhiraja, Pranav Deepak Tanna, Lovenya Jain, Mihika Deshpande, Arjun Srivastava, Srinivas Rallapalli, Harshal D Akolekar, Jagat Sesh Challa, and Dhruv Kumar. 2023. "With Great Power Comes Great Respon- sibility!": Student and Instructor Perspectives on th...

  20. [32]

    Ericson, David Weintrop, and Tovi Grossman

    Majeed Kazemitabaar, Justin Chow, Carl Ka To Ma, Barbara J. Ericson, David Weintrop, and Tovi Grossman. 2023. Studying the Effect of AI Code Generators on Supporting Novice Learners in Introductory Programming. InProceedings of the 2023 CHI Conference on Human Factors in Compu...

  21. [33]

    Tyson Kendon, Leanne Wu, and John Aycock. 2023. AI-Generated Code Not Considered Harmful. In Proceedings of the 25th Western Canadian Conference on Computing Education (Vancouver, BC, Canada) (WCCCE ’23). Association for Computing Machinery, New York, NY, USA, Article 3, 7 pag...

  22. [34]

    Hieke Keuning, Johan Jeuring, and Bastiaan Heeren. 2018. A Systematic Literature Review of Automated Feedback Generation for Programming Exercises. ACM Trans. Comput. Educ. 19, 1, Article 3 (Sept. 2018), 43 pages. https: //doi.org/10.1145/3231711

  23. [35]

    Natalie Kiesler, Dominic Lohr, and Hieke Keuning. 2023. Exploring the Potential of Large Language Models to Generate Formative Programming Feedback. In 2023 IEEE Frontiers in Education Conference (FIE) . 1–5. https://doi.org/10.1109/ FIE58773.2023.10343457

  24. [36]

    Natalie Kiesler and Daniel Schiffner. 2023. Large Language Models in Introductory Programming Education: ChatGPT’s Performance and Implications for Assessments. arXiv:2308.08572 [cs.SE] https://arxiv.org/abs/2308.08572

  25. [37]

    JP Kincaid. 1975. Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel. Chief of Naval Technical Training (1975)

  26. [38]

    MJ Kolen and RL Brennan. 2016. ‘No More Marking’: An online tool for comparative judgement. ISSN 1756-509X (2016), 12

  27. [39]

    Charles Koutcheme, Nicola Dainese, Arto Hellas, Sami Sarsa, Juho Leinonen, Syed Ashraf, and Paul Denny. 2024. Evaluating Language Models for Generating and Judging Programming Feedback. arXiv:2407.04873 [cs.AI] https: //arxiv.org/abs/2407.04873

  28. [41]

    Ban It Till We Understand It

    Sam Lau and Philip Guo. 2023. From "Ban It Till We Understand It" to "Resistance is Futile": How University Programming Instructors Plan to Adapt as More Students Use AI Code Generation and Explanation Tools Such as ChatGPT and GitHub Copilot. In Proceedings of the 2023 ACM Co...

  29. [42]

    Juho Leinonen, Paul Denny, Stephen MacNeil, Sami Sarsa, Seth Bernstein, Joanne Kim, Andrew Tran, and Arto Hellas

  30. [44]

    Mark Liffiton, Brad E Sheese, Jaromir Savelka, and Paul Denny. 2024. CodeHelp: Using Large Language Models with Guardrails for Scalable Support in Programming Classes. , 11 pages. https://doi.org/10.1145/3631802.3631830

  31. [45]

    Steffen Lippert, Anna Dreber, Magnus Johannesson, Warren Tierney, Wilson Cyrus-Lai, Eric Luis Uhlmann, Emo- tion Expression Collaboration, and Thomas Pfeiffer. 2024. Can large language models help predict results from a complex behavioural science study? Royal Society Open Sci...

  32. [46]

    Let Them Try to Figure It Out First

    Dominic Lohr, Natalie Kiesler, Hieke Keuning, and Johan Jeuring. 2024. "Let Them Try to Figure It Out First" - Reasons Why Experts (Do Not) Provide Feedback to Novice Programmers. In Proceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1 (Milan...

  33. [47]

    Stephen MacNeil, Andrew Tran, Arto Hellas, Joanne Kim, Sami Sarsa, Paul Denny, Seth Bernstein, and Juho Leinonen

  34. [48]

    Joyce Mahon, Brian Mac Namee, and Brett A. Becker. 2023. No More Pencils No More Books: Capabilities of Generative AI on Irish and UK Computer Science School Leaving Examinations. In Proceedings of the 2023 Conference on United Kingdom & Ireland Computing Education Research (S...

  35. [49]

    Ok Pal, We Have to Code That Now

    Alina Mailach, Dominik Gorgosch, Norbert Siegmund, and Janet Siegmund. 2024. “Ok Pal, We Have to Code That Now”: Interaction Patterns of Programming Beginners with a Conversational Chatbot. Empirical Software Engineering (EMSE) (2024), to appear

  36. [50]

    In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V

    Experiences from Using Code Explanations Generated by Large Language Models in a Web Software Development E-Book. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1 (Toronto ON, Canada) (SIGCSE 2023). Association for Computing Machinery, New ...

  37. [51]

    Samiha Marwan, Nicholas Lytle, Joseph Jay Williams, and Thomas Price. 2019. The Impact of Adding Textual Explanations to Next-Step Hints in a Novice Programming Environment. In Proceedings of the 2019 ACM Conference on Innovation and Technology in Computer Science Education (A...

  38. [52]

    Samiha Marwan and Thomas W. Price. 2023. ISnap: Evolution and Evaluation of a Data-Driven Hint System for Block- Based Programming. IEEE Trans. Learn. Technol. 16, 3.2 (June 2023), 399–413. https://doi.org/10.1109/TLT.2022.3223577

  39. [53]

    Samiha Marwan, Joseph Jay Williams, and Thomas Price. 2019. An Evaluation of the Impact of Automated Programming Hints on Performance and Learning. In Proceedings of the 2019 ACM Conference on International Computing Education Research (Toronto ON, Canada) (ICER ’19) . Associa...

  40. [54]

    Angela J McLean, Carol H Bond, and Helen D Nicholson. 2015. An anatomy of feedback: a phenomenographic investigation of undergraduate students’ conceptions of feedback. Studies in Higher Education 40, 5 (2015), 921–932

  41. [55]

    Daye Nam, Andrew Macvean, Vincent Hellendoorn, Bogdan Vasilescu, and Brad Myers. 2024. Using an LLM to Help With Code Understanding. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering (Lisbon, Portugal) (ICSE ’24) . Association for Computing M...

  42. [56]

    Jessica McBroom, Irena Koprinska, and Kalina Yacef. 2021. A Survey of Automated Programming Hint Generation: The HINTS Framework. ACM Comput. Surv. 54, 8, Article 172 (oct 2021), 27 pages. https://doi.org/10.1145/3469885

  43. [57]

    Sydney Nguyen, Hannah McLean Babe, Yangtian Zi, Arjun Guha, Carolyn Jane Anderson, and Molly Q Feldman. 2024. How Beginning Programmers and Code LLMs (Mis)read Each Other. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ...

  44. [58]

    Florian Obermüller, Ute Heuer, and Gordon Fraser. 2021. Guiding Next-Step Hint Generation Using Automated Tests. In Proceedings of the 26th ACM Conference on Innovation and Technology in Computer Science Education V. 1 (Virtual Event, Germany) (ITiCSE ’21). Association for Com...

  45. [59]

    Ha Nguyen and Vicki Allan. 2024. Using GPT-4 to Provide Tiered, Formative Code Feedback. In Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1 (Portland, OR, USA) (SIGCSE 2024). Association for Computing Machinery, New York, NY, USA, 958–964. ht...

  46. [60]

    Phitchaya Mangpo Phothilimthana and Sumukh Sridhara. 2017. High-Coverage Hint Generation for Massive Courses: Do Automated Hints Help CS1 Students?. In Proceedings of the 2017 ACM Conference on Innovation and Technology in Computer Science Education (Bologna, Italy) (ITiCSE ’1...

  47. [61]

    Alastair Pollitt. 2012. The method of Adaptive Comparative Judgement. Assessment in Educa- tion: Principles, Policy & Practice 19, 3 (2012), 281–300. https://doi.org/10.1080/0969594X.2012.665354 arXiv:https://doi.org/10.1080/0969594X.2012.665354

  48. [62]

    Maciej Pankiewicz and Ryan S. Baker. 2024. Navigating Compiler Errors with AI Assistance - A Study of GPT Hints in an Introductory Programming Course. In Proceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1 (Milan, Italy) (ITiCSE 2024). Assoc...

  49. [63]

    Becker, Robert Nix, Brent N

    James Prather, Paul Denny, Brett A. Becker, Robert Nix, Brent N. Reeves, Arisoa S. Randrianasolo, and Garrett Powell

  50. [64]

    James Prather, Raymond Pettit, Kayla McMurry, Alani Peters, John Homer, and Maxine Cohen. 2018. Metacognitive Difficulties Faced by Novice Programmers in Automated Assessment Tools. In Proceedings of the 2018 ACM Conference on International Computing Education Research (Espoo,...

  51. [65]

    Stanislav Pozdniakov, Jonathan Brazil, Solmaz Abdi, Aneesha Bakharia, Shazia Sadiq, Dragan Gašević, Paul Denny, and Hassan Khosravi. 2024. Large language models meet user interfaces: The case of provisioning feedback. Computers and Education: Artificial Intelligence 7 (2024), ...

  52. [66]

    It’s Weird That it Knows What I Want

    James Prather, Brent N. Reeves, Paul Denny, Brett A. Becker, Juho Leinonen, Andrew Luxton-Reilly, Garrett Powell, James Finnie-Ansley, and Eddie Antonio Santos. 2023. “It’s Weird That it Knows What I Want”: Usability and Interactions with Copilot for Novice Programmers. ACM Tr...

  53. [67]

    In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V

    First Steps Towards Predicting the Readability of Programming Error Messages. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1 (Toronto ON, Canada) (SIGCSE 2023). Association for Computing Machinery, New York, NY, USA, 549–555. https://doi....

  54. [68]

    Price, Samiha Marwan, and Joseph Jay Williams

    Thomas W. Price, Samiha Marwan, and Joseph Jay Williams. 2021. Exploring Design Choices in Data-Driven Hints for Python Programming Homework. In Proceedings of the Eighth ACM Conference on Learning @ Scale (Virtual Event, Germany) (L@S ’21). Association for Computing Machinery...

  55. [69]

    James Prather, Raymond Pettit, Kayla Holcomb McMurry, Alani Peters, John Homer, Nevan Simone, and Maxine Cohen. 2017. On Novices’ Interaction with Compiler Error Messages: A Human Factors Approach. In Proceedings of the 2017 ACM Conference on International Computing Education ...

  56. [70]

    Arun Raman and Viraj Kumar. 2022. Programming Pedagogy and Assessment in the Era of AI/ML: A Position Paper. In Proceedings of the 15th Annual ACM India Compute Conference (Jaipur, India) (COMPUTE ’22) . Association for Computing Machinery, New York, NY, USA, 29–34. https://do...

  57. [71]

    Becker, Bailey Kimmel, Jared Wright, and Ben Briggs

    James Prather, Brent N Reeves, Juho Leinonen, Stephen MacNeil, Arisoa S Randrianasolo, Brett A. Becker, Bailey Kimmel, Jared Wright, and Ben Briggs. 2024. The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers. In Proceedings of the 2024 ACM Conferenc...

  58. [73]

    Price, Samiha Marwan, Michael Winters, and Joseph Jay Williams

    Thomas W. Price, Samiha Marwan, Michael Winters, and Joseph Jay Williams. 2020. An Evaluation of Data-Driven Programming Hints in a Classroom Setting. In Artificial Intelligence in Education, Ig Ibert Bittencourt, Mutlu Cukurova, Kasia Muldner, Rose Luckin, and Eva Millán (Eds...

  59. [74]

    Andreas Scholl and Natalie Kiesler. 2024. How Novice Programmers Use and Experience ChatGPT when Solving Programming Exercises in an Introductory Course. arXiv:2407.20792 [cs.AI] https://arxiv.org/abs/2407.20792 35 Brown et al

  60. [75]

    Eric F Rietzschel, Bernard A Nijstad, and Wolfgang Stroebe. 2006. Productivity is not enough: A comparison of interactive and nominal brainstorming groups on idea generation and selection. Journal of Experimental Social Psychology 42, 2 (2006), 244–251

  61. [76]

    Brad Sheese, Mark Liffiton, Jaromir Savelka, and Paul Denny. 2024. Patterns of Student Help-Seeking When Using a Large Language Model-Powered Programming Assistant. In Proceedings of the 26th Australasian Computing Education Conference (Sydney, NSW, Australia)(ACE ’24). Associ...

  62. [77]

    Vincent Donche San Verhavert, Renske Bouwer and Sven De Maeyer. 2019. A meta-analysis on the reliability of comparative judgement. Assessment in Education: Principles, Policy & Practice 26, 5 (2019), 541–562. https: //doi.org/10.1080/0969594X.2019.1602027 arXiv:https://doi.org...

  63. [78]

    Ryo Suzuki, Gustavo Soares, Elena Glassman, Andrew Head, Loris D’Antoni, and Björn Hartmann. 2017. Exploring the Design Space of Automatically Synthesized Hints for Introductory Programming Assignments. In Proceedings of the 2017 CHI Conference Extended Abstracts on Human Fact...

  64. [79]

    Judy Sheard, Paul Denny, Arto Hellas, Juho Leinonen, Lauri Malmi, and Simon. 2024. Instructor Perceptions of AI Code Generation Tools - A Multi-Institutional Interview Study. In Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1 (Portland, OR, U...

  65. [80]

    Wiggins, Fahmid M

    Joseph B. Wiggins, Fahmid M. Fahid, Andrew Emerson, Madeline Hinckle, Andy Smith, Kristy Elizabeth Boyer, Bradford Mott, Eric Wiebe, and James Lester. 2021. Exploring Novice Programmers’ Hint Requests in an Intelligent Block-Based Coding Environment. In Proceedings of the 52nd...

  66. [81]

    Rebecca Smith and Scott Rixner. 2019. The Error Landscape: Characterizing the Mistakes of Novice Programmers. In Proceedings of the 50th ACM Technical Symposium on Computer Science Education (Minneapolis, MN, USA) (SIGCSE ’19). Association for Computing Machinery, New York, NY...

  67. [82]

    Bai, Robert Tairas, and Yu Huang

    Yuankai Xue, Hanlin Chen, Gina R. Bai, Robert Tairas, and Yu Huang. 2024. Does ChatGPT Help With Introductory Programming?An Experiment of Students Using ChatGPT in CS1. In Proceedings of the 46th International Conference on Software Engineering: Software Engineering Education...

  68. [83]

    Jacqueline Whalley, Amber Settle, and Andrew Luxton-Reilly. 2023. A Think-Aloud Study of Novice Debugging. ACM Trans. Comput. Educ. 23, 2, Article 28 (jun 2023), 38 pages. https://doi.org/10.1145/3589004

  69. [85]

    Ruiwei Xiao, Xinying Hou, and John Stamper. 2024. Exploring How Multiple Levels of GPT-Generated Programming Hints Support or Disappoint Novices. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI EA ’24). Association for...

  70. [87]

    Zamfirescu-Pereira, Richmond Y

    J.D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, and Qian Yang. 2023. Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Asso...

  71. [2021]

    In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21)

    On Designing Programming Error Messages for Novices: Readability and Its Constituent Factors. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 55, 15 pag...

  72. [2023]

    In Proceedings of the 2023 33 Brown et al

    Comparing Code Explanations Created by Students and Large Language Models. In Proceedings of the 2023 33 Brown et al. Conference on Innovation and Technology in Computer Science Education V. 1 (Turku, Finland) (ITiCSE 2023). Association for Computing Machinery, New York, NY, U...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.