REVIEW 3 major objections 5 minor 72 references
Breaking the Programming Language Barrier: Multilingual Prompting to Empower Non-Native English Learners
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper claims that non-native English speakers can solve programming problems by prompting generative AI in their native languages, and provides first classroom evidence from Arabic, Chinese, and Portuguese learners.
desk verdict First student-facing multilingual prompting study, but the headline claim overreaches: the data don't actually show that successful solutions came from native-language prompts. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Prompt Problem, a programming exercise in which a student is shown a computational task visually and must write a natural-language prompt for an LLM so that the generated code passes a suite of test cases. The tool that runs this loop, Promptly, executes the generated code and reports pass or fail, giving an objective measure of whether a prompt succeeded. In this paper the same exercise format is used across three language groups, and prompts are categorized by how much English, native language, or code they contain. The Prompt Problem mechanism matters because it isolates the student's ability to specify a solution in natural language from their ability to write syntax, which is exactly the skill that English-centric programming instruction tends to suppress.
What would settle it
Conduct the same three Prompt Problems with matched groups of the same educational level, same programming language (say Python), and same institution, prompting in Arabic, Chinese, and Portuguese through the same model; if Arabic pass rates match Portuguese and Chinese, the paper's low-resource-language explanation is false.
Extended reading notes
Core claim
The central claim is that non-native English speakers can successfully use their native language as the interface to generative AI for programming, solving Prompt Problems by writing prompts in Arabic, Chinese, or Portuguese rather than in English. Success is measured by whether the code the model generates passes a hidden test suite; by that measure, a majority of students in all three groups completed at least the first problem, and the Portuguese and Chinese groups showed high success rates overall. Arabic speakers completed problems at a noticeably lower rate, and the paper argues this reflects the small share of Arabic in the training data of large language models. Across all groups, students reported that native-language prompting felt more expressive and natural, but that they had to supply English terms for programming concepts, and many believed English prompts gave more accurate results. The paper concludes that native-language prompting is viable today for some languages, while acknowledging that the AI's language support and the student's familiarity with English coding terms shape how well it works.
Load-bearing premise
The paper's explanation of the Arabic results assumes the three student groups are similar enough in educational level, programming language, and institution that the lower Arabic success rate should be attributed to the AI model's weaker Arabic support rather than to those background differences.
Editorial extensions
If this is right
- Instructors can assign Prompt Problems in learners' native languages and still grade them automatically with test cases, creating accessible entry points for non-native English speakers.
- Students who understand programming concepts but struggle with English can demonstrate that understanding through native-language prompts, at least in high-resource languages.
- Arabic-speaking learners currently face an uneven experience: the paper's data suggest their prompts are more likely to be misunderstood, so educators should plan extra iteration time or pair native-language prompting with an English fallback.
- Because the model's English-centric training drives many of the difficulties, improvements in multilingual training data should directly improve native-language success rates.
- The finding that students code-switch to English technical terms implies that curricula still need to teach programming vocabulary in English, even if problem-solving happens in the native language.
Reading between the lines
- If the Arabic gap is truly a training-data effect, then rerunning the same problems with a newer model should shrink it; this is a cheap, direct test of the paper's main explanation.
- The same logic suggests a policy implication the authors only gesture at: investment in low-resource language data for code-generating models could be one of the most effective ways to widen access to programming education worldwide.
- The 'thinking in English' reports hint that the deeper barrier is the English-centric design of programming languages themselves; GenAI may lower, not remove, that barrier until languages or interfaces are localized.
- A testable extension would measure learning outcomes, not just pass rates: does native-language Prompt Problem practice transfer to better performance on later English-only programming tasks, or does it mainly build prompting skill?
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an exploratory study in which three groups of non-native English speaking students (Arabic, Chinese, and Portuguese) solved 'Prompt Problems' by writing prompts to a generative AI model in their native language. The authors present per-problem success rates (Table 2), a categorization of prompting strategies by language (Table 3), and qualitative themes from post-activity surveys. The central claim is that students can successfully use their native language to solve programming problems, with particular difficulty observed for Arabic, which the authors attribute partly to low-resource language support in LLMs. The paper also reports student perceptions about expressivity, model performance, and the entrenchment of English in programming.
Significance. If the central claim were soundly established, the result would be significant for computing education: it would offer evidence that LLM-based prompting can lower English-language barriers for non-native speakers, with implications for inclusive pedagogy and global access to programming instruction. The study is timely and addresses a genuinely under-explored question. The use of an external success criterion (passing hidden test cases) is a strength, as is the collection of authentic classroom data from three institutions and the inclusion of student voices via qualitative analysis. However, as analyzed, the data do not demonstrate that successful submissions were actually produced by native-language prompting, and the cross-language comparisons are confounded. The paper's contribution is better described as a descriptive feasibility observation than as a demonstrated result.
major comments (3)
- [Section 4.1, Tables 2 and 3] The inference that 'students were able to successfully solve the Prompt Problems in their native language' is not supported by the data as presented. Table 2 reports pass rates for all students who attempted each problem, but it does not indicate the language of the successful prompts. Table 3 reports per-prompt counts, not per-submission-stream outcomes. The only stream-level statement in the paper, 'Of 172 successful submission streams, 152 used only 1 state, and 127 of those were Native,' is internally inconsistent: 172 equals the total number of submission streams (82 Arabic + 28 Chinese + 62 Portuguese from Table 2), not the number of successful streams, which is 132 (27+15+12+12+5+2+21+20+18). The sentence therefore appears to refer to all submission streams, not successful ones. The paper never reports the pass rate of streams that stayed exclusively in the native language. For instance, among Chinese students only 6 of 19 correct prompts were categorized as 'Native,' with 6 'English' and 5 'Mixed,' so successful Chinese streams may have involved switching to English. To establish the central claim, the authors need to provide a per-stream crosstabulation of prompting strategy and success, or explicitly weaken the claim to 'students can pass some tasks using prompts that include their native language.'
- [Section 3.2.1] The data cleaning step removes all prompts by students who only prompted in English. This selection biases the sample toward students who attempted native-language prompting and removes any possibility of an English-only baseline. RQ1 asks 'How successful are students at solving Prompt Problems in their non-English native languages?' but the reported success rates are computed only for students who used at least some non-English words in their prompts. Without an English-prompting control group or at least a report of the number and success rates of excluded students, the paper cannot distinguish native-language feasibility from general LLM problem-solving ability. The authors should report the exclusion counts and, if possible, compare success rates between native-language and English-prompting students within the same cohorts.
- [Section 5.3, Section 6] The cross-language comparison, particularly the claim that Arabic students experienced greater difficulty partially due to limited training data, is confounded. Section 5.3 admits that the Arabic group were undergraduates using Java at a different institution, the Chinese group were postgraduates using Python, and the Portuguese group were undergraduates using Python. These differences alone could explain the observed lower Arabic success rates. The conclusion in Section 6 states that Arabic-speaking students 'appeared to face greater challenges, with lower accuracy and higher model misinterpretations' and attributes this to training-data limitations, but the data do not support a causal attribution. The authors should either present within-group evidence (e.g., same problems, same programming language, comparable educational level) or rephrase the discussion to present the training-data explanation as one of several plausible hypotheses. As written, the conclusion overreaches what the confounded design can support.
minor comments (5)
- [Section 3.1] The manuscript contains a typographical error: 'a Portugese university' should be 'a Portuguese university.'
- [Section 3.2.1] The category labels are inconsistent between the text and Table 3: the text defines categories O, E, N, M, C, but Table 3 labels the columns 'Other,' 'English,' 'Native,' 'Mixed,' 'Code.' Please align the terminology.
- [Section 4.1] The sentence 'Native was the most frequent starting state as well as the most frequent ending state' would be more informative if accompanied by counts or a transition table, especially given the paper's emphasis on state changes.
- [Section 4.2.1] The phrase 'Anecdotally, this theme was much more prevalent in responses in the Arabic-speaking sample' is odd because the authors are reporting on their own coded data; consider replacing 'anecdotally' with 'in our coded data' or a similar phrase.
- [Section 5.2] The sentence 'It's possible that the student of the future does not think programmatically in English because of these advances' is speculative; consider flagging it more explicitly as a forward-looking hypothesis rather than an implication of the data.
Circularity Check
No significant circularity: the central feasibility claim rests on externally evaluated test-case pass rates, and the self-citations are background context rather than load-bearing premises.
full rationale
This paper makes no mathematical derivation and fits no parameters, so the main circularity patterns do not apply. The central RQ1 result is measured by whether code generated from student prompts passes a predefined test suite, an external criterion that is independent of the authors' assumptions; success is not defined in terms of the paper's own categories or conclusions. The authors' prior Prompt Problems work [13] is cited to introduce the activity and to explain why fewer students attempt later problems, but that self-citation does not establish the empirical outcome and is not used to forbid alternatives. The inference concern raised by the skeptical reader, namely that Table 3 reports per-prompt strategy counts rather than a per-submission-stream crosstabulation of strategy and success, is a validity or evidentiary limitation, not circularity: it does not make the conclusion true by construction. The paper also openly acknowledges cohort differences and small samples in Section 5.3, further showing that its claims are presented as observed empirical findings rather than as consequences of its own definitions. Therefore no circular step is present, and the score is 0.
Assumptions & free parameters
free parameters (1)
- Prompt category definitions
assumptions (4)
- domain assumption Passing the hidden test suite indicates successful problem solving
- domain assumption The three cohorts are comparable for cross-language comparison
- domain assumption Human categorization of prompts into linguistic strategies is reliable
- domain assumption The LLM's difficulty with Arabic stems from training-data imbalance
Cite this review
Pith. "Pith review of Breaking the Programming Language Barrier: Multilingual Prompting to Empower Non-Native English Learners." pith.science (2026). https://pith.science/paper/HZ5BX5PK
@misc{pith2026241212800,
author = {Pith},
title = {Pith review of: Breaking the Programming Language Barrier: Multilingual Prompting to Empower Non-Native English Learners},
year = {2026},
howpublished = {\url{https://pith.science/paper/HZ5BX5PK}},
note = {Machine review of arXiv:2412.12800}
}
read the original abstract
Non-native English speakers (NNES) face multiple barriers to learning programming. These barriers can be obvious, such as the fact that programming language syntax and instruction are often in English, or more subtle, such as being afraid to ask for help in a classroom full of native English speakers. However, these barriers are frustrating because many NNES students know more about programming than they can articulate in English. Advances in generative AI (GenAI) have the potential to break down these barriers because state of the art models can support interactions in multiple languages. Moreover, recent work has shown that GenAI can be highly accurate at code generation and explanation. In this paper, we provide the first exploration of NNES students prompting in their native languages (Arabic, Chinese, and Portuguese) to generate code to solve programming problems. Our results show that students are able to successfully use their native language to solve programming problems, but not without some difficulty specifying programming terminology and concepts. We discuss the challenges they faced, the implications for practice in the short term, and how this might transform computing education globally in the long term.
Reference graph
Works this paper leans on
-
[1]
Arav Agarwal, Karthik Mittal, Aidan Doyle, Pragnya Sridhar, Zipiao Wan, Ja- cob Arthur Doughty, Jaromir Savelka, and Majd Sakr. 2024. Understanding the Role of Temperature in Diverse Question Generation by GPT-4. In Proc. of the 55th ACM Technical Symposium on Computer Science Education V. 2 . 1550–1551
work page 2024
-
[2]
Nimisha Agarwal, Viraj Kumar, Arun Raman, and Amey Karkare. 2023. A Bug’s New Life: Creating Refute Questions from Filtered CS1 Student Code Snapshots. In Proc. of the ACM Conf. on Global Computing Education Vol 1 . ACM, NY, NY, USA, 7–14
work page 2023
-
[3]
Vardhan Agarwal, Yada Chuengsatiansup, Elise Kim, Yuzi LYu, and Adalbert Ger- ald Soosai Raj. 2022. An Analysis of Stress and Sense of Belonging Among Native and Non-native English Speakers Learning Computer Science. In Proc. of the 53rd ACM Technical Symposium on Computer Science Education - Volume 1 . ACM, NY, NY, USA, 376–382
work page 2022
-
[4]
Brett A. Becker. 2019. Parlez-vous Java? Bonjour La Monde != Hello World: Barriers to Programming Language Acquisition for Non-Native English Speakers. In 30th Workshop of the Psychology of Programming Interest Group - PPIG ’19
work page 2019
-
[5]
Brett A. Becker, Daniel Gallagher, Paul Denny, James Prather, Colleen Gostomski, Kelli Norris, and Garrett Powell. 2022. From the Horse’s Mouth: The Words We Use to Teach Diverse Student Groups Across Three Continents. In Proc. of the 53rd ACM Technical Symposium on Computer Science Education - Volume 1 . ACM, NY, NY, USA, 71–77
work page 2022
-
[6]
Seth Bernstein, Paul Denny, Juho Leinonen, Lauren Kan, Arto Hellas, Matt Little- field, Sami Sarsa, and Stephen Macneil. 2024. "Like a Nesting Doll": Analyzing Recursion Analogies Generated by CS Students Using Large Language Models. In Proc. of the 2024 on Innovation and Technology in Computer Science Education V. 1. ACM, NY, NY, USA, 122–128
work page 2024
-
[7]
Virginia Braun and Victoria Clarke. 2019. Reflecting on Reflexive Thematic Analysis. Qualitative Research in Sport, Exercise and Health 11, 4 (2019), 589–597
work page 2019
-
[8]
Merijke Coenraad, Jen Palmer, Donna Eatinger, David Weintrop, and Diana Franklin. 2022. Using participatory design to integrate stakeholder voices in the creation of a culturally relevant computing curriculum. Int. J. of Child-Computer Interaction 31 (2022), 100353
work page 2022
Show all 72 references
-
[9]
Malcolm Corney, Sue Fitzgerald, Brian Hanks, Raymond Lister, Renee McCauley, and Laurie Murphy. 2014. ‘Explain in Plain English’ Questions Revisited: Data Structures Problems. In Proc. of the 45th ACM Technical Symposium on Computer Science Education. ACM, NY, NY, USA, 591–596
2014
-
[10]
Raena Cota, Enrico Pontelli, Paige Prescott, Lauren Curry, Lisa Hufstedler, Francis Vigil, Yolanda Lozano, and David Rutledge. 2022. Culturally Responsive Pedagogy in Computer Science (CR in CS)- K-12 Teacher Professional Development- Needs and Challenges. In Proc. of the 53rd...
2022
-
[11]
Andre Del Carpio Gutierrez, Paul Denny, and Andrew Luxton-Reilly. 2024. Eval- uating Automatically Generated Contextualised Programming Exercises. In Proc. of the 55th ACM Technical Symposium on Computer Science Education V. 1 . ACM, NY, NY, USA, 289–295
2024
-
[12]
Paul Denny, Viraj Kumar, and Nasser Giacaman. 2023. Conversing with Copilot: Exploring Prompt Engineering for Solving CS1 Problems Using Natural Language. In Proc. of the 54th ACM Technical Symposium on Computer Science Education V. 1 . ACM, NY, USA, 1136–1142
2023
-
[13]
Becker, and Brent N
Paul Denny, Juho Leinonen, James Prather, Andrew Luxton-Reilly, Thezyrie Amarouche, Brett A. Becker, and Brent N. Reeves. 2024. Prompt Problems: A New Programming Exercise for the Generative AI Era. In Proc. of the 55th ACM Technical Symposium on Computer Science Education V. ...
2024
-
[14]
Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N
Paul Denny, James Prather, Brett A. Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N. Reeves, Eddie Antonio Santos, and Sami Sarsa. 2024. Computing Education in the Era of Generative AI.Commun. ACM 67, 2 (jan 2024), 56–67
2024
-
[15]
Smith, Max Fowler, James Prather, Brett A
Paul Denny, David H. Smith, Max Fowler, James Prather, Brett A. Becker, and Juho Leinonen. 2024. Explaining Code with a Purpose: An Integrated Approach for Developing Code Comprehension and Prompting Skills. In Proc. of the 2024 on Innovation and Technology in Computer Science...
2024
-
[16]
Arid Hasan, Imran Razzak, and Usman Naseem
Krishno Dey, Prerona Tarannum, Md. Arid Hasan, Imran Razzak, and Usman Naseem. 2024. Better to Ask in English: Evaluation of Large Language Models on English, Low-resource and Cross-Lingual Settings. arXiv:2410.13153 [cs.CL]
2024 arXiv
-
[17]
Jacob Doughty, Zipiao Wan, Anishka Bompelli, Jubahed Qayum, Taozhi Wang, Juran Zhang, Yujia Zheng, Aidan Doyle, Pragnya Sridhar, Arav Agarwal, et al
-
[18]
Tony Haoran Feng, Paul Denny, Burkhard C Wünsche, Andrew Luxton-Reilly, and Jacqueline Whalley. 2024. An Eye for an AI: Evaluating GPT-4o’s Visual Perception Skills and Geometric Reasoning Skills Using Computer Graphics Questions. In SIGGRAPH Asia 2024 Educator’s Forum(Tokyo, ...
2024
-
[19]
Becker, Andrew Luxton-Reilly, and James Prather
James Finnie-Ansley, Paul Denny, Brett A. Becker, Andrew Luxton-Reilly, and James Prather. 2022. The Robots Are Coming: Exploring the Implications of Ope- nAI Codex on Introductory Programming. In Australasian Computing Education Conf. ACM, NY, NY, USA, 10–19
2022
-
[20]
James Finnie-Ansley, Paul Denny, Andrew Luxton-Reilly, Eddie Antonio Santos, James Prather, and Brett A. Becker. 2023. My AI Wants to Know If This Will Be on the Exam: Testing OpenAI’s Codex on CS2 Programming Exercises. In Proc. of the 25th Australasian Computing Education Co...
2023
-
[21]
Fusch and Lawrence R
Patricia I. Fusch and Lawrence R. Ness. 2015. Are we there yet? Data saturation in qualitative research. (2015)
2015
-
[22]
Philip J. Guo. 2018. Non-Native English Speakers Learning Computer Program- ming: Barriers, Desires, and Design Opportunities. In Proc. of the 2018 CHI Conf. on Human Factors in Computing Systems . ACM, NY, NY, USA, 1–14
2018
-
[23]
Carmen Nayeli Guzman, Anne Xu, and Adalbert Gerald Soosai Raj. 2021. Ex- periences of Non-Native English Speakers Learning Computer Science in a US University. In Proc. of the 52nd ACM Technical Symposium on Computer Science Education. ACM, NY, NY, USA, 633–639
2021
-
[24]
Scott Hanselman. 2008. Do You Have To Know English To Be A Program- mer? https://www.hanselman.com/blog/do-you-have-to-know-english-to-be- a-programmer Accessed: 2024-10-10
2008
-
[25]
Arto Hellas, Juho Leinonen, Sami Sarsa, Charles Koutcheme, Lilja Kujanpää, and Juha Sorva. 2023. Exploring the Responses of Large Language Models to Beginner Programmers’ Help Requests. In Proc. of the 2023 ACM Conf. on Int. Computing Education Research - Volume 1. ACM, NY, NY...
2023
-
[26]
Muntasir Hoq, Yang Shi, Juho Leinonen, Damilola Babalola, Collin Lynch, Thomas Price, and Bita Akram. 2024. Detecting ChatGPT-generated code submissions in a CS1 course using machine learning models. In Proc. of the 55th ACM Technical Symposium on Computer Science Education V....
2024
-
[27]
Irene Hou, Owen Man, Sophia Mettille, Sebastian Gutierrez, Kenneth Angelikas, and Stephen MacNeil. 2024. More Robots are Coming: Large Multimodal Models (ChatGPT) can Solve Visually Diverse Images of Parsons Problems. In Proc. of the 26th Australasian Computing Education Conf....
2024
-
[28]
Irene Hou, Sophia Mettille, Owen Man, Zhuo Li, Cynthia Zastudil, and Stephen MacNeil. 2024. The Effects of Generative AI on Computing Students’ Help- Seeking Preferences. In Proc. of the 26th Australasian Computing Education Conf. Australasian Computing Education Conference, F...
2024
-
[29]
Smith IV, Viraj Kumar, and Paul Denny
David H. Smith IV, Viraj Kumar, and Paul Denny. 2024. Explain in Plain Language Questions with Indic Languages: Drawbacks, Affordances, and Opportunities. arXiv:2409.20297 [cs.CY]
2024 arXiv
-
[30]
Sharin Jacob, Leiny Garcia, and Mark Warschauer. 2020. Leveraging Multilingual Identities in Computer Science Education . Springer Int. Publishing, 309–331
2020
-
[31]
Mollie Jordan, Kevin Ly, and Adalbert Gerald Soosai Raj. 2024. Need a Program- ming Exercise Generated in Your Native Language? ChatGPT’s Got Your Back: Automatic Generation of Non-English Programming Exercises Using OpenAI GPT-3.5. In Proc. of the 55th ACM Technical Symposium...
2024
-
[32]
Breanna Jury, Angela Lorusso, Juho Leinonen, Paul Denny, and Andrew Luxton- Reilly. 2024. Evaluating LLM-generated Worked Examples in an Introductory Programming Course. In Proceedings of the 26th Australasian Computing Educa- tion Conference (Sydney, NSW, Australia)(ACE ’24)....
2024
-
[33]
Smith IV, James Prather, Juho Leinonen, Andrew Luxton-Reilly, and Stephen MacNeil
Chris Kerslake, Paul Denny, David H. Smith IV, James Prather, Juho Leinonen, Andrew Luxton-Reilly, and Stephen MacNeil. 2024. Integrating Natural Language Prompting Tasks in Introductory Programming Courses. In SIGCSE Virtual 2024
2024
-
[34]
Natalie Kiesler, Dominic Lohr, and Hieke Keuning. 2023. Exploring the Potential of Large Language Models to Generate Formative Programming Feedback. arXiv preprint arXiv:2309.00029 (2023)
2023 arXiv
-
[35]
Charles Koutcheme, Nicola Dainese, Sami Sarsa, Arto Hellas, Juho Leinonen, and Paul Denny. 2024. Open Source Language Models Can Provide Feedback: Evaluating LLMs’ Ability to Help Students Using GPT-4-As-A-Judge. In Proc. of the 2024 on Innovation and Technology in Computer Sc...
2024
-
[36]
Viraj Kumar. 2021. Refute: An Alternative to ‘Explain in Plain English’ Questions. In Proc. of the 17th ACM Conf. on Int. Computing Education Research . ACM, NY, NY, USA, 438–440
2021
-
[37]
Juho Leinonen, Paul Denny, Stephen MacNeil, Sami Sarsa, Seth Bernstein, Joanne Kim, Andrew Tran, and Arto Hellas. 2023. Comparing Code Explanations Created by Students and Large Language Models. In Proc. of the 2023 Conf. on Innovation and Technology in Computer Science Educat...
2023
-
[38]
Juho Leinonen, Arto Hellas, Sami Sarsa, Brent Reeves, Paul Denny, James Prather, and Brett A. Becker. 2023. Using Large Language Models to Enhance Program- ming Error Messages. In Proc. of the 54th ACM Technical Symposium on Computer Science Education V. 1. ACM, NY, NY, USA, 563–569
2023
-
[39]
Leonard and Sue Sentance
Hayley C. Leonard and Sue Sentance. 2021. Culturally-relevant and responsive pedagogy in computing: A Quick Scoping Review. Int. J. of Computer Science Education in Schools 5, 2 (2021), 3–13
2021
-
[40]
Raymond Lister, Colin Fidge, and Donna Teague. 2009. Further Evidence of a Relationship between Explaining, Tracing and Writing Skills in Introductory Programming. In Proc. of the 14th Annual ACM SIGCSE Conf. on Innovation and Technology in Computer Science Education . ACM, NY...
2009
-
[41]
Chaoqun Liu, Wenxuan Zhang, Yiran Zhao, Anh Tuan Luu, and Lidong Bing
-
[42]
Evanfiya Logacheva, Arto Hellas, James Prather, Sami Sarsa, and Juho Leinonen
-
[43]
arXiv:2403.10258 [cs.CL]
Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models. arXiv:2403.10258 [cs.CL]
-
[44]
Stephen MacNeil, Andrew Tran, Arto Hellas, Joanne Kim, Sami Sarsa, Paul Denny, Seth Bernstein, and Juho Leinonen. 2023. Experiences from Using Code Expla- nations Generated by Large Language Models in a Web Software Development E-Book. In Proc. of the 54th ACM Technical Sympos...
2023
-
[45]
Evaluating Contextually Personalized Programming Exercises Created with Generative AI. In Proc. of the 2024 ACM Conf. on Int. Computing Education Research - Volume 1, Vol. 1. ACM, NY, NY, USA, 95–113
2024
-
[46]
Mike Lopez, Jacqueline Whalley, Phil Robbins, and Raymond Lister. 2008. Re- lationships between Reading, Tracing and Writing Skills in Introductory Pro- gramming. In Proc. of the Fourth Int. Workshop on Computing Education Research . ACM, NY, NY, USA, 101–112
2008
-
[47]
Ismael Villegas Molina, Audria Montalvo, Benjamin Ochoa, Paul Denny, and Leo Porter. 2024. Leveraging LLM Tutoring Systems for Non-Native English Speakers in Introductory CS Courses. arXiv:2411.02725 [cs.HC] https://arxiv.org/abs/ 2411.02725
2024 arXiv
-
[48]
Alexandre Magueresse, Vincent Carles, and Evan Heetderks. 2020. Low- resource Languages: A Review of Past Work and Future Challenges. arXiv:2006.07264 [cs.CL]
2020 arXiv
-
[49]
Gretchen McCulloch. 2019. Coding Is for Everyone–as Long as You Speak English. https://www.wired.com/story/coding-is-for-everyoneas-long-as-you- speak-english/ Accessed: 2024-10-10
2019
-
[50]
James Prather, Raymond Pettit, Kayla McMurry, Alani Peters, John Homer, and Maxine Cohen. 2018. Metacognitive Difficulties Faced by Novice Programmers in Automated Assessment Tools. InProc. of the 2018 ACM Conf. on Int. Computing Education Research. ACM, NY, NY, USA, 41–50
2018
-
[51]
Michael Sheinman Orenstrakh, Oscar Karnalim, Carlos Anibal Suarez, and Michael Liut. 2024. Detecting LLM-generated text in computing education: Comparative study for ChatGPT cases. In 2024 IEEE 48th Annual Computers, Software, and Applications Conf. IEEE, 121–126
2024
-
[52]
Becker, Ibrahim Albluwi, Michelle Craig, Hieke Keuning, Natalie Kiesler, Tobias Kohn, Andrew Luxton- Reilly, Stephen MacNeil, Andrew Petersen, Raymond Pettit, Brent N
James Prather, Paul Denny, Juho Leinonen, Brett A. Becker, Ibrahim Albluwi, Michelle Craig, Hieke Keuning, Natalie Kiesler, Tobias Kohn, Andrew Luxton- Reilly, Stephen MacNeil, Andrew Petersen, Raymond Pettit, Brent N. Reeves, and Jaromir Savelka. 2023. The Robots Are Here: Na...
2023
-
[53]
Becker, Arto Hellas, Bailey Kimmel, Garrett Powell, and Juho Leinonen
Brent Reeves, Sami Sarsa, James Prather, Paul Denny, Brett A. Becker, Arto Hellas, Bailey Kimmel, Garrett Powell, and Juho Leinonen. 2023. Evaluating the Performance of Code Generation Models for Solving Parsons Problems With Small Prompt Variations. In Proc. of the 2023 Conf....
2023
-
[54]
Becker, Bailey Kimmel, Jared Wright, and Ben Briggs
James Prather, Brent N Reeves, Juho Leinonen, Stephen MacNeil, Arisoa S Ran- drianasolo, Brett A. Becker, Bailey Kimmel, Jared Wright, and Ben Briggs. 2024. The Widening Gap: The Benefits and Harms of Generative AI for Novice Pro- grammers. In Proc. of the 2024 ACM Conf. on In...
2024
-
[55]
Patel, and Richard Halverson
Adalbert Gerald Soosai Raj, Kasama Ketsuriyonk, Jignesh M. Patel, and Richard Halverson. 2017. What Do Students Feel about Learning Programming Using Both English and Their Native Language?. In Int. Conf. on Learning and Teaching in Computing and Engineering . 1–8
2017
-
[56]
Jaromir Savelka, Arav Agarwal, Marshall An, Chris Bogart, and Majd Sakr. 2023. Thrilled by Your Progress! Large Language Models (GPT-4) No Longer Struggle to Pass Assessments in Higher Education Programming Courses. The 19th ACM Conf. on Int. Computing Education Research (2023)
2023
-
[57]
Eddie Antonio Santos and Brett A Becker. 2024. Not the Silver Bullet: LLM- enhanced Programming Error Messages are Ineffective in Practice.arXiv preprint arXiv:2409.18661 (2024)
2024 arXiv
-
[58]
Sami Sarsa, Paul Denny, Arto Hellas, and Juho Leinonen. 2022. Automatic Gen- eration of Programming Exercises and Code Explanations Using Large Language Models. In Proc. of the 2022 ACM Conf. on Int. Computing Education Research - Volume 1. ACM, NY, NY, USA, 27–43
2022
-
[59]
Patel, and Richard Halverson
Adalbert Gerald Soosai Raj, Kasama Ketsuriyonk, Jignesh M. Patel, and Richard Halverson. 2018. Does Native Language Play a Role in Learning a Programming Language?. In Proc. of the 49th ACM Technical Symposium on Computer Science Education. ACM, NY, NY, USA, 417–422
2018
-
[60]
Smith, Paul Denny, and Max Fowler
David H. Smith, Paul Denny, and Max Fowler. 2024. Prompting for Compre- hension: Exploring the Intersection of Explain in Plain English Questions and Prompt Writing. In Proc. of the Eleventh ACM Conf. on Learning @ Scale . ACM, NY, NY, USA, 39–50
2024
-
[61]
Explain-in-Plain-English
David H. Smith and Craig Zilles. 2024. Code Generation Based Grading: Evalu- ating an Auto-grading Mechanism for “Explain-in-Plain-English” Questions. In Proc. of the 2024 on Innovation and Technology in Computer Science Education V. 1 . ACM, NY, NY, USA, 171–177
2024
-
[62]
Smith, and Stephen MacNeil
Andrew Tran, Kenneth Angelikas, Egi Rama, Chiku Okechukwu, David H. Smith, and Stephen MacNeil. 2023. Generating multiple choice questions for computing courses using large language models. In 2023 IEEE Frontiers in Education Conf. IEEE, 1–8
2023
-
[63]
Chris Stokel-Walker. 2024. AI chatbot models ‘think’ in English even when us- ing other languages. https://www.newscientist.com/article/2420973-ai-chatbot- models-think-in-english-even-when-using-other-languages/
2024
-
[64]
Andrew Taylor, Alexandra Vassar, Jake Renzella, and Hammond Pearce. 2024. dcc --help: Transforming the Role of the Compiler by Generating Context-Aware Error Explanations with Large Language Models. In Proc. of the 55th ACM Technical Symposium on Computer Science Education V. ...
2024
-
[65]
Sierra Wang, John Mitchell, and Chris Piech. 2024. A large scale RCT on effective error messages in CS1. InProc. of the 55th ACM Technical Symposium on Computer Science Education V. 1. 1395–1401
2024
-
[66]
Marcella Veldthuis and Felienne Hermans. 2024. A Word about Programming: Applying a Natural Language Vocabulary Acquisition Model to Programming Education. In Proc. of the 2024 ACM SIGPLAN Int. Symposium on SPLASH-E . ACM, NY, NY, USA, 56–65
2024
-
[67]
Anne Venables, Grace Tan, and Raymond Lister. 2009. A Closer Look at Tracing, Explaining and Code Writing Skills in the Novice Programmer. In Proc. of the Fifth Int. Workshop on Computing Education Research Workshop . ACM, NY, NY, USA, 117–128
2009
-
[69]
Whalley, Raymond Lister, Errol Thompson, Tony Clear, Phil Robbins, P
Jacqueline L. Whalley, Raymond Lister, Errol Thompson, Tony Clear, Phil Robbins, P. K. Ajith Kumar, and Christine Prasad. 2006. An Australasian study of reading and comprehension skills in novice programmers, using the bloom and SOLO taxonomies. In Proc. of the 8th Australasia...
2006
-
[70]
Australian Computer Society, Inc., AUS, 243–252
-
[71]
Jun Zhao, Zhihao Zhang, Luhui Gao, Qi Zhang, Tao Gui, and Xuanjing Huang
-
[72]
arXiv:2401.01055 [cs.CL]
LLaMA Beyond English: An Empirical Study on Language Capability Transfer. arXiv:2401.01055 [cs.CL]
-
[2024]
A comparative study of AI-generated (GPT-4) and human-crafted MCQs in programming education. In Proc. of the 26th Australasian Computing Education Conf. 114–123
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.