CS1 students correctly predicted the output of LLM-generated Python code in 32.5% of tasks, versus 59.4% for natural-language prompts, and still failed code prediction 58% of the time when they understood the prompt.
Duarte, Carlos Ferreira, Joao Duraes, Henrique Madeira, and Miguel Castelo-Branco
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
"I Would Have Written My Code Differently'': Beginners Struggle to Understand LLM-Generated Code
CS1 students correctly predicted the output of LLM-generated Python code in 32.5% of tasks, versus 59.4% for natural-language prompts, and still failed code prediction 58% of the time when they understood the prompt.