REVIEW 2 cited by
Substance Beats Style: Why Beginning Students Fail to Code with LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Although LLMs are increasing the productivity of professional programmers, existing work shows that beginners struggle to prompt LLMs to solve text-to-code tasks. Why is this the case? This paper explores two competing hypotheses about the cause of student-LLM miscommunication: (1) students simply lack the technical vocabulary needed to write good prompts, and (2) students do not understand the extent of information that LLMs need to solve code generation tasks. We study (1) with a causal intervention experiment on technical vocabulary and (2) by analyzing graphs that abstract how students edit prompts and the different failures that they encounter. We find that substance beats style: a poor grasp of technical vocabulary is merely correlated with prompt failure; that the information content of prompts predicts success; that students get stuck making trivial edits; and more. Our findings have implications for the use of LLMs in programming education, and for efforts to make computing more accessible with LLMs.
Forward citations
Cited by 2 Pith papers
-
Who's the Leader? Analyzing Novice Workflows in LLM-Assisted Debugging of Machine Learning Code
In an eight-person formative study, novice ML engineers who actively led the ChatGPT debugging conversation outperformed those who followed it, with patterns of over- and under-reliance.
-
"I Would Have Written My Code Differently'': Beginners Struggle to Understand LLM-Generated Code
CS1 students correctly predicted the output of LLM-generated Python code in 32.5% of tasks, versus 59.4% for natural-language prompts, and still failed code prediction 58% of the time when they understood the prompt.
Discussion (0). Continue with ORCID to comment.