On 72 Python tasks, GPT-4 code passed 87.3% of tests versus 54.9% for one student's code, but was more complex and showed more severe security issues.
Devlin et al., ”BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of NAACL-HLT, 2019
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SE 1years
2025 1verdicts
REJECT 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Comparing Human and LLM Generated Code: The Jury is Still Out!
On 72 Python tasks, GPT-4 code passed 87.3% of tests versus 54.9% for one student's code, but was more complex and showed more severe security issues.