GPT-4o and Gemini 2.0 Flash correctly judged code correctness in about 64-68% of cases and corrected 54-68% of faulty code, with better results when given problem descriptions.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Evaluating Large Language Models for Code Review
GPT-4o and Gemini 2.0 Flash correctly judged code correctness in about 64-68% of cases and corrected 54-68% of faulty code, with better results when given problem descriptions.