CodeQUEST, a GPT-4o-based evaluator-optimizer loop, reports a 52.6% mean relative improvement in code quality, but the improvement is measured by the same model that enforces monotonic score increases.
Is gpt-4 a reliable rater? evaluating consistency in gpt-4’s text ratings
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
On Iterative Evaluation and Enhancement of Code Quality Using GPT-4o
CodeQUEST, a GPT-4o-based evaluator-optimizer loop, reports a 52.6% mean relative improvement in code quality, but the improvement is measured by the same model that enforces monotonic score increases.