A review of 114 studies creates taxonomies for code and data quality issues, formalizes 18 propagation mechanisms from training data defects to LLM-generated code defects, and synthesizes detection and mitigation techniques.
Codejudge: Evaluating code generation with large language models.arXiv preprint arXiv:2410.02184,
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.SE 3years
2026 3roles
background 2polarities
background 2representative citing papers
LLM judges for code tasks show high sensitivity to prompt biases that systematically favor certain options, changing accuracy and model rankings even when code is unchanged.
Empirical study of Stack Overflow logging posts identifies 11 topics where containerized environments show the highest difficulty via unanswered rates and resolution time.
citing papers explorer
-
Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code
A review of 114 studies creates taxonomies for code and data quality issues, formalizes 18 propagation mechanisms from training data defects to LLM-generated code defects, and synthesizes detection and mitigation techniques.
-
Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering
LLM judges for code tasks show high sensitivity to prompt biases that systematically favor certain options, changing accuracy and model rankings even when code is unchanged.
-
An Empirical Study on Logging Evolution On Stack Overflow: Trends, Topics, and Challenges
Empirical study of Stack Overflow logging posts identifies 11 topics where containerized environments show the highest difficulty via unanswered rates and resolution time.