A comparative study of Python test suites generated by GPT-4o, Amazon Q, and LLama 3.3 found 151 execution errors and 512 test smells, with assertion failures and low-cohesion tests most common.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Quality Assessment of Python Tests Generated by Large Language Models
A comparative study of Python test suites generated by GPT-4o, Amazon Q, and LLama 3.3 found 151 execution errors and 512 test smells, with assertion failures and low-cohesion tests most common.