A seven-function longitudinal industrial case study reports LLM test generation scores improving from 32.44% to over 91% in nine months, but the improvement is partly driven by iterative prompt engineering on the same evaluation set.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Tracking the Moving Target: A Framework for Continuous Evaluation of LLM Test Generation in Industry
A seven-function longitudinal industrial case study reports LLM test generation scores improving from 32.44% to over 91% in nine months, but the improvement is partly driven by iterative prompt engineering on the same evaluation set.