For LLM-generated Java tests, coverage and mutation predict real-bug detection mainly when comparing models on bug-free code, not when the code under test is buggy, and suite size is not a major confounder.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study)
For LLM-generated Java tests, coverage and mutation predict real-bug detection mainly when comparing models on bug-free code, not when the code under test is buggy, and suite size is not a major confounder.