LLM-generated probes with execution feedback find counterexamples that unit tests miss, and semantic clustering based on these probes improves code-generation evaluation.
Automated testing of refactoring engines
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Disproving Program Equivalence with LLMs
LLM-generated probes with execution feedback find counterexamples that unit tests miss, and semantic clustering based on these probes improves code-generation evaluation.