Specification grounding, not test quantity, is the main reason AI-written tests improve LLM code: +38 pp correctness and 0% false alarms over a strong edge-prompted baseline.
Judging LLM -as-a-judge with MT -bench and chatbot arena
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Specification Grounding Drives Test Effectiveness for LLM Code
Specification grounding, not test quantity, is the main reason AI-written tests improve LLM code: +38 pp correctness and 0% false alarms over a strong edge-prompted baseline.