Specification grounding, not test quantity, is the main reason AI-written tests improve LLM code: +38 pp correctness and 0% false alarms over a strong edge-prompted baseline.
Teaching large language models to self-debug
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
CONDITIONAL 2representative citing papers
On hard multi-table text-to-SQL, verification-based LLM judges beat self-consistency and log-probability for predicting execution correctness, and fine-tuned verifiers fail to transfer across schemas.
citing papers explorer
-
Specification Grounding Drives Test Effectiveness for LLM Code
Specification grounding, not test quantity, is the main reason AI-written tests improve LLM code: +38 pp correctness and 0% false alarms over a strong edge-prompted baseline.
-
What Predicts Correctness in Text-to-SQL? A Selective-Prediction Study
On hard multi-table text-to-SQL, verification-based LLM judges beat self-consistency and log-probability for predicting execution correctness, and fine-tuned verifiers fail to transfer across schemas.