An LLM agent can improve itself at test time by rewriting its surrounding executable harness from unlabeled traces, using only proxy signals and a frozen model.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
CONDITIONAL 2representative citing papers
On hard multi-table text-to-SQL, verification-based LLM judges beat self-consistency and log-probability for predicting execution correctness, and fine-tuned verifiers fail to transfer across schemas.
citing papers explorer
-
TTHE: Test-Time Harness Evolution
An LLM agent can improve itself at test time by rewriting its surrounding executable harness from unlabeled traces, using only proxy signals and a frozen model.
-
What Predicts Correctness in Text-to-SQL? A Selective-Prediction Study
On hard multi-table text-to-SQL, verification-based LLM judges beat self-consistency and log-probability for predicting execution correctness, and fine-tuned verifiers fail to transfer across schemas.