On hard multi-table text-to-SQL, verification-based LLM judges beat self-consistency and log-probability for predicting execution correctness, and fine-tuned verifiers fail to transfer across schemas.
ICLR , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
LLM OOD detectors are length-confounded; a two-pathway embedding-plus-trajectory framework detects covert OOD inputs at 0.721 average AUROC and 0.850 on jailbreaks.
citing papers explorer
-
What Predicts Correctness in Text-to-SQL? A Selective-Prediction Study
On hard multi-table text-to-SQL, verification-based LLM judges beat self-consistency and log-probability for predicting execution correctness, and fine-tuned verifiers fail to transfer across schemas.
-
How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework
LLM OOD detectors are length-confounded; a two-pathway embedding-plus-trajectory framework detects covert OOD inputs at 0.721 average AUROC and 0.850 on jailbreaks.