REVIEW 4 cited by
Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Neural text-to-SQL models have achieved remarkable performance in translating natural language questions into SQL queries. However, recent studies reveal that text-to-SQL models are vulnerable to task-specific perturbations. Previous curated robustness test sets usually focus on individual phenomena. In this paper, we propose a comprehensive robustness benchmark based on Spider, a cross-domain text-to-SQL benchmark, to diagnose the model robustness. We design 17 perturbations on databases, natural language questions, and SQL queries to measure the robustness from different angles. In order to collect more diversified natural question perturbations, we utilize large pretrained language models (PLMs) to simulate human behaviors in creating natural questions. We conduct a diagnostic study of the state-of-the-art models on the robustness set. Experimental results reveal that even the most robust model suffers from a 14.0% performance drop overall and a 50.7% performance drop on the most challenging perturbation. We also present a breakdown analysis regarding text-to-SQL model designs and provide insights for improving model robustness.
Forward citations
Cited by 4 Pith papers
-
AuthentiCity: A Multi-Source Provenance-Aware Knowledge Graph and Benchmark for 3D City Models
AuthentiCity provides provenance-aware 3D city knowledge graphs for five cities and two benchmark families showing current LLMs and GNNs still struggle with provenance, coverage, and spatial reasoning.
-
Beyond Static Rules: Automated Discovery of Latent Vulnerabilities in Text-to-SQL
An evolving Vulnerability Codex plus hypothesis-driven perturbations exposes latent Text-to-SQL failures in LLMs far better than fixed expert rules, with transferable patterns and early remediation gains.
-
ExeSQL: Self-Taught Text-to-SQL Models with Execution-Driven Bootstrapping for SQL Dialects
Execution-driven bootstrapping, where a model generates SQL, executes it, and keeps only queries that run, lets a 7B model outperform GPT-4o on PostgreSQL, MySQL, and Oracle text-to-SQL benchmarks.
-
Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges
A structured review of table understanding with LLMs that proposes a taxonomy of input representations and identifies three research gaps.
Discussion (0). Continue with ORCID to comment.