Spider 2.0-AIFunc is a 465-instance benchmark for evaluating text-to-SQL systems on queries that incorporate Snowflake Cortex AI functions, with evaluations of ten models showing proprietary models reach 67-70% accuracy.
arXiv preprint arXiv:2511.21402 , year=
5 Pith papers cite this work. Polarity classification is still indexing.
years
2026 5representative citing papers
ProSPy proposes a profiling-driven SQL-Python agentic framework that achieves 60.15-60.51% execution accuracy on Spider 2.0 benchmarks for enterprise Text-to-SQL using Claude-4.5-Opus.
EviLink combines multi-hypothesis schema grounding with uncertainty-guided evidence acquisition, reporting 90.15% field-level recall and 123.30K average tokens on Spider2-Snow while improving downstream SQL generation.
AV-SQL uses a pipeline of LLM agents to generate intermediate CTE views that decompose complex Text-to-SQL queries, reaching 70.38% execution accuracy on Spider 2.0.
citing papers explorer
-
Spider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL Workflows
Spider 2.0-AIFunc is a 465-instance benchmark for evaluating text-to-SQL systems on queries that incorporate Snowflake Cortex AI functions, with evaluations of ten models showing proprietary models reach 67-70% accuracy.
-
ProSPy: A Profiling-Driven SQL-Python Agentic Framework for Enterprise Text-to-SQL
ProSPy proposes a profiling-driven SQL-Python agentic framework that achieves 60.15-60.51% execution accuracy on Spider 2.0 benchmarks for enterprise Text-to-SQL using Claude-4.5-Opus.
-
EviLink: Multi-Path Schema Linking with Uncertainty-Guided Evidence Acquisition for Large-Scale Text-to-SQL
EviLink combines multi-hypothesis schema grounding with uncertainty-guided evidence acquisition, reporting 90.15% field-level recall and 123.30K average tokens on Spider2-Snow while improving downstream SQL generation.
-
AV-SQL: Decomposing Complex Text-to-SQL Queries with Agentic Views
AV-SQL uses a pipeline of LLM agents to generate intermediate CTE views that decompose complex Text-to-SQL queries, reaching 70.38% execution accuracy on Spider 2.0.
- FlexSQL: Flexible Exploration and Execution Make Better Text-to-SQL Agents