Pith. sign in

REVIEW 4 cited by

A Survey on Employing Large Language Models for Text-to-SQL Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.15186 v5 pith:PJNQJQD3 submitted 2024-07-21 cs.CL

classification cs.CL
keywords largemethodsmodelscomprehensivelanguagellm-basedsurveytext-to-sql
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the development of the Large Language Models (LLMs), a large range of LLM-based Text-to-SQL(Text2SQL) methods have emerged. This survey provides a comprehensive review of LLM-based Text2SQL studies. We first enumerate classic benchmarks and evaluation metrics. For the two mainstream methods, prompt engineering and finetuning, we introduce a comprehensive taxonomy and offer practical insights into each subcategory. We present an overall analysis of the above methods and various models evaluated on well-known datasets and extract some characteristics. Finally, we discuss the challenges and future directions in this field.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SPOT: Bridging Natural Language and Geospatial Search for Investigative Journalists

    cs.IR 2025-06 conditional novelty 6.0 of 10

    SPOT converts natural-language scene descriptions into structured OpenStreetMap queries via a fine-tuned LLaMA 3 model and semantic tag bundles, reporting state-of-the-art query-interpretation accuracy on a 195-query ...

  2. SEED: Enhancing Text-to-SQL Performance and Practical Usability Through Automatic Evidence Generation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    SEED automatically generates evidence from database schemas, descriptions, and sampled values, improving text-to-SQL accuracy in no-evidence settings.

  3. Bootstrapping Learned Cost Models with Synthetic SQL Queries

    cs.DB 2025-08 conditional novelty 5.0 of 10

    LLM-based synthetic SQL generation can train a learned cost model with fewer, more diverse queries than mechanical generation, though the measured accuracy gains are small and the comparison is not matched by training size.

  4. Meta-aware Learning in text-to-SQL Large Language Model

    cs.AI 2025-05 conditional novelty 4.0 of 10

    Combining schema, chain-of-thought, metadata knowledge, and tokenized prompt structures during fine-tuning improves text-to-SQL execution accuracy on private business databases compared to schema-only fine-tuning.

Pith tools