REVIEW 6 cited by
How to Prompt LLMs for Text-to-SQL: A Study in Zero-shot, Single-domain, and Cross-domain Settings
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) with in-context learning have demonstrated remarkable capability in the text-to-SQL task. Previous research has prompted LLMs with various demonstration-retrieval strategies and intermediate reasoning steps to enhance the performance of LLMs. However, those works often employ varied strategies when constructing the prompt text for text-to-SQL inputs, such as databases and demonstration examples. This leads to a lack of comparability in both the prompt constructions and their primary contributions. Furthermore, selecting an effective prompt construction has emerged as a persistent problem for future research. To address this limitation, we comprehensively investigate the impact of prompt constructions across various settings and provide insights into prompt constructions for future text-to-SQL studies.
Forward citations
Cited by 6 Pith papers
-
ABISS: Evaluating Text-to-SQL Systems Through Agent Interaction
A new benchmark shows that text-to-SQL models detect problematic questions but fail to pinpoint the exact problem type and to resolve the question after a useful clarification.
-
Knowledge Base Construction for Knowledge-Augmented Text-to-SQL
KAT-SQL constructs a reusable knowledge base for text-to-SQL by expanding training data with LLM-generated knowledge and retrieving/refining the best entries for each query.
-
DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph
DCG-SQL retrieves text-to-SQL demonstrations by embedding a question-to-schema link graph, improving execution accuracy on Spider by up to about 10 points over random demonstrations on small LLMs.
-
Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions
In a leader/subordinate discussion simulation, most LLM agents do not reproduce the human pronoun pattern, and knowing the pattern does not help them show it.
-
What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects
By independently varying base models and training data across 12 models, this study shows that base model choice influences out-of-domain table task performance more than the instruction-tuning dataset does.
-
Rethinking Table Instruction Tuning
A systematic hyperparameter study of table instruction tuning finds that small learning rates and 2,600 examples suffice, yielding TAMA, an 8B model competitive with GPT-3.5/GPT-4 on table benchmarks while keeping gen...
Discussion (0). Continue with ORCID to comment.