REVIEW 5 cited by
Text-to-SQL in the Wild: A Naturally-Occurring Dataset Based on Stack Exchange Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Most available semantic parsing datasets, comprising of pairs of natural utterances and logical forms, were collected solely for the purpose of training and evaluation of natural language understanding systems. As a result, they do not contain any of the richness and variety of natural-occurring utterances, where humans ask about data they need or are curious about. In this work, we release SEDE, a dataset with 12,023 pairs of utterances and SQL queries collected from real usage on the Stack Exchange website. We show that these pairs contain a variety of real-world challenges which were rarely reflected so far in any other semantic parsing dataset, propose an evaluation metric based on comparison of partial query clauses that is more suitable for real-world queries, and conduct experiments with strong baselines, showing a large gap between the performance on SEDE compared to other common datasets.
Forward citations
Cited by 5 Pith papers
-
Towards LLM Agents for Earth Observation
On a new 140-question Earth observation benchmark, the best LLM agent scores 33% accuracy with Google Earth Engine access because generated code fails to run over 58% of the time.
-
STRuCT-LLM: Unifying Tabular and Graph Reasoning with Reinforcement Learning for Semantic Parsing
Jointly reinforcing LLMs on SQL and Cypher with a graph-edit-distance reward improves structured parsing performance and transfers to table and graph QA tasks.
-
Sparks of Tabular Reasoning via Text2SQL Reinforcement Learning
Training LLMs on Text-to-SQL with chain-of-thought supervision and GRPO reinforcement learning is reported to improve zero-shot accuracy on tabular question answering, though the gains are measured by an LLM judge rat...
-
Large-Scale Analysis of Discussions by CS Educators Across the Stack Exchange Network
Across Stack Exchange, CS educators' posts are mostly technical and dominated by programming topics, while non-IT topics such as mathematics grew steadily between 2013 and 2018.
-
Exploring React Library Related Questions on Stack Overflow: Answered vs. Unanswered
React questions on Stack Overflow are more likely to be answered when they have more views, code snippets, more code lines, and higher-reputation askers, while comment count, length, and images reduce answerability.
Discussion (0). Continue with ORCID to comment.