REVIEW 3 cited by
Text2Analysis: A Benchmark of Table Question Answering with Advanced Data Analysis and Unclear Queries
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Tabular data analysis is crucial in various fields, and large language models show promise in this area. However, current research mostly focuses on rudimentary tasks like Text2SQL and TableQA, neglecting advanced analysis like forecasting and chart generation. To address this gap, we developed the Text2Analysis benchmark, incorporating advanced analysis tasks that go beyond the SQL-compatible operations and require more in-depth analysis. We also develop five innovative and effective annotation methods, harnessing the capabilities of large language models to enhance data quality and quantity. Additionally, we include unclear queries that resemble real-world user questions to test how well models can understand and tackle such challenges. Finally, we collect 2249 query-result pairs with 347 tables. We evaluate five state-of-the-art models using three different metrics and the results show that our benchmark presents introduces considerable challenge in the field of tabular data analysis, paving the way for more advanced research opportunities.
Forward citations
Cited by 3 Pith papers
-
Large Language Models for Predictive Analysis: How Far Are They?
Existing LLMs perform poorly on predictive analysis, with the best model scoring 24.11/28 and most models failing to generate executable code.
-
Evaluating and Enhancing LLMs for Multi-turn Text-to-SQL with Multiple Question Types
MMSQL is a multi-turn text-to-SQL benchmark with four question types, and a multi-agent framework with a Question Detector improves LLM performance on it.
-
MDSF: Context-Aware Multi-Dimensional Data Storytelling Framework based on Large language Model
MDSF is an LLM-based framework for automated data insight ranking and storytelling that, by its own reported results, does not outperform GPT-4 on ranking and most narrative metrics.
Discussion (0). Continue with ORCID to comment.