Pith. sign in

REVIEW 16 cited by

Large Language Models(LLMs) on Tabular Data: Prediction, Generation, and Understanding -- A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.17944 v4 pith:5KKMFGDX submitted 2024-02-27 cs.CL

classification cs.CL
keywords datadatasetssurveytabularaddresschallengescomprehensivefield
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent breakthroughs in large language modeling have facilitated rigorous exploration of their application in diverse tasks related to tabular data modeling, such as prediction, tabular data synthesis, question answering, and table understanding. Each task presents unique challenges and opportunities. However, there is currently a lack of comprehensive review that summarizes and compares the key techniques, metrics, datasets, models, and optimization approaches in this research domain. This survey aims to address this gap by consolidating recent progress in these areas, offering a thorough survey and taxonomy of the datasets, metrics, and methodologies utilized. It identifies strengths, limitations, unexplored territories, and gaps in the existing literature, while providing some insights for future research directions in this vital and rapidly evolving field. It also provides relevant code and datasets references. Through this comprehensive review, we hope to provide interested readers with pertinent references and insightful perspectives, empowering them with the necessary tools and knowledge to effectively navigate and address the prevailing challenges in the field.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 22 citations worldwide. Full citation record

  1. Watermarking Large Language Model-based Time Series Forecasting

    cs.IR 2025-07 conditional novelty 7.0 of 10

    Waltz embeds watermarks into LLM-based time series forecasts by nudging a few patch embeddings toward 'cold' LLM tokens, and detects them with a z-score test.

  2. Orthogonal Hierarchical Decomposition for Structure-Aware Table Understanding with Large Language Models

    cs.CL 2026-02 conditional novelty 6.0 of 10

    A framework that splits tables into row-tree and column-tree representations and uses an LLM to arbitrate both achieves large gains on complex-table QA benchmarks, though one backbone underperforms on HiTab.

  3. MetaRank: Task-Aware Metric Selection for Model Transferability Estimation

    cs.CV 2025-11 conditional novelty 6.0 of 10

    A meta-learner ranks MTE metrics for a target dataset from text descriptions, improving average rank over fixed metrics on 11 datasets.

  4. Limited Reference, Reliable Generation: A Two-Component Framework for Tabular Data Generation in Low-Data Regimes

    cs.LG 2025-09 conditional novelty 6.0 of 10

    ReFine combines rule-guided prompting and dual-granularity filtering to improve LLM-based tabular data generation when only 30 to 90 labeled rows exist, achieving top average rank over baselines.

  5. Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning

    cs.AI 2025-09 conditional novelty 6.0 of 10

    An LLM agent using retrieval and summary uncertainty as training rewards and inference filters produces more factual, useful multi-omics summaries and better downstream survival predictions.

  6. AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data

    cs.CL 2025-07 conditional novelty 6.0 of 10

    AraTable is the first Arabic tabular QA benchmark; its experiments show LLMs are much weaker at reasoning over Arabic tables than at direct lookup.

  7. Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMs

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Tab-MIA shows LLMs fine-tuned on tabular data are vulnerable to membership inference attacks, with AUROC up to 97.7% after three epochs and encoding format strongly affecting leakage.

  8. Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new benchmark, TableEval, with 3017 tables in five formats, shows LLMs are robust to table representation but perform worse on scientific tables, with the caveat that the domain gap is confounded by task difficulty.

  9. An LLM-Based Automatic Sportscast Solution for Robot Soccer Matches

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A neuro-symbolic pipeline turns RoboCup robot-soccer video into real-time statistics and LLM-generated commentary, demonstrated on three German Open 2026 clips.

  10. A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data

    cs.AI 2026-01 conditional novelty 5.0 of 10

    A metric-oriented survey that classifies intrinsic quality and trustworthiness metrics for LLM-generated data across six modalities and documents systematic evaluation gaps in the current literature.

  11. Multimodal Behavioral Patterns Analysis with Eye-Tracking and LLM-Based Reasoning

    cs.HC 2025-07 conditional novelty 5.0 of 10

    A human-AI framework uses horizontal and vertical segmentation with LLMs, expert co-scoring, and LSTM anomaly detection to extract behavioral patterns from eye-tracking data.

  12. Large Language Models as Unified Multimodal Learners for Clinical Prediction

    cs.CL 2026-07 conditional novelty 4.0 of 10

    Serializing all patient data — notes, vitals, labs — into one text sequence and fine-tuning an LLM matches or beats task-specific multimodal fusion baselines on mortality, graft-failure, and triage prediction.

  13. A Generalised Exponentiated Gradient Approach to Enhance Fairness in Binary and Multi-class Classification Tasks

    cs.LG 2026-03 conditional novelty 4.0 of 10

    A generalized Exponentiated Gradient method trains binary and multi-class classifiers under multiple linear fairness constraints, achieving fairness gains up to 92% at an accuracy cost up to 14 points.

  14. Agentic LLMs for Question Answering over Tabular Data

    cs.CL 2025-09 conditional novelty 4.0 of 10

    A five-stage NL-to-SQL pipeline with GPT-4o achieves 70.5% on DataBench QA and 71.6% on DataBench Lite QA, beating baselines of 26% and 27%.

  15. Towards High Supervised Learning Utility Training Data Generation: Data Pruning and Column Reordering

    cs.LG 2025-07 reject novelty 4.0 of 10

    PRRO combines signal-based data pruning and column reordering to improve the supervised learning utility of synthetic tabular data, but its evaluation is undermined by data manipulation and an ill-defined correlation measure.

  16. Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges

    cs.CL 2025-07 conditional novelty 3.0 of 10

    A structured review of table understanding with LLMs that proposes a taxonomy of input representations and identifies three research gaps.

Pith tools