Pith. sign in

REVIEW 6 cited by

TabLib: A Dataset of 627M Tables with Context

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.07875 v1 pith:POZY6KWQ submitted 2023-10-11 cs.CL cs.AIcs.DBcs.LG

classification cs.CLcs.AIcs.DBcs.LG
keywords tablibdatasetstextcontextdiversityimagespromisesize
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

It is well-established that large, diverse datasets play a pivotal role in the performance of modern AI systems for text and image modalities. However, there are no datasets for tabular data of comparable size and diversity to those available for text and images. Thus we present "TabLib'', a compilation of 627 million tables totaling 69 TiB, along with 867B tokens of context. TabLib was extracted from numerous file formats, including CSV, HTML, SQLite, PDF, Excel, and others, sourced from GitHub and Common Crawl. The size and diversity of TabLib offer considerable promise in the table modality, reminiscent of the original promise of foundational datasets for text and images, such as The Pile and LAION.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Large Language Models are Good Relational Learners

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Rel-LLM combines a GNN encoder with a frozen LLM via soft prompts and masked attribute pretraining, reporting improved average performance on RelBench relational database tasks.

  2. Representation Learning for Tabular Data: A Comprehensive Survey

    cs.LG 2025-04 conditional novelty 6.0 of 10

    A comprehensive survey that categorizes deep tabular representation learning into specialized, transferable, and general models, with a feature/sample/objective taxonomy for specialized methods.

  3. SALT: Sales Autocompletion Linked Business Tables Dataset

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A dataset of authentic, anonymized, linked ERP sales tables (2.3M item rows) with eight multiclass autocompletion targets and baseline evaluations.

  4. An Efficient and Effective Evaluator for Text2SQL Models on Unseen and Unlabeled Data

    cs.CL 2026-03 conditional novelty 5.0 of 10

    FusionSQL predicts a Text2SQL model's execution accuracy on unseen unlabeled workloads from embedding-distance shift descriptors, reaching about 4-point MAE on benchmark transfers.

  5. Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Continuing the pre-training of TabPFN on 71 curated real-world tables raises its average normalized ROC-AUC from 0.954 to 0.976 on 29 AutoML Benchmark datasets.

  6. Tackling prediction tasks in relational databases with LLMs

    cs.LG 2024-11 conditional novelty 5.0 of 10

    Pre-trained LLMs, fed serialized relational rows with related examples, achieve competitive AUROC/MAE on RelBench without fine-tuning, but the headline comparison is weakened by pretraining contamination on Formula 1 tasks.

Pith tools