Pith. sign in

REVIEW 7 cited by

TabLib: A Dataset of 627M Tables with Context

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.07875 v1 pith:POZY6KWQ submitted 2023-10-11 cs.CL cs.AIcs.DBcs.LG

TabLib: A Dataset of 627M Tables with Context

classification cs.CL cs.AIcs.DBcs.LG
keywords tablibdatasetstextcontextdiversityimagespromisesize
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

It is well-established that large, diverse datasets play a pivotal role in the performance of modern AI systems for text and image modalities. However, there are no datasets for tabular data of comparable size and diversity to those available for text and images. Thus we present "TabLib'', a compilation of 627 million tables totaling 69 TiB, along with 867B tokens of context. TabLib was extracted from numerous file formats, including CSV, HTML, SQLite, PDF, Excel, and others, sourced from GitHub and Common Crawl. The size and diversity of TabLib offer considerable promise in the table modality, reminiscent of the original promise of foundational datasets for text and images, such as The Pile and LAION.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MacrOData: New Benchmarks of Thousands of Datasets for Tabular Outlier Detection

    cs.LG 2026-02 accept novelty 8.0

    MacrOData supplies three large, curated benchmark suites totaling 2,446 datasets for tabular outlier detection, complete with standardized splits, metadata, and a public leaderboard.

  2. Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-Learning

    cs.LG 2026-05 unverdicted novelty 7.0

    MetaEvaluator meta-learns an initialization from reference models to enable accurate, label-free performance estimation for unseen models across architectures and modalities.

  3. MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

    cs.LG 2026-05 unverdicted novelty 7.0

    MulTaBench is a new collection of 40 image-tabular and text-tabular datasets designed to test target-aware representation tuning in multimodal tabular models.

  4. TabEmbed: Benchmarking and Learning Generalist Embeddings for Tabular Understanding

    cs.CL 2026-05 unverdicted novelty 7.0

    TabEmbed is the first generalist embedding model for tabular data that unifies classification and retrieval in one space via contrastive learning and outperforms text embedding models on the new TabBench benchmark.

  5. Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-Learning

    cs.LG 2026-05 unverdicted novelty 6.0

    MetaEvaluator applies meta-learning over reference models to deliver label-free performance estimates for unseen models across architectures and modalities on unlabeled datasets.

  6. Mind the Gap? A Distributional Comparison of Real and Synthetic Priors for Tabular Foundation Models

    cs.AI 2026-05 unverdicted novelty 5.0

    The synthetic prior for tabular foundation models covers only a narrow part of real table distributions, but this mismatch does not degrade model generalization.

  7. An Efficient and Effective Evaluator for Text2SQL Models on Unseen and Unlabeled Data

    cs.CL 2026-03 conditional novelty 5.0

    FusionSQL predicts a Text2SQL model's execution accuracy on unseen unlabeled workloads from embedding-distance shift descriptors, reaching about 4-point MAE on benchmark transfers.