Pith. sign in

REVIEW 16 cited by

Why Tabular Foundation Models Should Be a Research Priority

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.01147 v2 pith:OSFM3QZF submitted 2024-05-02 cs.LG

Why Tabular Foundation Models Should Be a Research Priority

classification cs.LG
keywords tabularmodelsdataresearchfoundationdatasetslargemodality
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recent text and image foundation models are incredibly impressive, and these models are attracting an ever-increasing portion of research resources. In this position piece we aim to shift the ML research community's priorities ever so slightly to a different modality: tabular data. Tabular data is the dominant modality in many fields, yet it is given hardly any research attention and significantly lags behind in terms of scale and power. We believe the time is now to start developing tabular foundation models, or what we coin a Large Tabular Model (LTM). LTMs could revolutionise the way science and ML use tabular data: not as single datasets that are analyzed in a vacuum, but contextualized with respect to related datasets. The potential impact is far-reaching: from few-shot tabular models to automating data science; from out-of-distribution synthetic data to empowering multidisciplinary scientific discovery. We intend to excite reflections on the modalities we study, and convince some researchers to study large tabular models.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Probing Memorization of Tabular In-Context Learning

    cs.LG 2026-06 unverdicted novelty 7.0

    A new probing framework detects moderate parametric memorization signals in tabular in-context learning models under single-task fine-tuning, strongest on low-cardinality tasks, but signals largely disappear under rea...

  2. Data Language Models: A New Foundation Model Class for Tabular Data

    cs.AI 2026-05 unverdicted novelty 7.0

    Schema-1 is the first Data Language Model that natively understands raw tabular data and outperforms gradient-boosted ensembles, AutoML, and prior tabular foundation models on row-level prediction and imputation tasks.

  3. TFM-Retouche: A Lightweight Input-Space Adapter for Tabular Foundation Models

    cs.LG 2026-05 unverdicted novelty 7.0

    TFM-Retouche is an architecture-agnostic input-space residual adapter that improves tabular foundation model accuracy on 51 datasets by learning input corrections through the frozen backbone, with an identity guard to...

  4. Self-Improving Tabular Language Models via Iterative Reward-Guided Post-Training

    cs.LG 2026-04 unverdicted novelty 7.0

    TabGRAA enables self-improving tabular language models through iterative group-relative advantage alignment using modular automated quality signals like distinguishability classifiers.

  5. Tables Guide Vision: Learning to See the Heart through Tabular Data

    cs.CV 2025-03 unverdicted novelty 7.0

    Tabular clinical data guides contrastive learning on cardiac MR images to build better visual representations by identifying patient similarities, outperforming image-only augmentation on downstream disease prediction tasks.

  6. LLM-TabLogic: Preserving Inter-Column Logical Relationships in Synthetic Tabular Data via Prompt-Guided Latent Diffusion

    cs.LG 2025-03 unverdicted novelty 7.0

    LLM-TabLogic extracts inter-column logical constraints using LLMs and conditions a score-based latent diffusion model on them to generate synthetic tabular data that preserves those relationships.

  7. Topological Signatures of Context-Level Reliability in TabPFN

    cs.LG 2026-07 conditional novelty 6.0

    Fragmentation of TabPFN's internal representation topology (H0 zigzag homology) strongly tracks calibration error and Bayes-label disagreement across a six-family synthetic benchmark, with a scale-invariant 'scissors'...

  8. CRUMB: Efficient Prior Fitted Network Inference via Distributionally Matched Context Batching

    cs.LG 2026-06 unverdicted novelty 6.0

    CRUMB speeds up PFN inference on large tabular datasets by clustering queries and selecting MMD-matched context subsets, outperforming prior selection methods on the 51-dataset TabArena benchmark across three architec...

  9. Trajectory-Based Difficulty Scoring for Reliable Learning on Tabular Data

    cs.LG 2026-05 unverdicted novelty 6.0

    TDS uses per-tree prediction trajectories to derive instance difficulty scores that rank errors better than prior hardness measures and improve active learning, selective prediction, and Mondrian conformal prediction ...

  10. TFM-Retouche: A Lightweight Input-Space Adapter for Tabular Foundation Models

    cs.LG 2026-05 unverdicted novelty 6.0

    TFM-Retouche is an input-space residual adapter that lifts TabICLv2 performance by 56 Elo points on 51 tabular datasets while remaining architecture-agnostic and computationally light.

  11. SQuARE: Structured Query & Adaptive Retrieval Engine For Tabular Formats

    cs.CL 2025-12 unverdicted novelty 6.0

    SQuARE is a hybrid retrieval system that uses a complexity score to route tabular queries between chunk-based and SQL-based paths, outperforming single-strategy baselines and GPT-4o on precision and accuracy for compl...

  12. Exploring Differences Between Tabular Enterprise Data and Public Benchmarks

    cs.LG 2026-06 unverdicted novelty 5.0

    Enterprise tabular data differs from public benchmarks in ways that prevent good generalization of models like TabPFN, TabICL, and ConTextTab between the two domains.

  13. Self-Improving Tabular Language Models via Iterative Reward-Guided Post-Training

    cs.LG 2026-04 unverdicted novelty 5.0

    TabGRAA applies group-relative advantage alignment in an iterative reward-guided post-training loop to improve tabular language model generators on fidelity, utility, and privacy trade-offs across five benchmarks.

  14. TREASURE: The Visa Payment Foundation Model for High-Volume Transaction Understanding

    cs.LG 2025-11 unverdicted novelty 5.0

    TREASURE is a transformer model for payment transactions that boosts abnormal behavior detection performance by 111% over production systems and improves recommendation models by 104% when used as an embedding provider.

  15. Noise Immunity in In-Context Tabular Learning: An Empirical Robustness Analysis of TabPFN's Attention Mechanisms

    cs.LG 2026-04 unverdicted novelty 4.0

    TabPFN maintains high ROC-AUC and structured attention under controlled additions of irrelevant features, nonlinear correlations, and mislabeled targets in binary classification.

  16. Creating Artificial Students that Never Existed: Leveraging Large Language Models and CTGANs for Synthetic Data Generation

    cs.LG 2025-01 unverdicted novelty 3.0

    CTGAN and LLMs generate synthetic student data that passes statistical and predictive utility checks for learning analytics.