Pith. sign in

hub Baseline reference

TabICLv2: A better, faster, scalable, and open tabular foundation model.arXiv preprint arXiv:2602.11139

Baseline reference. 57% of citing Pith papers use this work as a benchmark or comparison.

41 Pith papers citing it
Baseline 57% of classified citations

hub tools

citation-role summary

baseline 3 background 2 dataset 1 method 1

citation-polarity summary

years

2026 41

representative citing papers

STRABLE: Benchmarking Tabular Machine Learning with Strings

cs.LG · 2026-05-12 · unverdicted · novelty 8.0

A new corpus of 108 mixed string-numeric tables shows that advanced tabular learners with basic string embeddings perform well on most real-world data, while large LLM encoders help on free-text heavy tables.

Beyond IID: How General Are Tabular Foundation Models, Really?

cs.LG · 2026-06-29 · unverdicted · novelty 7.0

Tabular foundation models excel on tiny- to medium-sized IID data but are outperformed by traditional tree-based and deep learning models on non-IID, large, and high-dimensional datasets, based on evaluations across 11 models and 142 datasets in the new BeyondArena benchmark.

Computational Identifiability

cs.LG · 2026-06-08 · unverdicted · novelty 7.0

The paper defines computational identifiability as success of a finite search procedure in finding an empirical estimator for a causal query within error tolerance, conditional on the search assumptions and procedure.

Speedrunning Tabular Foundation Model Pretraining

cs.LG · 2026-06-02 · unverdicted · novelty 7.0

A speedrun benchmark for nanoTabPFN pretraining reports a record of 0.92 minutes to target performance, an 81x speedup over the 74.32-minute baseline using 22x fewer synthetic datasets.

TabPrep: Closing the Feature Engineering Gap in Tabular Benchmarks

cs.LG · 2026-06-01 · unverdicted · novelty 7.0

TabPrep is a new feature engineering pipeline that targets three data patterns and improves performance of tree-based, neural, linear, and foundation models on tabular benchmarks, often more than model architecture changes.

CalArena: A Large-Scale Post-Hoc Calibration Benchmark

cs.LG · 2026-05-28 · conditional · novelty 7.0

CalArena is a large-scale benchmark that evaluates dozens of post-hoc calibration methods using Post-Hoc Improvement (PHI) in proper scoring rules and finds that smooth functions outperform binning while dedicated multiclass methods are required in high-dimensional settings.

Data Language Models: A New Foundation Model Class for Tabular Data

cs.AI · 2026-05-07 · unverdicted · novelty 7.0

Schema-1 is the first Data Language Model that natively understands raw tabular data and outperforms gradient-boosted ensembles, AutoML, and prior tabular foundation models on row-level prediction and imputation tasks.

TFM-Retouche: A Lightweight Input-Space Adapter for Tabular Foundation Models

cs.LG · 2026-05-07 · unverdicted · novelty 7.0 · 2 refs

TFM-Retouche is an architecture-agnostic input-space residual adapter that improves tabular foundation model accuracy on 51 datasets by learning input corrections through the frozen backbone, with an identity guard to fall back to the original model.

TimEE: End-to-end Time Series Classification via In-Context Learning

cs.LG · 2026-07-08 · conditional · novelty 6.0

A 4.5M-parameter transformer meta-trained on synthetic VARX-generated classification tasks achieves state-of-the-art ROC AUC on the UCR time series classification benchmark via in-context learning with no per-dataset training.

Causal Foundation Models with Continuous Treatments

cs.LG · 2026-05-14 · conditional · novelty 6.0

A transformer meta-trained on a novel continuous-treatment data-generating prior reconstructs individual treatment-response curves from observational data via in-context learning and reports SOTA versus task-specific causal models.

citing papers explorer

Showing 41 of 41 citing papers.