REVIEW 5 cited by
In-Context Data Distillation with TabPFN
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Foundation models have revolutionized tasks in computer vision and natural language processing. However, in the realm of tabular data, tree-based models like XGBoost continue to dominate. TabPFN, a transformer model tailored for tabular data, mirrors recent foundation models in its exceptional in-context learning capability, being competitive with XGBoost's performance without the need for task-specific training or hyperparameter tuning. Despite its promise, TabPFN's applicability is hindered by its data size constraint, limiting its use in real-world scenarios. To address this, we present in-context data distillation (ICD), a novel methodology that effectively eliminates these constraints by optimizing TabPFN's context. ICD efficiently enables TabPFN to handle significantly larger datasets with a fixed memory budget, improving TabPFN's quadratic memory complexity but at the cost of a linear number of tuning steps. Notably, TabPFN, enhanced with ICD, demonstrates very strong performance against established tree-based models and modern deep learning methods on 48 large tabular datasets from OpenML.
Forward citations
Cited by 5 Pith papers
-
On Finetuning Tabular Foundation Models
Full finetuning of TabPFNv2 outperforms in-context learning and partial finetuning on medium tabular datasets, and its gains come from sharper query-key attention that better reflects target similarity.
-
TabFlex: Scaling Tabular Learning to Millions with Linear Attention
Linear attention lets a TabPFN-style model process millions of tabular samples in seconds with near-identical accuracy on small datasets.
-
Representation Learning for Tabular Data: A Comprehensive Survey
A comprehensive survey that categorizes deep tabular representation learning into specialized, transferable, and general models, with a feature/sample/objective taxonomy for specialized methods.
-
TabPFN Unleashed: A Scalable and Effective Solution to Tabular Classification Problems
BETA augments TabPFN with encoder fine-tuning and bagging to reduce both bias and variance, achieving SOTA accuracy on 200+ tabular benchmarks while scaling to larger and higher-dimensional data.
-
Drift-Resilient TabPFN: In-Context Learning Temporal Distribution Shifts on Tabular Data
Drift-Resilient TabPFN learns to predict under temporal distribution shifts by pre-training on structural causal models whose edge weights drift over time, improving OOD accuracy and calibration on small tabular datasets.
Discussion (0). Continue with ORCID to comment.