Pith. sign in

REVIEW 1 cited by

Rethinking Pre-Training in Tabular Data: A Neighborhood Embedding Perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.00055 v2 pith:RKVLTAGR submitted 2023-10-31 cs.LG

classification cs.LG
keywords datadatasetstabularlabelspre-trainingtabptmtasksdiverse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-training is prevalent in deep learning for vision and text data, leveraging knowledge from other datasets to enhance downstream tasks. However, for tabular data, the inherent heterogeneity in attribute and label spaces across datasets complicates the learning of shareable knowledge. We propose Tabular data Pre-Training via Meta-representation (TabPTM), aiming to pre-train a general tabular model over diverse datasets. The core idea is to embed data instances into a shared feature space, where each instance is represented by its distance to a fixed number of nearest neighbors and their labels. This ''meta-representation'' transforms heterogeneous tasks into homogeneous local prediction problems, enabling the model to infer labels (or scores for each label) based on neighborhood information. As a result, the pre-trained TabPTM can be applied directly to new datasets, regardless of their diverse attributes and labels, without further fine-tuning. Extensive experiments on 101 datasets confirm TabPTM's effectiveness in both classification and regression tasks, with and without fine-tuning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TabPFN Unleashed: A Scalable and Effective Solution to Tabular Classification Problems

    cs.LG 2025-02 conditional novelty 6.0 of 10

    BETA augments TabPFN with encoder fine-tuning and bagging to reduce both bias and variance, achieving SOTA accuracy on 200+ tabular benchmarks while scaling to larger and higher-dimensional data.

Pith tools