REVIEW 8 cited by
ReConTab: Regularized Contrastive Representation Learning for Tabular Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Representation learning stands as one of the critical machine learning techniques across various domains. Through the acquisition of high-quality features, pre-trained embeddings significantly reduce input space redundancy, benefiting downstream pattern recognition tasks such as classification, regression, or detection. Nonetheless, in the domain of tabular data, feature engineering and selection still heavily rely on manual intervention, leading to time-consuming processes and necessitating domain expertise. In response to this challenge, we introduce ReConTab, a deep automatic representation learning framework with regularized contrastive learning. Agnostic to any type of modeling task, ReConTab constructs an asymmetric autoencoder based on the same raw features from model inputs, producing low-dimensional representative embeddings. Specifically, regularization techniques are applied for raw feature selection. Meanwhile, ReConTab leverages contrastive learning to distill the most pertinent information for downstream tasks. Experiments conducted on extensive real-world datasets substantiate the framework's capacity to yield substantial and robust performance improvements. Furthermore, we empirically demonstrate that pre-trained embeddings can seamlessly integrate as easily adaptable features, enhancing the performance of various traditional methods such as XGBoost and Random Forest.
Forward citations
Cited by 8 Pith papers
-
Representation Learning for Tabular Data: A Comprehensive Survey
A comprehensive survey that categorizes deep tabular representation learning into specialized, transferable, and general models, with a feature/sample/objective taxonomy for specialized methods.
-
APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-Tuning
APAR pre-trains a tabular transformer on arithmetic combinations of target labels and fine-tunes it with adaptive feature masking, beating GBDT and neural baselines on 10 regression datasets.
-
MIRRAMS: Learning Robust Tabular Models under Unseen Missingness Shifts
A training objective built on mutual-information robustness conditions plus extra random masking improves tabular model accuracy under missingness shifts between train and test, with gains also in fully observed settings.
-
Latte: Transfering LLMs` Latent-level Knowledge for Few-shot Tabular Learning
Latte transfers LLM latent-state knowledge via a knowledge adapter and unsupervised meta-learning, claiming SOTA few-shot tabular performance, though its own evaluation contradicts that claim on Diabetes.
-
Harnessing LLMs Explanations to Boost Surrogate Models in Tabular Data Classification
LLM-generated post hoc explanations improve demonstration selection and few-shot accuracy of a smaller surrogate language model on four tabular classification datasets.
-
TabDeco: A Comprehensive Contrastive Framework for Decoupled Representations in Tabular Data
TabDeco pairs SAINT-style attention with SwitchTab-style feature decoupling and six contrastive losses, but its claim of consistently beating gradient boosting is contradicted by its own results.
-
Random at First, Fast at Last: NTK-Guided Fourier Pre-Processing for Tabular DL
Fixed random Fourier projections on tabular inputs are claimed to bound the NTK, speed up gradient descent, and improve accuracy across four architectures and eight benchmarks.
-
Optimized Coordination Strategy for Multi-Aerospace Systems in Pick-and-Place Tasks By Deep Neural Network
A deep RL policy for multi-agent space debris pick-and-place claims 16% efficiency gains in simulation, but lacks the details needed to evaluate or reproduce the result.
Discussion (0). Continue with ORCID to comment.