Pith. sign in

REVIEW 8 cited by

ReConTab: Regularized Contrastive Representation Learning for Tabular Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.18541 v2 pith:E2GD7PES submitted 2023-10-28 cs.LG cs.AI

classification cs.LGcs.AI
keywords learningrecontabcontrastiveembeddingsfeaturesrepresentationdatadomain
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Representation learning stands as one of the critical machine learning techniques across various domains. Through the acquisition of high-quality features, pre-trained embeddings significantly reduce input space redundancy, benefiting downstream pattern recognition tasks such as classification, regression, or detection. Nonetheless, in the domain of tabular data, feature engineering and selection still heavily rely on manual intervention, leading to time-consuming processes and necessitating domain expertise. In response to this challenge, we introduce ReConTab, a deep automatic representation learning framework with regularized contrastive learning. Agnostic to any type of modeling task, ReConTab constructs an asymmetric autoencoder based on the same raw features from model inputs, producing low-dimensional representative embeddings. Specifically, regularization techniques are applied for raw feature selection. Meanwhile, ReConTab leverages contrastive learning to distill the most pertinent information for downstream tasks. Experiments conducted on extensive real-world datasets substantiate the framework's capacity to yield substantial and robust performance improvements. Furthermore, we empirically demonstrate that pre-trained embeddings can seamlessly integrate as easily adaptable features, enhancing the performance of various traditional methods such as XGBoost and Random Forest.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Representation Learning for Tabular Data: A Comprehensive Survey

    cs.LG 2025-04 conditional novelty 6.0 of 10

    A comprehensive survey that categorizes deep tabular representation learning into specialized, transferable, and general models, with a feature/sample/objective taxonomy for specialized methods.

  2. APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-Tuning

    cs.LG 2024-12 conditional novelty 6.0 of 10

    APAR pre-trains a tabular transformer on arithmetic combinations of target labels and fine-tunes it with adaptive feature masking, beating GBDT and neural baselines on 10 regression datasets.

  3. MIRRAMS: Learning Robust Tabular Models under Unseen Missingness Shifts

    stat.ML 2025-07 conditional novelty 5.0 of 10

    A training objective built on mutual-information robustness conditions plus extra random masking improves tabular model accuracy under missingness shifts between train and test, with gains also in fully observed settings.

  4. Latte: Transfering LLMs` Latent-level Knowledge for Few-shot Tabular Learning

    cs.LG 2025-05 reject novelty 5.0 of 10

    Latte transfers LLM latent-state knowledge via a knowledge adapter and unsupervised meta-learning, claiming SOTA few-shot tabular performance, though its own evaluation contradicts that claim on Diabetes.

  5. Harnessing LLMs Explanations to Boost Surrogate Models in Tabular Data Classification

    cs.LG 2025-05 conditional novelty 4.0 of 10

    LLM-generated post hoc explanations improve demonstration selection and few-shot accuracy of a smaller surrogate language model on four tabular classification datasets.

  6. TabDeco: A Comprehensive Contrastive Framework for Decoupled Representations in Tabular Data

    cs.LG 2024-11 reject novelty 4.0 of 10

    TabDeco pairs SAINT-style attention with SwitchTab-style feature decoupling and six contrastive losses, but its claim of consistently beating gradient boosting is contradicted by its own results.

  7. Random at First, Fast at Last: NTK-Guided Fourier Pre-Processing for Tabular DL

    cs.LG 2025-06 conditional novelty 3.0 of 10

    Fixed random Fourier projections on tabular inputs are claimed to bound the NTK, speed up gradient descent, and improve accuracy across four architectures and eight benchmarks.

  8. Optimized Coordination Strategy for Multi-Aerospace Systems in Pick-and-Place Tasks By Deep Neural Network

    cs.RO 2024-12 reject novelty 2.0 of 10

    A deep RL policy for multi-agent space debris pick-and-place claims 16% efficiency gains in simulation, but lacks the details needed to evaluate or reproduce the result.

Pith tools