Pith. sign in

REVIEW 3 cited by

Tabular Data: Deep Learning is Not All You Need

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.03253 v2 pith:I7YNGLH6 submitted 2021-06-06 cs.LG

classification cs.LG
keywords modelsdeepxgboostdatadatasetstabularcomparingensemble
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A key element in solving real-life data science problems is selecting the types of models to use. Tree ensemble models (such as XGBoost) are usually recommended for classification and regression problems with tabular data. However, several deep learning models for tabular data have recently been proposed, claiming to outperform XGBoost for some use cases. This paper explores whether these deep models should be a recommended option for tabular data by rigorously comparing the new deep models to XGBoost on various datasets. In addition to systematically comparing their performance, we consider the tuning and computation they require. Our study shows that XGBoost outperforms these deep models across the datasets, including the datasets used in the papers that proposed the deep models. We also demonstrate that XGBoost requires much less tuning. On the positive side, we show that an ensemble of deep models and XGBoost performs better on these datasets than XGBoost alone.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VisTabNet: Adapting Vision Transformers for Tabular Data

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A pre-trained image ViT encoder, fed with learned projections of tabular rows, beats tree ensembles and tabular deep learning baselines on average across 23 small datasets.

  2. PECKER: A Precisely Efficient Critical Knowledge Erasure Recipe For Machine Unlearning in Diffusion Models

    cs.AI 2026-04 unverdicted novelty 5.0 of 10

    Spline numerical encodings can match or beat standard scaling on tabular nets, but PLE is most robust for classification and learnable knots add substantial training cost.

  3. MARBLE: A Multi-Agent Rule-Based LLM Reasoning Engine for Accident Severity Prediction

    cs.AI 2025-07 reject novelty 4.0 of 10

    MARBLE claims near-90% accuracy for accident severity prediction by combining a machine learning model with specialized small language model agents and rule-based coordination, though the comparison to baselines is suspect.

Pith tools