Pith. sign in

ReMasker: Imputing Tabular Data with Masked Autoencoding

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We present ReMasker, a new method of imputing missing values in tabular data by extending the masked autoencoding framework. Compared with prior work, ReMasker is both simple -- besides the missing values (i.e., naturally masked), we randomly ``re-mask'' another set of values, optimize the autoencoder by reconstructing this re-masked set, and apply the trained model to predict the missing values; and effective -- with extensive evaluation on benchmark datasets, we show that ReMasker performs on par with or outperforms state-of-the-art methods in terms of both imputation fidelity and utility under various missingness settings, while its performance advantage often increases with the ratio of missing data. We further explore theoretical justification for its effectiveness, showing that ReMasker tends to learn missingness-invariant representations of tabular data. Our findings indicate that masked modeling represents a promising direction for further research on tabular data imputation. The code is publicly available.

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

LSM-2: Learning from Incomplete Wearable Sensor Data

cs.LG · 2025-06-05 · conditional · novelty 6.0

LSM-2 with Adaptive and Inherited Masking learns usable representations directly from incomplete day-long wearable data, outperforming imputation-based baselines on most tasks.

citing papers explorer

Showing 1 of 1 citing paper.

  • LSM-2: Learning from Incomplete Wearable Sensor Data cs.LG · 2025-06-05 · conditional · none · ref 23 · internal anchor

    LSM-2 with Adaptive and Inherited Masking learns usable representations directly from incomplete day-long wearable data, outperforming imputation-based baselines on most tasks.