Pith. sign in

REVIEW 2 cited by

ReMasker: Imputing Tabular Data with Masked Autoencoding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.13793 v1 pith:NDBI5IWT submitted 2023-09-25 cs.LG

classification cs.LG
keywords dataremaskermaskedmissingtabularvaluesautoencodingfurther
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present ReMasker, a new method of imputing missing values in tabular data by extending the masked autoencoding framework. Compared with prior work, ReMasker is both simple -- besides the missing values (i.e., naturally masked), we randomly ``re-mask'' another set of values, optimize the autoencoder by reconstructing this re-masked set, and apply the trained model to predict the missing values; and effective -- with extensive evaluation on benchmark datasets, we show that ReMasker performs on par with or outperforms state-of-the-art methods in terms of both imputation fidelity and utility under various missingness settings, while its performance advantage often increases with the ratio of missing data. We further explore theoretical justification for its effectiveness, showing that ReMasker tends to learn missingness-invariant representations of tabular data. Our findings indicate that masked modeling represents a promising direction for further research on tabular data imputation. The code is publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LSM-2: Learning from Incomplete Wearable Sensor Data

    cs.LG 2025-06 conditional novelty 6.0 of 10

    LSM-2 with Adaptive and Inherited Masking learns usable representations directly from incomplete day-long wearable data, outperforming imputation-based baselines on most tasks.

  2. Handling Missing Data in Downstream Tasks With Distribution-Preserving Guarantees

    cs.LG 2025-01 conditional novelty 6.0 of 10

    F3I learns neighbor weights for KNN imputation by maximizing a concave density-ratio objective and comes with high-probability bounds on imputation error and cumulative regret.

Pith tools