Pith. sign in

REVIEW 2 cited by

On the Performance of Imputation Techniques for Missing Values on Healthcare Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.14687 v1 pith:YNQW7LCH submitted 2024-03-13 cs.LG cs.AI

classification cs.LGcs.AI
keywords imputationmissingvaluesdatadatasetshealthcaremeanperform
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Missing values or data is one popular characteristic of real-world datasets, especially healthcare data. This could be frustrating when using machine learning algorithms on such datasets, simply because most machine learning models perform poorly in the presence of missing values. The aim of this study is to compare the performance of seven imputation techniques, namely Mean imputation, Median Imputation, Last Observation carried Forward (LOCF) imputation, K-Nearest Neighbor (KNN) imputation, Interpolation imputation, Missforest imputation, and Multiple imputation by Chained Equations (MICE), on three healthcare datasets. Some percentage of missing values - 10\%, 15\%, 20\% and 25\% - were introduced into the dataset, and the imputation techniques were employed to impute these missing values. The comparison of their performance was evaluated by using root mean squared error (RMSE) and mean absolute error (MAE). The results show that Missforest imputation performs the best followed by MICE imputation. Additionally, we try to determine whether it is better to perform feature selection before imputation or vice versa by using the following metrics - the recall, precision, f1-score and accuracy. Due to the fact that there are few literature on this and some debate on the subject among researchers, we hope that the results from this experiment will encourage data scientists and researchers to perform imputation first before feature selection when dealing with data containing missing values.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How to rank imputation methods?

    stat.ME 2025-07 conditional novelty 6.0 of 10

    The energy-I-Score ranks imputation methods by comparing observed values with repeated imputations using the energy score, and is claimed to be proper under a new condition, CIMAR_j.

  2. Handling Missing Data in Downstream Tasks With Distribution-Preserving Guarantees

    cs.LG 2025-01 conditional novelty 6.0 of 10

    F3I learns neighbor weights for KNN imputation by maximizing a concave density-ratio objective and comes with high-probability bounds on imputation error and cumulative regret.

Pith tools