REVIEW 4 cited by
Review for Handling Missing Data with special missing mechanism
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Missing data poses a significant challenge in data science, affecting decision-making processes and outcomes. Understanding what missing data is, how it occurs, and why it is crucial to handle it appropriately is paramount when working with real-world data, especially in tabular data, one of the most commonly used data types in the real world. Three missing mechanisms are defined in the literature: Missing Completely At Random (MCAR), Missing At Random (MAR), and Missing Not At Random (MNAR), each presenting unique challenges in imputation. Most existing work are focused on MCAR that is relatively easy to handle. The special missing mechanisms of MNAR and MAR are less explored and understood. This article reviews existing literature on handling missing values. It compares and contrasts existing methods in terms of their ability to handle different missing mechanisms and data types. It identifies research gap in the existing literature and lays out potential directions for future research in the field. The information in this review will help data analysts and researchers to adopt and promote good practices for handling missing data in real-world problems.
Forward citations
Cited by 4 Pith papers
-
SHIFT: Survival Prediction from Incomplete and Heterogeneous Genomic Data
A masked-attention survival model trained with variable-rate feature masking predicts from incomplete genomic panels without imputation and transfers across institutions with mismatched gene sets.
-
GPUs, CPUs, and... NICs: Rethinking the Network's Role in Serving Complex AI Pipelines
The paper proposes offloading AI pipeline data processing tasks to SmartNICs and sketches designs for normalization, bilinear interpolation, and tokenization, without implementing them.
-
MissMecha: An All-in-One Python Package for Studying Missing Data Mechanisms
MissMecha is a Python toolkit combining simulation, visualization, statistical testing, and evaluation of missing data mechanisms for mixed-type tabular data.
-
Navigating Data Corruption in Machine Learning: Balancing Quality, Quantity, and Imputation Strategies
Performance under data corruption follows an exponential diminishing-return curve, noise harms more than missingness, imputation helps only when accurate, and adding data cannot fully compensate.
Discussion (0). Continue with ORCID to comment.