Pith. sign in

REVIEW 4 cited by

Review for Handling Missing Data with special missing mechanism

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.04905 v1 pith:CF7L5XU3 submitted 2024-04-07 stat.ME cs.AIcs.LG

classification stat.MEcs.AIcs.LG
keywords missingdataexistinghandlehandlingliteraturemechanismsrandom
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Missing data poses a significant challenge in data science, affecting decision-making processes and outcomes. Understanding what missing data is, how it occurs, and why it is crucial to handle it appropriately is paramount when working with real-world data, especially in tabular data, one of the most commonly used data types in the real world. Three missing mechanisms are defined in the literature: Missing Completely At Random (MCAR), Missing At Random (MAR), and Missing Not At Random (MNAR), each presenting unique challenges in imputation. Most existing work are focused on MCAR that is relatively easy to handle. The special missing mechanisms of MNAR and MAR are less explored and understood. This article reviews existing literature on handling missing values. It compares and contrasts existing methods in terms of their ability to handle different missing mechanisms and data types. It identifies research gap in the existing literature and lays out potential directions for future research in the field. The information in this review will help data analysts and researchers to adopt and promote good practices for handling missing data in real-world problems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 8 citations worldwide. Full citation record

  1. SHIFT: Survival Prediction from Incomplete and Heterogeneous Genomic Data

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A masked-attention survival model trained with variable-rate feature masking predicts from incomplete genomic panels without imputation and transfers across institutions with mismatched gene sets.

  2. GPUs, CPUs, and... NICs: Rethinking the Network's Role in Serving Complex AI Pipelines

    cs.NI 2025-01 conditional novelty 5.0 of 10

    The paper proposes offloading AI pipeline data processing tasks to SmartNICs and sketches designs for normalization, bilinear interpolation, and tokenization, without implementing them.

  3. MissMecha: An All-in-One Python Package for Studying Missing Data Mechanisms

    cs.LG 2025-08 conditional novelty 4.0 of 10

    MissMecha is a Python toolkit combining simulation, visualization, statistical testing, and evaluation of missing data mechanisms for mixed-type tabular data.

  4. Navigating Data Corruption in Machine Learning: Balancing Quality, Quantity, and Imputation Strategies

    cs.LG 2024-12 reject novelty 4.0 of 10

    Performance under data corruption follows an exponential diminishing-return curve, noise harms more than missingness, imputation helps only when accurate, and adding data cannot fully compensate.

Pith tools