REVIEW 4 cited by
Improving Missing Data Imputation with Deep Generative Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Datasets with missing values are very common on industry applications, and they can have a negative impact on machine learning models. Recent studies introduced solutions to the problem of imputing missing values based on deep generative models. Previous experiments with Generative Adversarial Networks and Variational Autoencoders showed interesting results in this domain, but it is not clear which method is preferable for different use cases. The goal of this work is twofold: we present a comparison between missing data imputation solutions based on deep generative models, and we propose improvements over those methodologies. We run our experiments using known real life datasets with different characteristics, removing values at random and reconstructing them with several imputation techniques. Our results show that the presence or absence of categorical variables can alter the selection of the best model, and that some models are more stable than others after similar runs with different random number generator seeds.
Forward citations
Cited by 4 Pith papers
-
TimeCHEAT: A Channel Harmony Strategy for Irregularly Sampled Multivariate Time Series Analysis
TimeCHEAT combines local channel-dependent embedding via bipartite graph learning with global channel-independent Transformer encoding, beating or matching prior models on several irregularly sampled multivariate time...
-
MuSiCNet: A Gradual Coarse-to-Fine Framework for Irregularly Sampled Multivariate Time Series Analysis
MuSiCNet combines multi-scale attention with Lomb-Scargle periodograms and dynamic time warping to represent irregularly sampled multivariate time series, reporting strong results across three tasks.
-
From Classical Machine Learning to Emerging Foundation Models: Review on Multimodal Data Integration for Cancer Research
A review mapping the transition from classical machine learning to foundation models for multimodal data integration in cancer research.
-
Efficacy Analysis in Clinical Trials: A Comprehensive Review of Statistical and Machine Learning Approaches
A review summarizing parametric, nonparametric, Bayesian, and machine learning methods for efficacy analysis in clinical trials and identifying gaps such as high-dimensional data and missingness.
Discussion (0). Continue with ORCID to comment.