REVIEW 4 cited by
Handling Incomplete Heterogeneous Data using VAEs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Variational autoencoders (VAEs), as well as other generative models, have been shown to be efficient and accurate for capturing the latent structure of vast amounts of complex high-dimensional data. However, existing VAEs can still not directly handle data that are heterogenous (mixed continuous and discrete) or incomplete (with missing data at random), which is indeed common in real-world applications. In this paper, we propose a general framework to design VAEs suitable for fitting incomplete heterogenous data. The proposed HI-VAE includes likelihood models for real-valued, positive real valued, interval, categorical, ordinal and count data, and allows accurate estimation (and potentially imputation) of missing data. Furthermore, HI-VAE presents competitive predictive performance in supervised tasks, outperforming supervised models when trained on incomplete data.
Forward citations
Cited by 4 Pith papers
-
CACTI: Leveraging Copy Masking and Contextual Information to Improve Tabular Data Imputation
CACTI combines median-truncated copy masking with language-model column embeddings to improve tabular imputation accuracy across MCAR, MAR, and MNAR missingness.
-
Neighbour-Driven Gaussian Process Variational Autoencoders for Scalable Structured Latent Modelling
Two neighbour-based approximations, HPA and SPA, make Gaussian Process variational autoencoders scalable while preserving local latent correlations.
-
Icebreaker: Element-wise Active Information Acquisition with Bayesian Deep Latent Gaussian Model
A Bayesian deep latent variable model with active, element-wise data acquisition lets machine learning systems learn from far fewer costly measurements.
-
Scalable Machine Learning Algorithms using Path Signatures
Path signatures can be embedded in Gaussian process, deep learning, kernel, and graph diffusion models to match or beat established baselines on time series and graph benchmarks.
Discussion (0). Continue with ORCID to comment.