Pith. sign in

REVIEW 4 cited by

Handling Incomplete Heterogeneous Data using VAEs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1807.03653 v4 pith:YDQNK3NN submitted 2018-07-10 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords dataincompletevaesmodelsaccurateheterogenoushi-vaemissing
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Variational autoencoders (VAEs), as well as other generative models, have been shown to be efficient and accurate for capturing the latent structure of vast amounts of complex high-dimensional data. However, existing VAEs can still not directly handle data that are heterogenous (mixed continuous and discrete) or incomplete (with missing data at random), which is indeed common in real-world applications. In this paper, we propose a general framework to design VAEs suitable for fitting incomplete heterogenous data. The proposed HI-VAE includes likelihood models for real-valued, positive real valued, interval, categorical, ordinal and count data, and allows accurate estimation (and potentially imputation) of missing data. Furthermore, HI-VAE presents competitive predictive performance in supervised tasks, outperforming supervised models when trained on incomplete data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CACTI: Leveraging Copy Masking and Contextual Information to Improve Tabular Data Imputation

    cs.LG 2025-06 conditional novelty 6.0 of 10

    CACTI combines median-truncated copy masking with language-model column embeddings to improve tabular imputation accuracy across MCAR, MAR, and MNAR missingness.

  2. Neighbour-Driven Gaussian Process Variational Autoencoders for Scalable Structured Latent Modelling

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Two neighbour-based approximations, HPA and SPA, make Gaussian Process variational autoencoders scalable while preserving local latent correlations.

  3. Icebreaker: Element-wise Active Information Acquisition with Bayesian Deep Latent Gaussian Model

    cs.LG 2019-08 conditional novelty 6.0 of 10

    A Bayesian deep latent variable model with active, element-wise data acquisition lets machine learning systems learn from far fewer costly measurements.

  4. Scalable Machine Learning Algorithms using Path Signatures

    stat.ML 2025-06 conditional novelty 4.0 of 10

    Path signatures can be embedded in Gaussian process, deep learning, kernel, and graph diffusion models to match or beat established baselines on time series and graph benchmarks.

Pith tools