Pith. sign in

REVIEW 1 cited by

Efficient EM Training of Gaussian Mixtures with Missing Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1209.0521 v2 pith:7T5ZZJFN submitted 2012-09-04 cs.LG stat.ML

classification cs.LGstat.ML
keywords missingtrainingalgorithmdatadiscriminantgaussiangenerativelearning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In data-mining applications, we are frequently faced with a large fraction of missing entries in the data matrix, which is problematic for most discriminant machine learning algorithms. A solution that we explore in this paper is the use of a generative model (a mixture of Gaussians) to compute the conditional expectation of the missing variables given the observed variables. Since training a Gaussian mixture with many different patterns of missing values can be computationally very expensive, we introduce a spanning-tree based algorithm that significantly speeds up training in these conditions. We also observe that good results can be obtained by using the generative model to fill-in the missing values for a separate discriminant learning algorithm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mixture-based Multiple Imputation Model for Clinical Data with a Temporal Dimension

    cs.LG 2019-08 conditional novelty 6.0 of 10

    MixMI, a mixture of Gaussian-process and linear-regression imputers with individualized mixing weights, reports lower mean absolute scaled error than six benchmarks on all four datasets tested.

Pith tools