Pith. sign in

REVIEW 1 cited by

What to Expect of Classifiers? Reasoning about Logistic Regression with Missing Features

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1903.01620 v2 pith:BMCVFJB7 submitted 2019-03-05 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords featuresmissinglogisticregressionclassifiersdistributionexpectedprediction
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

While discriminative classifiers often yield strong predictive performance, missing feature values at prediction time can still be a challenge. Classifiers may not behave as expected under certain ways of substituting the missing values, since they inherently make assumptions about the data distribution they were trained on. In this paper, we propose a novel framework that classifies examples with missing features by computing the expected prediction with respect to a feature distribution. Moreover, we use geometric programming to learn a naive Bayes distribution that embeds a given logistic regression classifier and can efficiently take its expected predictions. Empirical evaluations show that our model achieves the same performance as the logistic regression with all features observed, and outperforms standard imputation techniques when features go missing during prediction time. Furthermore, we demonstrate that our method can be used to generate "sufficient explanations" of logistic regression classifications, by removing features that do not affect the classification.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning High-dimensional Gaussians from Censored Data

    cs.LG 2025-04 conditional novelty 7.0 of 10

    Efficient algorithms recover the mean and covariance of a high-dimensional Gaussian from samples censored by known self-censoring or linear-thresholding missingness rules.

Pith tools