Pith. sign in

REVIEW 4 cited by

Ensemble Distribution Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.00076 v3 pith:DSMBZM4K submitted 2019-04-30 stat.ML cs.LG

classification stat.MLcs.LG
keywords ensembledistillationmodelsingledistributionuncertaintyemphperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Ensembles of models often yield improvements in system performance. These ensemble approaches have also been empirically shown to yield robust measures of uncertainty, and are capable of distinguishing between different \emph{forms} of uncertainty. However, ensembles come at a computational and memory cost which may be prohibitive for many applications. There has been significant work done on the distillation of an ensemble into a single model. Such approaches decrease computational cost and allow a single model to achieve an accuracy comparable to that of an ensemble. However, information about the \emph{diversity} of the ensemble, which can yield estimates of different forms of uncertainty, is lost. This work considers the novel task of \emph{Ensemble Distribution Distillation} (EnD$^2$) --- distilling the distribution of the predictions from an ensemble, rather than just the average prediction, into a single model. EnD$^2$ enables a single model to retain both the improved classification performance of ensemble distillation as well as information about the diversity of the ensemble, which is useful for uncertainty estimation. A solution for EnD$^2$ based on Prior Networks, a class of models which allow a single neural network to explicitly model a distribution over output distributions, is proposed in this work. The properties of EnD$^2$ are investigated on both an artificial dataset, and on the CIFAR-10, CIFAR-100 and TinyImageNet datasets, where it is shown that EnD$^2$ can approach the classification performance of an ensemble, and outperforms both standard DNNs and Ensemble Distillation on the tasks of misclassification and out-of-distribution input detection.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Factor-Informed Uncertainty Distillation for Gaze Estimation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    FIUD distills image-quality-based error predictions into a single-pass gaze uncertainty head, improving Spearman rank correlation and selective prediction versus ensembles, MC dropout, and heteroscedastic NLL baselines.

  2. Density-Informed Pseudo-Counts for Calibrated Evidential Deep Learning

    stat.ML 2026-02 conditional novelty 5.0 of 10

    DIP-EDL sets Dirichlet pseudo-counts to the product of marginal input density and learned class probabilities, concentrating on the true label distribution while sending OOD inputs to the prior.

  3. Ensemble Distribution Distillation for Self-Supervised Human Activity Recognition

    cs.LG 2025-09 conditional novelty 5.0 of 10

    A single prior network distilled from a 50-member self-supervised ensemble matches ensemble accuracy and robustness on HAR benchmarks at single-model inference cost.

  4. Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Two probability-only confidence metrics, a top-to-kth beam ratio and a tail-thinness score, improve quality correlation for BART and Flan-T5 on several summarization, translation, and QA datasets.

Pith tools