Pith. sign in

REVIEW 2 cited by

Hydra: Preserving Ensemble Diversity for Model Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2001.04694 v2 pith:FJ4ALNRD submitted 2020-01-14 cs.LG stat.ML

classification cs.LGstat.ML
keywords ensembledistillationbehaviordiversityhydramodelpredictiveuncertainty
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Ensembles of models have been empirically shown to improve predictive performance and to yield robust measures of uncertainty. However, they are expensive in computation and memory. Therefore, recent research has focused on distilling ensembles into a single compact model, reducing the computational and memory burden of the ensemble while trying to preserve its predictive behavior. Most existing distillation formulations summarize the ensemble by capturing its average predictions. As a result, the diversity of the ensemble predictions, stemming from each member, is lost. Thus, the distilled model cannot provide a measure of uncertainty comparable to that of the original ensemble. To retain more faithfully the diversity of the ensemble, we propose a distillation method based on a single multi-headed neural network, which we refer to as Hydra. The shared body network learns a joint feature representation that enables each head to capture the predictive behavior of each ensemble member. We demonstrate that with a slight increase in parameter count, Hydra improves distillation performance on classification and regression settings while capturing the uncertainty behavior of the original ensemble over both in-domain and out-of-distribution tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generative Data Mining with Longtail-Guided Diffusion

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A diffusion model guided by a classifier's uncertainty signal generates hard but in-distribution images, and fine-tuning on them improves accuracy, especially on rare classes.

  2. Function Space Diversity for Uncertainty Prediction via Repulsive Last-Layer Ensembles

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A repulsive last-layer ensemble trained with function-space diversity on OOD or augmented samples gives competitive uncertainty estimates at a fraction of deep-ensemble cost.

Pith tools