Pith. sign in

REVIEW 1 cited by

Few-shot learning of neural networks from scratch by pseudo example optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1802.03039 v3 pith:VSW3IZNS submitted 2018-02-08 stat.ML cs.LGcs.NE

classification stat.MLcs.LGcs.NE
keywords trainingknowledgemethodmodelamountdatadistillationproposed
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we propose a simple but effective method for training neural networks with a limited amount of training data. Our approach inherits the idea of knowledge distillation that transfers knowledge from a deep or wide reference model to a shallow or narrow target model. The proposed method employs this idea to mimic predictions of reference estimators that are more robust against overfitting than the network we want to train. Different from almost all the previous work for knowledge distillation that requires a large amount of labeled training data, the proposed method requires only a small amount of training data. Instead, we introduce pseudo training examples that are optimized as a part of model parameters. Experimental results for several benchmark datasets demonstrate that the proposed method outperformed all the other baselines, such as naive training of the target model and standard knowledge distillation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

    cs.CL 2019-08 conditional novelty 7.0 of 10

    Pre-training compact BERT models before distillation, called Pre-trained Distillation, outperforms pre-training plus fine-tuning, plain distillation, and more complex compression baselines.

Pith tools