{"work":{"id":"d927ab1f-17b8-4002-9d09-c3d55764fbad","openalex_id":"https://openalex.org/W4402727638","doi":"10.1109/cvpr52733.2024.01515","arxiv_id":"1503.02531","raw_key":null,"title":"Distilling the Knowledge in a Neural Network","authors":null,"authors_text":"Geoffrey Hinton, Oriol Vinyals, and Jeff Dean","year":2015,"venue":"stat.ML","abstract":"A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions. Unfortunately, making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users, especially if the individual models are large neural nets. Caruana and his collaborators have shown that it is possible to compress the knowledge in an ensemble into a single model which is much easier to deploy and we develop this approach further using a different compression technique. We achieve some surprising results on MNIST and we show that we can significantly improve the acoustic model of a heavily used commercial system by distilling the knowledge in an ensemble of models into a single model. We also introduce a new type of ensemble composed of one or more full models and many specialist models which learn to distinguish fine-grained classes that the full models confuse. Unlike a mixture of experts, these specialist models can be trained rapidly and in parallel.","external_url":"https://arxiv.org/abs/1503.02531","cited_by_count":28,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"1503.02531","created_at":"2026-05-08T16:53:30.132726+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Distilling the Knowledge in a Neural Network","render_title":"Distilling the Knowledge in a Neural Network"},"hub":{"state":{"work_id":"d927ab1f-17b8-4002-9d09-c3d55764fbad","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":747,"external_cited_by_count":28,"distinct_field_count":38,"first_pith_cited_at":"2016-06-15T08:20:51+00:00","last_pith_cited_at":"2026-07-09T09:10:49+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T05:09:21.352242+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":72},{"context_role":"method","n":14},{"context_role":"other","n":2},{"context_role":"dataset","n":1}],"polarity_counts":[{"context_polarity":"background","n":71},{"context_polarity":"use_method","n":13},{"context_polarity":"unclear","n":3},{"context_polarity":"support","n":1},{"context_polarity":"use_dataset","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Distilling the Knowledge in a Neural Network","claims":[{"claim_text":"A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions. Unfortunately, making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users, especially if the individual models are large neural nets. Caruana and his collaborators have shown that it is possible to compress the knowledge in an ensemble into a single model which is much easier to deploy and we develop this approach further using","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Distilling the Knowledge in a Neural Network because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-13T19:33:27.111749+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"8b2cc23b-b0d5-44c1-9db1-00fbb88c798c","orcid":null,"display_name":"Geoffrey Hinton"},{"id":"4398f623-cc78-4ab8-ad6b-a4676a455680","orcid":null,"display_name":"Oriol Vinyals"},{"id":"c8675e7b-8d91-4ff7-98cd-e41a9fe4b32d","orcid":null,"display_name":"and Jeff Dean"}]},"error":null,"updated_at":"2026-05-13T19:33:27.110026+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-13T19:23:32.881405+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":29},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":29},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":23},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":22},{"title":"DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter","work_id":"756f9764-ecd6-4672-8043-b37c698c7ad2","shared_citers":22},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":19},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":18},{"title":"Scaling Laws for Neural Language Models","work_id":"b7dd8749-9c45-4977-ab9b-64478dce1ae8","shared_citers":18},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":17},{"title":"MiniLLM: On-Policy Distillation of Large Language Models","work_id":"16edb291-dd18-41c5-8486-c6c715ec5311","shared_citers":16},{"title":"Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge","work_id":"28ea1282-d657-4c61-a83c-f1249be6d6b1","shared_citers":16},{"title":"FitNets: Hints for thin deep nets","work_id":"057e16e2-a9fd-4a16-8fa3-8806ac7d3165","shared_citers":15},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":15},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":14},{"title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","work_id":"41fe12c4-e538-4890-a244-480650ed3078","shared_citers":14},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":13},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":13},{"title":"Self-Distillation Enables Continual Learning","work_id":"e9aa25e3-870c-46c8-8270-e4e5948d09f0","shared_citers":13},{"title":"Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models","work_id":"bae00e84-9b0d-433d-a066-20b951f0b4d0","shared_citers":13},{"title":"Program Synthesis with Large Language Models","work_id":"fd241a05-03b9-4de2-9588-9d77ce176125","shared_citers":12},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":11},{"title":"Reinforcement Learning via Self-Distillation","work_id":"b193541d-5853-4ea4-8e4b-8e4c08617eb6","shared_citers":11},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":10},{"title":"A Survey of On-Policy Distillation for Large Language Models","work_id":"f6aaea8e-1f0d-43e3-b28f-6066d3e0a66b","shared_citers":10}],"time_series":[{"n":1,"year":2016},{"n":1,"year":2017},{"n":3,"year":2019},{"n":2,"year":2020},{"n":1,"year":2021},{"n":2,"year":2022},{"n":3,"year":2023},{"n":3,"year":2024},{"n":2,"year":2025},{"n":208,"year":2026}]},"error":null,"updated_at":"2026-05-13T19:33:26.690493+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"fixed":1,"items":[{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-13T19:23:32.200230+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Distilling the Knowledge in a Neural Network","claims":[{"claim_text":"A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions. Unfortunately, making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users, especially if the individual models are large neural nets. Caruana and his collaborators have shown that it is possible to compress the knowledge in an ensemble into a single model which is much easier to deploy and we develop this approach further using","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Distilling the Knowledge in a Neural Network because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-13T19:33:26.619929+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Distilling the Knowledge in a Neural Network","claims":[{"claim_text":"A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions. Unfortunately, making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users, especially if the individual models are large neural nets. Caruana and his collaborators have shown that it is possible to compress the knowledge in an ensemble into a single model which is much easier to deploy and we develop this approach further using","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Distilling the Knowledge in a Neural Network because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-13T19:23:32.883222+00:00"}},"summary":{"title":"Distilling the Knowledge in a Neural Network","claims":[{"claim_text":"A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions. Unfortunately, making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users, especially if the individual models are large neural nets. Caruana and his collaborators have shown that it is possible to compress the knowledge in an ensemble into a single model which is much easier to deploy and we develop this approach further using","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Distilling the Knowledge in a Neural Network because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":29},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":29},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":23},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":22},{"title":"DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter","work_id":"756f9764-ecd6-4672-8043-b37c698c7ad2","shared_citers":22},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":19},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":18},{"title":"Scaling Laws for Neural Language Models","work_id":"b7dd8749-9c45-4977-ab9b-64478dce1ae8","shared_citers":18},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":17},{"title":"MiniLLM: On-Policy Distillation of Large Language Models","work_id":"16edb291-dd18-41c5-8486-c6c715ec5311","shared_citers":16},{"title":"Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge","work_id":"28ea1282-d657-4c61-a83c-f1249be6d6b1","shared_citers":16},{"title":"FitNets: Hints for thin deep nets","work_id":"057e16e2-a9fd-4a16-8fa3-8806ac7d3165","shared_citers":15},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":15},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":14},{"title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","work_id":"41fe12c4-e538-4890-a244-480650ed3078","shared_citers":14},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":13},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":13},{"title":"Self-Distillation Enables Continual Learning","work_id":"e9aa25e3-870c-46c8-8270-e4e5948d09f0","shared_citers":13},{"title":"Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models","work_id":"bae00e84-9b0d-433d-a066-20b951f0b4d0","shared_citers":13},{"title":"Program Synthesis with Large Language Models","work_id":"fd241a05-03b9-4de2-9588-9d77ce176125","shared_citers":12},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":11},{"title":"Reinforcement Learning via Self-Distillation","work_id":"b193541d-5853-4ea4-8e4b-8e4c08617eb6","shared_citers":11},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":10},{"title":"A Survey of On-Policy Distillation for Large Language Models","work_id":"f6aaea8e-1f0d-43e3-b28f-6066d3e0a66b","shared_citers":10}],"time_series":[{"n":1,"year":2016},{"n":1,"year":2017},{"n":3,"year":2019},{"n":2,"year":2020},{"n":1,"year":2021},{"n":2,"year":2022},{"n":3,"year":2023},{"n":3,"year":2024},{"n":2,"year":2025},{"n":208,"year":2026}]},"authors":[{"id":"c8675e7b-8d91-4ff7-98cd-e41a9fe4b32d","orcid":null,"display_name":"and Jeff Dean","source":"manual","import_confidence":0.72},{"id":"8b2cc23b-b0d5-44c1-9db1-00fbb88c798c","orcid":null,"display_name":"Geoffrey Hinton","source":"manual","import_confidence":0.72},{"id":"4398f623-cc78-4ab8-ad6b-a4676a455680","orcid":null,"display_name":"Oriol Vinyals","source":"manual","import_confidence":0.72}]}}