Distilling the Knowledge in a Neural Network
A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions. Unfortunately, making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users, especially if the individual models are large neural nets.
Knowledge Distillation has no earlier papers in this dataset.
Led to
- MobileNets2017 · cited 3דAnother method for training small networks is distillation hinton2015distilling which uses a larger network to teach a smaller network.”From MobileNets · §Prior Work
- JFT-300M (unreasonable effectiveness)2017 · cited 2דJFT-300M is a follow up version of the dataset introduced by [17, 7].”From JFT-300M (unreasonable effectiveness) · §The JFT-300M Dataset
- Instance discrimination2018 · cited 2×, 1 in Method“Fig. 1 shows that an image from class leopard is rated much higher by class jaguar rather than by class bookcase hinton2015distilling.”From Instance discrimination · §Introduction
- Invariant & spreading instance features2019 · cited 1×, 1 in Method“where τ\tau is the temperature parameter controlling the concentration level of the sample distribution arxiv15temp. 𝐯iT𝐟jv_{i}^{T}{f_{j}} measures the cosine similarity between the feature 𝐟jf_{j} and the ii-th memo…”From Invariant & spreading instance features · §Proposed Method
- Noisy Student2019 · cited 5דThe algorithm is an improved version of self-training, a method in semi-supervised learning (e.g., scudder1965probability; yarowsky1995unsupervised), and distillation hinton2015distilling.”From Noisy Student · §Noisy Student Training
- GPT-32020 · cited 2דThis approach includes ALBERT [62] as well as general [44] and task-specific [121, 52, 59] approaches to distillation of language models.”From GPT-3 · §Related Work
- DeiT2020 · cited 2ד(KD), introduced by Hinton et al. [24], refers to the training paradigm in which a student model leverages “soft” labels coming from a strong teacher network.”From DeiT · §Related work
- Switch Transformer2021 · cited 2דFurther, our large sparse models can be distilled (Hinton et al. 2015) into small dense versions while preserving 30% of the sparse model quality gain.”From Switch Transformer · §Introduction
- Conceptual 12M2021 · cited 1×, 1 in Method“We train a Faster-RCNN [68] on Visual Genome [43], with a ResNet101 [29] backbone trained on JFT [30] and fine-tuned on ImageNet [69].”From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
- ALBEF2021 · cited 2דKnowledge distillation [28] aims to improve a student model’s performance by distilling knowledge from a teacher model, usually through matching the student’s prediction with the teacher’s.”From ALBEF · §Related Work
Abstract
A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions. Unfortunately, making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users, especially if the individual models are large neural nets. Caruana and his collaborators have shown that it is possible to compress the knowledge in an ensemble into a single model which is much easier to deploy and we develop this approach further using a different compression technique. We achieve some surprising results on MNIST and we show that we can significantly improve the acoustic model of a heavily used commercial system by distilling the knowledge in an ensemble of models into a single model. We also introduce a new type of ensemble composed of one or more full models and many specialist models which learn to distinguish fine-grained classes that the full models confuse. Unlike a mixture of experts, these specialist models can be trained rapidly and in parallel.