Paper Lineage
Esc
MethodMar 2015arXiv 1503.02531stat.ML

Distilling the Knowledge in a Neural Network

Geoffrey Hinton, Oriol Vinyals, Jeff Dean

A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions. Unfortunately, making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users, especially if the individual models are large neural nets.

From the abstract

Built on

0 papers · 0 verifiedSee as graph

Knowledge Distillation has no earlier papers in this dataset.

Led to

  • MobileNets2017 · cited 3×
    “Another method for training small networks is distillation hinton2015distilling which uses a larger network to teach a smaller network.”
    From MobileNets · §Prior Work
  • “JFT-300M is a follow up version of the dataset introduced by [17, 7].”
    From JFT-300M (unreasonable effectiveness) · §The JFT-300M Dataset
  • Instance discrimination2018 · cited 2×, 1 in Method
    “Fig. 1 shows that an image from class leopard is rated much higher by class jaguar rather than by class bookcase hinton2015distilling.”
    From Instance discrimination · §Introduction
  • Invariant & spreading instance features2019 · cited 1×, 1 in Method
    “where τ\tau is the temperature parameter controlling the concentration level of the sample distribution arxiv15temp. 𝐯iT​𝐟jv_{i}^{T}{f_{j}} measures the cosine similarity between the feature 𝐟jf_{j} and the ii-th memo…”
    From Invariant & spreading instance features · §Proposed Method
  • Noisy Student2019 · cited 5×
    “The algorithm is an improved version of self-training, a method in semi-supervised learning (e.g., scudder1965probability; yarowsky1995unsupervised), and distillation hinton2015distilling.”
    From Noisy Student · §Noisy Student Training
  • GPT-32020 · cited 2×
    “This approach includes ALBERT [62] as well as general [44] and task-specific [121, 52, 59] approaches to distillation of language models.”
    From GPT-3 · §Related Work
  • DeiT2020 · cited 2×
    “(KD), introduced by Hinton et al. [24], refers to the training paradigm in which a student model leverages “soft” labels coming from a strong teacher network.”
    From DeiT · §Related work
  • Switch Transformer2021 · cited 2×
    “Further, our large sparse models can be distilled (Hinton et al. 2015) into small dense versions while preserving 30% of the sparse model quality gain.”
    From Switch Transformer · §Introduction
  • Conceptual 12M2021 · cited 1×, 1 in Method
    “We train a Faster-RCNN [68] on Visual Genome [43], with a ResNet101 [29] backbone trained on JFT [30] and fine-tuned on ImageNet [69].”
    From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
  • ALBEF2021 · cited 2×
    “Knowledge distillation [28] aims to improve a student model’s performance by distilling knowledge from a teacher model, usually through matching the student’s prediction with the teacher’s.”
    From ALBEF · §Related Work
Abstract

A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions. Unfortunately, making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users, especially if the individual models are large neural nets. Caruana and his collaborators have shown that it is possible to compress the knowledge in an ensemble into a single model which is much easier to deploy and we develop this approach further using a different compression technique. We achieve some surprising results on MNIST and we show that we can significantly improve the acoustic model of a heavily used commercial system by distilling the knowledge in an ensemble of models into a single model. We also introduce a new type of ensemble composed of one or more full models and many specialist models which learn to distinguish fine-grained classes that the full models confuse. Unlike a mixture of experts, these specialist models can be trained rapidly and in parallel.