Paper Lineage
Esc
training techniqueIntroduced by DeiT · 2020

Distillation token (data-efficient ViT)

Train ViTs on ImageNet alone by adding a token that learns from a CNN teacher.

Drafted by AI · not yet reviewed

How this idea evolved

Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.

  1. 2014

    Let a decoder look back at the most relevant input positions instead of one fixed vector.

    Cites · not yet reviewedcited 5× · §Model Architecture
    “Most competitive neural sequence transduction models have an encoder-decoder structure [5, 2, 35].”
    From Transformer · §Model Architecture
  2. 2017

    A sequence model built only from attention and feed-forward layers, with no recurrence.

    Also draws on: Encoder–decoder seq2seq (Seq2Seq)

    Cites · not yet reviewedcited 4× · §Method
    “In model design we follow the original Transformer (Vaswani et al. 2017) as closely as possible.”
    From ViT · §Method
  3. 2020

    Split an image into patches and feed them to a plain Transformer as tokens.

    Cites · not yet reviewedcited 12× · §Introduction
    “We build upon the visual transformer architecture from Dosovitskiy et al. [15] and improvements included in the timm library [55].”
    From DeiT · §Introduction
  4. 2020

    Train ViTs on ImageNet alone by adding a token that learns from a CNN teacher.

    Also draws on: Knowledge distillation (Knowledge Distillation)

Papers using this