Paper Lineage
Esc

Vision Transformers

Transformers applied to images and video, and how to train and scale them.

Concepts in Vision Transformers
ConceptIntroduced byYearsPapers
Distillation token (data-efficient ViT)
Train ViTs on ImageNet alone by adding a token that learns from a CNN teacher.
DeiT2020–20213
Vision Transformer (ViT)
Split an image into patches and feed them to a plain Transformer as tokens.
ViT2020–202327
Convolution–attention hybrids
Stack convolution stages before attention stages to get both inductive bias and capacity.
CoAtNet20211
Local video Transformers
Extend windowed attention across space and time for video recognition.
Video Swin20210
Residual post-norm & cosine attention
Stabilise very large vision Transformers and transfer them across resolutions.
Swin V220210
Scaling laws for vision Transformers
How ViT error falls with model size, data and compute, up to billions of parameters.
Scaling ViTs (ViT-G)20212
Shifted-window hierarchical ViT
Attention inside local windows that shift between layers, with multi-scale feature maps.
Swin2021–20225