Vision Transformers
Transformers applied to images and video, and how to train and scale them.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Distillation token (data-efficient ViT) Train ViTs on ImageNet alone by adding a token that learns from a CNN teacher. | DeiT | 2020–2021 | 3 |
| Vision Transformer (ViT) Split an image into patches and feed them to a plain Transformer as tokens. | ViT | 2020–2023 | 27 |
| Convolution–attention hybrids Stack convolution stages before attention stages to get both inductive bias and capacity. | CoAtNet | 2021 | 1 |
| Local video Transformers Extend windowed attention across space and time for video recognition. | Video Swin | 2021 | 0 |
| Residual post-norm & cosine attention Stabilise very large vision Transformers and transfer them across resolutions. | Swin V2 | 2021 | 0 |
| Scaling laws for vision Transformers How ViT error falls with model size, data and compute, up to billions of parameters. | Scaling ViTs (ViT-G) | 2021 | 2 |
| Shifted-window hierarchical ViT Attention inside local windows that shift between layers, with multi-scale feature maps. | Swin | 2021–2022 | 5 |