architectureIntroduced by Swin · 2021
Shifted-window hierarchical ViT
Attention inside local windows that shift between layers, with multi-scale feature maps.
Drafted by AI · not yet reviewed
How this idea evolved
Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.
- 2014
Let a decoder look back at the most relevant input positions instead of one fixed vector.
Cites · not yet reviewedcited 5× · §Model Architecture“Most competitive neural sequence transduction models have an encoder-decoder structure [5, 2, 35].”
From Transformer · §Model Architecture - 2017
A sequence model built only from attention and feed-forward layers, with no recurrence.
Also draws on: Encoder–decoder seq2seq (Seq2Seq)
Cites · not yet reviewedcited 4× · §Method“In model design we follow the original Transformer (Vaswani et al. 2017) as closely as possible.”
From ViT · §Method - 2020
Split an image into patches and feed them to a plain Transformer as tokens.
Cites · not yet reviewedcited 8× · §Related Work“Most related to our work is the Vision Transformer (ViT) [20] and its follow-ups [63, 72, 15, 28, 66].”
From Swin · §Related Work - 2021
Attention inside local windows that shift between layers, with multi-scale feature maps.
Papers using this
- 2021Video Swin
- 2021METER
- 2021Swin V2
- 2021Florence
- 2022UniCL