architectureIntroduced by Transformer · 2017
Self-attention
Every token attends to every other token, weighted by learned query–key similarity.
Drafted by AI · not yet reviewed
How this idea evolved
Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.
- 2014
Let a decoder look back at the most relevant input positions instead of one fixed vector.
Cites · not yet reviewedcited 5× · §Model Architecture“Most competitive neural sequence transduction models have an encoder-decoder structure [5, 2, 35].”
From Transformer · §Model Architecture - 2017
Every token attends to every other token, weighted by learned query–key similarity.
Papers using this
- 2018Set Transformer
- 2018BERT
- 2019UniLM
- 2019VisualBERT
- 2019LXMERT
- 2019Unified VLP
- 2019BART
- 2020GLU variants
- 2020Oscar
- 2020ViT
- 2020VL-BERT meta-analysis
- 2021VinVL
- 2021Conceptual 12M
- 2021Swin
- 2021RoFormer (RoPE)
- 2021CoAtNet
- 2021Dynamic Head
- 2021Video Swin
- 2021VLMo
- 2023BLIP-2