Attention & Transformers
Attention mechanisms and the Transformer family, plus their positional and feed-forward refinements.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Attention mechanism Let a decoder look back at the most relevant input positions instead of one fixed vector. | Bahdanau attention | 2014–2021 | 12 |
| Self-attention Every token attends to every other token, weighted by learned query–key similarity. | Transformer | 2017–2023 | 20 |
| Sinusoidal positional encoding Add fixed sine/cosine signals so an order-blind Transformer knows token positions. | Transformer | 2017–2021 | 6 |
| Transformer A sequence model built only from attention and feed-forward layers, with no recurrence. | Transformer | 2017–2023 | 78 |
| Inducing-point attention (learned latents) A few learned vectors attend to a large set and summarise it: the idea behind later resamplers and query transformers. | Set Transformer | 2018–2021 | 1 |
| Segment-level recurrence & relative positions Reuse hidden states from earlier segments so attention can reach beyond a fixed context. | Transformer-XL | 2019–2021 | 2 |
| Gated feed-forward variants (SwiGLU) Replace the Transformer's ReLU feed-forward with gated linear units for better quality. | GLU variants | 2020–2022 | 1 |
| Latent-array cross-attention (Perceiver) A small latent array cross-attends to huge inputs, decoupling depth from input size. | Perceiver | 2021–2022 | 1 |
| Rotary position embedding (RoPE) Encode position by rotating query/key vectors, giving relative-position awareness inside attention. | RoFormer (RoPE) | 2021–2022 | 2 |