Paper Lineage
Esc

Attention & Transformers

Attention mechanisms and the Transformer family, plus their positional and feed-forward refinements.

Concepts in Attention & Transformers
ConceptIntroduced byYearsPapers
Attention mechanism
Let a decoder look back at the most relevant input positions instead of one fixed vector.
Bahdanau attention2014–202112
Self-attention
Every token attends to every other token, weighted by learned query–key similarity.
Transformer2017–202320
Sinusoidal positional encoding
Add fixed sine/cosine signals so an order-blind Transformer knows token positions.
Transformer2017–20216
Transformer
A sequence model built only from attention and feed-forward layers, with no recurrence.
Transformer2017–202378
Inducing-point attention (learned latents)
A few learned vectors attend to a large set and summarise it: the idea behind later resamplers and query transformers.
Set Transformer2018–20211
Segment-level recurrence & relative positions
Reuse hidden states from earlier segments so attention can reach beyond a fixed context.
Transformer-XL2019–20212
Gated feed-forward variants (SwiGLU)
Replace the Transformer's ReLU feed-forward with gated linear units for better quality.
GLU variants2020–20221
Latent-array cross-attention (Perceiver)
A small latent array cross-attends to huge inputs, decoupling depth from input size.
Perceiver2021–20221
Rotary position embedding (RoPE)
Encode position by rotating query/key vectors, giving relative-position awareness inside attention.
RoFormer (RoPE)2021–20222