architectureIntroduced by RoFormer (RoPE) · 2021
Rotary position embedding (RoPE)
Encode position by rotating query/key vectors, giving relative-position awareness inside attention.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that RoFormer (RoPE) cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Every token attends to every other token, weighted by learned query–key similarity.
Add fixed sine/cosine signals so an order-blind Transformer knows token positions.
A sequence model built only from attention and feed-forward layers, with no recurrence.
Papers using this
- 2022PaLM
- 2022GPT-NeoX-20B