Paper Lineage
Esc
MethodApr 2021arXiv 2104.09864cs.CL

RoFormer: Enhanced Transformer with Rotary Position Embedding

Jianlin Su, Yu Lu, Shengfeng Pan and 3 others

Position encoding recently has shown effective in the transformer architecture. It enables valuable supervision for dependency modeling between elements at different positions of the sequence.

From the abstract

Built on

6 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (6)

  • Transformer2017 · cited 8×, 3 in Method
    “Generally speaking, all these approaches attempt to modify Equation 6 based on the decomposition of Equation 3 under the self-attention settings in Equation 2, which was originally proposed in Vaswani et al. 2017.”
    From this paper · §Background and Related Work
  • T52019 · cited 4×, 3 in Method
    “Later work Raffel et al. 2020; He et al. 2020; Ke et al. 2020; Huang et al. 2020 followed these settings by only encoding the relative position information into the attention weights.”
    From this paper · §Background and Related Work
  • BERT2018 · cited 7×, 1 in Method
    “Previous work Devlin et al. 2019; Lan et al. 2020; Clark et al. 2020; Radford et al. 2019; Radford and Narasimhan 2018 introduced the use of a set of trainable vectors 𝒑i∈{𝒑t}t=1L{\boldsymbol{p}}_{i}\in\{{\boldsymbol{p…”
    From this paper · §Background and Related Work
  • Transformer-XL2019 · cited 2×, 1 in Method
    “Keeping the form of Equation 3, the authors Dai et al. 2019 have proposed to decompose 𝒒m⊺​𝒌n{\boldsymbol{q}}_{m}^{\intercal}{\boldsymbol{k}}_{n} of Equation 2 as”
    From this paper · §Background and Related Work
Show 2 more
  • ALBERT2019 · cited 2×, 1 in Method
    “Previous work Devlin et al. 2019; Lan et al. 2020; Clark et al. 2020; Radford et al. 2019; Radford and Narasimhan 2018 introduced the use of a set of trainable vectors 𝒑i∈{𝒑t}t=1L{\boldsymbol{p}}_{i}\in\{{\boldsymbol{p…”
    From this paper · §Background and Related Work
  • ConvS2S2017 · cited 2×
    “Convolution neural networks (CNNs) based models (CNNs) Gehring et al. 2017 were typically considered position-agnostic, but recent work Islam et al. 2020 has shown that the commonly used padding operation can implicitly…”
    From this paper · §Introduction

Led to

  • PaLM2022 · cited 1×, 1 in Method
    “RoPE Embeddings – We use RoPE embeddings (Su et al. 2021) rather than absolute or relative position embeddings, since RoPE embeddings have been shown to have better performance on long sequence lengths.”
    From PaLM · §Model Architecture
  • GPT-NeoX-20B2022 · cited 2×, 2 in Method
    “We use rotary embeddings (Su et al. 2021) instead of the learned positional embeddings used in GPT models (Radford et al. 2018), based on our positive prior experiences using it in training LLMs.”
    From GPT-NeoX-20B · §Model Design and Implementation
Abstract

Position encoding recently has shown effective in the transformer architecture. It enables valuable supervision for dependency modeling between elements at different positions of the sequence. In this paper, we first investigate various methods to integrate positional information into the learning process of transformer-based language models. Then, we propose a novel method named Rotary Position Embedding(RoPE) to effectively leverage the positional information. Specifically, the proposed RoPE encodes the absolute position with a rotation matrix and meanwhile incorporates the explicit relative position dependency in self-attention formulation. Notably, RoPE enables valuable properties, including the flexibility of sequence length, decaying inter-token dependency with increasing relative distances, and the capability of equipping the linear self-attention with relative position encoding. Finally, we evaluate the enhanced transformer with rotary position embedding, also called RoFormer, on various long text classification benchmark datasets. Our experiments show that it consistently overcomes its alternatives. Furthermore, we provide a theoretical analysis to explain some experimental results. RoFormer is already integrated into Huggingface: \url{https://huggingface.co/docs/transformers/model_doc/roformer}.