Paper Lineage
Esc
MethodSep 2019arXiv 1909.11942cs.CL

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Zhenzhong Lan, Mingda Chen, Sebastian Goodman and 3 others

Increasing model size when pretraining natural language representations often results in improved performance on downstream tasks. However, at some point further model increases become harder due to GPU/TPU memory limitations and longer training times.

From the abstract

Built on

5 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (5)

  • BERT2018 · cited 14×
    “BERT (Devlin et al. 2019) uses a loss based on predicting whether the second segment in a pair has been swapped with a segment from another document.”
    From this paper · §Related work
  • RoBERTa2019 · cited 13×
    “One of the most compelling signs of these breakthroughs is the evolution of machine performance on a reading comprehension task designed for middle and high-school English exams in China, the RACE test (Lai et al. 2017):…”
    From this paper · §Introduction
  • XLNet2019 · cited 9×
    “Like BERT, we use a vocabulary size of 30,000, tokenized using SentencePiece (Kudo & Richardson 2018) as in XLNet (Yang et al. 2019).”
    From this paper · §Experimental Results
  • Transformer2017 · cited 2×
    “The backbone of the ALBERT architecture is similar to BERT in that it uses a transformer encoder (Vaswani et al. 2017) with GELU nonlinearities (Hendrycks & Gimpel 2016).”
    From this paper · §The Elements of ALBERT
Show 1 more
  • ELMo2018 · cited 2×
    “Learning representations of natural language has been shown to be useful for a wide range of NLP tasks and has been widely adopted (Mikolov et al. 2013; Le & Mikolov 2014; Dai & Le 2015; Peters et al. 2018; Devlin et al.…”
    From this paper · §Related work

Led to

  • Megatron-LM2019 · cited 9×, 4 in Method
    “For BERT models, we largely follow the training process described in (Lan et al. 2019).”
    From Megatron-LM · §Setup
  • FreeLB2019 · cited 2×
    “To further explore its ability to improve more sophisticated language models, we apply FreeLB to the fine-tuning stage of ALBERT-xxlarge-v2 (Lan et al. 2020) model on the dev set of GLUE.”
    From FreeLB · §Experiments
  • T52019 · cited 6×
    “Recent results suggest that this may hold true for transfer learning in NLP (Liu et al. 2019c; Radford et al. 2019; Yang et al. 2019; Lan et al. 2019), i.e. it has repeatedly been shown that scaling up produces improved…”
    From T5 · §Experiments
  • Knowledge in LM parameters2020 · cited 2×, 1 in Method
    “Big, deep neural language models that have been pre-trained on unlabeled text have proven to be extremely performant when fine-tuned on downstream Natural Language Processing (NLP) tasks Devlin et al. 2018; Yang et al. 2…”
    From Knowledge in LM parameters · §Introduction
  • GPT-32020 · cited 3×
    “This last paradigm has led to substantial progress on many challenging NLP tasks such as reading comprehension, question answering, textual entailment, and many others, and has continued to advance based on new architect…”
    From GPT-3 · §Introduction
  • Perceiver2021 · cited 1×, 1 in Method
    “We note that weight sharing has been used for similar goals in Transformers (Dehghani et al. 2019; Lan et al. 2020).”
    From Perceiver · §Methods
  • RoFormer (RoPE)2021 · cited 2×, 1 in Method
    “Previous work Devlin et al. 2019; Lan et al. 2020; Clark et al. 2020; Radford et al. 2019; Radford and Narasimhan 2018 introduced the use of a set of trainable vectors 𝒑i∈{𝒑t}t=1L{\boldsymbol{p}}_{i}\in\{{\boldsymbol{p…”
    From RoFormer (RoPE) · §Background and Related Work
  • METER2021 · cited 1×, 1 in Method
    “In this work, we study the use of BERT devlin2018bert, RoBERTa liu2019roberta, ELECTRA clark2020electra, ALBERT lan2019albert, and DeBERTa he2020deberta for text encoding.”
    From METER · §The Meter Framework
Abstract

Increasing model size when pretraining natural language representations often results in improved performance on downstream tasks. However, at some point further model increases become harder due to GPU/TPU memory limitations and longer training times. To address these problems, we present two parameter-reduction techniques to lower memory consumption and increase the training speed of BERT. Comprehensive empirical evidence shows that our proposed methods lead to models that scale much better compared to the original BERT. We also use a self-supervised loss that focuses on modeling inter-sentence coherence, and show it consistently helps downstream tasks with multi-sentence inputs. As a result, our best model establishes new state-of-the-art results on the GLUE, RACE, and \squad benchmarks while having fewer parameters compared to BERT-large. The code and the pretrained models are available at https://github.com/google-research/ALBERT.