Paper Lineage
Esc
BenchmarkJun 2022arXiv 2206.04615cs.CL

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao and 448 others

Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabilities are as yet poorly characterized.

From the abstract

Built on

18 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (18)

  • PaLM2022 · cited 4×, 2 in Method
    “Recent results on large language models suggest that this brittleness to question phrasing may improve with further increases in scale (Chowdhery et al. 2022).”
    From this paper · §Behavior of language models and human raters on BIG-bench
  • LPAQA (What LMs know)2019 · cited 2×, 2 in Method
    “Furthermore, Jiang et al. 2020 studied calibration on generative language models (T5, BART, and GPT-2) and found that their predictive probabilities on question-answering tasks are not well calibrated.”
    From this paper · §Behavior of language models and human raters on BIG-bench
  • GPT-32020 · cited 26×, 1 in Method
    “We use OpenAI GPT models corresponding to the GPT-3 model series in Brown et al. 2020.”
    From this paper · §What is in BIG-bench?
  • SQuAD2016 · cited 8×
    “For instance, benchmarks often propose tasks that codify narrow subsets of areas, such as language understanding (Wang et al. 2019a), summarization (See et al. 2017; Hermann et al. 2015; Narayan et al. 2018; Koupaee & Wa…”
    From this paper · §Introduction
Show 14 more
  • MMLU2020 · cited 5×
    “New evaluations can be contributed as test suits on their project website.55 5 https://syntaxgym.org/ The Massive Multitask Language Understanding (MMLU) benchmark (Hendrycks et al. 2021b) is a collection of 57 diverse t…”
    From this paper · §Additional related work
  • BERT2018 · cited 4×
    “Task cites: (Devlin et al. 2018; Wu et al. 2016; Won et al. 2021; Xu et al. 2018; Edmiston & Stratos 2018; El-Kishky et al. 2019)”
    From this paper · §Author contributions
  • LaMDA2022 · cited 3×
    “We use 13 dense decoder-only Transformer models (Vaswani et al. 2017) with gated activation layers (Dauphin et al. 2017) and GELU activations based on the LaMDA architectures (Thoppilan et al. 2022).”
    From this paper · §What is in BIG-bench?
  • Transformer2017 · cited 2×
    “We use 13 dense decoder-only Transformer models (Vaswani et al. 2017) with gated activation layers (Dauphin et al. 2017) and GELU activations based on the LaMDA architectures (Thoppilan et al. 2022).”
    From this paper · §What is in BIG-bench?
  • CoQA2018 · cited 2×
    “Task cites: (Reddy et al. 2019; Radford et al. 2019; Brown et al. 2020)”
    From this paper · §Author contributions
  • RoBERTa2019 · cited 2×
    “Task cites: (Brown et al. 2020; Devlin et al. 2018; Lan et al. 2019; Liu et al. 2019; Radford et al. 2019; Annamoradnejad & Zoghi 2020; Chen & Soo 2018; Mao & Liu 2019; Weller & Seppi 2019; Khodak et al. 2017; Ghosh et a…”
    From this paper · §Author contributions
  • Kaplan scaling laws2020 · cited 2×
    “Their cross entropy on a test set scales as a power law with regards to model size, training data size, and the amount of compute used in training (Hestness et al. 2017; Hestness et al. 2019; Rosenfeld et al. 2019; Kapla…”
    From this paper · §Introduction
  • Switch Transformer2021 · cited 2×
    “Motivated by this predictable improvement, researchers have now scaled language models to more than one trillion parameters (Fedus et al. 2021), and we expect models to grow orders of magnitude larger over the next sever…”
    From this paper · §Introduction
  • Natural Instructions2021 · cited 2×
    “This project is an expansion of NATURAL INSTRUCTIONS (Mishra et al. 2021), which contained a collection of 61 tasks.”
    From this paper · §Additional related work
  • Codex2021 · cited 2×
    “For instance, they demonstrate nascent abilities in writing computer code (Hendrycks et al. 2021a; Chen et al. 2021; Austin et al. 2021; Schuster et al. 2021b; Biderman & Raff 2022), playing chess (Noever et al. 2020; St…”
    From this paper · §Introduction
  • T02021 · cited 2×
    “BIG-bench has also been partially evaluated on other models including Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), and T0 (Sanh et al. 2022).”
    From this paper · §What is in BIG-bench?
  • Gopher2021 · cited 2×
    “BIG-bench has also been partially evaluated on other models including Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), and T0 (Sanh et al. 2022).”
    From this paper · §What is in BIG-bench?
  • Chinchilla2022 · cited 2×
    “BIG-bench has also been partially evaluated on other models including Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), and T0 (Sanh et al. 2022).”
    From this paper · §What is in BIG-bench?
  • GPT-NeoX-20B2022 · cited 2×
    “The quantitative and qualitative changes that occur in language models as they become larger are potentially transformative (Bommasani et al. 2021; Black et al. 2022).”
    From this paper · §Introduction

Led to

  • Chinchilla2022 · cited 2×
    “We thus place more emphasis on other tasks for which leakage is less of a concern, such as MMLU (Hendrycks et al. 2020) and BIG-bench (BIG-bench collaboration 2021) along with various closed-book question answering and c…”
    From Chinchilla · §Chinchilla
  • PaLM2022 · cited 2×
    “In section 6.2 we additionally highlight breakthrough performance on BIG-bench (BIG-bench collaboration 2021), a recently released suite of 150150+ new language understanding and generation tasks, many of which are extre…”
    From PaLM · §Introduction
  • Emergent abilities2022 · cited 4×
    “Figure 2A–D depicts four emergent few-shot prompted tasks from BIG-Bench, a crowd-sourced suite of over 200 benchmarks for language model evaluation (BIG-Bench 2022).”
    From Emergent abilities · §Few-Shot Prompted Tasks
  • Flan-T5 / Flan-PaLM2022 · cited 4×
    “Instead, we use the following challenging benchmarks, for which current language models still perform well below expert human raters. (1) MMLU (Hendrycks et al. 2020) includes exam questions from 57 tasks such as mathema…”
    From Flan-T5 / Flan-PaLM · §unknown section
Abstract

Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabilities are as yet poorly characterized. In order to inform future research, prepare for disruptive new model capabilities, and ameliorate socially harmful effects, it is vital that we understand the present and near-future capabilities and limitations of language models. To address this challenge, we introduce the Beyond the Imitation Game benchmark (BIG-bench). BIG-bench currently consists of 204 tasks, contributed by 450 authors across 132 institutions. Task topics are diverse, drawing problems from linguistics, childhood development, math, common-sense reasoning, biology, physics, social bias, software development, and beyond. BIG-bench focuses on tasks that are believed to be beyond the capabilities of current language models. We evaluate the behavior of OpenAI's GPT models, Google-internal dense transformer architectures, and Switch-style sparse transformers on BIG-bench, across model sizes spanning millions to hundreds of billions of parameters. In addition, a team of human expert raters performed all tasks in order to provide a strong baseline. Findings include: model performance and calibration both improve with scale, but are poor in absolute terms (and when compared with rater performance); performance is remarkably similar across model classes, though with benefits from sparsity; tasks that improve gradually and predictably commonly involve a large knowledge or memorization component, whereas tasks that exhibit "breakthrough" behavior at a critical scale often involve multiple steps or components, or brittle metrics; social bias typically increases with scale in settings with ambiguous context, but this can be improved with prompting.