Paper Lineage
Esc
BenchmarkJul 2021arXiv 2107.03374cs.LG

Evaluating Large Language Models Trained on Code

Mark Chen, Jerry Tworek, Heewoo Jun and 55 others

We introduce Codex, a GPT language model fine-tuned on publicly available code from GitHub, and study its Python code-writing capabilities. A distinct production version of Codex powers GitHub Copilot.

From the abstract

Built on

5 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (5)

  • GPT-32020 · cited 4×
    “Following the success of large natural language models (Devlin et al. 2018; Radford et al. 2019; Liu et al. 2019; Raffel et al. 2020; Brown et al. 2020) large scale Transformers have also been applied towards program syn…”
    From this paper · §Related Work
  • BERT2018 · cited 3×
    “Following the success of large natural language models (Devlin et al. 2018; Radford et al. 2019; Liu et al. 2019; Raffel et al. 2020; Brown et al. 2020) large scale Transformers have also been applied towards program syn…”
    From this paper · §Related Work
  • T52019 · cited 2×
    “Following the success of large natural language models (Devlin et al. 2018; Radford et al. 2019; Liu et al. 2019; Raffel et al. 2020; Brown et al. 2020) large scale Transformers have also been applied towards program syn…”
    From this paper · §Related Work
  • The Pile2020 · cited 2×
    “More recently, language models have also fueled progress towards the longstanding challenge of program synthesis (Simon 1963; Manna & Waldinger 1971), spurred by the presence of code in large datasets (Husain et al. 2019…”
    From this paper · §Introduction
Show 1 more
  • DALL·E2021 · cited 2×
    “Further, just as text-conditional generative models in other modalities (Ramesh et al. 2021) have difficulty with binding attributes to objects, Codex can make mistakes binding operations to variables, especially when th…”
    From this paper · §Limitations

Led to

  • InstructGPT2022 · cited 1×, 1 in Method
    “The definition of alignment has historically been a vague and confusing topic, with various competing proposals (Chen et al., 2021; Leike et al., 2018; Gabriel, 2020).”
    From InstructGPT · §Methods and experimental details
  • PaLM2022 · cited 10×
    “Second, we compare to the early Codex model 12B described in Chen et al. 2021, which reports results only on the HumanEval dataset.”
    From PaLM · §Evaluation
  • BIG-bench2022 · cited 2×
    “For instance, they demonstrate nascent abilities in writing computer code (Hendrycks et al. 2021a; Chen et al. 2021; Austin et al. 2021; Schuster et al. 2021b; Biderman & Raff 2022), playing chess (Noever et al. 2020; St…”
    From BIG-bench · §Introduction
Abstract

We introduce Codex, a GPT language model fine-tuned on publicly available code from GitHub, and study its Python code-writing capabilities. A distinct production version of Codex powers GitHub Copilot. On HumanEval, a new evaluation set we release to measure functional correctness for synthesizing programs from docstrings, our model solves 28.8% of the problems, while GPT-3 solves 0% and GPT-J solves 11.4%. Furthermore, we find that repeated sampling from the model is a surprisingly effective strategy for producing working solutions to difficult prompts. Using this method, we solve 70.2% of our problems with 100 samples per problem. Careful investigation of our model reveals its limitations, including difficulty with docstrings describing long chains of operations and with binding operations to variables. Finally, we discuss the potential broader impacts of deploying powerful code generation technologies, covering safety, security, and economics.