Large language models & scaling
Very large autoregressive LMs, what scale buys you, and how to spend compute.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Autoregressive language modelling Predict the next token, so the model can generate text left to right. | — | 2017–2022 | 16 |
| Prompt mining & paraphrasing Search for better prompts to get a truer read on what an LM knows. | LPAQA (What LMs know) | 2019–2021 | 1 |
| In-context / few-shot learning A large model solves a new task from a few examples in its prompt, with no gradient updates. | GPT-3 | 2020–2022 | 25 |
| Knowledge stored in LM parameters Answer factual questions from model weights alone, with no retrieval. | Knowledge in LM parameters | 2020–2021 | 1 |
| Neural scaling laws Loss falls as a power law in model size, data and compute. | Kaplan scaling laws | 2020–2022 | 9 |
| Code LLMs & pass@k LMs fine-tuned on code, measured by whether sampled programs pass unit tests. | Codex | 2021–2022 | 1 |
| Trained verifiers Sample many solutions and let a trained verifier pick the right one. | GSM8K verifiers | 2021–2022 | 3 |
| Compute-optimal training For a fixed budget, grow parameters and training tokens together; most big LMs were undertrained. | Chinchilla | 2022 | 6 |
| Emergent abilities Abilities that are absent in small models and appear abruptly at scale. | Emergent abilities | 2022 | 2 |
| Knowledge-grounded dialogue LMs Fine-tune dialogue models to consult external tools and knowledge for factual answers. | LaMDA | 2022 | 4 |
| Open-weights LLMs Release large LM weights so researchers can study and build on them. | OPT | 2022–2023 | 2 |