Paper Lineage
Esc

Large language models & scaling

Very large autoregressive LMs, what scale buys you, and how to spend compute.

Concepts in Large language models & scaling
ConceptIntroduced byYearsPapers
Autoregressive language modelling
Predict the next token, so the model can generate text left to right.
—2017–202216
Prompt mining & paraphrasing
Search for better prompts to get a truer read on what an LM knows.
LPAQA (What LMs know)2019–20211
In-context / few-shot learning
A large model solves a new task from a few examples in its prompt, with no gradient updates.
GPT-32020–202225
Knowledge stored in LM parameters
Answer factual questions from model weights alone, with no retrieval.
Knowledge in LM parameters2020–20211
Neural scaling laws
Loss falls as a power law in model size, data and compute.
Kaplan scaling laws2020–20229
Code LLMs & pass@k
LMs fine-tuned on code, measured by whether sampled programs pass unit tests.
Codex2021–20221
Trained verifiers
Sample many solutions and let a trained verifier pick the right one.
GSM8K verifiers2021–20223
Compute-optimal training
For a fixed budget, grow parameters and training tokens together; most big LMs were undertrained.
Chinchilla20226
Emergent abilities
Abilities that are absent in small models and appear abruptly at scale.
Emergent abilities20222
Knowledge-grounded dialogue LMs
Fine-tune dialogue models to consult external tools and knowledge for factual answers.
LaMDA20224
Open-weights LLMs
Release large LM weights so researchers can study and build on them.
OPT2022–20232