LaMDA: Language Models for Dialog Applications
We present LaMDA: Language Models for Dialog Applications. 56T words of public dialog data and web text.
Also cited · not yet reviewed (6)
- T52019 · cited 2×, 1 in Method“The Transformer has 64 layers, dmodel=8192d_{model}=8192, dff=65536d_{ff}=65536, h=128h=128, dk=dv=128d_{k}=d_{v}=128, relative attention as described in T5 [11], and gated-GELU activation as described in Raffel et…”From this paper · §LaMDA pre-training
- Transformer2017 · cited 1×, 1 in Method“We use a decoder-only Transformer [92] language model as the model architecture for LaMDA.”From this paper · §LaMDA pre-training
- SentencePiece2018 · cited 1×, 1 in Method“Over 90% of the pre-training dataset is in the English language.”From this paper · §LaMDA pre-training
- GLU variants2020 · cited 1×, 1 in Method“The Transformer has 64 layers, dmodel=8192d_{model}=8192, dff=65536d_{ff}=65536, h=128h=128, dk=dv=128d_{k}=d_{v}=128, relative attention as described in T5 [11], and gated-GELU activation as described in Raffel et…”From this paper · §LaMDA pre-training
Show 2 more
- GPT-32020 · cited 7דSimilar to the concept of prompts in GPT-3 [12], we precondition LaMDA on a few turns of application-specific dialog to adapt LaMDA to the target applications.”From this paper · §Introduction
- Kaplan scaling laws2020 · cited 3דOur study of scaling laws with respect to model sizes is inspired by recent work on the scaling laws of neural language models [12, 13].”From this paper · §Related work
Led to
- InstructGPT2022 · cited 2×, 1 in Method“This is because the language modeling objective used for many recent large LMs—predicting the next token on a webpage from the internet—is different from the objective “follow the user’s instructions helpfully and safely…”From InstructGPT · §Introduction
- Chinchilla2022 · cited 4דThese include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”From Chinchilla · §Related Work
- PaLM2022 · cited 8דThis dataset is based on the datasets used to train LaMDA (Thoppilan et al. 2022) and GLaM (Du et al. 2021).”From PaLM · §Training Dataset
- OPT2022 · cited 6דSimilar to other LLMs, OPT-175B can produce factually incorrect statements Adiwardana et al. 2020; Brown et al. 2020; Roller et al. 2021; Rae et al. 2021; Chowdhery et al. 2022; Thoppilan et al. 2022.”From OPT · §Limitations
- BIG-bench2022 · cited 3דWe use 13 dense decoder-only Transformer models (Vaswani et al. 2017) with gated activation layers (Dauphin et al. 2017) and GELU activations based on the LaMDA architectures (Thoppilan et al. 2022).”From BIG-bench · §What is in BIG-bench?
Abstract
We present LaMDA: Language Models for Dialog Applications. LaMDA is a family of Transformer-based neural language models specialized for dialog, which have up to 137B parameters and are pre-trained on 1.56T words of public dialog data and web text. While model scaling alone can improve quality, it shows less improvements on safety and factual grounding. We demonstrate that fine-tuning with annotated data and enabling the model to consult external knowledge sources can lead to significant improvements towards the two key challenges of safety and factual grounding. The first challenge, safety, involves ensuring that the model's responses are consistent with a set of human values, such as preventing harmful suggestions and unfair bias. We quantify safety using a metric based on an illustrative set of human values, and we find that filtering candidate responses using a LaMDA classifier fine-tuned with a small amount of crowdworker-annotated data offers a promising approach to improving model safety. The second challenge, factual grounding, involves enabling the model to consult external knowledge sources, such as an information retrieval system, a language translator, and a calculator. We quantify factuality using a groundedness metric, and we find that our approach enables the model to generate responses grounded in known sources, rather than responses that merely sound plausible. Finally, we explore the use of LaMDA in the domains of education and content recommendations, and analyze their helpfulness and role consistency.