Decoupled Weight Decay Regularization
L$_2$ regularization and weight decay regularization are equivalent for standard stochastic gradient descent (when rescaled by the learning rate), but as we demonstrate this is \emph{not} the case for adaptive gradient algorithms, such as Adam. t.
AdamW has no earlier papers in this dataset.
Led to
- Megatron-LM2019 · cited 1×, 1 in Method“For our optimizer we utilize Adam (Kingma & Ba 2014) with weight decay (Loshchilov & Hutter 2019) λ=0.01\lambda=0.01.”From Megatron-LM · §Setup
- Prefix-Tuning2021 · cited 1×, 1 in Method“At training time, we use the AdamW optimizer Loshchilov and Hutter 2019 and a linear learning rate scheduler, as suggested by the Hugging Face default setup.”From Prefix-Tuning · §Experimental Setup
- VL-T52021 · cited 1×, 1 in Method“We use AdamW (Loshchilov & Hutter 2019) with (β1,β2)=(0.9,0.999)(\beta^{1},\beta^{2})=(0.9,0.999) and learning rate 1e-4 with 5% linear warmup schedule.”From VL-T5 · §Pretraining
- CLIP2021 · cited 1×, 1 in Method“We use the Adam optimizer (Kingma & Ba 2014) with decoupled weight decay regularization (Loshchilov & Hutter 2017) applied to all weights that are not gains or biases, and decay the learning rate using a cosine schedule…”From CLIP · §Approach
- BEiT2021 · cited 2×, 1 in Method“On ADE20K, we use Adam [24] as the optimizer.”From BEiT · §Experiments
- ALBEF2021 · cited 1×, 1 in Method“We use the AdamW [44] optimizer with a weight decay of 0.02.”From ALBEF · §ALBEF Pre-training
- METER2021 · cited 2×, 2 in Method“For example, PixelBERT huang2020pixel and CLIP-ViL shen2021much use AdamW loshchilov2018decoupled for transformer and SGD for CNN.”From METER · §Glossary of VLP Models
- Swin V22021 · cited 13דFollowing liu2021swin, we employ an AdamW loshchilov2017decoupled optimizer for 300 epochs using a cosine decay learning rate scheduler with 20 epochs of linear warm-up.”From Swin V2 · §A1 Experimental Settings for Ablation
- Simple end-to-end captioning2022 · cited 1×, 1 in Method“Implementation Details.For all cross-entropy-based experiments, we train our models with the AdamW optimization algorithm (Loshchilov and Hutter 2019), 4k batch size, mixed-precision training and FP16.”From Simple end-to-end captioning · §. Experiment Setup
- BEiT-32022 · cited 1×, 1 in Method“We use the AdamW [28] optimizer with β1=0.9\beta_{1}=0.9, β2=0.98\beta_{2}=0.98 and ϵ=\epsilon=1e-6 for optimization.”From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
- BLIP-22023 · cited 1×, 1 in Method“We use the AdamW (Loshchilov & Hutter 2017) optimizer with β1=0.9\beta_{1}=0.9, β1=0.98\beta_{1}=0.98, and a weight decay of 0.05.”From BLIP-2 · §Method
Abstract
L$_2$ regularization and weight decay regularization are equivalent for standard stochastic gradient descent (when rescaled by the learning rate), but as we demonstrate this is \emph{not} the case for adaptive gradient algorithms, such as Adam. While common implementations of these algorithms employ L$_2$ regularization (often calling it "weight decay" in what may be misleading due to the inequivalence we expose), we propose a simple modification to recover the original formulation of weight decay regularization by \emph{decoupling} the weight decay from the optimization steps taken w.r.t. the loss function. We provide empirical evidence that our proposed modification (i) decouples the optimal choice of weight decay factor from the setting of the learning rate for both standard SGD and Adam and (ii) substantially improves Adam's generalization performance, allowing it to compete with SGD with momentum on image classification datasets (on which it was previously typically outperformed by the latter). Our proposed decoupled weight decay has already been adopted by many researchers, and the community has implemented it in TensorFlow and PyTorch; the complete source code for our experiments is available at https://github.com/loshchil/AdamW-and-SGDW