BEiT: BERT Pre-Training of Image Transformers
We introduce a self-supervised vision representation model BEiT, which stands for Bidirectional Encoder representation from Image Transformers. Following BERT developed in the natural language processing area, we propose a masked image modeling task to pretrain vision Transformers.
Also cited · not yet reviewed (12)
- DALL·E2021 · cited 8×, 6 in Method“Following [34], we use the image tokenizer learned by discrete variational autoencoder (dVAE).”From this paper · §Methods
- ViT2020 · cited 7×, 3 in Method“Following ViT [12], we use the standard Transformer [42] as the backbone network.”From this paper · §Methods
- BERT2018 · cited 3×, 2 in Method“The image patches {𝒙ip}i=1N\{{\bm{x}}^{p}_{i}\}_{i=1}^{N} are flattened into vectors and are linearly projected, which is similar to word embeddings in BERT [13].”From this paper · §Methods
- Instance discrimination2018 · cited 3×, 1 in Method“Our augmentation policy includes random resized cropping, horizontal flipping, color jittering [43].”From this paper · §Methods
Show 8 more
- Transformer2017 · cited 2×, 1 in Method“Following ViT [12], we use the standard Transformer [42] as the backbone network.”From this paper · §Methods
- AdamW2017 · cited 2×, 1 in Method“On ADE20K, we use Adam [24] as the optimizer.”From this paper · §Experiments
- VQ-VAE2017 · cited 1×, 1 in Method“We learn the model following a two-stage procedure similar to [41, 36].”From this paper · §Methods
- T52019 · cited 1×, 1 in Method“Moreover, blockwise (or n-gram) masking is also widely applied in BERT-like models [19, 2, 35].”From this paper · §Methods
- DeiT2020 · cited 3דWe directly follow the most of hyperparameters of DeiT [38] in our fine-tuning experiments for a fair comparison.”From this paper · §Experiments
- MoCo2019 · cited 2דThe recent strand of research follows contrastive paradigm [43, 31, 16, 3, 17, 7, 5].”From this paper · §Related Work
- SimCLR2020 · cited 2דThe recent strand of research follows contrastive paradigm [43, 31, 16, 3, 17, 7, 5].”From this paper · §Related Work
- Scaling ViTs (ViT-G)2021 · cited 2דThe results suggest that BEiT tends to help more for extremely larger models (such as 1B, or 10B), especially when labeled data are insufficient22 2 [47] report that supervised pre-training of a 1.8B-size vision Transfor…”From this paper · §Experiments
Led to
- VLMo2021 · cited 6×, 2 in Method“Following vision Transformers [12, 44, 2], the 2D image 𝒗∈ℝH×W×C{\bm{v}}\in\mathbb{R}^{H\times W\times C} is split and reshaped into N=HW/P2N={HW}/{P^{2}} patches 𝒗p∈ℝN×(P2C){\bm{v}}^{p}\in\mathbb{R}^{N\times(P^{2}C)…”From VLMo · §Methods
- METER2021 · cited 4×, 3 in Method“Specifically, as shown in Figure 1, we dissect the model designs along multiple dimensions, including vision encoders (e.g., CLIP-ViT radford2021learning, Swin transformer liu2021swin), text encoders (e.g., RoBERTa liu20…”From METER · §Introduction
- MAE2021 · cited 8דMotivated by the success in NLP, related recent methods Chen2020c; Dosovitskiy2021; Bao2021 are based on Transformers Vaswani2017. iGPT Chen2020c operates on sequences of pixels and predicts unknown pixels.”From MAE · §Related Work
- Swin V22021 · cited 3דSpecifically, it obtains 84.0% top-1 accuracy on the ImageNet-V2 image classification validation set recht2019imagenet, 63.1 / 54.4 box / mask AP on the COCO test-dev set of object detection, 59.9 mIoU on ADE20K semantic…”From Swin V2 · §Introduction
- Simple end-to-end captioning2022 · cited 2×, 1 in Method“Visual Encoder.We follow the same setting as the visual transformer (ViT) which recently attracts much attention and shows remarkable performance in the computer vision area (Dosovitskiy et al. 2021; Radford et al. 2021;…”From Simple end-to-end captioning · §. Methodology
- VL-BEiT2022 · cited 6×, 1 in Method“Following BEiT [2], we apply block-wise masking strategy to mask 40% of image patches.”From VL-BEiT · §Methods
- BEiT v22022 · cited 10×, 2 in Method“BEiT v2 inherits the masked image modeling framework defined by BEiT (Bao et al. 2022), which uses a visual tokenizer to convert each image to a set of discrete visual tokens.”From BEiT v2 · §Methodology
- BEiT-32022 · cited 7×, 3 in Method“Second, the pretraining task based on masked data modeling has been successfully applied to various modalities, such as texts [14], images [3, 40], and image-text pairs [7].”From BEiT-3 · §Introduction: The Big Convergence
- EVA2022 · cited 10דRecently, masked image modeling (MIM) bao2021beit; xie2021simmim; he2021masked has boomed as a viable approach for vision model pre-training and scaling.”From EVA · §Introduction
Abstract
We introduce a self-supervised vision representation model BEiT, which stands for Bidirectional Encoder representation from Image Transformers. Following BERT developed in the natural language processing area, we propose a masked image modeling task to pretrain vision Transformers. Specifically, each image has two views in our pre-training, i.e, image patches (such as 16x16 pixels), and visual tokens (i.e., discrete tokens). We first "tokenize" the original image into visual tokens. Then we randomly mask some image patches and fed them into the backbone Transformer. The pre-training objective is to recover the original visual tokens based on the corrupted image patches. After pre-training BEiT, we directly fine-tune the model parameters on downstream tasks by appending task layers upon the pretrained encoder. Experimental results on image classification and semantic segmentation show that our model achieves competitive results with previous pre-training methods. For example, base-size BEiT achieves 83.2% top-1 accuracy on ImageNet-1K, significantly outperforming from-scratch DeiT training (81.8%) with the same setup. Moreover, large-size BEiT obtains 86.3% only using ImageNet-1K, even outperforming ViT-L with supervised pre-training on ImageNet-22K (85.2%). The code and pretrained models are available at https://aka.ms/beit.