Masked Autoencoders Are Scalable Vision Learners
This paper shows that masked autoencoders (MAE) are scalable self-supervised learners for computer vision. Our MAE approach is simple: we mask random patches of the input image and reconstruct the missing pixels.
Also cited · not yet reviewed (11)
- ViT2020 · cited 12×, 2 in Method“Motivated by the success in NLP, related recent methods Chen2020c; Dosovitskiy2021; Bao2021 are based on Transformers Vaswani2017. iGPT Chen2020c operates on sequences of pixels and predicts unknown pixels.”From this paper · §Related Work
- BERT2018 · cited 9×, 2 in Method“The solutions, based on autoregressive language modeling in GPT Radford2018; Radford2019; Brown2020 and masked autoencoding in BERT Devlin2019, are conceptually simple: they remove a portion of the data and learn to pred…”From this paper · §Introduction
- BEiT2021 · cited 8דMotivated by the success in NLP, related recent methods Chen2020c; Dosovitskiy2021; Bao2021 are based on Transformers Vaswani2017. iGPT Chen2020c operates on sequences of pixels and predicts unknown pixels.”From this paper · §Related Work
- GPT-32020 · cited 6דThe solutions, based on autoregressive language modeling in GPT Radford2018; Radford2019; Brown2020 and masked autoencoding in BERT Devlin2019, are conceptually simple: they remove a portion of the data and learn to pred…”From this paper · §Introduction
Show 7 more
- BYOL2020 · cited 5דRecently, contrastive learning Becker1992; Hadsell2006 has been popular, e.g., Wu2018a; Oord2018; He2020; Chen2020, which models image similarity and dissimilarity (or only similarity Grill2020; Chen2021) between two or…”From this paper · §Related Work
- SimCLR2020 · cited 4דRecently, contrastive learning Becker1992; Hadsell2006 has been popular, e.g., Wu2018a; Oord2018; He2020; Chen2020, which models image similarity and dissimilarity (or only similarity Grill2020; Chen2021) between two or…”From this paper · §Related Work
- ResNet2015 · cited 3דDeep learning has witnessed an explosion of architectures of continuously growing capability and capacity Krizhevsky2012; He2016; Vaswani2017.”From this paper · §Introduction
- Transformer2017 · cited 3דMotivated by the success in NLP, related recent methods Chen2020c; Dosovitskiy2021; Bao2021 are based on Transformers Vaswani2017. iGPT Chen2020c operates on sequences of pixels and predicts unknown pixels.”From this paper · §Related Work
- DALL·E2021 · cited 3דSpecifically for this variant, we use the DALLE pre-trained dVAE Ramesh2021 as the tokenizer, following Bao2021.”From this paper · §ImageNet Experiments
- Instance discrimination2018 · cited 2דRecently, contrastive learning Becker1992; Hadsell2006 has been popular, e.g., Wu2018a; Oord2018; He2020; Chen2020, which models image similarity and dissimilarity (or only similarity Grill2020; Chen2021) between two or…”From this paper · §Related Work
- MoCo2019 · cited 2דRecently, contrastive learning Becker1992; Hadsell2006 has been popular, e.g., Wu2018a; Oord2018; He2020; Chen2020, which models image similarity and dissimilarity (or only similarity Grill2020; Chen2021) between two or…”From this paper · §Related Work
Led to
- Simple end-to-end captioning2022 · cited 2×, 1 in Method“Visual Encoder.We follow the same setting as the visual transformer (ViT) which recently attracts much attention and shows remarkable performance in the computer vision area (Dosovitskiy et al. 2021; Radford et al. 2021;…”From Simple end-to-end captioning · §. Methodology
- CoCa2022 · cited 2×, 1 in Method“MAE [23] and SimMIM [24] remove the need for an image tokenizer and directly use a light-weight decoder or projection layer to regress pixel values.”From CoCa · §Related Work
- BEiT v22022 · cited 4דAs shown in Table 3, compared with MAE (He et al. 2022), BEiT v2 achieves dramatic gains across datasets, demonstrating the superiority of the proposed method in terms of model generalization.”From BEiT v2 · §Experiments
- EVA2022 · cited 4דRecently, masked image modeling (MIM) bao2021beit; xie2021simmim; he2021masked has boomed as a viable approach for vision model pre-training and scaling.”From EVA · §Introduction
Abstract
This paper shows that masked autoencoders (MAE) are scalable self-supervised learners for computer vision. Our MAE approach is simple: we mask random patches of the input image and reconstruct the missing pixels. It is based on two core designs. First, we develop an asymmetric encoder-decoder architecture, with an encoder that operates only on the visible subset of patches (without mask tokens), along with a lightweight decoder that reconstructs the original image from the latent representation and mask tokens. Second, we find that masking a high proportion of the input image, e.g., 75%, yields a nontrivial and meaningful self-supervisory task. Coupling these two designs enables us to train large models efficiently and effectively: we accelerate training (by 3x or more) and improve accuracy. Our scalable approach allows for learning high-capacity models that generalize well: e.g., a vanilla ViT-Huge model achieves the best accuracy (87.8%) among methods that use only ImageNet-1K data. Transfer performance in downstream tasks outperforms supervised pre-training and shows promising scaling behavior.