Paper Lineage
Esc
MethodNov 2021arXiv 2111.06377cs.CV

Masked Autoencoders Are Scalable Vision Learners

Kaiming He, Xinlei Chen, Saining Xie and 3 others

This paper shows that masked autoencoders (MAE) are scalable self-supervised learners for computer vision. Our MAE approach is simple: we mask random patches of the input image and reconstruct the missing pixels.

From the abstract

Built on

11 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (11)

  • ViT2020 · cited 12×, 2 in Method
    “Motivated by the success in NLP, related recent methods Chen2020c; Dosovitskiy2021; Bao2021 are based on Transformers Vaswani2017. iGPT Chen2020c operates on sequences of pixels and predicts unknown pixels.”
    From this paper · §Related Work
  • BERT2018 · cited 9×, 2 in Method
    “The solutions, based on autoregressive language modeling in GPT Radford2018; Radford2019; Brown2020 and masked autoencoding in BERT Devlin2019, are conceptually simple: they remove a portion of the data and learn to pred…”
    From this paper · §Introduction
  • BEiT2021 · cited 8×
    “Motivated by the success in NLP, related recent methods Chen2020c; Dosovitskiy2021; Bao2021 are based on Transformers Vaswani2017. iGPT Chen2020c operates on sequences of pixels and predicts unknown pixels.”
    From this paper · §Related Work
  • GPT-32020 · cited 6×
    “The solutions, based on autoregressive language modeling in GPT Radford2018; Radford2019; Brown2020 and masked autoencoding in BERT Devlin2019, are conceptually simple: they remove a portion of the data and learn to pred…”
    From this paper · §Introduction
Show 7 more
  • BYOL2020 · cited 5×
    “Recently, contrastive learning Becker1992; Hadsell2006 has been popular, e.g., Wu2018a; Oord2018; He2020; Chen2020, which models image similarity and dissimilarity (or only similarity Grill2020; Chen2021) between two or…”
    From this paper · §Related Work
  • SimCLR2020 · cited 4×
    “Recently, contrastive learning Becker1992; Hadsell2006 has been popular, e.g., Wu2018a; Oord2018; He2020; Chen2020, which models image similarity and dissimilarity (or only similarity Grill2020; Chen2021) between two or…”
    From this paper · §Related Work
  • ResNet2015 · cited 3×
    “Deep learning has witnessed an explosion of architectures of continuously growing capability and capacity Krizhevsky2012; He2016; Vaswani2017.”
    From this paper · §Introduction
  • Transformer2017 · cited 3×
    “Motivated by the success in NLP, related recent methods Chen2020c; Dosovitskiy2021; Bao2021 are based on Transformers Vaswani2017. iGPT Chen2020c operates on sequences of pixels and predicts unknown pixels.”
    From this paper · §Related Work
  • DALL·E2021 · cited 3×
    “Specifically for this variant, we use the DALLE pre-trained dVAE Ramesh2021 as the tokenizer, following Bao2021.”
    From this paper · §ImageNet Experiments
  • Instance discrimination2018 · cited 2×
    “Recently, contrastive learning Becker1992; Hadsell2006 has been popular, e.g., Wu2018a; Oord2018; He2020; Chen2020, which models image similarity and dissimilarity (or only similarity Grill2020; Chen2021) between two or…”
    From this paper · §Related Work
  • MoCo2019 · cited 2×
    “Recently, contrastive learning Becker1992; Hadsell2006 has been popular, e.g., Wu2018a; Oord2018; He2020; Chen2020, which models image similarity and dissimilarity (or only similarity Grill2020; Chen2021) between two or…”
    From this paper · §Related Work

Led to

  • Simple end-to-end captioning2022 · cited 2×, 1 in Method
    “Visual Encoder.We follow the same setting as the visual transformer (ViT) which recently attracts much attention and shows remarkable performance in the computer vision area (Dosovitskiy et al. 2021; Radford et al. 2021;…”
    From Simple end-to-end captioning · §. Methodology
  • CoCa2022 · cited 2×, 1 in Method
    “MAE [23] and SimMIM [24] remove the need for an image tokenizer and directly use a light-weight decoder or projection layer to regress pixel values.”
    From CoCa · §Related Work
  • BEiT v22022 · cited 4×
    “As shown in Table 3, compared with MAE (He et al. 2022), BEiT v2 achieves dramatic gains across datasets, demonstrating the superiority of the proposed method in terms of model generalization.”
    From BEiT v2 · §Experiments
  • EVA2022 · cited 4×
    “Recently, masked image modeling (MIM) bao2021beit; xie2021simmim; he2021masked has boomed as a viable approach for vision model pre-training and scaling.”
    From EVA · §Introduction
Abstract

This paper shows that masked autoencoders (MAE) are scalable self-supervised learners for computer vision. Our MAE approach is simple: we mask random patches of the input image and reconstruct the missing pixels. It is based on two core designs. First, we develop an asymmetric encoder-decoder architecture, with an encoder that operates only on the visible subset of patches (without mask tokens), along with a lightweight decoder that reconstructs the original image from the latent representation and mask tokens. Second, we find that masking a high proportion of the input image, e.g., 75%, yields a nontrivial and meaningful self-supervisory task. Coupling these two designs enables us to train large models efficiently and effectively: we accelerate training (by 3x or more) and improve accuracy. Our scalable approach allows for learning high-capacity models that generalize well: e.g., a vanilla ViT-Huge model achieves the best accuracy (87.8%) among methods that use only ImageNet-1K data. Transfer performance in downstream tasks outperforms supervised pre-training and shows promising scaling behavior.