Paper Lineage
Esc
objectiveIntroduced by MAE · 2021

Masked autoencoders

Mask 75% of patches, encode only the visible ones, and reconstruct pixels with a light decoder.

Drafted by AI · not yet reviewed

How this idea evolved

Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.

  1. 2014

    Let a decoder look back at the most relevant input positions instead of one fixed vector.

    Cites · not yet reviewedcited 5× · §Model Architecture
    “Most competitive neural sequence transduction models have an encoder-decoder structure [5, 2, 35].”
    From Transformer · §Model Architecture
  2. 2017

    A sequence model built only from attention and feed-forward layers, with no recurrence.

    Also draws on: Encoder–decoder seq2seq (Seq2Seq)

    Cites · not yet reviewedcited 4× · §BERT
    “BERT’s model architecture is a multi-layer bidirectional Transformer encoder based on the original implementation described in Vaswani et al. 2017 and released in the tensor2tensor library.11 1 https://github.com/tensorflow/tensor2tensor Because the use of Transformers has become common and our implementation is almost identical to the original, we will omit an exhaustive background description of the model architecture and refer readers to Vaswani et al. 2017 as well as excellent guides such as ‘‘The Annotated Transformer.’’22 2 http://nlp.seas.harvard.edu/2018/04/03/attention.html”
    From BERT · §BERT
  3. 2018

    Hide some tokens and predict them from context on both sides.

    Cites · not yet reviewedcited 3× · §Methods
    “The image patches {𝒙ip}i=1N\{{\bm{x}}^{p}_{i}\}_{i=1}^{N} are flattened into vectors and are linearly projected, which is similar to word embeddings in BERT [13].”
    From BEiT · §Methods
  4. 2021

    Hide image patches and predict them, the vision analogue of BERT.

    Also draws on: Discrete visual tokens (VQ-VAE) (VQ-VAE)

    Cites · not yet reviewedcited 8× · §Related Work
    “Motivated by the success in NLP, related recent methods Chen2020c; Dosovitskiy2021; Bao2021 are based on Transformers Vaswani2017. iGPT Chen2020c operates on sequences of pixels and predicts unknown pixels.”
    From MAE · §Related Work
  5. 2021

    Mask 75% of patches, encode only the visible ones, and reconstruct pixels with a light decoder.

Papers using this