architectureIntroduced by Flamingo · 2022
Gated cross-attention layers
New cross-attention layers inside a frozen LLM, gated to start as identity.
Drafted by AI · not yet reviewed
How this idea evolved
Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.
- 2014
Let a decoder look back at the most relevant input positions instead of one fixed vector.
Cites · not yet reviewedcited 5× · §Model Architecture“Most competitive neural sequence transduction models have an encoder-decoder structure [5, 2, 35].”
From Transformer · §Model Architecture - 2017
A sequence model built only from attention and feed-forward layers, with no recurrence.
Also draws on: Encoder–decoder seq2seq (Seq2Seq)
Cites · not yet reviewedcited 4× · §BERT“BERT’s model architecture is a multi-layer bidirectional Transformer encoder based on the original implementation described in Vaswani et al. 2017 and released in the tensor2tensor library.11 1 https://github.com/tensorflow/tensor2tensor Because the use of Transformers has become common and our implementation is almost identical to the original, we will omit an exhaustive background description of the model architecture and refer readers to Vaswani et al. 2017 as well as excellent guides such as ‘‘The Annotated Transformer.’’22 2 http://nlp.seas.harvard.edu/2018/04/03/attention.html”
From BERT · §BERT - 2018
Hide some tokens and predict them from context on both sides.
Cites · not yet reviewedcited 9× · §Introduction“In analogy to the training tasks in [12], we train our model on Conceptual Captions on two proxy tasks: predicting the semantics of masked words and image regions given the unmasked inputs, and predicting whether an image and text segment correspond.”
From ViLBERT · §Introduction - 2019
Separate image and text streams exchange information through co-attention layers.
Cites · not yet reviewedcited 2× · §Related work“In particular, BERT [23] inspired a large body of vision-language work [66, 106, 16, 38, 121, 61, 109, 151, 118, 59, 29, 28, 142, 143, 101, 107].”
From Flamingo · §Related work - 2022
New cross-attention layers inside a frozen LLM, gated to start as identity.
Papers using this
No other paper in this dataset is tagged with it yet.