architectureIntroduced by Transformer · 2017
Transformer
A sequence model built only from attention and feed-forward layers, with no recurrence.
Drafted by AI · not yet reviewed
How this idea evolved
Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.
- 2014
Let a decoder look back at the most relevant input positions instead of one fixed vector.
Cites · not yet reviewedcited 5× · §Model Architecture“Most competitive neural sequence transduction models have an encoder-decoder structure [5, 2, 35].”
From Transformer · §Model Architecture - 2017
A sequence model built only from attention and feed-forward layers, with no recurrence.
Also draws on: Encoder–decoder seq2seq (Seq2Seq)
Papers using this
- 2018Set Transformer
- 2018BERT
- 2018GPipe
- 2019Transformer-XL
- 2019Adapters
- 2019UniLM
- 2019XLNet
- 2019RoBERTa
- 2019ViLBERT
- 2019VisualBERT
- 2019Unicoder-VL
- 2019LXMERT
- 2019VL-BERT
- 2019Megatron-LM
- 2019RLHF for LMs (Ziegler)
- 2019Unified VLP
- 2019UNITER
- 2019FreeLB
- 2019ALBERT
- 2019T5
- 2019BART
- 2019Meshed-Memory Transformer
- 2020Knowledge in LM parameters
- 2020GLU variants
- 2020X-Linear attention
- 2020GPT-3
- 2020VILLA
- 2020VirTex
- 2020GShard
- 2020Learning to summarize from human feedbac
- 2020ViT
- 2020VL-BERT meta-analysis
- 2020DeiT
- 2021VinVL
- 2021Switch Transformer
- 2021VL-T5
- 2021ViLT
- 2021DALL·E
- 2021CLIP
- 2021Perceiver
- 2021Swin
- 2021RoFormer (RoPE)
- 2021Scaling ViTs (ViT-G)
- 2021CoAtNet
- 2021BEiT
- 2021Dynamic Head
- 2021Video Swin
- 2021Frozen
- 2021ALBEF
- 2021SimVLM
- 2021ViT-VQGAN
- 2021T0
- 2021GSM8K verifiers
- 2021VLMo
- 2021METER
- 2021LiT
- 2021Swin V2
- 2021Florence
- 2021Gopher
- 2021ViTCAP
- 2021GLaM
- 2021Fairseq MoE LMs
- 2022LaMDA
- 2022Megatron-Turing NLG
- 2022BLIP
- 2022Simple end-to-end captioning
- 2022Chinchilla
- 2022PaLM
- 2022UniCL
- 2022GPT-NeoX-20B
- 2022Super-NaturalInstructions
- 2022Flamingo
- 2022OPT
- 2022CoCa
- 2022VL-BEiT
- 2022BIG-bench
- 2022BEiT v2
- 2023BLIP-2