architectureIntroduced by CoAtNet · 2021
Convolution–attention hybrids
Stack convolution stages before attention stages to get both inductive bias and capacity.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that CoAtNet cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Attention inside local windows that shift between layers, with multi-scale feature maps.
Split an image into patches and feed them to a plain Transformer as tokens.
How ViT error falls with model size, data and compute, up to billions of parameters.
Papers using this
- 2021SimVLM