architectureIntroduced by ConvS2S · 2017
Convolutional seq2seq
A fully convolutional encoder–decoder with gated linear units and per-layer attention.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that ConvS2S cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Model long sequences with stacked dilated 1-D convolutions that run in linear time.
Encode an input sequence, then decode an output sequence from it.
Split rare words into frequent sub-word pieces so the vocabulary stays small and open.
Papers using this
- 2017Transformer
- 2020GLU variants