dataIntroduced by SentencePiece · 2018
Language-independent subword tokenizer
Train subword models directly on raw text, with no language-specific pre-tokenisation.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that SentencePiece cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Split rare words into frequent sub-word pieces so the vocabulary stays small and open.
Encode an input sequence, then decode an output sequence from it.