Paper Lineage
Esc
MethodAug 2018arXiv 1808.06226cs.CL

SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Taku Kudo, John Richardson

This paper describes SentencePiece, a language-independent subword tokenizer and detokenizer designed for Neural-based text processing, including Neural Machine Translation. It provides open-source C++ and Python implementations for subword units.

From the abstract

Built on

1 paper · 0 verifiedSee as graph

Also cited · not yet reviewed (1)

  • GNMT2016 · cited 3×
    “Neural machine translation (NMT) Bahdanau et al. 2014; Luong et al. 2015; Wu et al. 2016; Vaswani et al. 2017 has especially gained increasing popularity, as it can leverage neural networks to directly perform translatio…”
    From this paper · §Introduction

Led to

  • Frozen2021 · cited 1×, 1 in Method
    “Text is decomposed into a sequence of discrete tokens 𝐲=y1,y2,…,yLy=y_{1},y_{2},...,y_{L} by the SentencePiece tokenizer [17].”
    From Frozen · §The Frozen Method
  • Gopher2021 · cited 1×, 1 in Method
    “We tokenize the text using SentencePiece (Kudo and Richardson 2018) with a vocabulary of 32,000 and use a byte-level backoff to support open-vocabulary modelling.”
    From Gopher · §Method
  • GLaM2021 · cited 1×, 1 in Method
    “We use the SentencePiece (Kudo & Richardson 2018) subword tokenizer with a vocabulary of size of 256256K.”
    From GLaM · §Experiment Setup
  • LaMDA2022 · cited 1×, 1 in Method
    “Over 90% of the pre-training dataset is in the English language.”
    From LaMDA · §LaMDA pre-training
  • PaLM2022 · cited 1×, 1 in Method
    “Vocabulary – We use a SentencePiece (Kudo & Richardson 2018a) vocabulary with 256k tokens, which was chosen to support the large number of languages in the training corpus without excess tokenization.”
    From PaLM · §Model Architecture
  • BEiT-32022 · cited 2×, 2 in Method
    “Specifically, text data is tokenized by a SentencePiece tokenizer [25].”
    From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
Abstract

This paper describes SentencePiece, a language-independent subword tokenizer and detokenizer designed for Neural-based text processing, including Neural Machine Translation. It provides open-source C++ and Python implementations for subword units. While existing subword segmentation tools assume that the input is pre-tokenized into word sequences, SentencePiece can train subword models directly from raw sentences, which allows us to make a purely end-to-end and language independent system. We perform a validation experiment of NMT on English-Japanese machine translation, and find that it is possible to achieve comparable accuracy to direct subword training from raw sentences. We also compare the performance of subword training and segmentation with various configurations. SentencePiece is available under the Apache 2 license at https://github.com/google/sentencepiece.