about
SentencePiece: Neural Text Processing (Google Research) (arxiv.org)
2 points by Katydid on Aug 27, 2018 | hide | past | pdf | discuss on HN

In plain words: A tool that chops raw text into small word pieces and rebuilds it again, learning straight from unsplit sentences so it works for any language without word boundaries. On English-to-Japanese translation it matched the accuracy of the usual approach that first splits text into words.

Abstract · SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

This paper describes SentencePiece, a language-independent subword tokenizer and detokenizer designed for Neural-based text processing, including Neural Machine Translation. It provides open-source C++ and Python implementations for subword units. While existing subword segmentation tools assume that the input is pre-tokenized into word sequences, SentencePiece can train subword models directly from raw sentences, which allows us to make a purely end-to-end and language independent system. We perform a validation experiment of NMT on English-Japanese machine translation, and find that it is possible to achieve comparable accuracy to direct subword training from raw sentences. We also compare the performance of subword training and segmentation with various configurations. SentencePiece is available under the Apache 2 license at https://github.com/google/sentencepiece.

Taku Kudo, John Richardson
arXiv:1808.06226 · cs.CL · submitted Aug 19, 2018
abstract · pdf · html · Accepted as a demo paper at EMNLP2018

add comment on HN