Open-source project
google/sentencepiece avatar
google/sentencepiece

SentencePiece: a self-contained tokenizer you train from raw text

Unsupervised text tokenizer for Neural Network-based text generation.

12,104 stars1,385 forksC++Apache-2.0

At a glance

What is it?
SentencePiece trains BPE or unigram subword models directly on raw Unicode text and stores normalization, vocabulary and segmentation in one .model file. It is the right tool when you control the tokenizer, and the wrong one when you just need to load someone else's.
Who is it for?
Adopt SentencePiece when you are training a model from scratch or fine-tuning a tokenizer on a corpus you control, especially for Chinese, Japanese, Thai or other scripts without space delimiters, because the raw-Unicode input path removes the pre-tokenizer you would otherwise have to maintain. Do not adopt it when your model already ships a tokenizer: loading a Gemma or T5 tokenizer in Hugging Face is a few lines and gives you the same vocabulary without a second runtime.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The problem SentencePiece removes: pre-tokenization as a separate artifact

Most tokenizer pipelines are two programs glued together. A language-specific segmenter splits the sentence into words (Moses for European languages, MeCab for Japanese, KyTea for Japanese, and so on), and a subword model then splits those words further. The glue is where the bugs live. The segmenter has its own configuration, its own version, and its own idea of where a word ends, and detokenization has to invert all of it. Whitespace is usually the first casualty: the README points out that a traditional tokenizer treats Tokenize("World.") identically to Tokenize("World ."), which makes the reverse mapping ambiguous.

SentencePiece collapses that stack. It takes raw sentences as input and treats the text as a sequence of Unicode characters, with no pre-tokenizer in between. Whitespace is not thrown away; it is escaped with the meta-symbol ▁ (U+2581) and carried as a normal token. The README states the consequence plainly: detokenization is a string join, and the join is language-independent. For anyone building a pipeline that has to run the same code over English, Japanese and Thai, that is the whole argument. The project describes itself as not an official Google product, which matters if you are picking a dependency for a long-lived system.

Two algorithms, one model file: how the tokenization actually works

SentencePiece implements two subword algorithms. Byte-Pair-Encoding comes from Sennrich et al., and the unigram language model comes from Kudo. Both are trained unsupervised on raw sentences, and both end up in the same kind of artifact: a .model file plus a .vocab file, produced by a single training call.

The design decision worth noting is that the .model file is self-contained. The README says it holds the entire normalization rules, the vocabulary mapping and the segmentation model, and that this guarantees identical tokenization in any environment (C++, Python, Go) as long as the same model file is used. That is a stronger guarantee than it sounds. With a two-stage pipeline you have to pin the segmenter version, the subword model and the normalization tables together; here there is one file to version and one hash to record.

On top of the base algorithms SentencePiece offers sampling. For unigram, Subword Regularization draws different segmentations of the same sentence during training; for BPE, the equivalent is BPE-Dropout. The README frames this as augmenting the training data and making the model more resilient to spelling variation. Treat it as a training-time knob, not a serving-time one: sampling at inference makes outputs non-reproducible, which is usually the last thing you want in a production encoder.

Installing SentencePiece and training your first model

The Python module is on PyPI and installs with pip. The README gives this as the quick start:

bash
pip install sentencepiece

Once that completes you can import spm and train a model without leaving Python. The README's example trains on data/botchan.txt, a file that ships in the repository's data directory, writes the model as m.model with the vocabulary as m.vocab, and fixes the vocabulary size at training time:

python
import sentencepiece as spm

spm.SentencePieceTrainer.train(
    input='data/botchan.txt',
    model_prefix='m',
    vocab_size=1000
)

After training, load the model file and encode. The out_type argument decides whether you get subword strings or integer IDs, and the README shows both on the sentence "I saw a girl with a telescope." The string output begins ['▁I', '▁saw', '▁a', '▁girl', ...] and the integer output begins [9, 459, 11, 939, 44, ...]. Note the ▁ prefix on the first piece and on each word start: that is the whitespace escape doing its job.

python
sp = spm.SentencePieceProcessor(model_file='m.model')
pieces = sp.encode(text, out_type=str)
ids = sp.encode(text, out_type=int)
print(sp.decode(ids))

The README states that decode(ids) and decode(pieces) both return the original string exactly. That reversibility is the property to check first on your own corpus, because it is the one the rest of your data pipeline will depend on. If you want to see sampling in action, the README's example calls encode with enable_sampling=True, alpha=0.1 and nbest_size=-1 on the string 'New York' and shows three different segmentations across three calls.

The training-time vocabulary constraint, and where SentencePiece is the wrong tool

The README is explicit that the vocabulary size is fixed prior to training. That is not a footnote. It means the tokenizer is decided before the model exists, and changing your mind later means retraining the tokenizer and re-embedding the model. If your workflow is "try a bigger vocabulary next week," SentencePiece's training API will not accommodate you without a full rerun.

The second constraint is that SentencePiece is a tokenizer, not a model hub. It has no notion of downloading a pretrained checkpoint, no mapping from model names to tokenizer files, and no chat template handling. If you want T5's tokenizer, you are still going through the library that distributes T5. The README's own benchmark section compares against Hugging Face Fast, which tells you the intended relationship is coexistence, not replacement.

The third is sampling. Subword Regularization and BPE-Dropout change the segmentation of the same input across calls. That is the point during training. In an inference service it means two requests with identical text can produce different token sequences, which will break caching keyed on the input string and any test that asserts on exact IDs. The README presents sampling as a training aid; nothing in it suggests turning it on by default at serving time.

SentencePiece compared with BPE implementations and tiktoken

The comparison people actually make is against two things: a bare BPE implementation, and tiktoken.

Against a bare BPE implementation, the difference is the whitespace handling and the artifact. A textbook BPE trainer operates on pre-split words and drops the separators, so detokenization needs the original spacing recovered by a language-specific rule. SentencePiece folds the separator into the vocabulary as ▁ and produces one self-contained .model file. The README's own framing is that this makes detokenization a simple string join. If you are tokenizing a single language with clean spacing and you already control the pre-split, a plain BPE implementation is less machinery. If you are not, the escape symbol is doing real work.

Against tiktoken, the split is architectural rather than algorithmic. tiktoken is built around serving a fixed set of published vocabularies fast; it is not a trainer. SentencePiece is a trainer that also serves, and the model it serves is one you produced. The README's benchmark section pitches SentencePiece against Hugging Face Fast on throughput, reporting encoding in MB/s on a 60,720-sentence multilingual batch, with SentencePiece ahead at every thread count shown in both the unigram T5-base table and the BPE Gemma 3 table. The same section explains why scaling is not linear: the core tokenization runs in parallel, but converting the native results into Python objects is the bottleneck. That explanation is the useful part, because it tells you the Python binding, not the C++ core, is what you are measuring.

Maintenance, releases and the Apache-2.0 licence

The repository is not archived and the last push was on 2026-09-21. Releases are infrequent rather than continuous: v0.2.0 landed on 2024-02-19, v0.2.1 on 2025-08-12, and v0.2.2 on 2026-07-12. Between releases the master branch moves, so pinning to a tag is the safer default if you build from source rather than installing the wheel. The build files visible at the top level cover both Bazel (BUILD.bazel, MODULE.bazel, .bazelversion) and CMake (CMakeLists.txt, cmake/, config.h.in), so a source build has two supported paths. The repository also carries a lite/ directory and a contrib/ directory, which the README does not document; treat them as unverified until you read the source.

The licence is Apache-2.0, which is permissive and includes an explicit patent grant. Two implications are worth stating without giving legal advice: the grant is what most corporate legal reviews look for, and Apache-2.0 requires you to preserve notices and state changes when you redistribute. The README's note that this is not an official Google product is separate from the licence and affects support expectations, not your right to use the code. The training data you feed to SentencePieceTrainer.train is yours to account for; the licence covers the software, not the corpus.

Editorial conclusion

Adopt SentencePiece when you are training a model from scratch or fine-tuning a tokenizer on a corpus you control, especially for Chinese, Japanese, Thai or other scripts without space delimiters, because the raw-Unicode input path removes the pre-tokenizer you would otherwise have to maintain. Do not adopt it when your model already ships a tokenizer: loading a Gemma or T5 tokenizer in Hugging Face is a few lines and gives you the same vocabulary without a second runtime. Before committing, verify three things: that the vocab_size you pass to SentencePieceTrainer.train is the one your embedding matrix expects, that your chosen model_type (unigram or bpe) matches what the downstream checkpoint was trained with, and that decode(encode(text)) returns your input byte for byte on a sample that includes tabs and newlines, since the ▁ escape covers spaces and the README does not spell out what happens to other whitespace.

Frequently asked questions

What is SentencePiece and what is it used for?

It is an unsupervised text tokenizer and detokenizer for neural network based text generation, where the vocabulary size is fixed before training. It trains subword models, BPE or unigram, directly from raw sentences.

How do I install the SentencePiece library?

The Python module installs from PyPI with pip install sentencepiece. The repository also builds from source with Bazel or CMake, and ships wheels through its build workflow.

How do I use the SentencePiece tokenizer?

Train a model with spm.SentencePieceTrainer.train, passing input, model_prefix and vocab_size, then load it with spm.SentencePieceProcessor(model_file='m.model'). Encode with out_type=str for pieces or out_type=int for IDs, and decode returns the original string.

What is the difference between WordPiece and SentencePiece?

The README does not discuss WordPiece, so a direct comparison is not available. What it does state is that SentencePiece trains on raw Unicode text with no language-specific pre-tokenizer and escapes whitespace with the ▁ meta-symbol.

What does the SentencePiece tokenizer do with whitespace?

It treats the input as a raw sequence of Unicode characters and escapes whitespace with the meta-symbol ▁ (U+2581), which becomes part of the tokenization. The README states this makes detokenization a lossless string join independent of language.

Official sources

  1. google/sentencepiece on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/google-sentencepiece.svg)](https://hysenlabs.com/projects/google-sentencepiece)
Community notes

Community notes