# sugarme/tokenizer: Hugging Face tokenizer configs in pure Go

> This package reimplements the Hugging Face Tokenizers pipeline in Go, split into normalizer, pretokenizer, tokenizer and post-processing sub-packages, with word level, wordpiece and BPE models. The tree also contains a sentencepiece sub-package and a unigram example that the feature list never claims.

**sugarme/tokenizer** — NLP tokenizers written in Go language

- Repository: https://github.com/sugarme/tokenizer
- Stars: 339 · Forks: 69
- Language: Go
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/sugarme-tokenizer

## Four sub-packages, one pipeline

The architecture is the same one Hugging Face popularised, and the package is explicit about that inspiration and about the concepts it borrowed. The work is divided into four modules, each in its own sub-package: a normalizer, a pretokenizer, the tokenizer itself, and post-processing. That order matters more than it looks, because it mirrors where text actually gets mangled. Normalization cleans the string, pre-tokenization chops it into chunks that are cheap to model, the model turns chunks into ids, and post-processing assembles those ids into the encoding with the special tokens and offsets you asked for.

The root of the module holds the pieces that do not belong to any single stage: tokenizer.go, pretokenizer.go, encoding.go, config.go, added-vocabulary.go, file-util.go, init.go and util.go. Alongside them sit the sub-packages decoder/, model/, normalizer/, pretokenizer/, pretrained/, processor/, spm/ and util/. Tests live next to the code they cover rather than in one directory, with bpe_test.go, config_test.go, encoding_test.go, added-vocabulary_test.go and pretokenizer_test.go at the root.

## CachedPath is how a Hub model gets in

The load path runs through two calls, and understanding them tells you what format this package actually speaks. tokenizer.CachedPath takes a model name and a filename, here bert-base-uncased and tokenizer.json, and returns a local path to the downloaded file. pretrained.FromFile then parses that file into a usable tokenizer. The comment in the example is precise about the requirement: any model with a tokenizer.json available will work, and a large model is given as the example.

So the interchange format is one JSON file per model, the same artifact the Python library emits. Everything else in the package is downstream of parsing it. There is also a path that skips the download entirely, since models can be loaded from files manually and the API reference on pkg.go.dev is where the details live.

Encoding is a method on the tokenizer rather than a free function, which keeps the model, normalizer and post-processor together:

```go
configFile, err := tokenizer.CachedPath("bert-base-uncased", "tokenizer.json")
if err != nil {
	panic(err)
}

tk, err := pretrained.FromFile(configFile)
```

The sample sentence in the documentation contains a mask token, so the example doubles as a check that special tokens survive the round trip.

## Three model families are actually claimed

The feature list is a checklist with three boxes ticked: the word level model, the wordpiece model, and byte pair encoding. That is the honest scope, and it is worth reading as a claim about what is finished rather than as a menu of everything in the directory. Byte pair encoding is the one most Go ports skip, and it is here.

Both training paths are claimed. The package can train new models from scratch and can fine-tune existing ones, and the examples directory is pointed at for the details. The example folders give a better sense of the intended entry points than the prose does: basic/, bpe/, decode/, pretrained/, truncation/ and unigram/ each correspond to one thing you would otherwise have to read the API reference to discover.

The pipeline is described as usable for training, test and inference, and the surrounding repositories matter for that claim. The module is one of three in the same author's effort to bring deep learning tooling to Go, sitting alongside transformer and gotch. If you only need inference, this package is self-contained; if you need to train, you are using the first piece of a stack you will have to assemble.

## spm/ and unigram are in the tree, not on the list

Here is the gap between what the repository contains and what it advertises. The top level carries an spm/ sub-package, which is the conventional name for sentence piece support, and the examples include a unigram/ directory. Neither appears in the list of implemented models.

That is not necessarily a defect. A sub-package can exist, be partially ported, or work for the subset of configurations the author needed, and the README makes no claim either way. What it does mean is that the feature list is the only statement about model coverage you can rely on, so if your production model is a sentence piece or unigram model, read the code in spm/ before you plan around it. The same caution applies to anything not ticked: treat the checklist as the contract and the directory listing as a hint.

The same logic applies to decoder/, processor/ and util/, which exist as sub-packages without appearing in the four module list. The four modules describe the processing stages, not the full inventory of the repository.

## The dependency list tells you what it does at scale

go.mod is short and readable, which makes it a decent map of the implementation. It needs emirpasic/gods for generic data structures, patrickmn/go-cache for the download and file cache behind CachedPath, rivo/uniseg for grapheme cluster segmentation so a multi-byte character is not split in half, schollz/progressbar/v2 for download progress, golang.org/x/sync for concurrency, golang.org/x/text for Unicode tables, and a regex set package from the same author for the normalisation and pre-tokenisation patterns.

Two indirect dependencies are declared: a colour string helper and testify for assertions. That testify entry is a hint about how the test suite is written, and the presence of root level test files alongside a coverage.out file checked into the repository suggests coverage has been run locally and committed rather than reported through a service.

A cache library in the list is the part with a practical consequence. CachedPath will keep downloaded files around, so in a container that runs once per build you pay the download every time, and on a laptop you get a warm cache. If you need to pin or purge that cache in a deployment, the documentation says nothing about a cache directory or an eviction policy, so treat it as unspecified.

## Go 1.23 declared, 1.24.4 toolchain

The module declares go 1.23.0 and a toolchain line of go1.24.4. In practice that means the package will build on Go 1.23 and newer, and Go 1.23 installations will fetch the newer toolchain automatically to compile it. If you are on a machine where toolchain downloads are disabled by policy, that line is where the build fails first, and the fix is to install Go 1.24.4 rather than to edit the module.

The default branch is master rather than main, which matters for the example directory paths and for anyone following a raw file link. The repository has a single published release, v0.3.0, dated 2025-09-18, and the last push to the default branch is dated 2026-05-26. Version tags and commits are therefore far apart, so a go get of the latest tag and a clone of master are not the same code, and there is no changelog entry in the README explaining what v0.3.0 contained.

## Conclusion

This is the right package if your Go service has to load a pretrained tokenizer.json and you would rather not shell out to Python, particularly for inference paths where token counts affect billing or truncation limits. It is the wrong choice if you need a training loop, since the documentation points at the companion transformer and gotch repositories for that, and if you need a model family the feature list does not claim. Before adopting it, check that your target tokenizer ships a tokenizer.json at all, since that file is the only interchange format described here, and pin the Go version to 1.23 or newer because the module declares that floor.

## FAQ

### What does a tokenizer do?

In this package the work is split into four stages in separate sub-packages: a normalizer, a pretokenizer, the tokenizer model itself, and post-processing. A single sentence is turned into an encoding that can be decoded back.

### What algorithm is used for tokenization?

Three model families are marked as implemented: the word level model, the wordpiece model, and byte pair encoding. The repository also contains an spm sub-package and a unigram example, which the feature list does not claim.

### how tokenizer works

Download the model's tokenizer.json through tokenizer.CachedPath, parse it with pretrained.FromFile, then call the encoder, for example EncodeSingle on a sentence. The stages run in the order normalizer, pretokenizer, tokenizer, post-processing.

## Sources

- [Issues](https://github.com/sugarme/tokenizer/issues)
- [License: Apache-2.0](https://github.com/sugarme/tokenizer/blob/master/LICENSE)
- [README](https://github.com/sugarme/tokenizer/blob/master/README.md)
- [Releases](https://github.com/sugarme/tokenizer/releases)
- [sugarme/tokenizer on GitHub](https://github.com/sugarme/tokenizer)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sugarme-tokenizer
