Model or dataset
marcelroed/gigatoken avatar
marcelroed/gigatoken

Gigatoken: Language Model Tokenization at Gigabytes per Second

Language model tokenization at GB/s

4,116 stars220 forksRustMIT

At a glance

What is it?
Gigatoken is a Rust-backed Python tokenizer built for language model data pipelines that need to process large corpora fast. It provides a drop-in compatibility layer for HuggingFace Tokenizers and tiktoken as well as a native API that reads files directly in Rust for maximum throughput.
Who is it for?
Gigatoken fits ML engineers who preprocess multi-gigabyte or multi-terabyte text corpora for language model training and where tokenization throughput is the bottleneck. It is less useful when you tokenize single documents at inference time, since the per-call startup cost matters more there.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 27 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The Problem Gigatoken Solves

Preprocessing a raw text corpus for language model training involves tokenizing billions of words. HuggingFace Tokenizers and tiktoken are both already written in Rust and run multithreaded, as the README notes. Despite that, they process text at tens of megabytes per second on typical hardware. A dataset the size of the Common Crawl can take many hours or days to tokenize with those tools. Gigatoken targets this pipeline stage specifically, encoding the same corpora at gigabytes per second by exploiting wider parallelism and direct file I/O that bypasses Python entirely for the hot path.

The README benchmarks encoding throughput on owt_train.txt, an 11.9 GB file of OpenWebText data used as a representative proxy for Common Crawl content after extraction. The benchmarks run across three hardware classes: a dual-socket AMD EPYC 9565 with 144 cores total, an Apple M4 Max with 16 cores, and an AMD Ryzen 7 9800X3D with 16 cores. Results vary by tokenizer vocabulary due to algorithmic differences: the GPT-2 tokenizer reaches 24.53 GB/s on the EPYC and 8.79 GB/s on the M4 Max, while tokenizers using the SentencePiece algorithm (such as Gemma and Mistral variants) reach 3 to 4 GB/s on those same machines.

The intended users are ML infrastructure engineers and researchers running large-scale pretraining or fine-tuning data preparation who have already profiled their pipeline and identified tokenization as the bottleneck.

Architecture: Rust SIMD Parallelism and Direct File Reading

Gigatoken is implemented in Rust and compiled as a Python extension module using Maturin and PyO3. The Cargo.toml dependency list includes SIMD-accelerated libraries such as simdutf (for UTF-8 validation), winnow with the simd feature (for parsing), memchr and aho-corasick (for fast pattern matching), and rayon (for data-parallel execution across CPU cores). Memory mapping via memmap2 lets the Rust layer read files from disk without copying data through the Python interpreter.

The Cargo.toml also lists parquet support via the arrow and parquet crates, and SentencePiece precompiled vocabulary support via spm_precompiled, indicating that these input formats are handled within the Rust layer. The Cargo.toml is compiled with fat link-time optimization (lto = "fat") for maximum runtime performance in the release profile.

The native API exposes a TextFileSource that accepts a list of file paths and a byte separator token, then encodes the entire file set in one call from the Rust side. The compatibility layer wraps the same engine but translates its output format to match the behavior of HuggingFace Tokenizers or tiktoken exactly. The README states that this output-matching overhead is non-negligible: compatibility mode is faster than the original libraries but not as fast as the native API.

Installing Gigatoken and Choosing an API Mode

Gigatoken is available from PyPI and requires Python 3.10 or later:

bash
pip install gigatoken

The README describes two API modes. Compatibility mode minimizes changes to existing code: you wrap an existing HuggingFace or tiktoken tokenizer with gt.Tokenizer, then call .as_hf() or .as_tiktoken() to get an object that behaves like the original. The native API accepts a HuggingFace model name string directly and returns a tokenizer that reads from files with maximum parallelism.

The pyproject.toml marks the project as Beta (Development Status :: 4) and version 0.10.0. The repository has no tagged GitHub releases.

Compatibility Mode: Using Gigatoken with Existing Code

Compatibility mode requires the smallest code change for projects already using HuggingFace Tokenizers:

python
import gigatoken as gt

hf_tokenizer = ...
tokenizer = gt.Tokenizer(hf_tokenizer).as_hf()

tokens = tokenizer.encode_batch(["This is a test string", "And here is another"])

The README states that a substantial effort has been made to match HuggingFace output exactly in this mode. For tiktoken users, the same pattern applies with .as_tiktoken(). The same mode works with tiktoken:

python
import gigatoken as gt

tiktokenizer = ...
tokenizer = gt.Tokenizer(tiktokenizer).as_tiktoken()
tokens = tokenizer.encode_batch(["This is a test string", "And here is another"])

The README notes that exact output matching in compatibility mode carries a non-negligible performance cost compared to the native API, so switching to compatibility mode will improve throughput over the original libraries but not reach the ceiling of what the Rust implementation can do.

The Native API: Reading Files Directly from Rust

The native API skips the Python data structure overhead entirely by reading files from disk in Rust:

python
import gigatoken as gt

tokenizer = gt.Tokenizer("Qwen/Qwen3-8B")  # Accepts HF model names
file_source = gt.TextFileSource(["owt_train.txt"], separator=b"<|endoftext|>")
tokens = tokenizer.encode_files(file_source)

TextFileSource accepts a list of file paths and a byte separator that is inserted between documents. According to the benchmarks in the README, this path achieves the highest throughput. The README also notes that passing Python data structures through this API still incurs the overhead of reading from Python, so the full benefit requires providing file paths rather than in-memory strings.

Where Gigatoken Falls Short

Gigatoken's throughput advantage comes from batching and parallelism. For single-sentence inference, where you encode one string at a time and the per-call overhead dominates, the speed improvement over HuggingFace tokenizers will be smaller or absent. The README makes no claim about latency for short inputs, and it does not document any configuration for minimizing per-call overhead.

The package is classified as Beta (Development Status :: 4) at version 0.10.0 and carries no tagged GitHub releases. There is no published API stability guarantee, and the changelog format is not documented in the README. Engineers maintaining long-lived data pipelines that depend on specific output formats should pin the installed version explicitly and validate outputs against their existing tokenizer before switching.

Not all vocabulary types reach the same throughput. The README's benchmark table shows that BPE-based vocabularies (GPT-2, Llama, Qwen families) achieve the highest throughput, while SentencePiece-based vocabularies (Gemma, Mistral, CodeLlama, and TinyLlama) top out at roughly 1 to 4 GB/s on the same hardware. The speedup ratio over HuggingFace varies accordingly.

Gigatoken requires Python 3.10 or later, as stated in pyproject.toml. Environments locked to Python 3.9 or older cannot install it.

Tiktoken: Simpler, Slower, More Established

OpenAI's tiktoken is a Rust-backed tokenizer that covers the BPE vocabularies used by OpenAI's models, including GPT-2, GPT-4, and the cl100k_base vocabulary. It is widely used in production inference code because it is the reference implementation for those specific tokenizers and has a simple, stable Python API.

The key differences are scope, stability, and speed. Tiktoken handles the OpenAI vocabulary family and is optimized for correct, reliable encoding rather than maximum throughput on large datasets. It carries no Beta classification and has a multi-year track record in production inference pipelines. Gigatoken supports tiktoken vocabularies through the compatibility layer and reports higher throughput in its README benchmarks, but it also introduces a dependency on a Beta Rust extension and requires Python 3.10 or later. Teams that tokenize text at inference time on a per-request basis and already use tiktoken have little reason to switch. Teams running offline corpus preprocessing jobs that have outgrown tiktoken's throughput are the target audience for Gigatoken.

Editorial conclusion

Gigatoken fits ML engineers who preprocess multi-gigabyte or multi-terabyte text corpora for language model training and where tokenization throughput is the bottleneck. It is less useful when you tokenize single documents at inference time, since the per-call startup cost matters more there. Before adopting it, note that the PyPI package is classified as Beta (version 0.10.0), there are no GitHub release tags, and exact output matching with HuggingFace tokenizers in compatibility mode carries a performance penalty.

Frequently asked questions

What is Gigatoken?

Gigatoken is a Python tokenizer for language models, implemented in Rust, that processes text corpora at gigabytes per second. It provides a compatibility layer for existing HuggingFace Tokenizers and tiktoken code, and a native file-reading API for maximum throughput in data preprocessing pipelines.

How does Gigatoken's compatibility mode work?

You wrap an existing HuggingFace tokenizer with gt.Tokenizer(hf_tokenizer).as_hf(), and the resulting object behaves like the original tokenizer. The README states that output matching with HuggingFace is exact in this mode but carries a non-negligible performance cost compared to the native API.

Does Gigatoken require a GPU to achieve high throughput?

The README does not document GPU usage. The benchmarks in the README cover CPU hardware only (AMD EPYC with 144 cores, Apple M4 Max, and AMD Ryzen 7 9800X3D), and the architecture relies on SIMD and multi-core CPU parallelism rather than GPU acceleration.

Official sources

  1. Issues
  2. License: MIT
  3. marcelroed/gigatoken on GitHub
  4. README
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/marcelroed-gigatoken.svg)](https://hysenlabs.com/projects/marcelroed-gigatoken)
Community notes

Community notes