tiktoken: a BPE tokeniser you can call from Python and Rust
tiktoken is a fast BPE tokeniser for use with OpenAI's models.
At a glance
- What is it?
- tiktoken converts text to the token sequences OpenAI's models consume, with a Rust core behind a Python API. It is a good fit when you need exact token counts or a reversible encoding, and the wrong tool when you need a full generation stack.
- Who is it for?
- Adopt tiktoken when you need exact token counts, a reversible encoding, or your own special tokens on top of an existing encoding. Do not adopt it expecting a generation library, a training tokeniser, or a tool that works without a network on the first call.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 44 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem tiktoken solves: text in, token ids out
OpenAI's API bills and truncates in tokens, not characters. A Python string of 4,000 characters is not 4,000 tokens, and the ratio changes with language, punctuation and code. tiktoken exists to make that conversion explicit and local, so you can measure before you send a request instead of guessing from character counts.
The README describes it as "a fast BPE tokeniser for use with OpenAI's models" and gives the round trip as the first example: `enc.decode(enc.encode("hello world")) == "hello world"`. That reversibility is the property that separates it from a generic word splitter. BPE is lossless, works on arbitrary text including text outside the training data, and compresses: the README states that on average each token corresponds to about 4 bytes.
The audience is narrower than the name suggests. This is for engineers who are building on top of OpenAI models and need a token count in their own process: a context-budget check before a call, a chunking step in a retrieval pipeline, a cost estimate, or a custom encoding that adds chat control tokens. It is not a library that calls the API for you, and it does not train a tokeniser from your corpus.
How the Rust core and the Python registry fit together
The repository is a Python package with a compiled extension. `setup.py` declares a `RustExtension` named `tiktoken._tiktoken` with the PyO3 binding, and it is always built in release mode: the comment in the file says that between editable installs and wanting Rust for performance-sensitive code, "it makes sense to just always use --release". The Rust side lives in `src/` and is declared in `Cargo.toml`, which pulls in `fancy-regex`, `regex`, `rustc-hash` and `bstr`.
On the Python side, `tiktoken/core.py` holds the tokeniser API and `tiktoken/registry.py` holds the lookup. `tiktoken.get_encoding("o200k_base")` resolves a name to a constructor, and `tiktoken.encoding_for_model("gpt-4o")` maps a model name to the encoding that model uses. The constructors themselves live in `tiktoken_ext/openai_public.py`, which is the file the README points at for examples of the arguments a specific encoding needs.
That split matters for two reasons. First, the hot loop (splitting text by the pattern, then applying merges) runs in compiled code, which is where the README's performance claim comes from: tiktoken is described as between 3 and 6 times faster than a comparable open source tokeniser, measured on 1GB of text with the GPT-2 tokeniser against `GPT2TokenizerFast` from `tokenizers==0.13.2` and `transformers==4.24.0`. That is the project's own measurement, not an independent one, and it is tied to those versions. Second, the encoding definitions are data plus a constructor, not hard-coded logic, which is what makes the extension mechanism possible.
Installing tiktoken with pip and counting tokens
The README gives one install line and points at PyPI for the open source version. The package metadata in `pyproject.toml` sets `requires-python = ">=3.9"` and lists two runtime dependencies, `regex` and `requests`. There is also an optional extra, `blobfile`, declared as `blobfile = ["blobfile>=3"]`, which is relevant when encodings are loaded from remote storage rather than a local cache.
pip install tiktokenOnce installed, the shortest useful program resolves an encoding and counts tokens. The README's own first example uses `o200k_base` directly, and it shows `encoding_for_model` as the way to get the tokeniser for a specific model in the OpenAI API.
import tiktoken
enc = tiktoken.get_encoding("o200k_base")
assert enc.decode(enc.encode("hello world")) == "hello world"
# To get the tokeniser corresponding to a specific model in the OpenAI API:
enc = tiktoken.encoding_for_model("gpt-4o")You should see no assertion error, and `enc` should end up holding the encoding mapped to `gpt-4o`. If the model name is not in the mapping, that call fails rather than falling back silently, which is the behaviour you want in a cost check. To turn the encoding into a count, take the length of the list returned by `encode`.
The README also links the OpenAI Cookbook notebook "How_to_count_tokens_with_tiktoken.ipynb" for worked examples. That notebook is outside this repository, so treat it as a companion rather than part of the package.
If you want to see what BPE is doing rather than just count, the repository ships an educational submodule. The README shows training a small encoding and visualising how a real one splits text:
from tiktoken._educational import *
# Train a BPE tokeniser on a small amount of text
enc = train_simple_encoding()
# Visualise how the GPT-4 encoder encodes text
enc = SimpleBytePairEncoding.from_tiktoken("cl100k_base")
enc.encode("hello world aaaaaaaaaaaa")Note the leading underscore in `_educational`. It is a teaching aid, not a stable surface, and the README presents it as a way to learn the procedure.
Adding your own special tokens without forking the package
The README documents two extension paths and is explicit that the second is optional. The first is to construct an `Encoding` yourself and pass it around. The example copies the pattern string, the mergeable ranks and the special tokens from `cl100k_base`, then adds `<|im_start|>` and `<|im_end|>` at ids 100264 and 100265.
cl100k_base = tiktoken.get_encoding("cl100k_base")
# In production, load the arguments directly instead of accessing private attributes
# See openai_public.py for examples of arguments for specific encodings
enc = tiktoken.Encoding(
# If you're changing the set of special tokens, make sure to use a different name
# It should be clear from the name what behaviour to expect.
name="cl100k_im",
pat_str=cl100k_base._pat_str,
mergeable_ranks=cl100k_base._mergeable_ranks,
special_tokens={
**cl100k_base._special_tokens,
"<|im_start|>": 100264,
"<|im_end|>": 100265,
}
)The README carries two warnings next to this snippet, and both are worth taking literally. It tells you to load the arguments directly in production instead of touching private attributes, and it says that if you change the set of special tokens you must use a different name, because the name should make the expected behaviour clear. Renaming is not cosmetic: an encoding called `cl100k_base` that behaves differently from the real one is a bug waiting for the next reader.
The second path registers your encoding so that `get_encoding` can find it. You create a namespace package under `tiktoken_ext` with a module that defines `ENCODING_CONSTRUCTORS`, a dictionary from encoding name to a zero-argument function returning the arguments for `tiktoken.Encoding`. The README gives the layout and says to omit `tiktoken_ext/__init__.py`:
my_tiktoken_extension
├── tiktoken_ext
│ └── my_encodings.py
└── setup.pyThe `setup.py` uses `find_namespace_packages(include=['tiktoken_ext*'])` and depends on `tiktoken`. The README then installs it with `pip install ./my_tiktoken_extension` and adds one constraint in bold: do not use an editable install. That constraint is the kind of detail that costs an afternoon if you skip it, and it is the reason the registry mechanism is harder to debug than the first path.
Where tiktoken is the wrong tool
The most common mistake is treating a token count as a request count. tiktoken counts the tokens you hand it. It does not know about the chat template your endpoint applies, the system message, or any tokens the server adds around your payload. The README's mapping function gets you the right encoding for a model; it does not tell you what the API will do with your messages.
The second limitation is the encoding data. The package resolves encodings by name, and `requests` is a runtime dependency, which means the first lookup can involve fetching encoding files. The README itself does not document an offline workflow or an environment variable for pointing at a local cache directory, so if you are deploying into an air-gapped environment you should confirm the caching behaviour against the code in `tiktoken/registry.py` and the loader rather than assuming it works. The `blobfile` extra exists for the case where encodings are read from remote storage, which suggests the loading path is configurable, but the README leaves the details to the source.
Third, this is not a training tokeniser. There is no documented path for learning merges from your own corpus at production scale. The educational submodule trains a toy encoding on a small amount of text; that is a teaching example, not a corpus pipeline. If you need to fit a vocabulary to a domain corpus, tiktoken is not the component you are looking for.
Finally, the performance claim has a shape. It is stated against a specific comparison at specific versions, on 1GB of text with the GPT-2 tokeniser. It says nothing about your text distribution, your batch size, or how you call the API from Python. Counting tokens in a tight loop across millions of documents is a different workload from counting one prompt, and the README does not break the numbers down by that case.
tiktoken and Hugging Face tokenizers: different centres of gravity
The obvious alternative is the `tokenizers` library from Hugging Face, which is the comparison the README itself uses: `GPT2TokenizerFast` from `tokenizers==0.13.2` and `transformers==4.24.0` is the baseline for the 3 to 6 times figure.
The difference in approach is where the vocabulary comes from. tiktoken ships a fixed set of OpenAI encodings and a registry that maps model names to them; you can add special tokens or register your own constructor, but the mergeable ranks you start from are the ones OpenAI published. The Hugging Face stack is built around loading a tokenizer definition that travels with a model, which is the natural fit when the model is yours or comes from the Hub. If your pipeline already uses `transformers`, adding tiktoken means a second tokeniser dependency and a second source of truth about how text is split.
There is also a language split. tiktoken's Rust core is compiled into the Python extension through PyO3, and `Cargo.toml` declares both `cdylib` and `rlib` crate types, which means the same crate can be consumed as a Rust library. If you need tokenisation inside a Rust service that talks to OpenAI models, that is a real advantage over a Python-only tokeniser. If you need one tokeniser that covers dozens of non-OpenAI models, the Hugging Face route is shorter.
Maintenance, releases and what the MIT licence asks of you
The repository is not archived, and the last push was on 2026-08-17, which is the same date as the 0.14.0 release. Before that, 0.13.0 landed on 2026-05-15 and 0.12.0 on 2025-10-06. That is a release cadence measured in months, not weeks, and the CHANGELOG.md at the repository root is the place to read what changed between them.
The upgrade cost is mostly in the encoding names and the private attributes. The README's extension example reaches into `_pat_str`, `_mergeable_ranks` and `_special_tokens`. Those leading underscores are a signal: a minor release can change them, and the README already tells you to load arguments directly in production rather than depend on them. If you pin `tiktoken` and copy encoding arguments into your own module, upgrades are cheap. If you subclass or introspect the internals, they are not.
On licensing, `pyproject.toml` declares `license = { file = "LICENSE" }` and the repository is MIT. MIT is permissive, but the obligations are real: the copyright notice and permission notice need to travel with copies or substantial portions of the software. The encoding data is part of the package you install, so if you redistribute it, read LICENSE rather than assuming the permissive label covers every file. This is a description of what the repository states, not legal advice.
One more cost that is easy to miss: `setup.py` builds the Rust extension on install when no wheel matches your platform. The `pyproject.toml` build requirements include `setuptools-rust>=1.5.2`, and the wheel build skips `*-manylinux_i686`, `*-musllinux_i686` and `*-win32`, with macOS wheels built for `x86_64` and `arm64`. On a platform outside that matrix you are compiling Rust, which means a Rust toolchain in your build image.
Editorial conclusion
Adopt tiktoken when you need exact token counts, a reversible encoding, or your own special tokens on top of an existing encoding. Do not adopt it expecting a generation library, a training tokeniser, or a tool that works without a network on the first call. Before you commit, verify which encoding name your model maps to, whether your environment can reach the encoding files or needs a pre-populated cache, and whether the MIT licence text in LICENSE matches your redistribution plans.
Frequently asked questions
How do I install tiktoken?
The README gives a single command, `pip install tiktoken`, and points at PyPI for the open source version. The package requires Python 3.9 or newer according to pyproject.toml. If no wheel matches your platform, the install compiles the Rust extension declared in setup.py.
How do I use tiktoken to count tokens?
Get an encoding with `tiktoken.get_encoding("o200k_base")` or `tiktoken.encoding_for_model("gpt-4o")`, call `encode` on your text, and take the length of the result. The README's first example uses the same round trip to show that decoding restores the original string.
What is tiktoken used for?
It converts text into the token sequences OpenAI's models consume, using byte pair encoding. The README lists the properties that follow from BPE: the conversion is reversible and lossless, it works on arbitrary text, and it compresses, with each token averaging about 4 bytes.
What is the tiktoken cache?
The README does not document a cache directory or an environment variable for one. The package depends on `requests` and offers an optional `blobfile` extra for reading encodings from remote storage, so the loading path exists in the code, but the README leaves the details to the source files.
Can I use tiktoken offline?
The README does not document an offline mode or a cache location, so this is not something the README confirms. The presence of `requests` as a runtime dependency and the `blobfile` optional extra indicates encodings can be fetched rather than bundled, which you should verify in the loader before relying on an air-gapped deployment.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/openai-tiktoken)