tiktoken-rs packages OpenAI tokenizers with the error handling left in
Ready-made tokenizer library for working with GPT and tiktoken
At a glance
- What is it?
- A small Rust workspace whose only member is a crate that wraps the tiktoken library with ready-made OpenAI encodings, context sizes and chat completion budgets. Every fallible call returns a Result, and the GPT-6 context entry has no encoding behind it.
- Who is it for?
- Adopt tiktoken-rs if you are writing a Rust service that has to count tokens against an OpenAI model and would rather not shell out to Python. Do not adopt it if your stack needs Llama, Gemini or Mistral vocabularies, since the crate states its scope is OpenAI tokenizers and points elsewhere for the rest.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A workspace with one member and a vendored upstream
tiktoken-rs is a Rust library for tokenizing text with the OpenAI tokenizers. It sits on top of the `tiktoken` library rather than reimplementing byte pair encoding, and its pitch is convenience: ready-made tokenizer constructors instead of assembling rank data and special tokens by hand. The models named as supported are GPT-5.4, GPT-5, GPT-4.1, GPT-4o, o1, o3, o4-mini and the gpt-oss family.
The repository layout is small. `Cargo.toml` declares a workspace with `resolver = "1"` and exactly one member, `tiktoken-rs`, which is also where the `examples` directory lives. A `vendor/` directory and a `.gitmodules` entry at the top level suggest the upstream project is carried in-tree rather than resolved at build time, which matters if you build in a network-restricted CI job.
The rest of the tree is what a published crate accumulates: `Justfile`, `rustfmt.toml`, `renovate.json`, `scripts/`, plus `LICENSE`, `CONTRIBUTING.md`, `CODE_OF_CONDUCT.md` and `SECURITY.md`. Credit for the original code and the `.tiktoken` files goes to @spolu.
Counting tokens is three lines and one unwrap
Adding the crate is a single command, and the shortest useful program prints a token count:
cargo add tiktoken-rsuse tiktoken_rs::o200k_base;
let bpe = o200k_base().unwrap();
let tokens = bpe.encode_with_special_tokens(
"This is a sentence with spaces"
);
println!("Token count: {}", tokens.len());Two details in those lines matter. The constructor returns a `Result`, so the `unwrap` is load-bearing and the surrounding code has to decide what a failure means. And `encode_with_special_tokens` is the call to reach for when the input may contain special token text, since it treats them as tokenizable content rather than rejecting them.
The string in the example has three spaces in a row, which is there for a reason: runs of whitespace are where naive token counting disagrees with the model, and this is the case the example is set up to exercise.
The singleton exists to stop rebuilding the tokenizer
For code that encodes on every request, the project ships a second constructor with the same shape and no `Result`:
use tiktoken_rs::o200k_base_singleton;
let bpe = o200k_base_singleton();
let tokens = bpe.encode_with_special_tokens(
"This is a sentence with spaces"
);
println!("Token count: {}", tokens.len());The stated reason is to avoid re-initializing the tokenizer on repeated calls. That is the whole decision this crate forces on you: pay the parse cost per call and handle a `Result`, or pay it once and hand out a shared handle. In a server where every inbound request needs a token count, the second option is the one most code ends up using.
Encoding itself keeps upstream semantics. `CoreBPE::encode` mirrors the upstream `tiktoken` behaviour and returns a `Result`, and the project shows it used with the allowed set taken from `special_tokens()`:
let bpe = o200k_base().unwrap();
let allowed = bpe.special_tokens();
let (tokens, last_piece_token_len) = bpe.encode("hello ", &allowed).unwrap();The generic `encode_as` and `count` helpers return a `Result` as well, so there is no path through this crate that quietly yields a zero on failure.
Six encodings, and only two of them are current
Choosing the wrong encoding is the mistake that silently breaks a token budget, so the mapping is worth reading rather than guessing. `o200k_base` is the current one, used by the GPT-5 series, the o1, o3 and o4 series, gpt-4o, gpt-4.5, gpt-4.1 and the codex models. `o200k_harmony` is the exception, reserved for `gpt-oss-20b` and `gpt-oss-120b`.
The rest exist for older deployments. `cl100k_base` covers gpt-4, gpt-3.5-turbo and the text-embedding-ada-002 and text-embedding-3 models, so it is what an embedding pipeline wants. `p50k_base` is for the code models and text-davinci-002 and text-davinci-003, `p50k_edit` for edit models such as text-davinci-edit-001, and `r50k_base`, also known as `gpt2`, for GPT-3 era models like davinci.
Four of those six encode models nobody deploys today. They are kept because a tokenizer count has to match whatever model the request is actually going to, and a library that deleted the old tables would force you to guess on legacy traffic.
get_context_size answers for gpt-6 without an encoding
The context size table is where a rejected request gets avoided, and it holds one entry that deserves a second look. For any model name beginning with `gpt-6`, including variants and snapshots, `get_context_size` returns the GPT-6 family default of 1,050,000. The same figure is given for gpt-5.4 and gpt-5.4-pro.
That number and the tokenizer are two separate lookups, and only one of them works. GPT-6 tokenizer lookup and token-budget helpers are unsupported because no encoding mapping is available. So a budget calculation against a gpt-6 name returns a window size while the encoder for that family does not exist yet in the crate.
Elsewhere the table is a plain lookup: 1,047,576 for the gpt-4.1 family, 400,000 for gpt-5 and its mini and nano variants, 200,000 for o1, o3, o3-mini, o3-pro, o4-mini and codex-mini, 131,072 for gpt-oss, 128,000 for gpt-4o, o1-mini and gpt-5.3-codex-spark, 16,385 for gpt-3.5-turbo and 8,192 for gpt-4. Read it as the model's window, not as a promise about what this crate can encode.
Chat budgets need a typed message struct, or a feature flag
Counting the prompt of a chat request is a separate helper, `get_chat_completion_max_tokens`, and it does not take a bare string. It expects `ChatCompletionRequestMessage` values, each carrying a `content: Option<String>`, a `role: String` and the rest filled from `Default::default()`. Roles are plain strings, so the type does not stop you from writing a role the API will reject.
There is a second path for anyone already using the official async client, and it costs a Cargo feature. Enable `async-openai` and import `get_chat_completion_max_tokens` from `tiktoken_rs::async_openai` instead, then build the same conversation out of `async_openai` types such as `ChatCompletionRequestSystemMessage` and `ChatCompletionRequestUserMessage` with their content variants.
The feature matters because it decides which message type your counting helper accepts. A codebase that has already committed to the async client should enable it rather than maintain two parallel message structs, and the cost is a dependency the crate otherwise does not need.
Scope stops at OpenAI, and says so
The project states its boundary in one quoted note: it is focused on OpenAI tokenizers, and for non-OpenAI models such as Llama, Gemini and Mistral you should use the HuggingFace `tokenizers` crate instead. That sentence saves an afternoon, because a general-purpose tokenizer library would otherwise look like the obvious dependency.
The version floor is Rust 1.85 or newer, which is what the encode examples assume when they say the crate requires it. Full working examples for the supported features live in the `examples` directory inside the `tiktoken-rs` member, and for background on why the encodings differ the README points at the OpenAI Cookbook notebook on counting tokens with tiktoken.
The gap to be aware of is precision rather than coverage. This crate tells you how many tokens a model will see, which is what you need for a budget check. It does not tell you what the model will do with them, and a count that matches the API still leaves prompt design, output limits and retries entirely on you.
Three releases since April, all under 1.0
Version history is short and recent: v0.11.0 on 2026-04-08, v0.12.0 on 2026-06-02 and v0.12.1 on 2026-09-24, with the last push landing on the same day as that patch. The repository is not archived, and the licence is MIT.
Two details of that history are worth reading before you pin a version. Every release so far is below 1.0, which in practice means the API surface is still moving, and the project does not publish a compatibility promise for the 0.x line. And the gap between releases is two to four months, so an upgrade is a deliberate event rather than a weekly one.
`renovate.json` at the repository root suggests dependency updates are automated even when releases are not. Pin the minor version you test against, read CHANGELOG before moving, and remember that the constructor names and the encoding list are the parts you will actually notice changing.
Editorial conclusion
Adopt tiktoken-rs if you are writing a Rust service that has to count tokens against an OpenAI model and would rather not shell out to Python. Do not adopt it if your stack needs Llama, Gemini or Mistral vocabularies, since the crate states its scope is OpenAI tokenizers and points elsewhere for the rest. Verify two things before wiring it in: pick your constructor, o200k_base() or o200k_base_singleton(), on the basis of how often you will call it, and check that your model name is in the context size table rather than relying on the gpt-6 default, which resolves a window without an encoding.
Frequently asked questions
What does TikToken do?
It converts text into the integer token IDs that OpenAI models consume. tiktoken-rs packages those tokenizers for Rust, with ready-made constructors such as o200k_base() and helpers for counting tokens and for chat completion token budgets.
Which Rust version does tiktoken-rs require?
Rust 1.85 or newer. The crate is added with cargo add tiktoken-rs, and every fallible call returns a Result, including o200k_base(), CoreBPE::encode and the generic encode_as and count helpers.
Can tiktoken-rs tokenize Llama or Gemini models?
No. The crate states its scope is OpenAI tokenizers, and directs non-OpenAI models such as Llama, Gemini and Mistral to the HuggingFace tokenizers crate instead.
What context window does tiktoken-rs report for gpt-6 models?
get_context_size uses the GPT-6 family default of 1,050,000 for any name beginning with gpt-6, including variants and snapshots. Tokenizer lookup and token-budget helpers for that family remain unsupported because no encoding mapping is available.
How do I stop tiktoken-rs rebuilding the tokenizer on every call?
Use o200k_base_singleton() instead of o200k_base(). The project gives avoiding re-initialization of the tokenizer on repeated calls as the reason that constructor exists.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zurawiki-tiktoken-rs)