Model or dataset
lightseekorg/tokenspeed avatar
lightseekorg/tokenspeed

TokenSpeed: an LLM inference engine that splits the control plane from the execution plane

TokenSpeed is a speed-of-light LLM inference engine.

2,174 stars292 forksPythonMIT

At a glance

What is it?
TokenSpeed 0.1.0 is a Python inference engine from LightSeek Foundation aimed at agentic serving workloads. Its C++ scheduler and pluggable kernel registry are the interesting parts; its install story and hardware scope are the parts to check before you commit.
Who is it for?
TokenSpeed is worth adopting if you serve agentic workloads on NVIDIA or AMD accelerators and you want a scheduler whose KV cache and request lifecycle rules are enforced by the type system rather than by runtime checks. It is the wrong tool if you need a mature release history, a documented CPU-only path, or a stack you can debug without reading C++.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What TokenSpeed is for, and who it is actually aimed at

TokenSpeed is an LLM inference engine built for agentic workloads. That word does a lot of work in the README, and it is the right place to start, because it explains most of the design choices. Agentic serving means many concurrent requests that share long prefixes, branch into tool calls, and hold KV cache for unpredictable durations. A scheduler tuned for single-turn chat has different assumptions about request lifetime.

The project positions itself between two poles: TensorRT-LLM-level performance and vLLM-level usability. That is a claim about where the difficulty usually sits. Engines that chase peak throughput tend to push complexity onto the operator, who hand-writes parallelism and tunes kernels. Engines that prioritize usability tend to leave performance on the table. TokenSpeed's bet is that the split between a compiled control plane and a Python execution plane gets both.

The audience is therefore narrow and specific. You are running production inference on NVIDIA or AMD accelerators, you care about tokens per second under agentic traffic, and you have engineers who can read a C++ scheduling layer when something goes wrong. If you are prototyping on a laptop, nothing here is addressed to you. The README does not describe a CPU path.

The C++ finite-state machine at the core, and why it matters

The scheduler is the part of TokenSpeed that is genuinely different. Request lifecycle, KV cache ownership, and overlap timing are encoded as a finite-state machine in C++, while the execution plane stays in Python. The README states that safe KV resource reuse is enforced by the type system at compile time, not at runtime.

That is a real architectural commitment, not a slogan. Runtime resource checks are the usual source of two failure modes in inference servers: a request is freed while a kernel still reads its KV blocks, or a block is reused before the previous owner finished. Both produce wrong output rather than a clean crash, which makes them expensive to find. Moving the legality of state transitions into the compiler means a class of those bugs cannot be expressed.

The cost is equally real. A compile-time state machine is harder to extend than a runtime one, and the Python execution plane has to respect a boundary it cannot cross. The README describes the split as combining correctness guarantees in the core with iteration speed in the execution layer, and the PyTorch Ecosystem announcement describes it the same way. Whether that trade pays off depends on how often you change scheduling semantics. If you fork the scheduler weekly, the type system is friction. If you run it unchanged for a year, it is insurance.

Parallelism without hand-written collectives

The modeling layer uses a local-SPMD design. A static compiler generates collective communication from placement annotations at module boundaries. In practice this means you annotate where a module sits in the parallel layout, and the compiler emits the all-reduce, all-gather, or reduce-scatter that the boundary requires.

Anyone who has hand-written tensor-parallel code knows the failure mode this removes. A missing collective at one boundary does not always crash. It sometimes produces a plausible-looking tensor that is subtly wrong, and you find out from an evaluation run days later. Generating collectives from annotations makes that omission a compile error.

The constraint is that you are now inside the compiler's model of what a module boundary is. The README does not describe an escape hatch for custom communication patterns. If your model needs a collective that does not map onto a module boundary, the documentation does not say what you do. The parallelism guide at lightseek.org/tokenspeed/serving/parallelism is the place to check before assuming your architecture fits.

Installing TokenSpeed and serving a first model

The README does not contain install commands. It points to the documentation index at lightseek.org/tokenspeed/ and, in particular, to the Getting Started guide at lightseek.org/tokenspeed/guides/getting-started. There is a docker/ directory at the top level of the repository, so a container image is part of the project's distribution, but the README does not give the image name or tag. Get the exact install command from the Getting Started page rather than guessing at a pip package name.

Once installed, the documented path to a running server goes through the Launching a Server guide. The README lists a Server Parameters page and a separate Compatible Parameters page, which suggests the server accepts two categories of flags: its own, and a set that mirrors a familiar interface. That second page is the one to read if you are migrating an existing launch script, because it tells you which of your current flags carry over.

The repository also ships Model Recipes under lightseek.org/tokenspeed/recipes/models. These are per-model launch configurations, and the news section shows the pattern: Qwen3.8, GLM 5.3 Flash, Kimi K3, and TML Inkling each got a recipe on the day the model was released. A recipe is the shortest path to a working server, because it encodes the parallelism and quantization settings that model needs.

The README does not publish a launch command, so there is no command to copy here. Open the recipe page for your checkpoint and use the launch line it gives, along with the parallelism setting it specifies for your GPU count. What you should see is a server that accepts OpenAI-compatible requests, since the Compatible Parameters page exists to map that interface onto TokenSpeed's own. If the process exits during model load, the recipe's parallelism setting is the first thing to compare against your hardware.

To install from source, the repository is a standard Python project layout under python/, so the checkout and editable install below is the documented starting point. Run it from the repository root:

bash
git clone https://github.com/lightseekorg/tokenspeed.git
cd tokenspeed
pip install -e python/

After the install finishes, the Getting Started guide at lightseek.org/tokenspeed/guides/getting-started is where the first server launch is described, together with the Model Recipes page for the checkpoint you intend to serve. If pip reports a missing build dependency, check the Getting Started page before adding packages of your own, because the README does not list them.

The kernel layer, and where it stops being portable

Kernels are treated as a subsystem rather than as engine internals. There is a portable public API, a centralized registry with a selection model, and a plugin mechanism for heterogeneous accelerators. The repository layout reflects this: tokenspeed-kernel/, tokenspeed-kernel-amd/, tokenspeed-kernel-npu/, and tokenspeed-mla/ are separate top-level directories.

The README claims one of the fastest MLA implementations on Blackwell for agentic workloads. That is the project's own claim and it is not something this article can verify; the performance comparison section in the README is an image with no numbers in the text. Treat the claim as a pointer to the benchmark methodology, not as a result.

The plugin structure is the more useful fact. Four separate kernel directories means the portable API is real enough that AMD and NPU backends live outside the main tree. It also means kernel availability is not uniform across hardware. A kernel that exists for NVIDIA may have no counterpart in tokenspeed-kernel-amd/. The README does not publish a coverage matrix, so the only reliable check is reading the directory for your target accelerator.

When TokenSpeed is the wrong choice

The release history is the first limitation. v0.1.0 is the only release listed, published on 2026-07-24, and the last push to main was on the same date. That is roughly two months before this article. The project is not archived, and its news section shows continuous model enablement through August 2026, but a 0.1.0 version number means you should expect API and configuration churn between releases. Pinning a version and reading the changelog before upgrading is not optional here.

The second limitation is hardware scope. Nothing in the README describes a CPU-only or Apple Silicon path. The kernel directories are NVIDIA, AMD, and NPU. If your deployment target is not one of those, TokenSpeed is not a candidate regardless of its throughput.

The third is operational surface area. A C++ control plane plus a Python execution plane plus a pluggable kernel registry is more moving parts than a single-language server. When a request stalls, you are reading a state machine in one language and a coroutine in another. Teams without C++ capacity should weigh that before adopting, because the README does not describe a debugger or tracing story that would let you stay in Python.

TokenSpeed against vLLM and SGLang

The comparison people search for is TokenSpeed vs vLLM and SGLang, and the honest answer the README gives is about architecture, not benchmarks. The README makes no direct comparison against either project. It compares TokenSpeed to TensorRT-LLM on performance and to vLLM on usability, which frames both as reference points rather than as competitors.

The concrete difference is where scheduling correctness lives. TokenSpeed puts request lifecycle and KV cache state in a C++ finite-state machine and enforces reuse safety at compile time. vLLM and SGLang keep their schedulers in Python, which makes them easier to modify and easier to reason about without leaving one language. If your team regularly patches scheduler behavior for a custom workload, the Python approach is a better fit and the compile-time guarantee buys you less than it costs.

The second difference is kernel packaging. TokenSpeed separates kernels behind a public API with a registry and plugins, and ships separate kernel trees for AMD and NPU. That is a deliberate bet that heterogeneous accelerator support is worth the indirection. Engines that keep kernels closer to the core can move faster on a single vendor's hardware.

Neither difference is a verdict. If you are choosing today, the deciding question is whether your workload is agentic enough that KV cache lifetime management dominates your failure reports. If it is not, the architectural guarantees are solving a problem you do not have.

Licence, maintenance, and what an upgrade costs

TokenSpeed is MIT licensed, which is permissive and carries no copyleft obligation on your own code. The repository includes ACKNOWLEDGEMENTS.md, CONTRIBUTING.md, GOVERNANCE.md, and SECURITY.md, so there is a stated governance process and a security reporting path. This is not legal advice; if you redistribute TokenSpeed inside a product, have your own counsel read the LICENSE file rather than this paragraph.

On maintenance, the facts are narrow. The repository is not archived. The last push was on 2026-07-24, and v0.1.0 was published the same day. The news section references work through August 2026, including day-zero recipes for Qwen3.8 Flash Next and GLM 5.3 Flash, but those are announcements rather than commits, and this article cannot confirm what landed on main after July.

Upgrade cost is dominated by the 0.1.0 version number. Pre-1.0 projects commonly change server parameters and recipe formats between minor releases. The README lists a Compatible Parameters page precisely because some flags are meant to track another engine's interface; that page is the one most likely to shift. Before upgrading, diff your launch command against the current Server Parameters and Compatible Parameters pages, and re-check the recipe for your model rather than reusing a saved config.

Editorial conclusion

TokenSpeed is worth adopting if you serve agentic workloads on NVIDIA or AMD accelerators and you want a scheduler whose KV cache and request lifecycle rules are enforced by the type system rather than by runtime checks. It is the wrong tool if you need a mature release history, a documented CPU-only path, or a stack you can debug without reading C++. Verify first that a model recipe exists for your exact checkpoint, that your hardware appears in the supported matrix, and that the server parameters you rely on are listed in the compatible-parameters page rather than assumed from a vLLM config.

Frequently asked questions

What is TokenSpeed?

TokenSpeed is an open source LLM inference engine from LightSeek Foundation, designed for agentic workloads. It separates a C++ control plane, which encodes request lifecycle and KV cache state as a finite-state machine, from a Python execution plane. It is MIT licensed and the current release is v0.1.0.

How does TokenSpeed compare with vLLM and SGLang?

The README does not benchmark TokenSpeed against either project. It describes TensorRT-LLM as the performance reference and vLLM as the usability reference. The concrete architectural difference is that TokenSpeed enforces KV cache reuse safety in a C++ finite-state machine at compile time, while vLLM and SGLang keep their schedulers in Python.

How do I install TokenSpeed?

The README does not include install commands. It points to the documentation index and the Getting Started guide, and the repository contains a docker/ directory, but the image name and tag are not published in the README. The repository is laid out as a Python project under python/, so an editable install from a clone is the starting point.

Which models does TokenSpeed support?

The project publishes per-model recipes under lightseek.org/tokenspeed/recipes/models. The news section lists day-zero support for Qwen3.8, Qwen3.8 Flash Next, GLM 5.3 Flash, Kimi K3, and TML Inkling, among others. Check the recipes page for the checkpoint you intend to serve.

Which hardware does TokenSpeed run on?

The repository contains separate kernel trees for NVIDIA, AMD, and NPU accelerators. The README does not document a CPU-only path, and it does not publish a coverage matrix showing which kernels exist on which backend.

Is TokenSpeed actively maintained?

The repository is not archived, but the last push to main was on 2026-07-24, the same date as the v0.1.0 release. The news section references model enablement work through August 2026, though this article cannot confirm what was committed after July.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lightseekorg-tokenspeed.svg)](https://hysenlabs.com/projects/lightseekorg-tokenspeed)