maciejhirsz/logos: a Rust lexer generator that builds its state machine at compile time
Create ridiculously fast Lexers
At a glance
- What is it?
- Logos turns an enum of token attributes into a single deterministic state machine at compile time. It suits Rust projects that need a fast tokenizer and can accept the constraints of the derive macro.
- Who is it for?
- Adopt Logos if you are writing a Rust lexer and want token definitions expressed as enum variants with #[token] and #[regex] attributes, and you can live with the derive macro's compile-time work. Do not adopt it if you need a runtime-configurable grammar, a lexer for a language you cannot express as regular patterns, or a parser: Logos produces tokens, not syntax trees.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 21 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Logos solves and who it is for
Writing a lexer by hand in Rust means writing a loop, matching characters, tracking spans, and deciding what to do when no rule matches. Logos replaces that loop with a derive macro. You declare an enum, annotate variants with #[token] for literal strings and #[regex] for patterns, and the macro generates the lexer at compile time. The README states the two goals plainly: make it easy to create a lexer, and make the generated lexer faster than anything you would write by hand.
The audience is Rust developers building a tokenizer as the first stage of a parser, an interpreter, a configuration language, or a data format reader. The repository ships examples for a JSON lexer, a calculator, a Brainfuck interpreter, and a small array language, which indicates the intended scale: small to medium grammars where tokens are regular and the interesting work happens after tokenization. If your grammar needs context-sensitive tokenization, Logos is the wrong layer.
How the derive macro builds one deterministic state machine
The README lists the mechanics: all token definitions are combined into a single deterministic state machine, branches are optimized into lookup tables or jump tables, backtracking inside token definitions is prevented, loops are unwound, and reads are batched to minimize bounds checking. All of that happens at compile time, which is why the generated code is fast and why compile times grow with the number of patterns.
The data flow is visible in the README example. Token::lexer("Create ridiculously fast Lexers.") returns an iterator. Each call to next() returns Option<Result<Token, _>>, and the helper methods span() and slice() describe the matched region. In the example, the first next() returns Some(Ok(Token::Text)), span() returns 0..6, and slice() returns "Create". The lexer therefore carries the source and the current span, and the iterator yields either a token, an error, or None at the end of input.
The skip attribute is part of the same mechanism. In the README example, #[logos(skip r"[ \t\n\f]+")] tells the lexer to ignore that pattern between tokens, so whitespace never reaches the iterator. The repository layout separates the runtime crate in src/ from the code generation crates logos-codegen and logos-derive, with logos-cli as a separate workspace member.
Installing Logos and writing a first lexer
Logos is published on crates.io. The workspace Cargo.toml sets the minimum supported Rust version to 1.80.0, so check your toolchain before adding the dependency. The book replaces the version string in its getting-started page on release, which is where the current version number is maintained.
Add the crate to Cargo.toml:
[dependencies]
logos = "0.16.1"The default features are export_derive and std. The Cargo.toml comments explain export_derive as re-exporting the Logos derive macro so that a user only needs to import this crate and write use logos::Logos. The std feature controls whether the crate uses the standard library.
A minimal lexer follows the README example. The enum derives Logos, Debug, and PartialEq, and the skip attribute ignores whitespace:
use logos::Logos;
#[derive(Logos, Debug, PartialEq)]
#[logos(skip r"[ \t\n\f]+")]
enum Token {
#[token("fast")]
Fast,
#[token(".")]
Period,
#[regex("[a-zA-Z]+")]
Text,
}Construct the lexer with Token::lexer on a string and iterate. The README example asserts that the first item is Some(Ok(Token::Text)), that span() is 0..6, and that slice() is "Create". If you see those values, the derive macro compiled and the state machine is matching as expected.
Where Logos stops being the right tool
The design choice that makes Logos fast also limits it. Token definitions are combined into one deterministic state machine and backtracking inside token definitions is prevented, so a pattern that would require the lexer to try one interpretation, fail, and return to an earlier position is not expressible. If your language needs that, you will be fighting the tool.
Compile time is the other cost. The heavy lifting happens at compile time, and the README presents that as a benefit, which it is at runtime. It also means a large enum of patterns increases build time, and error messages from the derive macro can point at generated code rather than at your attribute. The repository has a fuzz/ directory excluded from the workspace, which suggests the maintainers treat unexpected input as a real concern, but the README does not document a rollback or recovery strategy for a partially consumed token.
Finally, Logos is a lexer generator. It does not build a syntax tree, does not resolve ambiguities between grammar rules, and does not give you a parser. If you need a parser, you need a parser on top of the tokens.
Logos compared with a hand-written lexer
The alternative most Rust developers reach for is a hand-written lexer: a loop over bytes with a match statement, manual span tracking, and manual whitespace skipping. That approach has no derive macro, no compile-time state machine construction, and no restriction on backtracking. You can write any control flow you want, including lookahead and context-sensitive rules.
The difference in approach is where the complexity lives. With a hand-written lexer, the complexity is in your source file and in the tests you write for it. With Logos, the complexity moves into the derive macro and the generated state machine, and you interact with it through attributes. The README claims the generated lexer is faster than anything you would write by hand, and it publishes benchmark lines for identifiers, keywords and punctuation, and strings. Those numbers come from the project's own benchmark suite, not from an independent comparison, so treat them as the project's claim rather than a neutral measurement. The practical trade-off is that Logos gives you less control over the matching strategy in exchange for not writing the matching loop.
Maintenance, releases, and licensing
The repository is not archived. The last push was on 2026-09-11, which is close to the date of writing, and the recent release history shows v0.16.1 on 2026-01-30, v0.16 on 2025-12-07, and v0.15.1 on 2025-08-08. The v0.16 release is described as a major overhaul of Logos, and v0.16.1 is described as fixing a no_std problem and clarifying edge cases and docs. The workspace version is 0.16.1, so the crate and the workspace are in step.
Upgrade cost is concentrated in the derive macro's attribute surface. A major overhaul between v0.15 and v0.16 means code written against v0.15 should be checked against the v0.16 notes before upgrading. The Cargo.toml has a release configuration that rewrites the version string in the book's getting-started page, so the documented dependency version tracks releases automatically.
On licensing, the README states the code is distributed under the terms of both the MIT license and the Apache License (Version 2.0), choose whatever works for you, and the repository contains LICENSE-MIT and LICENSE-APACHE. The Cargo.toml license field is MIT OR Apache-2.0. This is a permissive dual license, but whether it fits your distribution model is a question for your own legal review.
Editorial conclusion
Adopt Logos if you are writing a Rust lexer and want token definitions expressed as enum variants with #[token] and #[regex] attributes, and you can live with the derive macro's compile-time work. Do not adopt it if you need a runtime-configurable grammar, a lexer for a language you cannot express as regular patterns, or a parser: Logos produces tokens, not syntax trees. Before committing, check that the patterns you need fit the supported regex subset in the Logos handbook, confirm the crate version you pin against the release notes for v0.16 and v0.16.1, and verify that the crate builds on your minimum supported Rust version, which the workspace Cargo.toml sets at 1.80.0.
Frequently asked questions
How do I install maciejhirsz/logos in a Rust project?
Add logos to the dependencies section of Cargo.toml; the book's getting-started page tracks the current version string. The workspace Cargo.toml sets the minimum supported Rust version to 1.80.0.
Does Logos work without the standard library?
The crate has a std feature that controls whether it uses the standard library, and the v0.16.1 release notes mention fixing a no_std problem. The default features are export_derive and std, so no_std use requires adjusting features.
What does the #[logos(skip ...)] attribute do in Logos?
It tells the lexer to ignore the given pattern between tokens. In the README example, skip r"[ \t\n\f]+" ignores whitespace so those characters never appear in the token stream.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/maciejhirsz-logos)