Open-source project
rust-lang/regex avatar
rust-lang/regex

rust-lang/regex: linear time matching, and what it costs you

An implementation of regular expressions for Rust. This implementation uses finite automata and guarantees linear time matching on all inputs.

4,038 stars536 forksRustApache-2.0

At a glance

What is it?
The regex crate is the standard regular expression engine for Rust, built on finite automata so every search runs in worst case O(m * n) time. That guarantee is bought with look-around and backreferences, which the crate does not support.
Who is it for?
Adopt rust-lang/regex when you are writing Rust and want a regex engine whose running time you can reason about before you ship, especially on untrusted input. Do not adopt it if your patterns need look-around or backreferences, or if you are working outside Rust and only want a general purpose regex library.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 51 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What rust-lang/regex is for, and who should reach for it

This crate provides routines for searching strings for matches of a regular expression. The README states the central trade-off plainly: the syntax is similar to other regex engines, but it lacks several features that are not known how to implement efficiently, including look-around and backreferences. In exchange, all regex searches have worst case O(m * n) time complexity, where m is proportional to the size of the regex and n is proportional to the size of the string being searched.

That exchange is the whole product. If you are writing Rust and the input to your pattern comes from somewhere you do not control, a log line, a user field, a config file, a network payload, the linear time bound is the reason to pick this crate over a backtracking engine. A backtracking engine can be fast on typical input and then fall off a cliff on a crafted one. This crate's documentation does not promise fast; it promises bounded.

The audience follows from that. Rust developers who need a general purpose matching library and who are willing to give up the exotic constructs. People who need those constructs should look elsewhere rather than fight the syntax.

Finite automata, the O(m * n) promise, and where regex-automata fits

The description in Cargo.toml says the implementation uses finite automata. The README ties the complexity bound directly to that design choice: it is the absence of look-around and backreferences that makes the bound hold. Those two constructs are what force a backtracking engine to revisit positions, and revisiting is what turns a linear scan into an exponential one on adversarial input.

The crate does not expose one engine. The regex-automata directory contains a crate that exposes all the internal matching engines used by the regex crate. The README frames the split as intent: the regex crate exposes a simple API for 99% of use cases, while regex-automata exposes customizable behaviors. If you find yourself wanting to choose a matching strategy rather than accept the default, that subdirectory is where the choice lives, not in the top level crate.

There is a third crate in the workspace worth knowing about. regex-syntax provides a tested regular expression parser, abstract syntax and a high level intermediate representation for analysis. It does no compilation or execution. The README recommends it only if you are implementing your own engine or need to analyze pattern syntax, and explicitly says it is otherwise not recommended for general use. That is an unusually direct disclaimer, and it should be taken at face value.

Installing rust-lang/regex and matching your first date string

The README gives two equivalent ways to bring the crate into a repository: add regex to your Cargo.toml, or run cargo add regex. The second is the shorter path.

bash
cargo add regex

After that command, the dependency appears in your manifest and you can use the crate. The README's first example matches a date in YYYY-MM-DD format and pulls out the year, month and day using named capture groups and the x flag, which lets the pattern span lines with comments.

rust
use regex::Regex;

fn main() {
    let re = Regex::new(r"(?x)
(?P<year>\d{4})  # the year
-
(?P<month>\d{2}) # the month
-
(?P<day>\d{2})   # the day
").unwrap();

    let caps = re.captures("2010-03-14").unwrap();
    assert_eq!("2010", &caps["year"]);
    assert_eq!("03", &caps["month"]);
    assert_eq!("14", &caps["day"]);
}

Captures returns an Option, and the README's example unwraps it because the input is known to match. In real code the unwrap is the part you change. If you have many dates in one body of text, the README adapts the same pattern with captures_iter and extract, which yields the capture groups as an array rather than requiring named lookups.

rust
use regex::Regex;

fn main() {
    let re = Regex::new(r"(\d{4})-(\d{2})-(\d{2})").unwrap();
    let hay = "On 2010-03-14, foo happened. On 2014-10-14, bar happened.";

    let mut dates = vec![];
    for (_, [year, month, day]) in re.captures_iter(hay).map(|c| c.extract()) {
        dates.push((year, month, day));
    }
    assert_eq!(dates, vec![
      ("2010", "03", "14"),
      ("2014", "10", "14"),
    ]);
}

The README's own warning about this step is the one most likely to bite a new user: compiling the same regular expression in a loop is an anti-pattern, because compilation typically takes anywhere from a few microseconds to a few milliseconds depending on pattern size. It also prevents the reuse of allocations inside the matching engines. The recommended fix is std::sync::LazyLock, or the once_cell crate if you cannot use the standard library.

rust
use std::sync::LazyLock;

use regex::Regex;

fn some_helper_function(haystack: &str) -> bool {
    static RE: LazyLock<Regex> = LazyLock::new(|| Regex::new(r"...").unwrap());
    RE.is_match(haystack)
}

The static is initialized on first use and reused afterward. The README notes a regex! macro that handles lazy compilation as well.

When the &str API is the wrong door: the bytes module

regex::Regex requires the caller to pass a &str, and a &str must be valid UTF-8. That means the main API cannot search arbitrary bytes. If your data is a binary protocol, a file format with mixed encodings, or anything that may contain invalid UTF-8, the main API is not merely inconvenient, it is unusable for the task.

The README points to regex::bytes::Regex for this. The API is identical except that it takes an &[u8]. There is a second difference that matters more than the type change: the byte oriented APIs permit disabling Unicode mode even when the pattern would match invalid UTF-8. The README's example is (?-u:.), which is rejected by regex::Regex but accepted by regex::bytes::Regex, where it matches any byte except \n. By contrast, a plain . matches the UTF-8 encoding of any Unicode scalar value except \n.

rust
use regex::bytes::Regex;

let re = Regex::new(r"(?-u)(?<cstr>[^\x00]+)\x00").unwrap();
let text = b"foo\xFFbar\x00baz\x00";

let cstrs: Vec<&[u8]> =
    re.captures_iter(text)
      .map(|c| c.name("cstr").unwrap().as_bytes())
      .collect();

The pattern finds null terminated strings in a byte slice, and [^\x00]+ matches any byte except NUL, including \xFF, which is not valid UTF-8. The README notes the same pattern under the main API would instead match valid UTF-8 sequences only. This is the kind of thing that produces silently different results rather than an error, so it is worth deciding deliberately which module you are in.

RegexSet: many patterns, one pass over the text

A single regex is not always the shape of the problem. If you need to know which of several patterns appear in a body of text, running them one at a time scans the input repeatedly. RegexSet is the crate's answer: it matches multiple, possibly overlapping, regular expressions in a single scan of the search text.

rust
use regex::RegexSet;

let set = RegexSet::new(&[
    r"\w+",
    r"\d+",
    r"\pL+",
    r"foo",
    r"bar",
    r"barfoo",
    r"foobar",
]).unwrap();

let matches = set.matches("foobar");
assert!(!matches.matched(5));
assert!(matches.matched(6));

In the README's example the haystack foobar matches several of the patterns at once, and the result reports indices rather than capture groups: index 6, foobar, matches, while index 5, barfoo, does not. The matches value can be iterated to collect all matching indices. So the set answers which patterns matched, not where each one matched. If you need spans or capture groups per pattern, you are back to individual Regex values. That is a real boundary in the API, not a configuration detail.

Note also that the set is constructed from a slice of patterns and compiled once, which puts it on the same side of the compilation cost warning as a single Regex held in a static.

The missing features are the point, not an oversight

Look-around and backreferences are absent, and the README is explicit that this is because they are not known to be implementable efficiently. This is the clearest case in the crate of a limitation that is also the feature. A pattern that relies on a lookahead assertion to express a constraint will not compile here, and no amount of tuning changes that.

The practical consequence is that porting patterns from another engine is not a copy and paste operation. A pattern that worked in Python or JavaScript may need restructuring, or may need to move out of the regex layer entirely and into code that inspects the match afterward. If your team's patterns are full of lookahead, the migration cost is real and should be measured before adoption rather than discovered during it.

The second limitation is less obvious. The README describes the crate as exposing a simple API for 99% of use cases, with regex-automata for the rest. That phrasing is honest about the boundary but it also means the top level crate deliberately withholds control. If your requirement is a specific matching strategy rather than a correct answer within the time bound, you are in the 1% and should be reading regex-automata's documentation, not the main crate's.

What rust-lang/regex is not, compared with a backtracking engine

The natural alternative is a backtracking regex engine of the kind found in most languages: the engines behind Python's re module, JavaScript's RegExp, and Java's java.util.regex. Those engines support look-around and backreferences, which is exactly what this crate gives up, and they are the reason a pattern written for one of those languages may not work here.

The difference in approach is not cosmetic. A backtracking engine explores alternatives and can revisit the same position many times; on typical input that is fast, and on adversarial input it can blow up. A finite automata engine does not backtrack, which is what makes the O(m * n) bound hold, and it is also why the constructs that require backtracking cannot be supported. You are trading expressive power for a running time you can state in advance.

That trade is not always worth making. If your patterns are fixed, written by your own team, and run against input you generate, the expressive power of a backtracking engine may be the better deal. The bound matters most when the pattern is fixed but the input is not, or when the input is large enough that a pathological case would be noticed.

Within Rust there is also a lighter option in this same repository. The workspace members include regex-lite, a separate crate from the main one. If you want the Rust regex API without the full feature set, that member exists in the tree, and its own documentation is the place to check what it drops.

Licence and the cost of staying current

The repository carries both LICENSE-MIT and LICENSE-APACHE at the top level, and the Cargo.toml declares license = "MIT OR Apache-2.0". The dual licence is the conventional Rust arrangement: you choose one of the two. The workspace include list also ships LICENSE-UNICODE, which is separate from the two code licences and is worth reading if you redistribute the crate, because Unicode data carries its own terms. This is a description of what the files say, not legal advice; if the distinction matters to your organisation, have someone qualified read the actual licence texts.

The version in Cargo.toml is 1.13.1, and the crate declares rust-version = "1.65" and edition = "2021". Those two lines are the real upgrade constraint: a toolchain older than 1.65 cannot build this version. The default feature set in Cargo.toml is std, perf, unicode and regex-syntax/default, and the comment above the std feature states that removing it will prevent regex from compiling, so std is not currently optional in practice despite being listed as a feature.

The last push to the repository was on 2026-08-10, which is recent. The crate has been at 1.x since 1.0.0 in 2018, so the public API has had a long stable run, and the workspace layout with regex-automata, regex-syntax, regex-lite and regex-cli suggests the maintenance effort is spread across several crates rather than concentrated in one file.

Editorial conclusion

Adopt rust-lang/regex when you are writing Rust and want a regex engine whose running time you can reason about before you ship, especially on untrusted input. Do not adopt it if your patterns need look-around or backreferences, or if you are working outside Rust and only want a general purpose regex library. Before committing, verify two things: that your patterns compile under the documented syntax, and whether your use case needs the bytes API rather than the &str one. The regex-automata subdirectory is the escape hatch if the default engine's strategy choices do not fit.

Frequently asked questions

What does rust-lang/regex mean by linear time matching?

The README states that all regex searches in the crate have worst case O(m * n) time complexity, where m is proportional to the size of the regex and n is proportional to the size of the string being searched. That bound comes from the finite automata implementation described in Cargo.toml.

How do I write a simple regex with the Rust regex crate?

The README's first example builds a Regex from a raw string pattern, calls captures on the input, and reads named groups out of the result by name, such as caps["year"]. It also shows the x flag, which lets the pattern span multiple lines with comments.

How do I install rust-lang/regex in a Rust project?

The README says to either add regex to your Cargo.toml or run cargo add regex. The crate declares rust-version = "1.65" and edition = "2021" in Cargo.toml, so the toolchain has to meet that.

How do I use rust-lang/regex to find every match in a string?

The README shows captures_iter combined with extract to iterate over all matches and pull out the capture groups as an array, using a pattern for dates in YYYY-MM-DD form as the example.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. rust-lang/regex on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rust-lang-regex.svg)](https://hysenlabs.com/projects/rust-lang-regex)