Hyperscan: Intel's Multi-Regex Matching Library for DPI Stacks
High-performance regular expression matching library
At a glance
- What is it?
- Hyperscan is a standalone C library that matches tens of thousands of regular expressions at once, in block or streaming mode, using hybrid automata. It is aimed at deep packet inspection pipelines, not at general-purpose scripting.
- Who is it for?
- Hyperscan fits teams building deep packet inspection or similar high-volume scanning pipelines in C or C++, where thousands of patterns must be evaluated per packet or per stream and where the BSD licence terms are acceptable. It is the wrong tool for one-off text processing, for callers who need capture groups and backreferences, and for anyone who cannot take a C toolchain dependency.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Hyperscan was written to solve
A conventional regex engine evaluates patterns one at a time. If you have ten thousand signatures to check against every packet on a link, running ten thousand separate matches per packet is not viable. Hyperscan is built for that shape of workload: the README describes it as a high-performance multiple regex matching library that allows simultaneous matching of large numbers (up to tens of thousands) of regular expressions, and matching across streams of data.
The intended consumer is explicit in the README: Hyperscan is typically used in a DPI library stack. That framing matters. This is not a library you reach for to validate an email address in a web handler. It is infrastructure for intrusion detection, traffic classification and content filtering, where the pattern set is large, mostly fixed at compile time, and applied to a firehose of bytes. The syntax follows libpcre, which lowers the cost of porting an existing signature set, but Hyperscan is a standalone library with its own C API rather than a drop-in PCRE replacement.
How the matching engine is structured
The README states that Hyperscan uses hybrid automata techniques. That phrase is doing real work: rather than committing to a single automaton representation, the compiler picks among several and combines them, which is what lets one compiled database hold tens of thousands of patterns without the state blowup a pure DFA would suffer.
The API shape follows from that design. Patterns are compiled ahead of time into a database, and matching runs against that compiled artifact rather than against pattern strings. Two modes exist, and the README names both: matching regular expressions across streams of data, and the non-streaming case. Streaming mode is what a DPI stack needs, because a signature can straddle a packet boundary and the engine has to carry state between chunks.
This is also where the main constraint lives. A compiled database is tied to the platform and to the feature set it was built with, so the compile step is not something you casually do per request. The repository layout reflects the split: include/ holds the public headers, src/ the engine, hs.def and hs_runtime.def the exported symbol lists, and chimera/ a separate component. The examples/ directory contains simplegrep.c, pcapscan.cc and patbench.cc, which map onto the three things a newcomer usually wants: a minimal match, packet scanning, and pattern-set benchmarking.
Building Hyperscan and running a first match
The README does not carry build instructions. It points at the Developer Reference Guide at intel.github.io/hyperscan/dev-reference for information on building the library and using its API, and the repository root has a CMakeLists.txt plus a cmake/ directory, so the build is CMake-driven. Treat the guide as the authority for the exact toolchain versions and options of the release you intend to use.
The repository ships examples, and examples/README.md is where the project describes them. The example sources are simplegrep.c, pcapscan.cc and patbench.cc. The simplest of them is a grep-style tool over a single pattern, and its source file name tells you what to look for in the build tree after compiling.
What the example demonstrates is the two-phase flow the library is built around: a pattern is compiled into a database, and that database is then used for matching. If you want to see the library under a realistic load, pcapscan.cc scans a packet capture instead of a text file, and patbench.cc is the tool the project provides for measuring a pattern set. Read examples/README.md before running any of them, since the README here does not document their arguments.
Where Hyperscan is the wrong choice
The first limitation is syntactic. Hyperscan follows libpcre syntax, but following the syntax is not the same as implementing all of it. Constructs that require backtracking, capture groups and backreferences among them, are not part of the model the README describes. If your patterns depend on extracting what matched inside a group, this library is answering a different question than the one you are asking. You get whether a pattern matched, at scale, not a parse tree.
The second limitation is operational. Compiling a large database is a heavyweight step, and the resulting artifact is platform-specific. That pushes Hyperscan toward long-running services with a fixed signature set and away from anything that compiles patterns per request or ships them between machines. A Python or Node process that wants to match a user-supplied regex on demand is the wrong fit on both counts.
The third is the dependency itself. Hyperscan is a C++ library with a C API, and adopting it means adopting a native build step in your pipeline. For a team whose deployment target is a managed runtime, that cost can exceed the benefit. The README also does not describe any rollback or version-migration path for compiled databases, so pinning a version and rebuilding on upgrade is the safe assumption.
Hyperscan versus RE2 and PCRE2
The natural comparison is with RE2, the other well-known engine built around the guarantee that matching time stays linear in the input. Both avoid backtracking, and both will refuse pattern constructs that cannot be handled that way. The difference in approach is the workload they were designed around. RE2 is a general-purpose library you embed to match one pattern, or a small set, safely. Hyperscan is built for the multi-pattern case: the README's headline claim is simultaneous matching of up to tens of thousands of expressions, and it is the compiled-database model, not a single-pattern API, that makes that tractable.
Against PCRE2 the split is sharper. PCRE2 implements the full backtracking feature set, including capture and backreferences, and pays for it with worst-case behaviour that depends on the pattern and the input. Hyperscan gives up those features and buys predictable throughput on large pattern sets. If your signature set was written for PCRE and uses backreferences, porting it to Hyperscan is a rewrite, not a recompile.
The honest summary is that these are not interchangeable. Pick Hyperscan when the pattern count is the problem. Pick PCRE2 when the pattern features are the problem.
Maintenance, licensing and upgrade cost
The repository is not archived, and the last push was on 2026-09-17, so the tree is being touched. That is a weaker signal than it sounds. The most recent release listed is v5.4.2 from 2023-04-19, preceded by v5.4.1 in February 2023 and v5.4.0 in January 2021. Commits landing on master and a new tagged release are different events, and the release cadence here is slow. Plan for a stable, infrequently versioned dependency rather than one that tracks upstream fixes quickly.
The README describes the branch policy: master always contains the most recent release, and further development toward the next release happens on the develop branch. It also advises that users, rather than developers, should be on master. That is a clear instruction about which branch to pin, and it implies the develop branch is not a supported target for production.
Licensing is stated plainly in the README: Hyperscan is licensed under the BSD License, with the LICENSE file in the repository as the authority. The repository metadata carries a NOASSERTION licence identifier, which means the automated classifier did not recognize the file; the README's statement and the LICENSE file are what you should read. A permissive BSD licence is generally compatible with closed-source distribution, but the terms are yours to review, not something to take from a summary. Note also the COPYING file at the repository root alongside LICENSE.
Editorial conclusion
Hyperscan fits teams building deep packet inspection or similar high-volume scanning pipelines in C or C++, where thousands of patterns must be evaluated per packet or per stream and where the BSD licence terms are acceptable. It is the wrong tool for one-off text processing, for callers who need capture groups and backreferences, and for anyone who cannot take a C toolchain dependency. Before adopting it, verify that the pattern set you intend to compile stays within the constructs Hyperscan supports, and check the Developer Reference Guide for the exact build requirements of the version you plan to pin, since the README itself points there rather than listing them.
Frequently asked questions
What is Hyperscan?
Hyperscan is a high-performance multiple regex matching library from Intel. It follows libpcre regular expression syntax but is a standalone library with its own C API, and it is typically used in a DPI library stack.
How do I install Hyperscan?
The README does not contain build steps; it directs readers to the Developer Reference Guide for building the library and using its API. The repository root has a CMakeLists.txt and a cmake/ directory, so the build is driven by CMake.
How does Hyperscan compare with RE2?
Both avoid backtracking, but RE2 is a general-purpose engine for matching one pattern or a small set, while Hyperscan is built for the multi-pattern case, with the README claiming simultaneous matching of up to tens of thousands of regular expressions via a compiled database.
How does Hyperscan compare with PCRE2?
PCRE2 implements the full backtracking feature set, including capture and backreferences, while Hyperscan follows libpcre syntax but omits constructs that require backtracking. Choosing between them comes down to whether pattern count or pattern features are your constraint.
What are alternatives to Hyperscan?
RE2 and PCRE2 are the closest points of comparison. RE2 shares the no-backtracking approach but targets single-pattern or small-set matching, and PCRE2 offers the full backtracking feature set at the cost of input-dependent worst-case behaviour.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/intel-hyperscan)