StringZilla: one string kernel wearing eleven language bindings
Up to 100x faster strings for C, C++, CUDA, Python, Rust, Swift, JS, & Go, leveraging NEON, AVX2, AVX-512, SVE, GPGPU, & SWAR to accelerate search, hashing, sorting, edit distances, sketches, and memory ops π¦
At a glance
- What is it?
- A header-only C library that treats substring search, hashing, edit distance and sorting as vectorizable work, then ships the same kernel through CPython, Cargo, npm, cGo, SwiftPM, Maven and NuGet.
- Who is it for?
- StringZilla is an unusual project in that the interesting engineering sits below the language binding you actually touch. The C and C++ kernels, the runtime ISA dispatch and the vectorized edit distance families are the product; the eleven language packages exist so that a Python or Rust developer can reach them without writing an extension by hand.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Why a string library exists at all when CPUs have no string instruction
The opening of the README states the premise more plainly than any benchmark table can. Strings are the first fundamental data type every language implements in software rather than hardware. The closest thing a CPU offers is x86's `PCMPISTRI`, described as too slow and too narrow to build a library on, and nothing at all ships an instruction for computing a string hash. What you are left with is the familiar shape of string processing code, a loop with a branch and a per-character comparison, where the surrounding control flow frequently costs more than the character logic itself.
The second half of the argument is about register pressure. A modern CPU carries dozens of 16 to 64 byte architectural registers and hundreds of physical ones to feed out-of-order execution, and byte-at-a-time processing simply leaves that machinery idle. StringZilla reaches for SIMD and SWAR instructions directly, where SWAR means treating a machine word as a bundle of independent lanes and using arithmetic to test many bytes at once when a vector unit is unavailable. The library bills itself as one of the widest, fastest and most portable collections of text-processing primitives, with lazily evaluated iterators throughout and no allocation on the hot path.
The claim of breadth is the part worth believing first, because it is the part a competitor cannot copy by adding a kernel. Besides exact and fuzzy matching, hashing, edit distances, sorting, segmentation and random string generation, the package.json description names fingerprinting explicitly. The `Cargo.toml` categories add `no-std` and `hardware-support`, which tells you the same headers compile for a microcontroller and for a database storage engine.
Reading the benchmark table without believing it
The README reports throughput on two CPUs and one GPU, grouped by operation, and only the languages that ship a counterpart appear under each heading. The framing note is the most useful sentence in the section: `StringZilla.C` is the C kernel called directly while `StringZilla.Py` is the same kernel reached through the CPython binding, so the gap between the two columns is the cost of crossing the interpreter boundary.
A fragment of the table, reproduced exactly as the README lays it out:
Xeon4 M5 Pro H100
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Unicode case-insensitive substring search (GB/s)
Python icu.StringSearch 0.06 0.15
StringZilla.C sz_utf8_uncased_search 13.16 9.30
StringZilla.Py sz.utf8_uncased_search 12.40 9.19
Find the first occurrence of a random word, β
5 bytes (GB/s)
LibC strstr 21.3 3.5
STL C++ std::string::find 9.1 10.9
Python str.find 2.5 3.1
StringZilla.C sz_find 21.0 37.1
StringZilla.Py sz.find 18.6 33.4Two things stand out. On the Xeon, `sz_find` lands roughly level with `strstr` at 21.0 against 21.3 GB/s, while on the M5 Pro the system implementation drops to 3.5 and StringZilla climbs to 37.1. That asymmetry is the whole argument for the library on Apple hardware, and it explains the separate bullet claiming 3x over LibC on Arm servers and 9x on Apple Silicon where the system `strstr` is weaker. The second thing is the interpreter overhead, which is small for search, since 21.0 becomes 18.6. It is a much larger fraction of the edit distance work, which is the harder case for a binding to hide.
Treat these numbers as vendor benchmarks until you run them on your own data. String length distribution, needle frequency and cache residency all matter, and the README is upfront that this is throughput on inputs the author chose.
One header tree, eleven ways in
The repository tree is the clearest statement of scope. Directories include `c/`, `include/`, `forkunion/`, `rust/`, `python/`, `javascript/`, `csharp/`, `java/`, `swift/`, `golang/`, alongside `bench/`, `test/` and `probes/`. Build manifests sit at the root for nearly every one of them: `CMakeLists.txt`, `pyproject.toml`, `setup.py`, `package.json`, `Cargo.toml`, `Package.swift`, `StringZilla.csproj`, `binding.gyp` and `pom.xml`. A language gets a manifest only if there is a real package to distribute, so the directory list is the honest feature list.
The README positions each entry point as an upgrade rather than an addition. In C 99 you swap `<string.h>` for `<stringzilla/stringzilla.h>`, in C++ 11 the STL `<string>` for `<stringzilla/stringzilla.hpp>`, and bulk GPU work goes through `<stringzillas/stringzillas.cuh>` in CUDA C++ 17. Python gets a `Str` type, Rust a traits crate, Go a cGo module, Swift a `String+StringZilla` extension, C# a zero-copy view over `ReadOnlySpan<byte>` that is NativeAOT friendly, and Java a pure FFM API over `MemorySegment` with no JNI. There is even a shell entry point, an `sz-` prefix for accelerating common command line tools.
The C# and Java rows are the most architecturally interesting. Zero-copy over `ReadOnlySpan<byte>` avoids marshalling and stays compatible with ahead-of-time compilation, while the Java binding using the Foreign Function and Memory API avoids JNI entirely, which matters because JNI has historically been a poor neighbour for native libraries that want direct memory access. Both choices trade some ergonomics for control, and both are recent enough additions that you should read the binding directory before assuming parity with C.
What the packaging reveals about who this is for
The `pyproject.toml` is unusually well commented, and the comments explain the release matrix rather than restating the schema. One `SZ_TARGET` environment variable selects which of three packages a build produces, and the release workflow drives a different platform set for each. `stringzilla` covers CPython 3.10 through 3.14 including the 3.14t free-threaded build, on Linux manylinux and musllinux across x86_64, aarch64, i686, ppc64le, s390x, armv7l and riscv64, plus macOS on both architectures and Windows AMD64, x86 and ARM64. `stringzillas-cpus` narrows to x86_64 and aarch64 on Linux, with a stated reason: the C++ engines are too slow to build and test under QEMU. `stringzillas-cuda` narrows further because NVIDIA ships no CUDA Toolkit for Windows on Arm and no musl toolchain at all.
That matrix is a map of the intended audience, and the README lists it directly: data engineers parsing large corpora such as CommonCrawl, RedPajama or LAION, engineers optimizing string handling inside an application, bioinformaticians and search engineers looking for edit distances in the style of USearch, DBMS developers optimizing `LIKE`, `ORDER BY` and `GROUP BY`, hardware designers who need a SWAR baseline, and students studying SIMD and SWAR applied to non-data-parallel work. The `no-std` category in `Cargo.toml` and the claim about sandboxed browser, database and LLM environments through a custom WebAssembly backend line up with that last pair.
The npm side is compact by comparison. `package.json` declares ESM, requires Node 22 or newer, depends on `bindings`, `node-addon-api` and `node-gyp-build`, and lists six platform packages as optional dependencies pinned at the same 5.1.2 version. The scripts are `node-gyp rebuild` for building and `node-gyp-build` for installing, which means a prebuilt binary is preferred and compilation is the fallback.
Build-time hazards the maintainers wrote down
The comments in `setup.py` are worth reading even if you never build from source, because they document failures that are otherwise baffling. Two constants set the memory budget: `CPP_MEMORY_PER_WORKER_GB` at 2.0 and `CUDA_MEMORY_PER_WORKER_GB` at 4.0, the second annotated with the reason it exists. Four concurrent `.cu` passes exhausted the 15.6 GB of a four-core arm64 runner and the kernel killed the build.
The function that turns those numbers into a worker count is explicit about why core count alone is the wrong bound. A compiler killed by the out-of-memory killer takes the whole build down with no diagnostic, which on a small many-core box is a real hazard rather than a theoretical one. The helper reads `MemAvailable` and `MemTotal` from `/proc/meminfo`, returns `None` where that file is absent as on Windows, and lets `SZ_MAX_COMPILE_WORKERS` override both the core and memory bounds. If you hit a build failure on a memory-constrained machine, that environment variable is the first lever to reach for.
The Rust manifest carries constraints of a different kind. It requires Rust 1.83, with the comment explaining that the `Byteset` mutators are `const fn` taking `&mut self`, which needs `const_mut_refs`. It sets `links = "stringzilla"` so Cargo refuses a dependency graph holding two copies of the crate, since the bundled C sources export global `sz_*` symbols. That is exclusivity added to resolution only, with the actual linking left to `build.rs`. The default feature set is `std` plus `dynamic-dispatch`, and the `cpus` feature compiles every ISA tier and picks the best one at load via a dispatch table, mirroring the CMake shared library. Disabling default features resolves the tier at compile time instead, the analogue of the header-only build.
What the recent releases say about the state of the code
Release notes are the cheapest available evidence of maintenance quality, and v5.1.2 is unusually informative. It is a patch release dated 2026-08-12 and its list is almost entirely safety fixes: a silent offset wrap-around in `arrow_strings_tape`, rejection of single-pass iterators in `arrow_strings_tape::try_assign`, a heap overflow in `sz_string_reserve` when shrinking, and an over-read in the Haswell four-lane Myers kernel. The rest match CPython helper linkage to its declarations, bin benchmark durations in CPU cycles, complete the cipher aliasing contract in the docs, and decouple the aarch64 CUDA wheels from the x86 ones.
Each of those is a specific kind of bug. A silent offset wrap-around means a length computation could go negative without complaint. A heap overflow when shrinking a buffer is the kind of defect that only appears under a particular reallocation sequence. An over-read in a hand-written SIMD kernel is the classic tail-handling failure, and naming the microarchitecture means the fix landed in the ISA tier where it happened rather than being masked everywhere. The Arrow tape items suggest the library has grown a columnar-adjacent surface where iterators are assigned in bulk, which is exactly where single-pass iterator misuse becomes possible.
v5.1.1 from a week earlier shows the same pattern with a different flavor: snapshot caller sequences before walking them, cache Unicode data atomically for the parallel tests, guard the device scope in the NW and SW cross-products, and link the CUDA objects against the dynamic CRT. That one also added the memory bound on compile workers discussed above, so a build failure someone hit in CI turned into a guardrail in the packaging.
v5.1.0 is the substantive release underneath those patches, and its rationale is worth reading in full because it explains a direction rather than a fix. It argues that AES-256 and SHA-256 have been optimized for decades and that on x86 the implementation has largely converged on the same assembly snippets reused across most TLS and cryptography libraries, but that optimization axes remain. The first is parallelizing across input streams, computing checksums for more than one file simultaneously, which the author notes is attempted mainly by Intel's ISA-L Crypto and Multi-Buf Crypto and that neither is widely used, easy to integrate, or friendly to small inputs. On that axis, wide AVX-512 intrinsics beat relying on specialized SHA hardware. The second is using newer and wider specialized hardware. The library is expanding beyond strings into crypto primitives, which raises the stakes on the aliasing contract the patch release documents.
Where this fits against the usual alternatives
The comparison that actually decides adoption is against what you already have. For ASCII or byte-level substring search in C, `strstr` and `memmem` are well optimized and free, which the table shows, since `sz_find` only ties `strstr` on the Xeon. If your workload is short strings or you are not memory-bound, the case for a dependency is weak. The wins appear when text is long, when you need Unicode-aware matching, or when the platform's implementation is the weak one, as on Apple Silicon.
For Unicode work the comparison is ICU, and the README claims 10x to 70x over both ICU4C and its Rust successor ICU4X across UTF-8 handling, case folding, segmentation and tokenization. That is the range where the project's ambition is clearest. ICU is a very large library with a large surface, and the case for a leaner alternative is strongest for services that need matching and folding but not the full internationalization stack. The risk in that direction is correctness at the edges of Unicode, which is where most of the difficult bugs live.
For fuzzy matching, the reference points are RapidFuzz on the Python side and the GPU edit distance families on the C side. The README claims over 10x faster than NVIDIA's own libraries for on-GPU Levenshtein, NW and SW distances, with the H100 column showing `szs_levenshtein_distances` at 5,980,110 MCUPS for roughly 100 byte DNA strings on one core, which is a different measurement axis from the throughput numbers and should not be compared to them directly. On GPU the pitch is that the grid-parallel formulation beats the vendor-tuned serial kernel, which is a plausible outcome given how differently the two problems are shaped.
The honest summary is that StringZilla replaces several libraries you may already carry, but only in the operations it actually covers, and it expects you to own the correctness question. It is at 3,563 stars, 138 forks, 31 open issues, Apache-2.0 licensed, not archived, with the last recorded push on 2026-09-22, and a homepage essay at `ashvardanian.com` that the README points to for design decisions.
Editorial conclusion
StringZilla is an unusual project in that the interesting engineering sits below the language binding you actually touch. The C and C++ kernels, the runtime ISA dispatch and the vectorized edit distance families are the product; the eleven language packages exist so that a Python or Rust developer can reach them without writing an extension by hand. The tradeoffs follow from that. You get allocation-free lazy iterators and a single code path across languages, at the cost of a native build step, a wide platform matrix to reason about, and a header set that has to be right about memory alignment because it does its own vector loads. The release history is the most reassuring signal available: v5.1.2 is a patch full of fixes for a silent offset wrap-around in the Arrow tape, a heap overflow in `sz_string_reserve`, and an over-read in the Haswell four-lane Myers kernel. That is a maintainer who reads sanitizer output. If you want the numbers on your own data rather than the README table, use the `bench/` directory before you commit to anything, because the gap between the C kernel and the language binding is the part most likely to surprise you.
Frequently asked questions
Which languages can I actually use StringZilla from?
Ten language surfaces plus a shell prefix. C 99 uses `<stringzilla/stringzilla.h>`, C++ 11 uses `<stringzilla/stringzilla.hpp>`, CUDA C++ 17 uses `<stringzillas/stringzillas.cuh>`, Python exposes a `Str` type, Rust ships a traits crate, Go uses cGo, Swift adds a `String+StringZilla` extension, JavaScript uses the npm package, C# works zero-copy over `ReadOnlySpan<byte>` and is NativeAOT friendly, and Java uses a pure FFM API over `MemorySegment` without JNI. There is also an `sz-` prefix for accelerating command line tools. The npm package requires Node 22 or newer, the Cargo crate requires Rust 1.83, and the Python wheels target CPython 3.10 through 3.14 including the free-threaded build.
How much of the speedup survives crossing a language boundary?
For search, most of it. The README reports `sz_find` at 21.0 GB/s on the Xeon against 18.6 GB/s for the same kernel reached through the CPython binding, roughly a 12 percent cost for leaving Python. For heavier operations the ratio changes because there is more per-call work to amortize against, which is why the project publishes the two columns separately rather than one headline number. The README states the intent plainly: `StringZilla.C` is the kernel called directly and `StringZilla.Py` is that same kernel through the binding, so the gap between columns is the interpreter boundary cost. Batch work amortizes it well. One call per short string does not.
Why did my StringZilla build get killed instead of reporting an error?
Almost certainly the out-of-memory killer, which is why `setup.py` treats memory as a first-class bound alongside core count. `CPP_MEMORY_PER_WORKER_GB` is set to 2.0 and `CUDA_MEMORY_PER_WORKER_GB` to 4.0, annotated with the observation that four concurrent `.cu` passes exhausted the 15.6 GB of a four-core arm64 runner. The worker count reads `MemAvailable` and `MemTotal` from `/proc/meminfo` and returns `None` where that file is missing, as on Windows. Setting `SZ_MAX_COMPILE_WORKERS` overrides both bounds and is the first thing to try.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ashvardanian-stringzilla)