RapidFuzz: version 3.0 changed your scores, and a Windows runtime is the usual install failure
Rapid fuzzy string matching in Python using various string metrics
At a glance
- What is it?
- RapidFuzz is a permissively licensed fuzzy string matching library for Python and C++ that positions itself as a drop-in replacement for fuzzywuzzy, adding string metrics the original lacks and a compiled core for speed. Two things dominate the practical experience of adopting it: from version 3.0 the library stopped preprocessing your strings, and on Windows a missing runtime library is the documented reason the import fails.
- Who is it for?
- Use this if you need fuzzy matching with several measures available and a licence your project can live with, since the permissiveness is often the reason people migrate. Three things to know before you upgrade.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Version 3.0 stopped preprocessing your strings
The single most consequential line in the usage section is a note about a behaviour change. From version 3.0.0 onwards, strings are not preprocessed by default: non-alphanumeric characters are not removed, whitespace is not trimmed and case is not folded. The consequence is spelled out, that two strings containing the same characters in different cases can now produce different similarity scores, so someone upgrading will see numbers they did not see before and will assume something broke. The examples make the size of the change concrete. Two strings that differ only in case score in the low twenties with the default behaviour and a perfect score when a processor is supplied. So the fix is not a workaround but an explicit choice, and the README points at the weighted scorer's own section for examples of matching with preprocessing enabled. Everything else in the library is unchanged by this, which is why it belongs at the top of any upgrade note.
A missing runtime library is the number one install failure on Windows
The install section carries a dedicated callout for one error message, a dynamic link library load failure at import time, and gives the cause and the fix in the same paragraph. The reason is most likely that a Visual C++ redistributable is not installed, because the compiled extension cannot find the C++ libraries without it. The note adds a useful detail about that runtime: the 2019 release includes the earlier 2015 and 2017 releases as well, so installing the newest one satisfies all three. That is a rough edge of shipping a C++ extension into Python, and documenting it in the install guide rather than in an issue tracker is the difference between a five-minute fix and an afternoon. The same requirement is listed again in the requirements section at the top, which is mild duplication that costs nothing and saves a scroll.
Two tags for one version, seventeen minutes apart
The release history shows something a packaging pipeline does. There is a tag for version 3.14.6 and, seventeen minutes later, another tag for the same version with a suffix, both titled with the same release number. Then a gap back to the previous version from April. So the pattern is a quarterly-ish release with a packaging retry shortly after, which is the signature of a wheel or a source distribution that failed late and was rebuilt under a second tag. It is harmless, but it does mean that when someone pins a version they may find two tags claiming it and have to work out which one has the artefact they need. The version number in the package metadata is dynamic rather than written down, so it comes from the tag itself, which makes the tag choice matter rather than being cosmetic.
Drop-in for fuzzywuzzy, with the differences in their own file
The project positions itself as a successor rather than as a new idea, and it does so by naming its predecessor. The string similarity calculations come from fuzzywuzzy, and five things are claimed as differences. It is permissively licensed, so it can be dropped into a project under whatever terms that project already uses, which is often the decisive point for a library derived from one with a copyleft licence. It implements a range of string metrics that its predecessor does not, naming the Hamming and Jaro-Winkler families. It is mostly written in C++ and adds algorithmic work on top, with the claim that the results are unchanged and the benchmarks live in the documentation. It fixes several bugs in the partial ratio implementation. And it can largely be used as a drop-in replacement, with the caveat in the fifth point: the remaining API differences are documented in a file at the repository root rather than in the README, which is the right place for them and the wrong place to look first.
A Cython build system, submodules for externals and a recursive clone
The build configuration explains why a Python package has a Windows runtime requirement. The build backend is the scikit-build-core tool, and the build requirements are that tool plus Cython with a bounded version range, which means the Python layer of this package is generated rather than hand written, with the compiled core underneath. That is consistent with the claim about being mostly written in C++, and it is why a C++17-capable compiler is needed to install from a source checkout. The documented clone is recursive, and the repository's own files explain why:
git clone --recursive https://github.com/rapidfuzz/rapidfuzz.git
cd rapidfuzz
pip install .The repository carries a submodule configuration file and an externals directory, which is what the recursive flag is for: a plain clone leaves the vendored third-party sources empty and the build fails in a way that looks like a compiler problem. That is the most likely first-build failure for anyone installing from source, and the one line in the instructions prevents it.
An entry point from the PyInstaller four series
One line in the package metadata is quietly out of step with the rest of the file. Alongside the project URLs and the optional dependency group there is a declared entry point under a versioned group for a packaging tool, pointing at a test helper inside the package. The group name is the specification version for PyInstaller 4, from an era when the tool's plugin system was versioned that way. Everything else in the same file is current: the build backend is a modern one, the Cython bound is narrow, and the interpreter classifiers run from 3.11 up through 3.15, which is a longer forward-looking range than most projects commit to. So one compatibility shim has outlived the generation of packaging tools that needed it, in a file that otherwise reads as if it were written this year. It costs nothing at install time, since entry point groups nobody claims are ignored, but it is the kind of detail that tells you how a file has been edited over time.
The process module exists to beat calling scorers from Python
The final documented feature is a separate module, and the reason it exists is performance rather than ergonomics. Comparing one string against a list of strings through the scorers directly from Python means paying the interpreter's cost per candidate, so the module moves that loop into compiled code and is described as generally more performant. Its two entry points mirror what you would write by hand: one takes a query, a list of choices, a scorer and a limit, and returns matching tuples of the choice, the score and the index in the original list; the other returns a single best match with the same shape. The index in the tuple is the part worth noting, because it is what lets you map a result back to a record without searching the list again. The scorers themselves live in two modules by kind, fuzzy scorers in one and distance metrics in the other, so a caller choosing between them is choosing a measure rather than a function.
Editorial conclusion
Use this if you need fuzzy matching with several measures available and a licence your project can live with, since the permissiveness is often the reason people migrate. Three things to know before you upgrade. Scores will move, because preprocessing is off by default and case now matters, so pin the version and re-baseline any threshold you were tuning against older output. On Windows, install the runtime library first and the import failure goes away. And if you install from a clone, clone recursively, because the external sources are submodules and the missing-directory build error looks nothing like its cause. The metric you pick matters as much as the performance: token set scoring returns a perfect score for subset relationships, which is rarely what you want in a matcher.
Frequently asked questions
How do I install RapidFuzz?
With pip or conda, which are the recommended routes. Pre-built wheels cover macOS from 10.9 onward, Linux on x86_64 and Windows. Wheels for the Raspberry Pi architectures are hosted on a third-party wheel index rather than on the main package index. Installing from a clone needs a recursive checkout and a C++17-capable compiler.
Why do my RapidFuzz scores differ from earlier versions?
Since version 3.0.0 strings are not preprocessed by default: non-alphanumeric characters are no longer stripped, whitespace is no longer trimmed and case is no longer folded. Two strings with the same characters in different cases can therefore score differently. Passing a processor to a scorer restores the old behaviour for that call.
What does the import error on Windows mean?
The README attributes a dynamic link library load failure most likely to a missing Visual C++ 2019 redistributable, which the compiled extension needs in order to find the C++ libraries. It notes that that runtime release includes the 2015 and 2017 versions as well, so installing the newest satisfies all three.
Is RapidFuzz a drop-in replacement for fuzzywuzzy?
Largely, and the project says so while listing five differences: it is permissively licensed, it implements string metrics that fuzzywuzzy does not, it is mostly written in C++ with algorithmic changes that keep the same results, it fixes bugs in the partial ratio implementation, and the remaining API differences are described in a separate document in the repository.
What do the token set and weighted scorers do?
Token set scoring returns a perfect score when one string is a subset of the other regardless of extra content in the longer one, and reduces the score only where the two strings explicitly disagree. The weighted and quick scorers accept a processor, which is how you get pre-3.0 behaviour such as stripping punctuation or folding case for a specific comparison.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/rapidfuzz-rapidfuzz)