# datasketch 2.0 changed the hash values, so sketches persisted by 1.x stop matching

> A set of probabilistic data structures for cardinality and similarity at scale, distributed as a wheel with two decade-old dependency floors. Its current major changed the default permutation scheme for its most-used sketch, which invalidates anything already persisted, and its test extra still carries the runner that pytest replaced.

**ekzhu/datasketch** — MinHash, LSH, LSH Forest, Weighted MinHash, HyperLogLog, HyperLogLog++, LSH Ensemble and HNSW

- Repository: https://github.com/ekzhu/datasketch
- Website: https://ekzhu.github.io/datasketch
- Stars: 2,969 · Forks: 322
- Language: Python
- License: MIT
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/ekzhu-datasketch

## Version 2.0 changed the default hash scheme, so persisted sketches stop matching

A note near the top of the readme is the most consequential paragraph in the document, and it is easy to skim past.

The current major changes the default permutation scheme for the MinHash sketch to a 32-bit affine scheme. Three things are claimed for it: it fixes a bias that made similarity estimates come out too high on large sets, it halves the memory a sketch occupies, and it speeds up updates by roughly four times. The issue number for the bias is given, so the reasoning can be checked.

A wider 64-bit variant is offered for sets in the billions, where 32 bits would not leave enough distinct values.

Then comes the sentence that matters operationally. Hash values differ from earlier versions, and anyone with a sketch or an index on disk has two choices: rebuild the persisted data, or pass the legacy scheme explicitly to interoperate with what already exists.

That is a silent failure rather than a loud one. Nothing raises, nothing warns at load time, and the sketch still returns numbers. It just returns numbers computed from a different permutation, so a query against a newly built index will not find matches in an index built under the old scheme, and the symptom will look like a recall problem rather than a version problem.

The readme is upfront about it, which is more than most projects manage. The trap is for code that upgraded the package without re-running whatever produced its stored indexes.

## Four of the five indexes are MinHash-based, and one takes a custom metric

The package is described as eight things, but that count mixes sketches with indexes. Four sketches are listed separately: one for Jaccard similarity and cardinality, one for the weighted form of that similarity, and two for cardinality, the second of which is an improved variant of the first and is documented as an anchor within the same page.

The index table is where the structure shows. Five indexes are offered against those sketches, and they fall into four different query types.

Two of them answer Jaccard threshold queries, and both accept either form of the MinHash sketch. One answers Jaccard top-K queries and also accepts both. One answers containment threshold queries and accepts only the plain sketch, not the weighted one. And one accepts any sketch at all and answers top-K queries against a metric you supply.

Two things follow. Containment is supported by exactly one index, and that index is the only one that will not take the weighted sketch. Approximate nearest neighbour against a custom distance is supported by exactly one index too, and it is the only one not restricted to Jaccard at all, which makes it the escape hatch for anyone whose similarity is neither.

Only two of the five are described as supporting an external storage layer, Redis or Cassandra, and the documentation for running the threshold index at scale is a section within one index's own page rather than a separate topic. So the storage story is narrower than the index story, and the two live in different places.

## The dependency floors are more than a decade old

The requirement line says the package must be used with Python three point nine or above, an array library at one point eleven or above, and a scientific library. Those floors have not moved in a long time, and both numerical packages have had many major releases since.

A low floor is a deliberate choice with a real benefit: a library that accepts old environments can be adopted inside an organisation that cannot move everything at once. The cost is that nobody can tell from the floor which combinations are actually exercised.

The package metadata puts the same number in a second place. The declared Python requirement is an open-ended lower bound, and the classifier list then enumerates five specific versions, ending at three point thirteen. So the metadata claims support for everything from three point nine upward while advertising tested versions only up to three point thirteen.

The two runtime dependencies are declared with lower bounds only and no upper bounds, which for a numeric library is the right default: an upper bound on the array library would break users on newer releases for no benefit the library needs.

The licence is declared as inline text in the metadata rather than as a pointer to the file at the repository root, and both exist, so there is a licence file and a licence string and no mechanism tying them together.

## The test extra still carries the runner that replaced it

The optional dependency group for testing has not been modernised, and reading it is like looking at several years of the same file at once.

It includes two packages from the runner that the testing ecosystem moved away from, plus its exclusion helper. That runner does not work on recent Python versions, so on any modern interpreter those two entries are dead weight at best.

Alongside them sits the current runner, pinned twice with environment markers: an unconstrained version for older interpreters and a recent one for newer ones. The same two-tier pattern appears again for the natural language library in the benchmark group, which gets an old floor below a version boundary and a newer one above it.

Then there are the mock libraries. One mocks the cache server, and a document database driver is present for testing against the wide-column store, neither of which is a test framework. So the group mixes the runner, the coverage tooling, the mocks for two optional backends, and an asynchronous plugin.

Finally, the two optional backends appear in this group as well as in their own, so the cache driver and the wide-column driver are declared twice in the same file.

None of that stops the tests running. It does mean the group cannot be read as a description of what testing needs, and that a dependency audit on this group would flag the two obsolete entries every time.

## A benchmark dependency carries an unresolved note about installation

The benchmark group is the longest optional group in the file, and it carries two comments that explain more about the project than any documentation does.

The first is an unresolved note attached directly to one dependency: a comment reading that installation needs fixing. The dependency is an older hashing library, and the note sits on the same line as the version constraint, which is where a reader will see it when the install fails rather than when they are choosing to run benchmarks.

The second comment explains why several packages that nothing imports directly are declared anyway. It says they are transitive dependencies of the plotting library, listed to stop an automated dependency bot from producing pull requests that only touch the lockfile. That is a candid admission about a class of automated noise, and the fix chosen is to pin transitive dependencies directly so the bot has nothing to propose.

The group also contains an older lower bound on the scientific library than the one in the runtime requirements, which is harmless because a resolver takes the stricter of the two, and it declares the imaging library only for newer interpreters. So on the oldest supported interpreter the plotting stack is not installed at all.

The practical reading is that this group is for people running the comparisons rather than using the library, and it has known rough edges that are documented in comments rather than in prose.

## Three extras in the install steps, four in the manifest, two in the contributing steps

The package publishes five optional dependency groups. No single place in the documentation lists all five, and the three places that do list them each list a different subset.

The installation section, which is aimed at users, names three: the cache extra, the wide-column extra, and the bloom filter extra. Each is shown as its own install command with the group name in brackets, and the plain install is noted as also pulling in the array library.

The manifest defines those three plus the benchmark group and the testing group.

The contributing section, aimed at developers, then lists the testing group, the wide-column extra and the cache extra, and offers a flag that installs every extra at once. So bloom and benchmark appear in neither the user list nor the developer list, and the testing group appears in the developer list only.

The two lists also disagree about how to install. The user section uses the standard installer with bracketed extras. The developer section uses the newer environment tool, with a flag per group and a separate flag for all of them.

That last flag is the only route in the documentation that installs everything, and it is in the contributor instructions rather than the installation instructions. So a user who wants the bloom filter and is following the user-facing page will find it; a developer wanting the testing group and following the contributor page will find that; and anyone wanting one combination across both categories has to know that the tool's all-extras flag exists.

## One tool installs it and a different one develops it

The package is installed with the standard installer, which is what the installation section shows four times. It is developed with a newer environment tool, and the contributing section is explicit that the project uses that tool for package management.

The developer setup has five steps: install the tool, clone, set up, verify, and add optional groups. The setup step creates a virtual environment and syncs dependencies, and the comment beside the activation line notes that activating it is optional because commands run through the tool work without it. That is a smaller requirement than most projects impose, and it removes a step that would otherwise fail differently on each platform.

Verification is a single command that runs the test suite through the tool rather than through an activated environment.

Code quality is checked with a linter invoked in a way that deliberately does not touch the environment: a separate runner prefix fetches the linter and runs it against the tree, with a check mode and a format mode shown as two commands. So the linter is not a declared dependency of anything, not even the testing group, and a contributor's version of it is whatever that prefix resolves to.

The workflow that follows is conventional: fork, branch with one of two naming shapes, change, test, lint, commit with a described message, push. The only unusual element is that step four and step five are the same commands as steps four and three of the setup, repeated rather than referenced.

## Six example files, and none for the index that accepts any sketch

The examples directory holds six files, one per sketch and per index family: two for the MinHash sketch family, one each for the forest, the ensemble, the cardinality sketch, and the weighted sketch.

The gap is the graph-based index. It appears in the repository description, in the indexes table as the only entry accepting any sketch with a custom metric, and in the documentation link list, but it has no example file. So the one component with the widest input compatibility is the one you cannot learn from the examples directory.

There is also no example for the bloom filter index, for either external storage backend, or for the Redis and Cassandra paths, even though those are the features the readme calls out as what distinguishes the threshold index at scale.

The readme itself is written in reStructuredText rather than Markdown, which is why it uses directives for notes, images and code blocks and why the two tables are drawn as grid tables rather than pipe tables. The manifest points at that file by name for the package metadata, so the two are tied together and the format is a project choice rather than an accident.

Two badges sit above the prose. One is a monthly download counter from a third-party service, and one is an archival DOI, which means releases are deposited somewhere designed for citation rather than only hosted on a forge.

## Conclusion

datasketch suits someone deduplicating or counting at a scale where exact sets no longer fit in memory, and who wants one library covering cardinality, Jaccard and containment rather than several. Before adopting it, check four things. Decide which major version you are on, because the current one changed the default permutation scheme, so a sketch or index persisted by an earlier version will not match against newly built ones unless you pin the legacy scheme explicitly. Check the dependency floors against your own environment, because the required versions are old enough that a modern environment may be outside what has been exercised. Expect the only broadly maintained index to be the two Jaccard-threshold ones, since the forest, the ensemble and the graph index each take a different query type. And if you intend to run the benchmark extra, note that one of its dependencies carries an unresolved note about installation.

## FAQ

### What changed in datasketch version 2.0?

The default MinHash permutation scheme changed to a 32-bit affine scheme, which fixes a similarity over-estimation bias on large sets, halves sketch memory and speeds up updates by roughly four times. A 64-bit variant is available for billion-scale sets. Hash values differ from earlier versions, so persisted sketches and indexes must be rebuilt or the legacy scheme passed explicitly.

### Which data sketches does datasketch provide?

Four: one estimating Jaccard similarity and cardinality, one estimating weighted Jaccard similarity, and two for cardinality, the second being an improved variant of the first documented as an anchor within the same page.

### Which datasketch index supports a custom similarity metric?

Only the graph-based index. It is the one entry in the indexes table that accepts any sketch and answers top-K queries against a metric you supply. The other four are MinHash-based, and one of them additionally refuses the weighted sketch.

### Does datasketch support Redis and Cassandra storage?

Two of its indexes do: the MinHash threshold index and the ensemble index. The documentation for running the threshold index at scale is a section within that index's own page rather than a separate topic.

### What are datasketch's install extras?

Five are defined in the manifest: Redis, Cassandra, a bloom filter, a benchmark group and a testing group. The install section documents three, the contributing section names a different three, and only the environment tool's all-extras flag installs everything.

### How is datasketch developed?

With a newer environment tool rather than the standard installer used for installation. Setup creates a virtual environment and syncs dependencies, activation is optional because commands run through the tool regardless, and code quality is checked with a linter invoked outside the environment.

## Sources

- [ekzhu/datasketch on GitHub](https://github.com/ekzhu/datasketch)
- [License: MIT](https://github.com/ekzhu/datasketch/blob/master/LICENSE)
- [Project website](https://ekzhu.github.io/datasketch)
- [README](https://github.com/ekzhu/datasketch/blob/master/README.md)
- [Releases](https://github.com/ekzhu/datasketch/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ekzhu-datasketch
