Open-source project
aappleby/smhasher avatar
aappleby/smhasher

SMHasher: a test suite for non-cryptographic hash functions

Automatically exported from code.google.com/p/smhasher

2,891 stars481 forksC++License varies

At a glance

What is it?
SMHasher measures distribution, collision and performance properties of non-cryptographic hashes, and it is also the home of MurmurHash3. The repository is small, the build story is thin, and the suite is the point.
Who is it for?
Adopt SMHasher if you are choosing between non-cryptographic hashes for a hash table, a bloom filter or a checksum and you want collision and distribution data rather than a benchmark chart from a blog post. Do not adopt it if you need a cryptographic hash, or if you expect a packaged release: there are no releases, and the README gives no build or install steps beyond pointing at the src directory.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 98 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What SMHasher is for, and who ends up using it

SMHasher is a test suite, not a hash function. The README describes it as "a test suite designed to test the distribution, collision, and performance properties of non-cryptographic hash functions", and it says the project aims to be the DieHarder of hash testing. That comparison is the clearest statement of intent in the repository: DieHarder is a randomness battery, and SMHasher is the equivalent battery pointed at hash functions instead of random number generators.

The audience is narrow and specific. If you are picking a hash for a hash table, a bloom filter, a deduplication key or a checksum, and you care whether it distributes keys evenly and whether it collides on structured input, this is the tool that answers that question. If you are picking a hash for password storage, signatures or anything with an adversary in the threat model, this is the wrong tool entirely, and the README's own framing says so by scoping itself to non-cryptographic functions.

The repository also ships MurmurHash3 in src/MurmurHash3.cpp, which is the latest version in the MurmurHash series. That dual role matters in practice: the suite is the reason the repository exists, and MurmurHash3 is the reason a lot of people arrive at it. The README states that MurmurHash3 variants produce 32-bit and 128-bit hash values on both x86 and x64.

How the suite works: a battery of properties, not a single score

The mechanism is a collection of independent tests rather than one pass or fail verdict. The README names three families of property: distribution, collision and performance. Distribution tests ask whether output bits look independent across inputs that differ in structured ways. Collision tests feed the function keysets designed to expose weaknesses, including keys that vary in a single bit or in a single byte position, and count how often two inputs land on the same output. Performance tests time the function across input sizes.

The reason this matters is that a hash can pass a throughput benchmark and still be a poor choice. A function tuned for long inputs may behave badly on the short keys a hash table actually stores, and a function with good average distribution may still collide on the specific structured inputs a test suite constructs. Running the battery gives you evidence about which of those failure modes applies.

The repository layout is minimal: a .gitignore, README.md, build.hancho and a src directory. Hash functions under test live in src alongside the harness, so adding a candidate means adding an implementation and wiring it into the test list. The README does not document that wiring procedure, which is the first place a new contributor will have to read the source.

Getting the sources and finding the test harness

The README gives no install instructions. There is no package, no release, and no documented build command. What the repository does contain is a build.hancho file at the top level, which is a build descriptor rather than a Makefile, and a src directory holding the suite. Treat the build as something you assemble yourself from those files, and read the descriptor before you guess at a target.

What the repository does document is where the code lives. The README links to the src directory as the home of the SMHasher test suite and to src/MurmurHash3.cpp as the implementation of MurmurHash3. Those two paths are the entry points: the suite is in src, and so is the hash family it was written to verify.

The first real use is to run the suite and read the per-test output for the hash you are evaluating, rather than looking for a single number. The README's claim that the suite does "a pretty good job of finding flaws with a number of popular hashes" is a statement about that per-test output. If you are adding a candidate, you add its implementation to src and register it with the harness; the README does not describe that registration step, so budget time for reading the harness source before you assume the suite is pluggable. Because the README names no compiler, standard or target, the src directory is the authority on how the pieces fit together, and build.hancho is the only build description the repository ships.

Where SMHasher stops being useful

The first limitation is categorical. SMHasher tests non-cryptographic properties. Passing the suite says nothing about preimage resistance, collision resistance against a motivated attacker, or resistance to length extension. A function that scores well here can still be trivially broken as a cryptographic hash, and the README does not blur that line, which is to its credit. If your threat model includes an adversary choosing inputs, stop and use a cryptographic hash.

The second limitation is reproducibility. The README describes performance as one of the three properties under test, and performance results depend on the machine, the compiler and the flags. The repository ships no recorded baseline results, and the README does not publish a set of reference numbers. Any performance figure you produce is a figure about your machine.

The third is packaging. There are no releases, the README gives no version, and the licence statement is split: the README says SMHasher is released under the MIT license and that all MurmurHash versions are public domain with the author disclaiming copyright. The repository metadata does not carry a licence identifier, so the README is the only statement available and it covers two different components with two different terms. Anyone embedding this in a product should read that paragraph directly rather than relying on a licence field.

Finally, the README's most recent update is dated 1/8/2016 and is about the move to GitHub. The wiki pages were copied from code.google.com and, in the README's own words, have not been reformatted to Markdown. Documentation quality is therefore uneven, and the wiki is the place to look for anything the README omits.

SMHasher against a plain benchmark harness

The obvious alternative is not another test suite but the thing most people reach for first: a small timing harness that hashes a buffer in a loop and prints throughput. The difference in approach is what gets measured. A timing harness answers one question, how fast, and it answers it on the inputs you chose. SMHasher answers several questions at once, including whether the function collides on adversarial or structured keys, and it answers them with a fixed battery so two functions can be compared on the same terms.

That is the trade-off. A hand-rolled benchmark takes minutes to write and gives you a number you already understand. SMHasher takes longer to build, produces output you have to interpret test by test, and its results are tied to your machine. What you get in return is the collision and distribution evidence that a throughput loop never produces, which is the part that predicts whether a hash will behave badly in a hash table six months after you ship.

For the MurmurHash family specifically, the repository gives you both sides: the functions in src and the suite that tests them. If you only need a fast non-cryptographic hash and you trust the published analysis, you can take MurmurHash3 and skip the suite. If you are choosing between candidates, or you have written your own, the suite is the reason to be here.

Maintenance, upgrades and what the licence actually covers

The last push to the repository was on 2026-06-25. There are no releases, so there is no version to upgrade to and no changelog to read. Upgrading means pulling master and rebuilding, and because the README documents no build system, a pull that touches the harness can change how you compile the suite. Pin a commit if you depend on results staying comparable across runs.

The maintenance cost that matters is not the pull, it is the interpretation. Test output from a battery this broad is only useful if you keep the machine and compiler consistent between runs, because the performance portion moves with both. A result recorded on one workstation does not transfer to another.

On licensing, the README states that SMHasher is released under the MIT license and that all MurmurHash versions are public domain software with the author disclaiming all copyright to their code. The repository metadata carries no licence identifier, so the README is the source of record and it describes two components under two sets of terms. If you are embedding MurmurHash3 or shipping the suite, read that paragraph and confirm which component you are using. This is not legal advice, and the terms as stated are the terms you have to work from.

Editorial conclusion

Adopt SMHasher if you are choosing between non-cryptographic hashes for a hash table, a bloom filter or a checksum and you want collision and distribution data rather than a benchmark chart from a blog post. Do not adopt it if you need a cryptographic hash, or if you expect a packaged release: there are no releases, and the README gives no build or install steps beyond pointing at the src directory. Verify first that the tree compiles on your toolchain, because the README does not document one, and read the wiki before trusting any single result, since the suite's output depends on the machine and compiler you run it on.

Frequently asked questions

Which hash algorithm is best?

SMHasher does not produce a single ranking. It reports distribution, collision and performance results per test, and the README does not publish reference numbers, so the answer depends on the properties you measure on your own machine.

How do SHA hashes work?

The README does not cover SHA or any cryptographic hash. SMHasher is scoped to non-cryptographic hash functions, so passing its battery says nothing about the properties SHA is designed to provide.

Are SHA1 and SHA256 the same?

The README says nothing about SHA1 or SHA256. The repository covers the MurmurHash family and the SMHasher test suite, both aimed at non-cryptographic hashing.

Should I use MD5 or SHA256?

Neither is discussed in the README. SMHasher tests distribution, collision and performance properties of non-cryptographic hash functions, and it does not evaluate cryptographic choices such as MD5 or SHA256.

Official sources

  1. aappleby/smhasher on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/aappleby-smhasher.svg)](https://hysenlabs.com/projects/aappleby-smhasher)