CLI tool
Cyan4973/xxHash avatar
Cyan4973/xxHash

xxHash: the non-cryptographic hash you build into a table or a file format

Extremely fast non-cryptographic hash algorithm

11,279 stars916 forksCNOASSERTION

At a glance

What is it?
xxHash ships XXH32, XXH64 and the vectorized XXH3 family behind one C header, with a make target that also builds the xxhsum CLI. It is a checksum and hash-table primitive, not a security tool, and the README says so plainly.
Who is it for?
Adopt xxHash when you need a fast, portable, deterministic checksum or hash-table function and you have already decided that adversarial resistance is not part of the threat model. Do not adopt it for signatures, password storage or any protocol where an attacker chooses the input.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 10 days ago.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What xxHash is for, and who should reach for it

xxHash solves a narrow problem well: producing a fixed-width fingerprint of a byte buffer fast enough that hashing stops being the bottleneck. The README describes it as an extremely fast non-cryptographic hash algorithm, working at RAM speed limits, and lists three families. XXH32 uses 32-bit arithmetic, XXH64 uses 64-bit arithmetic, and XXH3, available since v0.8.0, produces 64-bit or 128-bit values using vectorized arithmetic. The 128-bit variant is called XXH128.

The intended users are people building hash tables, bloom filters, checksums and file formats. The README makes that explicit: hashing is useful in constructions like hash tables and bloom filters, where many small inputs are hashed, sometimes only a few bytes long. That is a different workload from hashing a 4 GB archive once, and it is the workload XXH3 was designed around.

The boundary is equally explicit. The README states in bold that xxHash is not a cryptographic hash function, and that it must not be used for signatures, password storage, or any other purpose that requires resistance to attacks. If your input comes from an adversary who benefits from a collision, this library is the wrong tool regardless of how fast it is.

The three algorithm families and how the code is laid out

The repository is deliberately flat at the top level. xxhash.h is the API header and the documented reference; xxhash.c is the single translation unit you can compile directly; xxh3.h and xxh_x86dispatch.c sit alongside for the XXH3 implementation and x86 dispatch. The cli/ directory holds xxhsum, tests/ holds the benchmark and collision testers, and build/cmake/ holds the CMake integration guide.

The public surface is small. The simplest call hashes a contiguous block in one shot, and the README gives this example:

c
#include <stddef.h>
#include "xxhash.h"

XXH64_hash_t hash_buffer(const void* buffer, size_t size)
{
    return XXH3_64bits(buffer, size);
}

For streams of unknown size, the API supports incremental hashing; the README points to complete single-shot and streaming examples inside xxhash.h itself rather than in the README. That is a real documentation choice worth knowing before you start: the header is the manual.

Portability is a first-class property here. The README states that xxHash produces identical hashes on all platforms, including little- and big-endian systems, and that once finalized, algorithm outputs remain stable across xxHash releases. That stability guarantee is what lets you write an XXH3 value into a file format or a database column and still verify it years later.

Installing xxHash and running a first checksum

The default make target builds both the library and the xxhsum command line utility. From a checkout of the repository, run make and then hash a file with the XXH3 algorithm selected by -H3:

bash
make
./xxhsum -H3 README.md

The README gives exactly this pair of commands as the getting-started path, so the output you should expect is a hexadecimal XXH3 digest printed next to the file name. Without -H3, xxhsum falls back to its default algorithm; the cli/xxhsum.1.md manual documents checksum generation, verification, benchmarking and the advanced command line options.

You do not have to build from source. xxHash is available from many package managers, and the README shows the vcpkg route:

bash
vcpkg install xxhash
vcpkg install "xxhash[xxhsum]"

The second form adds the xxhsum feature so the command line utility is installed alongside the library. The README points to Repology for the current package versions across distributions, which is the right place to check before assuming your distro carries a recent release.

For embedding, there are two integration styles. Link against libxxhash, or compile xxhash.c directly into your project. For a header-only integration, define XXH_INLINE_ALL before including xxhash.h:

c
#define XXH_INLINE_ALL
#include "xxhash.h"

CMake users are directed to build/cmake/README.md rather than the top-level README.

How fast, and why the README's table needs reading carefully

The README publishes a benchmark table measured on an Intel i7-9700K running Ubuntu x64 20.04, compiled with Clang v10.0 at -O3. Read as a ranking rather than as a promise for your hardware: XXH3 with AVX2 is listed at 59.4 GB/s and XXH128 at 57.9 GB/s, against memcpy from RAM at 28.0 GB/s. XXH3 with SSE2 drops to 31.5 GB/s, XXH64 to 19.4 GB/s and XXH32 to 9.7 GB/s. For comparison, the same table lists City64 at 22.0 GB/s, Murmur3 at 3.9 GB/s, SipHash at 3.0 GB/s and FNV64 at 1.2 GB/s.

The AVX2 and SSE2 rows are the interesting part. The same logical algorithm runs at roughly half the bandwidth when the wider vector path is unavailable, so a build that silently falls back to SSE2 will not match the headline number. The repository carries xxh_x86dispatch.c and xxh_x86dispatch.h for exactly this dispatch problem on x86.

The README also flags that some algorithms run faster than RAM speed, and can only reach their full potential when the input is already in CPU cache (L3 or better). Otherwise they are limited by RAM speed. That caveat matters for anyone benchmarking a cold file read and expecting 59 GB/s.

Small inputs are a separate regime. The README notes that initialization and finalization become fixed costs and branch misprediction has a much greater impact, and that XXH3 was designed for both long and small inputs. The table's small data velocity column is described by the README itself as a rough evaluation, not a precise measurement.

Quality claims, the SMHasher suite and the collision tester

Speed without distribution quality produces a bad hash table. The README addresses this directly: for non-adversarial inputs, xxHash aims to produce a uniform distribution so that any subset of the output bits can spread entries evenly in a table or index. That is a stronger claim than merely passing a test suite, because it covers truncated and masked uses of the digest.

On testing, the README states that all variants successfully complete Austin Appleby's SMHasher test suite, described as a baseline measure of statistical quality. The repository also carries its own massive collision tester under tests/collisions, which the README says is able to generate and compare billions of hashes to test the limits of 64-bit hash algorithms. Additional speed and collision tests live in tests/.

The honest framing is right there in the README: like any fixed-width hash, xxHash is still subject to collisions and the birthday paradox. A 64-bit digest has a finite collision space, and no amount of throughput changes that. If your design assumes collisions will never occur at your data volume, the design is wrong, not the hash.

Where xxHash is the wrong choice

The failure mode is not subtle. Anywhere an attacker can choose the input and benefit from a collision, xxHash is unsuitable. That covers signatures, password storage and the README's catch-all of any other purpose that requires resistance to attacks. A hash table keyed by attacker-supplied strings and protected only by a non-cryptographic hash is a denial-of-service surface; the README's own comparison table lists SipHash, a keyed hash designed for that situation, at 3.0 GB/s, roughly a fifth of XXH64's throughput. That gap is the price of the property xxHash deliberately does not provide.

There is a second, quieter failure mode: version and variant drift. XXH32, XXH64, XXH3-64 and XXH128 are four different outputs for the same bytes. If one service writes XXH64 digests and another verifies with XXH3_64bits(), every check fails, and the bug looks like corruption rather than a mismatch of algorithms. The README's stability guarantee applies to a finalized algorithm, not across algorithms, so the variant has to be pinned in your format documentation.

Finally, if you need a cryptographic-strength digest and cannot tolerate a non-cryptographic one, the README's own table puts Blake2 at 1.1 GB/s and SHA1 at 0.8 GB/s. Choosing xxHash there is not a performance decision, it is a security regression.

xxHash against crc32, MD5 and SHA256

The comparison that matters most is against crc32, because both are checksums and neither is cryptographic. CRC32 is built into almost every language runtime and needs no dependency, which is its entire advantage. xxHash gives you a wider output (XXH32 at 32 bits, XXH64 and XXH3 at 64, XXH128 at 128) and, per the README's table, much higher throughput: XXH32 at 9.7 GB/s against City32 at 9.1 GB/s, with CRC32 not listed in that table at all. If your only requirement is detecting accidental corruption and you already have crc32 available, adding a dependency buys you speed and a wider digest, not a new capability.

Against MD5 and SHA256 the difference is categorical, not incremental. Both are cryptographic hashes, and the README's table lists MD5 at 0.6 GB/s and SHA1 at 0.8 GB/s, with Blake2 at 1.1 GB/s. XXH64 at 19.4 GB/s is roughly an order of magnitude faster than the cryptographic options. The README's own annotation on MD5 and SHA1 reads cryptographic but broken, which is a reminder that a slow hash is not automatically a safe one. If you need collision resistance against an adversary, use a current cryptographic hash; if you need to detect accidental corruption or index a table, the cryptographic hashes are paying for a property you are not using.

Licence, packaging and the cost of upgrading

The repository's LICENSE file is the authority, and the metadata for this project reports the licence as NOASSERTION, meaning no standard identifier was detected automatically. The Makefile header, however, carries a GPL v2 notice covering the Makefile itself and points to the Free Software Foundation's text. Do not infer the library's licence from the Makefile's banner: read LICENSE directly before you ship, and treat this as a question for your own legal review rather than something to settle from a build file.

Upgrade cost is low by design. The README states that once finalized, algorithm outputs remain stable across xxHash releases, so upgrading the library does not invalidate digests you have already stored or written into files. The release cadence visible in the repository supports a slow-moving API: v0.8.2, v0.8.3 and v0.8.4 are the recent releases, and the last push to the repository was on 2026-09-20.

What you do pay for is the build matrix. XXH3's vectorized paths mean an AVX2 build and an SSE2 build of the same source produce the same digests at different speeds, so a container image built on one machine may not match the throughput of a binary built on another. Package availability also varies; the README defers to Repology rather than promising a version, so pinning through vcpkg or your distro is the practical route.

Editorial conclusion

Adopt xxHash when you need a fast, portable, deterministic checksum or hash-table function and you have already decided that adversarial resistance is not part of the threat model. Do not adopt it for signatures, password storage or any protocol where an attacker chooses the input. Before you commit, verify that the algorithm you pick is the one your format pins down: XXH3 is the README's recommendation for new applications, and the README states that once finalized, algorithm outputs remain stable across xxHash releases, so the choice is hard to reverse after data is written.

Frequently asked questions

What is xxHash used for?

The README describes it as a non-cryptographic hash algorithm used in constructions like hash tables and bloom filters, and as a checksum for files through the xxhsum utility. It is intended for non-adversarial inputs and for producing uniform distributions across output bits.

What are the key differences between xxHash and SHA256?

SHA256 is a cryptographic hash and the README lists SHA1 at 0.8 GB/s and MD5 at 0.6 GB/s, roughly an order of magnitude below XXH64 at 19.4 GB/s in the same table. xxHash is explicitly not a cryptographic hash function and the README says it must not be used for signatures or password storage.

What are the differences between XXH32 and XXH64?

XXH32 generates 32-bit hashes using 32-bit arithmetic, while XXH64 generates 64-bit hashes using 64-bit arithmetic. The README recommends XXH3_64bits() as the default for new applications, with XXH3_128bits() when a 128-bit hash is required.

How fast is xxHash?

The README's benchmark table, measured on an Intel i7-9700K running Ubuntu x64 20.04 compiled with Clang v10.0 at -O3, lists XXH3 with AVX2 at 59.4 GB/s and XXH64 at 19.4 GB/s. The same README notes that algorithms running faster than RAM speed only reach full potential when input is already in CPU cache.

How do I install xxHash?

The default make target builds both the library and the xxhsum command line utility from a source checkout. The README also shows vcpkg install xxhash, with the xxhsum feature adding the command line utility, and points to Repology for versions across distributions.

Is xxHash deterministic across platforms?

The README states that xxHash is highly portable and produces identical hashes on all platforms, including little- and big-endian systems. It also states that once finalized, algorithm outputs remain stable across xxHash releases.

Official sources

  1. Cyan4973/xxHash on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/cyan4973-xxhash.svg)](https://hysenlabs.com/projects/cyan4973-xxhash)