Hypothesis: property-based testing for Python that shrinks failures to the smallest case
The property-based testing library for Python
At a glance
- What is it?
- Hypothesis is a Python library that generates inputs for your tests instead of you writing them by hand, then reduces any failure to a minimal example. It is aimed at Python developers who already have a test suite and want coverage of edge cases they did not enumerate.
- Who is it for?
- Adopt Hypothesis if you write Python tests and can state a property that should hold for a whole range of inputs, such as sorted output matching a reference implementation. Do not adopt it if your test needs a fixed, hand-written fixture or if you cannot describe the input space as a strategy, because the library will generate values you have to reason about.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Hypothesis solves for Python test suites
Hand-written unit tests check the inputs you thought of. The README makes this the central claim: with Hypothesis you "write tests which should pass for all inputs in whatever range you describe, and let Hypothesis randomly choose which of those inputs to check - including edge cases you might not have thought about." That is the whole pitch, and it is a narrow one. Hypothesis is not a test runner. It plugs into pytest and unittest style test functions and supplies arguments to them.
The audience is Python developers who can express a rule about their code rather than a single example. Sorting is the README's own illustration: whatever list goes in, the sorted result should match a reference sort. That is a property. A test that asserts a specific function returns 42 for input 7 is not a property, and Hypothesis adds nothing to it beyond overhead.
Because the library chooses inputs, the value it provides scales with how well you describe the input space. A loose strategy produces mostly uninteresting values. A tight one produces adversarial values near boundaries. The library's usefulness is therefore a function of the strategies you write, not of the library alone.
How generation and shrinking work in practice
A test is decorated with @given, which takes one or more strategies. Strategies are the description of the input space: st.lists(st.integers()) means lists of integers, with no length bound stated in the README example. At run time Hypothesis draws values from those strategies and calls the test function with them.
The second half of the mechanism is shrinking, and it is the part that distinguishes Hypothesis from a plain random fuzzer. The README states that when Hypothesis finds a bug, "it doesn't just report any failing test case - it reports the simplest possible one." The README's own example makes this concrete. A my_sort implementation defined as sorted(set(ls)) drops duplicates, so it disagrees with the built-in sort. Hypothesis reports:
Failing test case: test_matches_builtin(ls=[0, 0])Two equal integers. Not a thousand-element list with a duplicate buried in it. That reduction is what makes the output usable for debugging: the failing case is small enough to read and reason about.
The trade-off is that shrinking is search, and search costs time. Every failing test is re-run against progressively smaller candidates. For a fast pure function this is invisible. For a test that touches a database or a network service, the cost compounds, and the README does not discuss that case.
Installing Hypothesis and writing a first property test
Installation is a single pip command. The README gives it directly, and notes that optional extras exist and are documented separately. If you need the extras, check that page rather than guessing at extra names.
pip install hypothesisAfter that, the smallest useful test imports given and the strategies module and decorates a function. This is the README's example verbatim, with a deliberately wrong implementation so you can see the failure output.
from hypothesis import given, strategies as st
def my_sort(ls):
return sorted(set(ls))
@given(st.lists(st.integers()))
def test_matches_builtin(ls):
assert sorted(ls) == my_sort(ls)Run it with pytest as you would any other test. It should fail, and the reported failing case should be a list of two equal integers, matching the README's output. If you see a larger list, the shrinker is still working or your strategy is producing values it cannot reduce further.
The first real use beyond a toy is replacing an existing example-based test. Take a test that asserts one input produces one output, identify the rule behind it, and rewrite the assertion to compare against a reference implementation over a strategy. That is exactly the shape of the sorting example.
Where Hypothesis is the wrong tool
Hypothesis needs a property. If your code's correct behaviour is defined by a specification table, a golden file, or a set of approved fixtures, there is no property to state, and generating inputs only tells you that your generated inputs are not in the table. That is noise, not signal.
The second limitation is that generated failures are reproducible only through the library's own machinery. The README does not document rollback, persistence of the failing database, or how a shrunk example is replayed in a later run. If your workflow requires a failing case to be pasted into a bug tracker as a literal input that anyone can re-run without Hypothesis installed, the shrunk example is a starting point, not the artifact you need.
The third is cost on slow tests. Shrinking re-runs the test body repeatedly. A property test wrapped around an integration path with real I/O multiplies that cost. The library does not know which parts of your test are expensive; you do.
Finally, Hypothesis is for Python. The repository's primary language is Python and the README describes it as the property-based testing library for Python. If your system is polyglot, the properties you can express here cover only the Python portion.
Hypothesis compared with example-based testing and plain fuzzing
The nearest alternative is the example-based test you already write: pick inputs, assert outputs, commit. The difference in approach is who chooses the inputs. With example-based tests you choose, which means your tests encode your assumptions about what matters. With Hypothesis the library chooses from the range you describe, which means the tests can contradict your assumptions. That is the entire value, and it is also the reason the results can be surprising in ways that take time to triage.
The second alternative is a hand-rolled fuzzer: a loop that calls random.randrange or random.choice and asserts something. The difference is shrinking. A random loop reports whatever input happened to fail, which for a list of a thousand integers is close to useless. Hypothesis reduces that input before reporting it. If you have written a random-input loop and found yourself unable to reproduce or minimise the failure, that gap is what the library fills.
Neither comparison is a reason to delete your example-based tests. A property test and a specific regression test for a known bug serve different purposes, and the README's framing is additive: Hypothesis catches things you would not have found.
Maintenance cadence, licence and upgrade cost
The repository is not archived, and the last push was on 2026-09-20, one day before this article's reference point. Releases are frequent: v6.168.0 on 2026-09-08, v6.167.1 on 2026-08-30, and v6.167.0 on 2026-08-30. A minor version every week or two is the pattern the release list shows, which means pinning a version and upgrading deliberately is the realistic approach rather than tracking the tip.
The upgrade cost is the usual one for a test-time dependency: a new version can change generation behaviour, and a test that passed by luck on one version may fail on the next. That is not a defect, it is the point of the library, but it means a version bump can surface a real bug at an inconvenient moment. Budget for that.
The licence is listed as NOASSERTION, which means the repository's licence metadata could not be classified automatically. The repository contains a LICENSE.txt file at the top level, so the terms are available there. Read that file before distributing anything that depends on Hypothesis; this article does not interpret it and is not legal advice.
Editorial conclusion
Adopt Hypothesis if you write Python tests and can state a property that should hold for a whole range of inputs, such as sorted output matching a reference implementation. Do not adopt it if your test needs a fixed, hand-written fixture or if you cannot describe the input space as a strategy, because the library will generate values you have to reason about. Before committing, verify which optional extras you need against the extras page of the documentation, and confirm the current version on PyPI, since the repository publishes releases on a frequent cadence.
Frequently asked questions
How do I install Hypothesis for Python?
The README gives a single command: pip install hypothesis. The README also notes that optional extras are available and links to a separate extras page in the documentation.
How do I use Hypothesis in a Python test?
Import given and strategies from hypothesis, decorate a test function with @given passing one or more strategies, and take the generated values as arguments. The README's example uses @given(st.lists(st.integers())) on a function that asserts sorted output matches a reference sort.
What does Hypothesis report when a property test fails?
According to the README, it reports the simplest possible failing test case rather than an arbitrary one. In the README's sorting example the reported case is ls=[0, 0].
Does Hypothesis work with pytest?
The README does not name a test runner. It shows a plain function decorated with @given, which is the form pytest and unittest-style runners collect, but the README itself does not state a runner requirement.
What licence does Hypothesis use?
The repository's licence is reported as NOASSERTION, meaning it could not be classified automatically, and a LICENSE.txt file is present at the top level of the repository. The README does not restate the licence terms.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hypothesisworks-hypothesis)