# Schemathesis builds its own CPython and disables the GIL

> Schemathesis generates test inputs from an OpenAPI or GraphQL schema, adapts them to what the server returns, and chains operations into workflows. Its repository also compiles an interpreter from source for the container image, which is where the unusual engineering decisions sit.

**schemathesis/schemathesis** — Catch API bugs before your users do

- Repository: https://github.com/schemathesis/schemathesis
- Website: https://schemathesis.readthedocs.io
- Stars: 3,649 · Forks: 225
- Language: Python
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/schemathesis-schemathesis

## pytest is a runtime dependency, not a test extra

Installing the command line tool does not spare you pytest. pyproject.toml puts pytest>=8.4,<10 in the runtime dependency list, alongside click, requests, werkzeug, rich, PyYAML, anyio, harfile, pyrate-limiter, jsonschema_rs and hypothesis_graphql, with hypothesis>=6.130.12,<7 among them. The documented way to use the project is either uv tool install schemathesis followed by schemathesis run, or uvx schemathesis run with no install step at all, and neither of those is a test suite. So a user who only intends to point the CLI at a schema URL still pulls in a test runner and its plugin machinery. The heavier frameworks are kept out of that list on purpose: coverage, Django, fastapi, Flask, httpx and httpx2 sit in the optional tests extra, which has to be requested deliberately.

## requires-python says 3.10 while the classifiers reach 3.14

The supported interpreter range is written twice in pyproject.toml. requires-python is >=3.10, and the classifiers list 3.10, 3.11, 3.12, 3.13 and 3.14, together with CPython and PyPy implementations. One dependency is conditional inside the same file: tomli>=2.2.1 carries the marker python_version < '3.11', which means the standard library tomllib only covers 3.11 and upward and anyone on 3.10 gets the backport instead. The same file marks the project as Development Status :: 5 - Production/Stable and includes src/schemathesis/py.typed in the package, so the distribution is typed and declared stable while it still carries a version-conditional shim for its oldest supported interpreter. Anything older than 3.10 is outside the declared range with no fallback path written into the metadata.

## The Dockerfile compiles CPython 3.14.3 with --disable-gil

The container image does not take a base image's interpreter. The Dockerfile starts from alpine:3.21, installs a compiler toolchain with apk, downloads the interpreter source and builds it:

```
ARG PYTHON_VERSION=3.14.3
RUN wget https://www.python.org/ftp/python/${PYTHON_VERSION}/Python-${PYTHON_VERSION}.tar.xz && \
    tar -xf Python-${PYTHON_VERSION}.tar.xz

# PGO (--enable-optimizations) disabled: hangs on Alpine/musl in test_datetime
# Still optimized with: -O2, LTO, stripped binaries
# `--disable-gil`: `--workers` are threads, so the GIL caps them. Measured 2.85x at 8 workers;
# with the GIL on, adding workers only makes runs slower.
```

The configure step that follows enables LTO, builds a shared library and strips the result, and a later layer deletes the test packages, tkinter, turtledemo and idlelib from the installed prefix. The three comments carry the reasoning a reader would otherwise have to reconstruct. Profile-guided optimization is switched off because it hangs on Alpine with musl inside test_datetime. The GIL is switched off deliberately, and the comment gives a figure for the reason: the workers are threads, the GIL caps them, and the measured speedup is 2.85x at eight workers. A second build file, Dockerfile.trixie, sits next to the first, so the image is defined twice.

## rate-limit = auto follows Retry-After instead of stopping

A fuzzer pointed at an HTTP API needs a throttle, and the throttle can be handed to the server. rate-limit = "auto" follows Retry-After on a 429, and that exact setting is what the documented configuration sample uses. The alternative is a fixed cap on the request rate. The behaviour worth understanding is what auto does not do: it does not abandon the run when a server asks it to slow down, it slows down. The dependency implementing the cap is pyrate-limiter>=4.0,<5.0. For a service you own on a shared cluster, and more so for a shared or third-party endpoint, the value in that file decides whether a run finishes usefully or turns into a self-inflicted incident, and the safe default is a number the operator picked rather than a header the operator's own service sent.

## Adaptive and stateful modes write real records

Two capabilities change what a run does to the target rather than only what it observes. Adaptive testing learns constraints, ids, and auth from responses and reuses them mid-run, so generation is not fixed when the run starts: a value the server rejected shapes what is sent next, and a value the server accepted can return as a live identifier. Stateful testing infers operation links from the schema, with no manual wiring, and the Python sample converts that into a state machine that builds a test class for pytest or unittest. Chaining the inferred links is what exercises a sequence such as create user, then get user, then delete user. The consequence for anyone running this against a real environment is that records are created, read and deleted, and the sample configuration contains nothing that narrows a run to a sandbox.

## Baseline mode lets known failures keep passing CI

The reporting formats are numerous: JUnit, VCR, HAR, NDJSON and JSON, plus Allure. VCR and HAR are the two that replay outside the tool, so a failure can be handed to someone else and inspected in another program. Coverage is reported at keyword level, which is a narrower claim than line coverage: the report shows which schema constraints the generated cases exercised, so an operation that documents three constraints is not counted as covered because one of them was hit. Baseline mode changes the verdict rather than the report. Past failures can be replayed, and a baseline lets a pipeline fail only on new ones, which is the right setting for an existing service with a backlog and the wrong one while you are trying to see whether anything is broken. A green build under baseline means no new breakage, not no breakage.

## The config file interpolates ${API_TOKEN} instead of storing it

One of the stated reasons for schemathesis.toml is that auth, phases and per-operation overrides can be set without writing Python, and the sample in the README is three lines long:

```toml
headers = { Authorization = "Bearer ${API_TOKEN}" }
generation.max-examples = 500
rate-limit = "auto"
```

The Authorization value is an interpolation, so the token is read from the environment and the file itself carries no credential. generation.max-examples = 500 is the ceiling on generated cases for an operation, which is the number to lower first when a target is fragile. Python users reach the same ground through the library instead:

```python
import schemathesis

schema = schemathesis.openapi.from_url("https://your-api.com/openapi.json")

@schema.parametrize()
def test_api(case):
    # Tests with random data, edge cases, and invalid inputs
    case.call_and_validate()

# Stateful testing: Tests workflows like: create user -> get user -> delete user
APIWorkflow = schema.as_state_machine()
# Creates a test class for pytest/unittest
TestAPI = APIWorkflow.TestCase
```

For CI there is a third door, a GitHub Action that takes the schema URL as an input:

```yaml
- uses: schemathesis/action@v3
  with:
    schema: "https://your-api.com/openapi.json"
```

## Live benchmarks run against real-world APIs without naming them

One section of the README advertises testing other people's systems and does not say whose. See it in action points at Live Benchmarks, described as continuous testing results from real-world APIs, and names three things the page displays: code and API schema coverage achieved, issues found with detailed categorization, and performance across different fuzzing strategies. No file in the repository states which APIs are involved, who operates them, or on what basis they are tested. The Try it now block makes the same move at a smaller scale, pointing at a hosted demo schema at example.schemathesis.io with a comment claiming it finds real bugs in 30 seconds. The who uses it section is the part with checkable names: Spotify, WordPress, JetBrains and Red Hat, with the first two links pointing at the backstage and openverse repositories rather than at company pages.

## Conclusion

Schemathesis earns a place in a pipeline that has a schema and an endpoint it is allowed to hammer, and its adaptive and stateful modes are the reason it finds things hand-written cases miss. Verify four things first: that the throttle in schemathesis.toml suits the target, because rate-limit = "auto" follows Retry-After instead of stopping; whether a pytest install is acceptable for a CLI-only user, since pytest is a runtime dependency; whether baseline mode is hiding known failures from your build; and whether the environment can absorb records that a stateful run creates and deletes.

## FAQ

### How do you install Schemathesis?

Two paths are documented. Install the tool with `uv tool install schemathesis` and then run `schemathesis run https://your-api.com/openapi.json`, or skip the install and use `uvx schemathesis run` against the same schema URL.

### How do you use Schemathesis against an API?

Point it at a schema URL. From the command line that is `uvx schemathesis run https://your-api.com/openapi.json`, in pytest it is `schemathesis.openapi.from_url` with `@schema.parametrize()` and `case.call_and_validate()`, and in a pipeline it is the `schemathesis/action@v3` GitHub Action with a schema input.

### Which API formats does Schemathesis support?

OpenAPI and GraphQL. The project description in pyproject.toml is Adaptive API testing for OpenAPI and GraphQL, GraphQL support arrives through the hypothesis_graphql dependency, and the CLI takes a schema URL such as https://your-api.com/openapi.json.

### What reports does a Schemathesis run produce?

JUnit, VCR, HAR, NDJSON and JSON, plus Allure. There is also a keyword-level schema coverage report showing which constraints the generated tests exercised, and a baseline mode that lets a pipeline fail only on new failures.

### How is Schemathesis related to Hypothesis?

Schemathesis is built on top of Hypothesis, described as a property-based testing library for Python, and pyproject.toml pins hypothesis>=6.130.12,<7 among its runtime dependencies alongside hypothesis_graphql for GraphQL. The README does not compare the two tools.

### How often does Schemathesis release new versions?

Three releases landed in four days: v4.29.0 on 2026-10-01, v4.29.1 on 2026-10-03 and v4.29.2 on 2026-10-04. The version field in pyproject.toml reads 4.29.2, matching the newest tag.

## Sources

- [License: MIT](https://github.com/schemathesis/schemathesis/blob/master/LICENSE)
- [Project website](https://schemathesis.readthedocs.io)
- [README](https://github.com/schemathesis/schemathesis/blob/master/README.md)
- [Releases](https://github.com/schemathesis/schemathesis/releases)
- [schemathesis/schemathesis on GitHub](https://github.com/schemathesis/schemathesis)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/schemathesis-schemathesis
