# Bernstein calls itself beta, warns that the version counts releases, and ships a changeme default

> Bernstein is a Python governance layer for AI agents that schedules work deterministically and records a replayable trail. Its own files disagree in useful places: the env template marks secrets required while compose fills them in, the adapter count is 50 in one place and 40 in another, and a VOLUME line was deleted on purpose with a warning not to restore it.

**sipyourdrink-ltd/bernstein** — The open‑source AI Agents Governance & Orchestration framework: write the rules declaratively, Bernstein enforces them and produces the verifiable, replayable record. Free, Apache-2.0. https://bernstein.run

- Repository: https://github.com/sipyourdrink-ltd/bernstein
- Website: https://bernstein.run
- Stars: 1,347 · Forks: 177
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/sipyourdrink-ltd-bernstein

## The status note says the version number counts releases, not maturity

The first line under the logo is a status block, and it is unusually specific. Bernstein is called beta and solo-maintained, with the instruction that the version number counts releases rather than maturity, that minor versions may change interfaces, and that anything you depend on should be pinned. The block ends with an invitation to file regressions rather than a promise about response times.

The numbers line up with that warning. The newest tag is v3.20.0 from 2026-09-20, the version field in pyproject.toml reads 3.20.0, and two patch tags sit behind it at v3.19.2 and v3.19.1. The branch itself was pushed on 2026-10-02, so commits on main are newer than the newest release, which is the state the pinning instruction is written for.

The project name explains the other thing a reader notices first. The README hangs a quote on the name, attributed on the page to Leonard Bernstein: to achieve great things you need a plan and not quite enough time.

## The env template marks secrets required, and compose supplies the defaults

Copy .env.example to .env and the file tells you which values are mandatory. The header states that every variable below is read by docker-compose.yaml and that lines marked REQUIRED must hold a real value, otherwise the service refuses to start. Several are marked that way, including ANTHROPIC_API_KEY, OPENAI_API_KEY and BERNSTEIN_AUTH_TOKEN, with the note that without a provider key the orchestrator and workers cannot reach any model and every task fails with an auth error.

The compose file does not enforce any of it. Each of those variables is written with a shell style default:

```yaml
    BERNSTEIN_AUTH_TOKEN: ${BERNSTEIN_AUTH_TOKEN:-changeme}
```

The provider keys follow the same pattern with empty fallbacks, and the Grafana login is admin/admin, described in the template as cosmetic and fine to leave. So the cluster starts cleanly with no .env at all, and the auth token it starts with is the literal word changeme. The template itself says that default is not safe for any network-reachable deployment.

## The compose file ships a database password and calls it a development default

The cluster definition names six roles: a task server with a web dashboard on port 8052, an orchestrator that reads the backlog and spawns workers over HTTP, workers you scale yourself, Prometheus scraping metrics on 9090, Grafana on 3000 with admin/admin, plus Postgres and Redis for state and locks. The database URL is written inline:

```yaml
    BERNSTEIN_DATABASE_URL: postgresql://bernstein:bernstein@postgres:5432/bernstein  # NOSONAR - development-only default, not used in production
```

The password is in the URL, and the trailing comment is a static analysis suppression marker rather than a runtime check, so nothing stops that stack from being started anywhere. Two smaller details travel with it: the file still opens with `version: "3.9"`, a key modern Compose ignores, and the worker count is left to you, which is why the header gives two commands.

```bash
#   docker compose up -d
#   docker compose up --scale bernstein-worker=4
```

## A VOLUME line was deleted on purpose, and the Dockerfile asks you not to restore it

The runtime stage runs as a user created in the image, `bernstein` at uid 1000, with state living under `/workspace/.sdd`. The comment block above that explains a scar: an anonymous volume declared at that path is created root-owned and 0755 by the daemon while the container runs as uid 1000, so the first run died with Permission denied while creating the worktree. The fix was to delete the declaration, and the file says in as many words not to re-add a VOLUME line.

The rest of the image is pinned tighter than most. Both stages use `python:3.13-slim` with a digest, the builder installs `hatchling==1.29.0` explicitly, and the wheel is installed as the uid 1000 user rather than switched to afterwards. Two ports are published for the task server, 8052 for HTTP and 50051 for gRPC.

Python versions spread wider than the image does. `requires-python` asks for 3.12 or newer and the classifiers name 3.12, 3.13 and 3.14, so a developer on the oldest supported interpreter is not running what the container runs.

## The adapter count is 50 in the README and 40 in the package metadata

The README says more than 50 selectable CLI agent adapters, naming Claude Code, Codex and Gemini CLI as the ones that work out of the box, plus a generic `--prompt` wrapper for anything else. The package description in pyproject.toml runs the same sentence with 40+, and the OCI label baked into the image repeats the 40+ figure while quoting the same three tools.

So the headline number for the project's main selling point is stated two ways in files that ship together. One of them is the metadata every installer reads, and the other is the page a reader lands on.

The tree also explains where the truth lives: ADOPTERS.md sits at the root next to the feature and capability references, so the count is meant to be checked rather than remembered. That same root holds instructions aimed at a dozen tools, from AGENTS.md and CLAUDE.md to .cursor/, .qwen/, .goosehints, .aider.conf.yml, .mcp.json, bernstein-skills.toml and context7.json.

## The HMAC audit chain is switched on by one environment variable

Three layers of record keeping are described, and they are not always on. The replay journal is described as recording every run, and the lineage spine as always on, recording every step that carries lineage. The third is opt-in: set `BERNSTEIN_AUDIT=1` and you get an HMAC chained audit log whose receipts a reviewer verifies offline, without rerunning anything. Non-determinism is designed to show up as a hash mismatch at the exact step rather than as a flaky repeat.

The same mechanism reaches past code. A task can declare an artifact contract, covering a report, dataset, action log or ops result, and completes on a signed lineage receipt rather than a git commit, so non-code deliverables get the same verification treatment as a diff.

The scheduler is the part that makes any of this reproducible. It is plain Python with no model in the coordination loop, so replaying yesterday's plan is supposed to reproduce yesterday's task graph, and the run is declared in one YAML file holding phases, roles, dependencies and the conditions under which a node runs at all.

## Turn worktrees off and every task runs in the shared checkout

Isolation is the default and it is per task. Each coding task gets its own git worktree placed behind merge gates, and artifact-mode tasks get a working directory under `.sdd/workspaces/` instead. Agents share no mutable workspace by default, and the one thing they do share, the task backlog, is claimed atomically so two agents cannot take the same task.

Stricter filesystem enforcement is available but not enabled by default, and it lives in the sandbox backends documentation. That is the setting to read before pointing an adapter at a real repository rather than a scratch checkout.

The switch has a blunt failure mode, stated in the README's own summary of the design: disable worktrees and every task runs in the shared checkout. With fifty plus adapters writing into one directory, that setting converts the isolation story into a coordination problem, so it is worth knowing which side of the line a given configuration sits on.

## Conclusion

Bernstein fits a team that wants an agent run to be checkable after the fact, since the design puts replay and lineage ahead of model freedom and keeps state in files. Check three things first: that you pin a version, because the project's own status note says minor releases may change interfaces; that you replace the `changeme` auth token, the in-repo Postgres password and the admin/admin Grafana login before anything listens on a network; and whether cluster mode is what you want, since the compose file turns on four provider keys and two backing services by default. Solo-maintained beta software that says so in its own README should be read as a tool to pin and file issues against, not a dependency to track on its default branch.

## FAQ

### What does the Bernstein framework do for AI agents?

It is a governance layer that runs agents from a declarative policy file and records what happened. Scheduling is plain Python with no model in the coordination loop, so replaying a plan reproduces its task graph, and run receipts can be checked offline from the artifacts.

### Why do search results about Bernstein not match this repository?

The name is borrowed from a quote attributed on the README to Leonard Bernstein, which collides with the conductor, the research firm and the mathematician. This project is the Python agent governance layer published under Apache-2.0 with its documentation at bernstein.run.

### How do you enable Bernstein's audit chain?

Set `BERNSTEIN_AUDIT=1`. The replay journal and the always-on lineage spine are described as always recording, while the HMAC chained audit log is the opt-in layer, producing receipts a reviewer verifies offline without rerunning the run.

### What does Bernstein require to run?

Python 3.12 or newer, per `requires-python`, with classifiers for 3.12, 3.13 and 3.14. The image builds on python:3.13-slim pinned by digest and installs hatchling 1.29.0 explicitly in the build stage.

### How does Bernstein deploy as a cluster?

docker-compose.yaml defines a server with a dashboard on port 8052, an orchestrator, workers you scale with `docker compose up --scale bernstein-worker=4`, plus Prometheus, Grafana, Postgres and Redis. The image exposes 8052 for HTTP and 50051 for gRPC.

### Is the Bernstein auth token safe at its default value?

No. docker-compose.yaml fills in `${BERNSTEIN_AUTH_TOKEN:-changeme}`, and .env.example says the changeme default lets the cluster start but is not safe for any network-reachable deployment. Grafana ships with admin/admin in the same file.

## Sources

- [License: Apache-2.0](https://github.com/sipyourdrink-ltd/bernstein/blob/main/LICENSE)
- [Project website](https://bernstein.run)
- [README](https://github.com/sipyourdrink-ltd/bernstein/blob/main/README.md)
- [Releases](https://github.com/sipyourdrink-ltd/bernstein/releases)
- [sipyourdrink-ltd/bernstein on GitHub](https://github.com/sipyourdrink-ltd/bernstein)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sipyourdrink-ltd-bernstein
