Bernstein: policy-as-code governance for AI agent fleets
The open‑source AI Agents Governance & Orchestration framework: write the rules declaratively, Bernstein enforces them and produces the verifiable, replayable record. Free, Apache-2.0. https://bernstein.run
At a glance
- What is it?
- Bernstein is an Apache-2.0 Python framework that enforces declarative rules over parallel CLI agents and produces a replayable, tamper-evident record. It is a beta project with a single maintainer, and the README says so.
- Who is it for?
- Adopt Bernstein if you already run several CLI coding agents and need a verifiable record of what each one did, and if you can pin the version and tolerate a beta interface. Skip it if you want a hosted control plane, a stable API contract, or a project with more than one maintainer.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Bernstein targets: parallel agents with no record
Running one coding agent is a session. Running eight of them against the same repository is an operations problem. Who decided which agent got which file? Which agent's output was accepted, and on whose authority? If a run produced a bad diff three days ago, can you reconstruct the task graph that led there?
Bernstein's answer is to move those questions out of the agent and into a layer that sits above it. The README describes it as "the open-source governance layer for AI agents", with policy as code: you write who may do what, what needs approval, and what must be recorded, and the framework enforces it. The audience is not a solo developer running one assistant. It is a platform or security engineer who has to answer for a fleet, and who needs the answer to survive an audit rather than a code review.
The scope claim is broader than code. The README states the same layer governs any agent workload, where the deliverable can be a diff, a research report, a dataset, or an audit evidence pack. That framing matters, because it means the unit of governance is a task with a declared artifact contract, not a git commit.
No model in the coordination loop, and what that buys
The design decision Bernstein leans on hardest is the absence of an LLM from scheduling. The README says scheduling is plain Python, so a run is reproducible end to end, and replaying yesterday's plan yields yesterday's task graph. This is the part that separates it from orchestrators that ask a model to decompose work and route it. If a model decides the task graph, the graph is a sample from a distribution, and two runs of the same input can differ. Bernstein pushes that variability into the agents, where it is expected, and keeps the coordination deterministic.
The record has three layers, per the README. A replay journal records every run. An always-on lineage spine records every lineage-bearing step. An opt-in HMAC-chained audit log, switched on with BERNSTEIN_AUDIT=1, adds receipts that a reviewer verifies offline. The failure mode is stated plainly: non-determinism surfaces as a hash mismatch at the exact step rather than a flaky re-run.
Two consequences follow. First, verification does not require re-executing the workload, which is the point of an offline check. Second, the audit chain is opt-in, so a deployment that never sets BERNSTEIN_AUDIT=1 does not get the receipts, only the journal and lineage. That is a deliberate cost trade-off, and it is also a footgun for anyone who assumes the strongest record is on by default.
Parallelism is handled with git worktrees, which the topic list and the Dockerfile comment both reference: the runtime installs git because it is "required for git_ops", and the container notes describe creating a worktree under /workspace/.sdd. Agents working in separate worktrees avoid the file-level collisions that make naive parallel edits unusable.
Installing Bernstein and running a first governed task
The package is on PyPI as bernstein and requires Python 3.12 or newer, per pyproject.toml. The README links to an install page and a first-run page rather than reproducing the steps, so the commands below are the ones the repository itself exposes.
The simplest path is pip, matching the project name in pyproject.toml:
pip install bernsteinFor a cluster deployment, docker-compose.yaml defines the roles and the ports. Copy the environment template first, because the compose file reads every variable from it:
cp .env.example .env
docker compose up -d
docker compose up --scale bernstein-worker=4The .env.example file marks at least one provider key as required, since without one the orchestrator and workers cannot reach a model and every task fails with an auth error. It also flags BERNSTEIN_AUTH_TOKEN, whose default value lets the cluster start but is not safe for a network-reachable deployment. The compose file lists the dashboard endpoints: the Bernstein web UI on port 8052 at /dashboard, Prometheus on 9090, and Grafana on 3000 with admin/admin as the documented default credentials.
A single-node run uses the repository's own example configuration. The examples directory contains simple.yaml and full.yaml, and the repository root holds bernstein.yaml. The README does not spell out a run command for those files, so check the first-run page before executing them.
What you should see is a task graph executed against your agents with a journal written under the state directory, which the Dockerfile identifies as /workspace/.sdd. The README does not document a rollback procedure for a run, so treat the journal as the recovery surface rather than expecting an undo command.
Where Bernstein is the wrong tool
The README is unusually direct about status: beta, solo-maintained, and "the version number counts releases, not maturity". It warns that minor versions may change interfaces and tells you to pin the version for anything you depend on. That is the single most important constraint on adoption. A framework whose own documentation says interfaces move between minor releases is not a foundation to build a product on without a pin and a test suite.
The audit chain being opt-in is a second limitation with real consequences. If your compliance story assumes every run is HMAC-chained, you have to verify that BERNSTEIN_AUDIT=1 is set in the environment that actually executes tasks, not just in your documentation.
Third, the coordination layer is deterministic but the work is not. Bernstein can tell you that a step diverged and where, by hash mismatch. It cannot make a model produce the same diff twice. Anyone hoping deterministic replay means reproducible agent output has misread the guarantee: replay reproduces the plan and the task graph, and the README is explicit that this is what it claims.
Finally, this is heavier than the problem many teams have. If you run one agent interactively and review its diffs yourself, the policy layer, the worktree isolation, the cluster roles and the Postgres and Redis dependencies in docker-compose.yaml are overhead with no matching risk. The project also does not present itself as a hosted service, so teams wanting a managed control plane will not find one here.
Bernstein compared with a general workflow orchestrator
The closest familiar alternative is a general-purpose workflow engine such as Airflow or Temporal, and the difference is where determinism lives. Those tools give you durable execution and retries around code you wrote; the agent, if there is one, is an opaque step inside the graph. Bernstein inverts the arrangement. The agents are the opaque part and the orchestrator is the part held to reproducibility, with the lineage spine and the audit chain attached to the agent steps specifically.
A second comparison is to the agent frameworks themselves. Projects in that space typically own the loop: they call the model, manage tools, and decide when the task is done. Bernstein explicitly does not sit in that position. It runs existing CLI agents, and the README names Claude Code, Codex and Gemini CLI among 40+ supported, with an MCP server identity declared in the Dockerfile. The governance layer wraps agents you already chose rather than replacing them.
That is the trade-off to weigh. You keep whatever agent you like and gain a record. You also inherit an extra process between you and the agent, a state directory, and a policy file to maintain. If your agents are already inside a platform that logs their actions, Bernstein may be a second record of the same events.
Maintenance, licensing and what an upgrade costs
The repository is not archived, and the last push was on 2026-09-10, the same day as the v3.19.2 release. Releases have been frequent: v3.19.0 on 2026-08-31, v3.19.1 on 2026-09-03, and v3.19.2 a week later. The README attributes the cadence to a solo maintainer and states that regressions get fixed fast and should be filed as issues. Frequent point releases from one person are a maintenance signal in both directions: responsive, and concentrated.
The licence is Apache-2.0, declared in both pyproject.toml and the README. That permits commercial use and modification, and it includes an explicit patent grant. Two project-specific notes matter more than the licence text. The repository ships a NOTICE file and a TRADEMARKS.md with a name policy, so redistributing under the Bernstein name has conditions separate from the code licence. And the DISCLAIMER.md file exists at the root; read it before treating any output as a compliance artifact. Nothing here is legal advice, and the trademark and notice terms are worth a lawyer's glance if you redistribute.
Upgrade cost is the practical concern. Because the README warns that minor versions may change interfaces, an upgrade should be treated as a change to your policy files and integration points, not a patch. Pin the version, read CHANGELOG.md between pins, and keep your bernstein.yaml under review alongside the dependency.
Frequently asked questions
The questions below cover the points that come up first when evaluating Bernstein: what it is, what it does not do, and how the record is verified.
Editorial conclusion
Adopt Bernstein if you already run several CLI coding agents and need a verifiable record of what each one did, and if you can pin the version and tolerate a beta interface. Skip it if you want a hosted control plane, a stable API contract, or a project with more than one maintainer. Before committing, read docs/reference/KNOWN_LIMITATIONS.md, check that your agent CLI appears in the supported list, and decide whether you need BERNSTEIN_AUDIT=1, since the audit chain is opt-in.
Frequently asked questions
What is Bernstein in this context?
It is an open-source governance and orchestration layer for AI agents, distributed as the Python package bernstein under Apache-2.0. It enforces declarative policy over agent runs and produces a replayable record.
Does Bernstein put a model in the coordination loop?
No. The README states that scheduling is plain Python and that there is no LLM in the coordination loop, so replaying a plan reproduces its task graph. Variability is pushed into the agents rather than the scheduler.
How do you install Bernstein?
The package is published on PyPI as bernstein and requires Python 3.12 or newer according to pyproject.toml. Cluster deployment uses the docker-compose.yaml file in the repository with an .env copied from .env.example.
Is the audit log enabled by default in Bernstein?
No. The README describes the HMAC-chained audit log as opt-in and switched on with the BERNSTEIN_AUDIT=1 environment variable. Without it you get the replay journal and the always-on lineage spine, but not the offline-verifiable receipts.
Is Bernstein production-ready?
The README labels the project beta, solo-maintained, and says the version number counts releases rather than maturity, with minor versions possibly changing interfaces. It advises pinning the version for anything you depend on.
Which agents does Bernstein support?
The README says CLI coding agents work out of the box, naming Claude Code, Codex and Gemini CLI among 40+ supported agents. The same layer is described as governing any agent workload, not only code.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sipyourdrink-ltd-bernstein)