Model or dataset
aaron-for-value/VeriRun avatar
aaron-for-value/VeriRun

VeriRun: six invariants, six releases, and an architecture diagram that stops mid string

Evidence-first infrastructure for reproducible, isolated executable evaluation and online rewards.

324 stars30 forksPythonApache-2.0

At a glance

What is it?
A Python runtime for executable code and agent evaluation that insists on provenance, structured failure classes and replay before claims. Reading its README against its own pyproject.toml and Makefile turns up a truncated mermaid block, a coverage gate behind four environment variables, and OpenTelemetry pinned as a base dependency while the observability work is still marked Planned.
Who is it for?
VeriRun is worth adopting if you are running code generation benchmarks and already have Postgres and S3 to hand, because the manifest, hashing and replay layers are the part you would otherwise write yourself. It is not ready for untrusted generated code, and the README says so itself.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The mermaid block ends inside an unterminated string literal

The target architecture section says the architecture is delivered incrementally and that components shown are not all implemented today. The diagram itself stops in the middle of a label:

mermaid
flowchart LR
    C["CLI / API / CI"] --> CP["Eval Control Plane"]
    CP --> DB["Run / Task / Attempt Store"]
    CP --> AB["Admission & Budget"]
    AB --> ORCH["Ray / KubeRay Orchestrator"]

    ORCH --> DATA["Versioned Dataset & Adapters"]
    ORCH --> MODEL["Async Model Gateway"]
    ORCH --> SBX["Sandbox Manager"]

    MODEL --> EP["OpenAI-compatible Endpoint"]
    SBX --> K8S["Kubernetes Job + gVisor"]

    DATA --> EV["EvalPlus / LiveCodeBench / Harbor"]
    EV --> SBX
    K8S --> COMMIT["Artifact & Idempotent Result Commit"]
    COMMIT --> DB
    COMMIT --> REPORT["Statistic

The last line opens `REPORT["Statistic` and never closes the bracket or the quote before the fence, so the node label and the whole diagram fail to render. What survives is a picture of the control plane and execution plane split, with the reporting node missing.

The one paragraph under the diagram is therefore the only prose description of the components: the control plane owns intent and durable state, the execution plane performs retryable attempts, and benchmark adapters preserve upstream workload semantics. That paragraph names four workloads, EvalPlus, LiveCodeBench, Harbor and Terminal-Bench. The dataset node in the diagram names three and omits Terminal-Bench, which the delivery table later lists as an optional v0.8 integration.

Three releases in fourteen days, and evidence the README disqualifies in writing

The release list runs v0.4.0 on 2026-09-02, v0.5.0 on 2026-09-15 and v0.6.0 on 2026-09-16, and the last push is 2026-09-16, seven minutes before the v0.6.0 tag. The cadence is fast and the tags track the package version, which `pyproject.toml` also carries at 0.6.0, so the manifest and the newest tag agree.

What does not agree is the status of the committed evidence. The M0 section states that the full gate is manually reproducible in GitHub Actions through a workflow named EvalPlus M0 full workload, and then says it must be regenerated from a clean revision before it is presented as merged or release evidence. So the report the README links, at `evidence/m0/evalplus/REPORT.md`, is by the project's own rule not yet something to quote as release evidence.

The same section also fences the claim itself. The M0 addendum is described as verifier and replay evidence, explicitly not a model score and not a new v0.1.0 release claim. And the subset rule is spelled out: a small frozen subset is used for development smoke tests, and any published HumanEval+ or MBPP+ claim requires the standard versioned workload, with subset results labelled as subset results. Both are sensible, and both mean a reader who skims the evidence directory will overestimate what is proven.

Every released row carries a local only qualifier, and the pre-alpha warning names the executor

The delivery table has six released rows and three planned ones, and every released row ends in a parenthetical that narrows it. v0.3.0 is released with local kind and gVisor evidence only. v0.4.0 is released with local PostgreSQL and MinIO evidence. v0.5.0 and v0.6.0 are released with a local CPU trusted fixture reference. Not one row claims evidence from the distributed environment the architecture diagram is drawn around.

The callout at the top of the file is blunter. It says VeriRun is pre-alpha, that v0.1's local executor remains for trusted fixtures only and is not a security boundary, and that it should not be used for model generated or otherwise untrusted code. The package metadata agrees on the version of pre-alpha, classifying the project as Development Status 2, while the classifier list pins only Python 3.12.

That leaves an odd shape. The second invariant is replay before claims, and the sixth says unreliable runs are marked partial or invalid instead of publishing misleading conclusions. The fourth says isolation is evidence rather than a checkbox, and that security claims require an explicit threat model plus an attack regression suite on the stated Linux, Kubernetes and gVisor environment. After the six invariants the file concedes that they are design targets until the corresponding roadmap gate is completed and linked to reproducible evidence. A threat model and an attack suite are not in the table at all, while `SECURITY.md` sits at the repository root.

OpenTelemetry is a pinned base dependency while the observability row reads Planned

The runtime dependencies list four packages. Two are ranged, `httpx>=0.28,<1` and `pydantic>=2.10,<3`. Two are pinned to the patch, `opentelemetry-api==1.44.0` and `opentelemetry-sdk==1.44.0`. Those telemetry packages are the only exact pins in the base install, so every consumer inherits them at that patch level whether or not it exports anything.

The delivery table then lists OpenTelemetry, capacity and chaos evidence, and statistically valid reports as a single v0.6 row with the status Planned. So the tracing SDK is a mandatory part of a plain install while the observability capability built on it has not shipped, and the two telemetry rows have to be read as the instrument, not the feature. The other two planned rows are veRL asynchronous reward integration at v0.7 and optional Harbor and Terminal-Bench agent workload integration at v0.8.

The remaining extras split the same way. The control plane extra carries ranged `minio>=7.2,<8` and `psycopg[binary]>=3.3,<4`, while distributed executor pins `ray[data,default]==2.50.1` exactly and evalplus pins `evalplus==0.3.1` exactly. The two extras that touch heavy infrastructure are frozen; the two that touch storage are ranged. Nothing in the file explains the split, so a conflict inside the Ray or EvalPlus pins has to be resolved by editing metadata rather than by resolving a range.

The coverage gate sits behind four environment variables and a running object store

The Makefile defines a `test` target that refuses to start without four variables, checked one at a time before pytest is ever invoked:

make
test:
	@test -n "$(VERIRUN_TEST_POSTGRES_DSN)" || (echo "VERIRUN_TEST_POSTGRES_DSN is required for the full coverage gate"; exit 1)
	@test -n "$(VERIRUN_TEST_S3_ENDPOINT)" || (echo "VERIRUN_TEST_S3_ENDPOINT is required for the full coverage gate"; exit 1)
	@test -n "$(VERIRUN_TEST_S3_ACCESS_KEY)" || (echo "VERIRUN_TEST_S3_ACCESS_KEY is required for the full coverage gate"; exit 1)
	@test -n "$(VERIRUN_TEST_S3_SECRET_KEY)" || (echo "VERIRUN_TEST_S3_SECRET_KEY is required for the full coverage gate"; exit 1)
	$(PYTHON) -m pytest --cov=verirun --cov-report=term-missing

The configuration that makes that gate meaningful lives in `pyproject.toml`: branch coverage on, `source` set to the `verirun` package, `show_missing` and `skip_covered` on, and `fail_under = 85`. So 85 percent branch coverage is enforced only on the path that needs a live PostgreSQL and a live S3 compatible endpoint.

The `test-unit` target runs plain pytest with the quiet flag and no environment checks, so on a laptop with no services the suite passes and the coverage number is never evaluated. The pytest configuration adds `--strict-config` and `--strict-markers` alongside `testpaths = ["tests"]`, meaning an unregistered marker or an unknown ini key is an error rather than a warning. That strictness applies to the cheap path; the expensive one is the one gated behind services.

Every target hardcodes .venv/bin/python and none of them use the declared console script

One variable sits at the top of the Makefile, `PYTHON := .venv/bin/python`, and every recipe in the file expands it. There is no target that creates that virtual environment, no `PYTHON ?=` override, and no fallback to `python3`, so the file only works in a checkout where someone has already made a `.venv` at exactly that path. Any other interpreter means editing the Makefile.

The same applies to how the tool is invoked. `pyproject.toml` declares a console script, `verirun = "verirun.cli:main"`, and the Makefile never uses it. Every recipe calls the module instead, such as `$(PYTHON) -m verirun smoke`, `$(PYTHON) -m verirun evalplus-m0`, and `$(PYTHON) -m verirun gateway-smoke`. The module form works without installing the package, which is convenient for development and means the declared entry point is exercised by nothing in the build tooling.

Scoping is inconsistent between targets as well. `typecheck` runs mypy over `src` only, in strict mode with `python_version` set to 3.12. `lint` runs ruff over the entire repository. `schemas` runs `scripts/export_schemas.py --check` against the `schemas/` directory. So the tests directory is linted but never type checked, which matters because the strict mypy pass covers the package and the fixtures that exercise it live on the other side of that boundary.

Two evidence directories, and the Makefile writes the one the README does not link

The README points readers at a committed report, `evidence/m0/evalplus/REPORT.md`, and at a guide, `docs/M0_EVALPLUS_EVIDENCE.md`. Both live under version control, and `evidence/` is listed among the top level directories.

The Makefile writes somewhere else. `smoke` writes to `.verirun/evidence/v0.1/synthetic`. `evalplus-smoke` writes to `.verirun/evidence/v0.1/evalplus`. `evalplus-m0` writes to `.verirun/evidence/m0/evalplus`. `gateway-smoke` writes to `.verirun/evidence/v0.2/gateway-smoke`. The path segment matches the report directory name exactly, but the root is a dot directory under the working tree rather than the committed `evidence/` tree.

So regenerating the M0 workload does not refresh the report the README links, and nothing in the visible portion of the file copies one to the other. The `.PHONY` line declares thirty nine targets, including six `evidence-` prefixed variants and named counterparts for the distributed, reliability, kubernetes and container runs, which suggests a copy step lives in the part of the Makefile that follows `container-smoke`. One detail is visible in what is shown: `evalplus-m0` is the only target that sets an inline environment override, `EVALPLUS_MAX_MEMORY_BYTES=-1`, which removes the memory ceiling for that one workload and is undocumented as to why.

Editorial conclusion

VeriRun is worth adopting if you are running code generation benchmarks and already have Postgres and S3 to hand, because the manifest, hashing and replay layers are the part you would otherwise write yourself. It is not ready for untrusted generated code, and the README says so itself. Before you rely on it, check which Python you are on, since the metadata refuses anything outside 3.12, decide whether the local only evidence rows are enough for your claim, and re-run the M0 workload from a clean revision rather than quoting the committed report.

Frequently asked questions

What Python version does VeriRun need?

The package metadata requires Python 3.12 and nothing else, with `requires-python` set to `>=3.12,<3.13`. The same version appears in the repository's `.python-version` file and in the mypy configuration, and the classifier list names only Python 3.12. An install on 3.13 will be refused.

Can I run VeriRun's tests without Postgres and S3?

The `test` target exits before pytest unless `VERIRUN_TEST_POSTGRES_DSN`, `VERIRUN_TEST_S3_ENDPOINT`, `VERIRUN_TEST_S3_ACCESS_KEY` and `VERIRUN_TEST_S3_SECRET_KEY` are all set, and that is the only path that enforces the 85 percent branch coverage floor. The `test-unit` target runs pytest with no such checks, so it passes without services and never evaluates coverage.

Is VeriRun safe to point at untrusted generated code?

No, by its own statement. The callout at the top of the README says VeriRun is pre-alpha, that v0.1's local executor remains for trusted fixtures only and is not a security boundary, and that it must not be used for model generated or otherwise untrusted code. The security invariant is listed among six that the file calls design targets.

Which benchmarks and agent workloads does VeriRun support?

Adapters are named for EvalPlus, LiveCodeBench and Harbor, and the prose adds Terminal-Bench to that list, while the diagram's dataset node names only the first three. The delivery table puts Harbor and Terminal-Bench agent workload integration at v0.8 with the status Optional.

Does VeriRun have tracing turned on out of the box?

The base install pins `opentelemetry-api` and `opentelemetry-sdk` at 1.44.0, so the SDK is always present. The delivery table still lists OpenTelemetry, capacity and chaos evidence, and statistically valid reports as Planned for v0.6, so the instrument ships before the capability it supports.

Official sources

  1. aaron-for-value/VeriRun on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/aaron-for-value-verirun.svg)](https://hysenlabs.com/projects/aaron-for-value-verirun)