# Reef: continual learning infrastructure for agents that keep their own history

> Reef is an Apache-2.0 Python framework that joins inference, feedback, training and versioned delivery in one loop. It suits teams already running a GPU training stack, and it is the wrong tool if you only need a served model.

**Human-Agent-Society/reef** — Continual learning infra for self-improving agents

- Repository: https://github.com/Human-Agent-Society/reef
- Website: reefinfra.ai
- Stars: 6,326 · Forks: 567
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/human-agent-society-reef

## The problem Reef targets: an agent that cannot keep what it learned

Most agent stacks are stateless in the ways that matter. An inference engine such as vLLM or SGLang serves traffic but does not train. An RL framework such as Slime, veRL or AReaL trains but does not stay live through an update and has no version history. Reef's own comparison table puts version management and staying live through updates in a column that only Reef fills, and adds a fifth row: evolving beyond weights, meaning skills and harness. That is the gap the project claims. The intended user is an engineer who already has a trainable model and an evaluation signal, and who wants the traffic their agent already serves to feed back into the next version. Reef's README frames the three entry points as model weight training, harness optimization (prompts, rules, skills) and test-time training for a measurable objective. The harness path is the cheapest to try because it needs a model endpoint and an evaluator rather than local training GPUs.

## Serve, observe, grow, commit: the four-step loop and the modules behind it

The README describes each learning cycle as four steps and names the directory that implements each one. Step one, Serve, handles agent requests and records interactions, split across reef/service and reef/runtime, where runtime also covers inference and artifact updates. Step two, Observe, matches feedback to recorded interactions, with stored records in reef/records.py and matching plus eligibility rules in reef/train/processors. Step three, Grow, produces an update from eligible records through reef/recipe and reef/train. Step four, Commit, applies the configured selection policy, evaluates candidates in reef/train/evaluation, writes version history in reef/artifact, and delivers artifacts through reef/surface.

The data flow that matters for a caller is the receipt. A request carries an x-reef-scenario header, and a new scenario name creates a scenario using the deployment's configured recipe. Requests do not select recipes; the deployment does. The response carries an x-reef-agent-record-id header, which the README calls the receipt. A later report to /reef/report names that receipt in a references list, alongside a numeric score and optional textual or structured feedback. Only eligible, scored records become training material, and the recipe decides what an update is.

## Installing reef-infra and sending a first scored request

Reef ships on PyPI as reef-infra while the import package stays reef. The README recommends uv. Artifact and checkpoint functionality requires the git-lfs system package, because Reef initializes Git LFS locally for its artifact repositories. The command below creates a virtual environment, installs the distribution, and prints the version so you can confirm the import resolves.

```bash
uv venv && source .venv/bin/activate
uv pip install reef-infra
python3 -c "import reef; print(reef.__version__)"
```

To work from a source checkout instead, the README installs git-lfs, clones the repository, and installs the package in editable mode.

## Starting a weight-training deployment from the SAO recipe

The README's worked example is the SAO deployment, run from a source checkout in an environment that satisfies the GPU requirements in the Evolve your model guide. It installs the slime extra and the runtime dependency group, exports a model path and a token, then starts the server on port 8900 with the recipe file recipes/sao/examples/sao/serve.yaml. The health check on /healthz is what tells you the deployment is ready to serve.

```bash
uv pip install -e ".[slime]" && uv pip install --no-deps --group runtime

export MODEL_PATH="Qwen/Qwen2.5-1.5B-Instruct"
export REEF_TOKEN="reef-local"

reef serve -c recipes/sao/examples/sao/serve.yaml \
  --reef.model_path "$MODEL_PATH" \
  --reef.port "8900"

curl -f http://127.0.0.1:8900/healthz
```

## A first real use: inference, receipt, and a report

The endpoint is OpenAI- and Anthropic-compatible: /v1/chat/completions and /v1/messages take the provider's own request body, and the response body uses the OpenAI-compatible format. The scenario header below creates a scenario named hello-reef on first use. After the call, read the x-reef-agent-record-id header, compare the answer to the expected string, and post the score with the receipt in references. That report is what the SAO recipe consumes to run a training step.

```python
import os
import httpx

reef = httpx.Client(
    base_url="http://127.0.0.1:8900",
    headers={"Authorization": f"Bearer {os.environ['REEF_TOKEN']}", "x-reef-scenario": "hello-reef"},
    timeout=300,
)

response = reef.post(
    "/v1/chat/completions",
    json={
        "model": os.environ["MODEL_PATH"],
        "messages": [{"role": "user", "content": "Return exactly: reef is ready"}],
    },
)

response.raise_for_status()
receipt = response.headers["x-reef-agent-record-id"]
answer = response.json()["choices"][0]["message"]["content"]

matched = answer.strip() == "reef is ready"

reef.post(
    "/reef/report",
    json={"score": float(matched), "feedback": "matched" if matched else "wrong answer", "references": [receipt]},
)
```

## Where Reef is the wrong tool, and what the docs leave open

Reef is not a serving layer you can adopt without a learning plan. The base install is deliberately free of training and GPU packages so that pip install reef-infra supports the base recipe configuration on CPU, which means the interesting paths arrive through extras: slime for the SAO example, sglang for the Python-side inference adapters (the README notes native SGLang belongs in its own GPU environment), tinker, postgres, terminus and wandb. If you have no trainable model, no evaluator and no feedback signal, the loop has nothing to close.

Two boundaries are worth stating plainly. The README does not document rollback, so the version history in reef/artifact is not described as a way to return to a previous artifact. And the release history is thin: v0.0.2 on 2026-09-02 is the only release listed. The project is not archived and the last push was on 2026-09-10, but a version number that low means the interfaces documented here can move. Reef also duplicates work you may already have: if your pipeline already trains weights and your harness never changes, Reef's version management and liveness rows solve problems you do not have.

## How Reef differs from an inference engine or an RL trainer alone

The honest alternative is the combination you probably already run: vLLM or SGLang in front, Slime or veRL behind, and your own glue for feedback matching and deployment. That stack can serve live traffic and train weights; what it does not do, per Reef's comparison, is version management, staying live through updates, or evolving skills and harness. The difference is not the training algorithm, since Reef integrates with Slime rather than replacing it. The difference is the record: Reef stores interactions and their feedback, decides eligibility, evaluates candidates, and publishes accepted updates through an artifact surface. If your team already built feedback matching and a promotion policy, adopting Reef means replacing that glue, not adding a model. If you have not built it, Reef is the argument that you should not start from scratch.

## Licence, maintenance and the cost of upgrading

The project is Apache-2.0, declared in both the repository licence file and pyproject.toml, with the distribution named reef-infra and the import package named reef. Apache-2.0 permits commercial use and modification with the usual notice obligations; that is a statement about the licence text, not legal advice, and you should read the LICENSE file for the terms that apply to you. On maintenance, the last push was on 2026-09-10 and the repository is not archived. Upgrade cost is dominated by the extras rather than the core: the tinker extra pins tinker==0.28.1 and tinker-cookbook==0.5.7 and requires Python 3.11 or newer, the terminus extra requires Python 3.12 or newer and pulls reef-eval[harbor], and the core requires Python 3.10 or newer. The core dependency list is short (aiohttp, alembic, huggingface_hub, pyyaml, reef-client, sqlalchemy, tomli-w), which keeps the base install cheap to move. Anything touching artifacts also depends on git-lfs being present on the machine.

## Conclusion

Adopt Reef if you already run a trainable model on a supported GPU stack and want feedback from live traffic to become versioned updates without stitching four systems together; the harness-optimization path needs only a model endpoint and an evaluator. Do not adopt it if you need a stable API surface: the newest release is v0.0.2, and the README does not document rollback. Before committing, verify that your GPU environment satisfies the requirements in the Evolve your model guide, that git-lfs is installed, and that recipes/sao/examples/sao/serve.yaml matches your model path and port.

## FAQ

### What is a human agent in AI?

The repository does not define the term directly; it is published by the Human-Agent-Society organization, and Reef itself is described as continual learning infrastructure for self-improving agents that learn from how you interact with them.

### How does an AI agent interact with its environment?

In Reef, the agent interacts through Reef's OpenAI- and Anthropic-compatible endpoints, /v1/chat/completions and /v1/messages, with an x-reef-scenario header. Reef records the interaction and returns a receipt in the x-reef-agent-record-id response header, which a later report uses to attach a score and feedback.

### Does Reef work with Claude Code?

The README describes Anthropic-compatible /v1/messages alongside OpenAI-compatible /v1/chat/completions, so the endpoint speaks the provider format. The README does not mention Claude Code by name, so integration with that specific client is not documented.

### What Python version does reef-infra require?

The core requires Python 3.10 or newer according to pyproject.toml. Some extras are stricter: tinker requires 3.11 or newer and terminus requires 3.12 or newer.

### Does Reef need a GPU to run?

The core dependency list is kept free of training and GPU packages so that pip install reef-infra supports the base recipe configuration on CPU. The weight-training example installs the slime extra and expects an environment that satisfies the GPU requirements in the Evolve your model guide; the harness optimization path is described as needing no local training GPUs.

## Sources

- [Human-Agent-Society/reef on GitHub](https://github.com/Human-Agent-Society/reef)
- [Issues](https://github.com/Human-Agent-Society/reef/issues)
- [License: Apache-2.0](https://github.com/Human-Agent-Society/reef/blob/main/LICENSE)
- [README](https://github.com/Human-Agent-Society/reef/blob/main/README.md)
- [Releases](https://github.com/Human-Agent-Society/reef/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/human-agent-society-reef
