# SIA: an agent framework that rewrites its own agent

> Hexo's SIA treats a benchmark task as the thing to optimise, and runs three agents in a loop where one writes the target agent and another rewrites it after reading its logs.

**hexo-ai/sia** — SIA is a Self Improving AI framework to autonomously improve the performance of any AI system (Model / Agent) on a benchmark task.

- Repository: https://github.com/hexo-ai/sia
- Website: https://hexolabs.com/
- Stars: 2,156 · Forks: 256
- Language: Python
- License: MIT
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/hexo-ai-sia

## Three agents in a loop, not one prompt

The architecture is best understood through the glossary in the README, which names three roles. The meta agent reads a task description and generates an initial target agent tailored to that task. The target agent, described as task specific, then attempts the task and records what it did and what came back. The feedback or improvement agent reads those logs, identifies what to change, and updates the target agent accordingly.

That is a closed loop with an artefact at each stage, and the artefact is a Python file. The README describes SIA as the official implementation of a paper titled SIA: Self Improving AI with Harness and Weight Updates, in which a language model agent updates both the harness and the weights of a task specific agent. Harness means the scaffolding around the model, so prompts, tool wiring and control flow; weights means fine tuning. The benchmark configurations in the results tables are labelled SIA-W and SIA-H accordingly, and the headline configuration updates both.

The practical difference from a prompt optimisation tool is scope. Most of those improve instructions. This one is set up to change the program.

The last push to the repository was on 2026-08-26, and the package version in `pyproject.toml` is 0.6.0 with a development status of alpha.

## Choosing an agent backend at install time

Installation is a two step affair, and the choice of extra decides which agent implementation you get. The Claude implementation uses the Claude Agent SDK and works with Claude models only.

```bash
python3 -m venv .venv && source .venv/bin/activate
pip install 'sia-agent[claude]'
export ANTHROPIC_API_KEY="..."
```

The OpenHands implementation is the multi-provider option, covering Gemini, OpenAI, Anthropic and others, and it takes whichever keys you need as environment variables: `ANTHROPIC_API_KEY` for anthropic models, `GEMINI_API_KEY` or `GOOGLE_API_KEY` for gemini models, `OPENAI_API_KEY` for openai models, and `OPENROUTER_API_KEY` for the openrouter provider.

```bash
python3 -m venv .venv && source .venv/bin/activate
pip install 'sia-agent[openhands]'
```

`pyproject.toml` fills in the rest. Python 3.11 or later is required, the base dependencies are deliberately ordinary, with `python-dotenv`, `numpy`, `pandas`, `scikit-learn`, `fastapi`, `uvicorn`, `pydantic` and `claude-agent-sdk`, and the optional extras add `openhands-ai`, `pydantic-ai`, `google-generativeai` for MLE-bench and a `dev` group with pytest, ruff and httpx. The `mlebench` and `dev` extras are not mentioned in the install section of the README, which is a small gap if you intend to reproduce that benchmark yourself.

## Running gpqa and reading the artefacts

The CLI has two subcommands, `sia run` for the improvement loop and `sia web` for the runs visualiser. A run against a bundled task looks like this.

```bash
sia run --task gpqa --max_gen 5 --run_id 1
```

Four tasks ship with the package: `gpqa`, `lawbench`, `longcot-chess` and `spaceship-titanic`. Swap the `--task` value for any of them. A shorter form without the subcommand still works and is treated as `sia run`.

What makes this worth doing rather than just reading about is that a run leaves evidence on disk. Output lands in `runs/run_{run_id}/gen_{n}/`, containing `target_agent.py`, the agent as it existed for that generation, `agent_execution.json`, the execution logs, and from generation two onward `improvement.md`, which is described as the diff rationale. So you can diff generation one against generation two, read the model's stated reason for the change, and check it against the logs.

Other flags cover the essentials: `--max_gen` defaults to 3, `--run_id` to 1, `--task_dir` points at an external task directory instead of a bundled one, and a live dashboard starts automatically at `http://127.0.0.1:8000` unless you pass `--no-web`. The port can be changed with `--web-port`, and the bind host with `--web-host`.

## One key for every model through OpenRouter

Model and provider selection happens through profiles rather than flags, and this is the part of the design that will save you time. The meta agent and the feedback agent share a profile called with `--meta-agent-profile`, defaulting to `default-meta`, and the target agent uses `--target-agent-profile`, defaulting to `default-target`. Each profile is a name or a path to a JSON file, and the packaged profiles ship inside the Python package under `defaults/profiles/`, with provider definitions under `defaults/providers/`.

The example that sells it is the OpenRouter path. Bundled profiles named `openrouter-meta` and `openrouter-target` route both agents through OpenRouter, so a single `OPENROUTER_API_KEY` covers more than 400 models without opening a separate account per provider.

```bash
sia run --task gpqa --meta-agent-profile openrouter-meta --target-agent-profile openrouter-target
```

The README also shows evaluating a specific model, here Kimi on Nebius as the target with the default meta agent alongside it.

```bash
export NEBIUS_API_KEY="..."
sia run --task gpqa --target-agent-profile kimi-nebius-target --max_gen 5 --run_id 2
```

The full agent, model and key reference is in `docs/configuration.md`, which is where to look when a provider is not listed in the README.

## Reading the benchmark claims with care

The README reports a 56.6% gain on LawBench, a 91.9% runtime reduction on GPU kernels and a 502% improvement on single-cell RNA denoising, and attributes them to the paper. The per-benchmark captions add detail: a top rank across all generations tested on OpenAI MLE-Bench Hard, 70.1% top-1 accuracy on LawBench against a prior state of the art of 45%, and a 14x speedup on an AlphaFold-3 TriMul Triton kernel while holding H100 latency targets.

One number deserves a second look rather than a repeat. The scRNA-seq denoising caption says the configuration scores 0.289 MSE normalised, which it describes as surpassing a prior state of the art of 0.240. For a squared error metric, lower is better, so the two figures cannot both be read the way the caption reads them. Either the metric is oriented differently, or the comparison is against a different baseline than the caption implies. If denoising is your task, read the paper before quoting either number.

More broadly, none of this is reproducible from the README. There is an `EVALUATION_GUIDE.md` in the repository root and an `mlebench` extra that installs `google-generativeai`, so the intent is there, but the README does not walk through reproducing any of it.

## Generated code runs with host access unless you stop it

The security note in the README is short and should not be skimmed. The default mode is `--sandbox none`, which runs agent generated code with access to the host. For untrusted tasks or untrusted models, the README directs you to `--sandbox docker`, which isolates execution in a container with no network access. The full model is documented in `SECURITY.md`.

That default is a deliberate choice for research convenience and it is the wrong default for anything else, because the loop executes code an LLM wrote, judges the result, and then writes more code, with no human in the middle. Combined with the fact that `improvement.md` is only a rationale rather than a gate, this is the part to design around: run it in a container, on a task directory you can throw away, with credentials scoped to what the run actually needs.

The repository structure supports that. Alongside `sia/` and `tests/` you get `docs/`, `environment.yml` for the Conda path, a `CONTRIBUTING.md`, a `CODE_OF_CONDUCT.md` and `tests/`, with no prebuilt model weights or data blobs checked in at the root. The licence is MIT, and the project is published by Hexo Labs, which also operates the homepage linked from the repository.

## Conclusion

SIA is interesting for a specific reason: it treats the improvement target as the whole agent rather than the prompt or the weights alone, and it makes the rewritten agent a file you can read. That makes it useful as a research tool and as a source of patterns for anyone building an automated improvement loop, since each generation leaves `target_agent.py`, `agent_execution.json` and `improvement.md` behind in the run directory. It is not a production framework, though. The package is at version 0.6.0 and classified as alpha, the default sandbox gives generated code full host access, and the headline benchmark numbers come from a paper rather than from anything you can check locally. Start with `gpqa`, the smallest bundled task, read `SECURITY.md` before pointing it at anything you care about, and run with `--sandbox docker` if the model is not one you control.

## FAQ

### What is SIA, the self-improving AI framework?

SIA is a Python framework from Hexo Labs that improves an agent on a benchmark task by running three agents in a loop. A meta agent writes a target agent for the task, the target agent attempts it and logs what happened, and a feedback agent reads those logs and rewrites the target agent, optionally updating weights as well as the surrounding harness.

### Which tasks does SIA ship with out of the box?

Four bundled tasks are named in the README: `gpqa`, `lawbench`, `longcot-chess` and `spaceship-titanic`. You select one with the `--task` flag, or point `--task_dir` at an external task directory of your own.

### What files does an SIA run leave behind?

Each generation is written to `runs/run_{run_id}/gen_{n}/` and contains `target_agent.py`, the agent as it stood for that generation, `agent_execution.json` with the execution logs, and from generation two onward `improvement.md`, which holds the diff rationale for the change.

### How do I run SIA safely against code it generates itself?

Do not rely on the default. The default `--sandbox none` mode runs agent generated code with host access, so for untrusted tasks or models the README directs you to `--sandbox docker`, which isolates execution in a container with no network access. The complete model is in `SECURITY.md`.

### What Python version and provider keys does SIA need?

Python 3.11 or later, with the version declared in `pyproject.toml` alongside the alpha development status. Keys depend on the extra you install: `pip install 'sia-agent[claude]'` needs an Anthropic key, `pip install 'sia-agent[openhands]'` works across providers, and the bundled `openrouter-meta` and `openrouter-target` profiles let a single OpenRouter key cover both agents.

## Sources

- [hexo-ai/sia on GitHub](https://github.com/hexo-ai/sia)
- [Issues](https://github.com/hexo-ai/sia/issues)
- [License: MIT](https://github.com/hexo-ai/sia/blob/main/LICENSE)
- [Project website](https://hexolabs.com/)
- [README](https://github.com/hexo-ai/sia/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hexo-ai-sia
