HyperAgents: agents that rewrite the code that runs them
Self-referential self-improving agents that can optimize for any computable task
At a glance
- What is it?
- A Facebook Research repository in which a foundation model produces an improved version of another agent, the improvement is scored, and the loop repeats across a population of candidates.
- Who is it for?
- HyperAgents is a research harness for search over agent code, not a library you would import into a product. It is at its best when a domain already has a numeric score and a cheap enough rollouts: symbolic manipulation with sympy, the Minigrid, Minihack and Baba Is AI environments in requirements.txt, or a GPU environment under Genesis.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 68 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A repository built around a loop rather than a model
The README describes the project in one sentence: self-referential self-improving agents that can optimize for any computable task. That phrasing is doing more work than marketing copy usually does. There is no training script here and no dataset of trajectories to replay. What exists is a mechanism for asking a foundation model to produce a modified version of an existing agent, scoring that modified agent, and repeating.
Two files carry the design. `meta_agent.py` is described as the main implementation of the meta agent, the component that writes a candidate rewrite. `task_agent.py` is the main implementation of the task agent, which is the thing being rewritten. `generate_loop.py` is the entry point for running the algorithm. Alongside those sit `run_meta_agent.py`, described in the README as a script to help run the meta agent and get the diffs, plus `select_next_parent.py` and `ensemble.py`.
The existence of a module called `select_next_parent.py` is the detail that tells you this is population-based search rather than a single chain of successive improvements. A chain would only need a loop variable. Choosing a parent implies several candidates are alive at once and each round decides which of them seeds the next generation.
Provenance is unusually visible for a research release. The README links arXiv 2603.19461, a blog post on ai.meta.com, and a BibTeX block naming Jenny Zhang, Bingchen Zhao, Wannan Yang, Jakob Foerster, Jeff Clune, Minqi Jiang, Sam Devlin and Tatiana Shavrina. The repository has 2,761 stars, 363 forks and 30 open issues, which is the shape of a paper drop rather than a slowly grown library. The last push was on 2026-07-31.
Three API keys and one routing library
Setup opens with the least interesting and most telling step:
OPENAI_API_KEY=...
ANTHROPIC_API_KEY=...
GEMINI_API_KEY=...Three providers are named with no preference expressed between them. The file structure section describes `agent/` as code for using foundation models, and `requirements.txt` pins `litellm==1.74.9`, the library that sits underneath all three keys. That means the choice of model is a runtime argument rather than a code edit, which is consistent with a project that wants to run the same meta agent against several backends.
Native packages come next, and the package manager is the first real clue about the reference environment:
sudo dnf install -y python3.12-devel
sudo dnf install -y graphviz graphviz-devel cmake ninja-build bzip2-devel zlib-devel ncurses-devel libffi-develdnf rather than apt places the original development on a Fedora or RHEL family machine. Graphviz next to cmake and ninja-build suggests that agent graphs get rendered, which fits a project where the structure of a generated agent is something a human wants to look at.
The environment step is a conventional virtualenv with a specific name:
python3.12 -m venv venv_nat
source venv_nat/bin/activate
pip install -r requirements.txt
pip install -r requirements_dev.txt
# To build the docker container
docker build --network=host -t hyperagents .The name `venv_nat` implies at least one other variant existed. The `--network=host` flag on the build is worth noting too, since it suggests the container is expected to reach model endpoints and whatever drives the environment directly rather than through a compose network.
How a run is actually invoked
Two commands stand between a fresh clone and a running search. The first initializes the starting population:
bash ./setup_initial.shThe second is the loop itself:
python generate_loop.py --domains <domain>The README adds that you should see the script for args and baseline selections. That is a candid admission that the command line surface is documented in code rather than in prose, and it has a practical consequence: the flags that exist, including what baseline selection accepts, are only discoverable by reading `generate_loop.py`.
Outputs are written to an `outputs/` directory by default. No schema is given for that directory, which is expected given that the interesting contents are whatever logs and score histories a particular domain produces.
The rest of the tree is organised by role. `domains/` holds code for each domain, which is the extension point the framework is designed around. `analysis/` holds scripts used for plotting and analysis, so the repository expects that you will want curves rather than just a final number. `baselines/` sits alongside `domains/`, implying that comparison runs are first-class rather than an afterthought. `utils/` collects code shared across the repo, and the top-level scripts sit next to the two agent implementations rather than inside a package directory.
There is no tests directory in the tree, and no examples either. For a project whose output is a search over generated code, that is a reasonable absence, but it does mean the README's demonstration is the setup script and nothing more.
What the dependency list reveals about the intended domains
The generic requirements are a short list of pinned versions: `requests==2.32.4`, `dotenv==0.9.9`, `tqdm==4.67.1`, `backoff==2.2.1`, `matplotlib==3.10.3`, `docker==7.1.0`, `datasets==3.6.0`, `GitPython==3.1.44`, `litellm==1.74.9`, `pandas==2.3.2` and `sympy==1.14.0`. Several of those are informative. GitPython alongside datasets suggests at least one domain involves code or data repositories as the environment. matplotlib is what makes the `analysis/` directory useful. The Docker client suggests some environments are containerised.
Then the requirements file changes register. Under a comment marked for balrog come `hydra-core==1.3.2`, `gym==0.23.0`, `gymnasium==1.2.0`, `setuptools==80.9.0`, `wheel==0.46.2`, and three git dependencies: a pinned fork of Minigrid, a pinned fork of minihack, and `git+https://github.com/nacloos/baba-is-ai`. Two different generations of the Gym API are present at once, gym and gymnasium, which is a small sign that support for older environments is being carried rather than dropped.
A second block marked for genesis adds `rsl-rl-lib==2.2.4` and `tensorboard==2.20.0` plus a git dependency on the Genesis physics engine. Right above it sits a comment explaining that torch and torchvision are installed separately in the Dockerfile with a CUDA index.
Put together, that is at least three embodied environments and one symbolic library, and it is the clearest available answer to what any computable task means here in practice. It means a task with a numeric score function and cheap rollouts, not an open-ended objective.
The safety warning is the most useful paragraph in the README
Before the citation block, the README puts a warning callout at the top of its own section, and it is unusually blunt. It states that the repository involves executing untrusted, model-generated code, advises users to be aware of the associated safety risks, and adds that while overtly malicious behaviour is considered unlikely under the current settings and with the models used, the code may still behave destructively due to limitations in model capability or alignment. Using the repository means accepting those risks.
Read that as an engineering constraint rather than a disclaimer. The loop executes code that a model wrote, in the same environment as the experiment. There is no sandbox named in the README, no mention of container isolation for generated code, and no discussion of filesystem permissions. The Dockerfile sets `PYOPENGL_PLATFORM=egl` and `DISPLAY=:99` and installs `osmesa`, which is headless rendering for simulated environments, not process isolation. So the honest reading is that generated code runs with the permissions of the user who started the loop.
That is a different posture from running a hosted agent framework where tool calls are mediated by an API. Here the agent's output is code, and code runs where you stand. Anyone reproducing these experiments is signing up to run generated programs, which is why the README asks for acknowledgement rather than burying the point.
The related concern is cost. `meta_agent.py` calling a foundation model once per candidate per generation means API spend scales with population size and generation count, and nothing in the README describes a budget, a token cap or a cache. The README points to a Google Drive folder for experiment logs instead, which is where you would look to see what a full run cost.
What the README leaves for the paper and the blog post
The gaps here are structural rather than accidental. There are no releases in the repository, no version tags, and no changelog. Results live in a linked Google Drive folder of experiment logs rather than in the README, which means the headline numbers behind the claim that these agents self-improve are documented somewhere other than where most readers start.
The credit for that is worth giving fairly. A paper at arXiv 2603.19461 and a Meta research blog post exist, and the citation block is complete enough to paste. A reader who wants the argument and the measurements has somewhere to go. What the README itself provides is the mechanism and a working entry point.
The licensing story is also worth reading carefully, because the two signals differ. The README carries a badge linking a file called `LICENSE.md` and labels it CC BY-NC-SA 4.0, and the tree does contain `LICENSE.md`. Repository metadata reports NOASSERTION. If the file is what the badge says, then the grant is non-commercial with a share-alike condition, which is a materially narrower permission than the permissive licences common on research code. Anyone planning to build on this should read `LICENSE.md` rather than infer the terms from the badge.
Two supporting files round out the repository: `CONTRIBUTING.md` and `CODE_OF_CONDUCT.md`. Their presence, alongside `setup_initial.sh` in the tree, suggests the project has a defined process for outside code, even if the README does not describe it.
Editorial conclusion
HyperAgents is a research harness for search over agent code, not a library you would import into a product. It is at its best when a domain already has a numeric score and a cheap enough rollouts: symbolic manipulation with sympy, the Minigrid, Minihack and Baba Is AI environments in requirements.txt, or a GPU environment under Genesis. The repository hands you the meta agent, the task agent and the selection loop, then leaves the scoring function and the cost of a rollout entirely to you, which is exactly the part that decides whether a run produces anything. Start with `setup_initial.sh`, run `generate_loop.py` on a single domain, and read the arXiv paper at 2603.19461 before scaling out, because the README explains the mechanism but not the results.
Frequently asked questions
What are hyperagents?
In the HyperAgents repository, a hyperagent is an agent whose own code or configuration can be rewritten by a second agent during a run. A meta agent produces a candidate modification, the modified task agent is scored on a domain, and `select_next_parent.py` decides which candidate continues. The claim that any computable task can be optimized depends on the domain supplying a score function and cheap enough rollouts.
Which model providers does HyperAgents support?
Setup asks for an OpenAI, an Anthropic and a Gemini key, and `requirements.txt` pins `litellm==1.74.9`, the routing library that covers those providers. The `agent/` directory holds the code for calling foundation models, so the backend is a configuration choice rather than something baked into the meta agent.
Is it safe to run HyperAgents on a workstation?
The README warns that the repository executes untrusted, model-generated code and that it may still behave destructively. It names no sandbox and no restricted permissions for generated code, so code from a model runs with the privileges of the user who launched the loop. Treat it the way you would treat any experiment that executes unreviewed generated programs.
What experimental domains does the repository ship with?
The requirements file points at Minigrid, Minihack and Baba Is AI under a block marked for balrog, at the Genesis physics engine under a block marked for genesis, and at sympy in the pinned base list. `domains/` holds one directory per domain and `baselines/` holds comparison runs, so adding a new domain means writing code against the score interface those directories expect.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/facebookresearch-hyperagents)