Model or dataset
Human-Agent-Society/CORAL avatar
Human-Agent-Society/CORAL

CORAL: autoresearch infrastructure that runs Claude Code, Codex and OpenCode in parallel git worktrees

Open-source autoresearch powered by autonomous coding agents. Run Claude Code, OpenCode, and Codex with grading, shared knowledge, and multi-agent evolution. Accepted at COLM 2026.

1,040 stars129 forksPythonApache-2.0

At a glance

What is it?
CORAL is an Apache-2.0 Python orchestration layer that puts multiple coding agents in isolated git worktrees, scores every commit with a grader daemon, and shares attempts and notes through a common state directory. It is for teams with a measurable objective, not for general code assistance.
Who is it for?
CORAL is for people who already have an objective a script can score: a benchmark harness, a kernel, a simulation, a Kaggle-style task. It is the wrong tool if you cannot write a grader, because the grader is the entire feedback signal and the project removed legacy auto-discovery of eval/grader.py.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 22 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What CORAL actually automates, and for whom

Most agent tooling assumes a human in the loop: you ask, the agent edits, you review. CORAL assumes the opposite. You supply a codebase and a grader, and the system runs a population of coding agents against that codebase until the score improves. The README puts it as infrastructure for "autonomous AI agent organizations" that run experiments, share knowledge, and continuously improve solutions.

The intended user is someone with a measurable objective and a budget for repeated model calls. The examples directory backs this up: circle_packing, dna_design, drug_design, erdos, kernel_engineering, mnist, spaceship_titanic, stanford_covid_vaccine, swebench-verified, terminal-bench. These are tasks where a script can return a number. Nothing in the repository suggests CORAL helps with ordinary feature work, refactoring, or code review, because those have no automatic score.

The paper is titled "CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery," and the project was accepted to COLM 2026. That framing matters for expectations: this is a research artifact that ships as a CLI, not a product with an opinionated onboarding path.

Worktrees, a shared .coral/public/ directory, and a grader daemon

The architecture is described in the README in three sentences, and the diagram is the clearest statement of it. Each agent runs in its own git worktree. Shared state (attempts, notes, skills) lives in .coral/public/ and is symlinked into every worktree, so agents see each other's work in real time. A grader daemon scores every commit. The manager interrupts agents with heartbeat prompts labeled reflect, consolidate and pivot.

That is a different shape from a message-passing multi-agent framework. There is no agent-to-agent protocol here; coordination happens through the filesystem and through the score. An agent writes an attempt, commits, gets scored, and reads what other agents wrote. The symlink is the mechanism, which also means concurrent writes to the same note file are a real hazard the documentation does not address.

The heartbeat prompts are the only steering signal from above. reflect, consolidate and pivot are not free-form instructions; they are fixed interruption types, and the manager decides when to fire them. If your task needs a human to redirect strategy mid-run, CORAL gives you a coarse lever rather than a conversational one.

v0.6.0 added multi-island runs, which partition agents into isolated islands with scoped attempts, notes, skills, heartbeat state, and migration between islands. The stated purpose is broader exploration. In practice this is the knob you reach for when a single population converges too early.

Installing CORAL and running a first task

The README's installation path is a shell script that installs the latest coral release globally via uv tool install. The script is fetched from the main branch of the repository.

bash
curl -fsSL https://raw.githubusercontent.com/Human-Agent-Society/CORAL/main/install.sh | sh

The README notes you can pin a specific release with CORAL_VERSION=<tag> if you need to. Manual install, dev setup and prerequisites are deferred to the installation docs rather than spelled out in the README. Python 3.11 or newer is required, and pyproject.toml caps it below 3.14.

The two-command first run is given directly in the README:

bash
coral init my-task
cd my-task && coral start -c task.yaml

coral init scaffolds a task directory. coral start -c task.yaml launches agents against that configuration. What the README does not show is a complete task.yaml, so the grader wiring is the part you will spend time on. The release notes for 2026-06-13 state that legacy eval/grader.py auto-discovery was deprecated and removed, and that graders must now be wired via grader.entrypoint pointing at a packaged grader. If you are following an older tutorial, that is why it fails.

Each agent runtime must be installed and authenticated separately. The supported values for agents.runtime are claude_code (the default), codex, dsh, cursor, kiro, opencode and pi. CORAL orchestrates these tools; it does not provide the models behind them.

Driving CORAL from Claude Code or Codex with the plugin

The plugin is the lower-friction path, and the README describes it as a skills-first bundle with no MCP. It teaches the workflow coral setup, then init or validate, then start, status and log, and it checks that coral is installed when a session starts.

code
/plugin marketplace add Human-Agent-Society/CORAL
/plugin install coral@coral-marketplace

For Codex, the README gives a parallel pair of commands and notes v0.117.0+ as the requirement:

code
codex plugin marketplace add Human-Agent-Society/CORAL
codex plugin add coral@coral-marketplace

The quickstart is aimed at code you already have. You open the repository you want to optimize and ask, in the README's own example phrasing, to make sample() in saga/decode.py faster without changing its output. The plugin then scaffolds a gitignored .coral_workspace/, copies your code into a seed/, writes a grader for your metric, and loops coral validate until the task is launch-ready, at which point it hands you the coral start command.

On Claude Code a coral-task-author subagent performs that work autonomously and a coral-run-doctor triages a stuck run. Both are named in the README; neither is documented in detail there. The bundled skills are coral-quickstart, setting-up-coral, creating-a-coral-task and running-coral-experiments. Note the plugin is not a substitute for the CLI: it authors and validates tasks, then hands control back.

The grader is the bottleneck, and it is yours to write

CORAL's only fitness signal is the grader. If the grader is slow, every commit costs that time. If it is noisy, agents optimize the noise. If it can be gamed, agents will game it, because that is what selection pressure does. The README's framing, "give it a codebase and a grader, and CORAL handles the rest," is accurate about where the burden sits.

The project does ship help here. Rubric judges, announced on 2026-04-24, are two reusable LLM-judge grader packages for open-ended tasks such as reports, memos and legal analysis. That is a meaningful concession that not every objective is a number, but an LLM judge is a noisier signal than a deterministic test, and CORAL does not claim otherwise.

There is also a security boundary to think about. The 2026-06-24 note says the Docker session now isolates the agent from the grader: each agent runs as an unprivileged user while manager and grader stay root, so agents can no longer read .coral/private/, which holds the grader venv and answer keys, not even via Bash. On the host this stays opt-in via agents.isolate_user. Read that carefully. If you run on the host without setting agents.isolate_user, the isolation described for Docker does not apply, and answer keys sitting in .coral/private/ are readable by the agents you are grading.

The wrong-tool case is any task without a defensible automatic score. Multi-agent evolution on an unmeasurable objective produces confident commits and no signal.

How CORAL differs from a plain agent harness

A single agent harness such as Claude Code or Codex is a general assistant: you describe a change, it edits files, you judge the result. The difference is not the model or the editing ability. It is that CORAL adds a population, a persistent shared memory, and an external score, and it removes the human from the inner loop.

Compared with running several agent sessions by hand in separate clones, the concrete differences are the symlinked .coral/public/ state, so attempts and notes are visible across agents without you copying anything, and the grader daemon, which scores every commit rather than only the ones you notice. Compared with a general workflow orchestrator, CORAL is narrower: it assumes git worktrees, a grader, and a coding-agent CLI that is already installed and authenticated.

Multi-island runs are the clearest example of a design choice with a cost. Partitioning agents into islands with scoped attempts, notes, skills and heartbeat state, plus migration between islands, buys exploration breadth. It also means knowledge written in one island does not reach another unless a migration happens, which is a coordination cost you accept deliberately or not at all.

Licence, release cadence and upgrade cost

CORAL is Apache-2.0, declared both in the repository's LICENSE file and in pyproject.toml. That is a permissive licence with an explicit patent grant, and it imposes no copyleft obligation on your own code. It says nothing about the licences or terms of the agent runtimes you point it at, which you install and authenticate yourself. This is not legal advice; if you redistribute CORAL inside a product, read the licence text rather than this paragraph.

The release cadence is fast. v0.7.19, v0.7.20 and v0.7.21 landed on 2026-08-20, 2026-08-30 and 2026-09-06, and the last push to main was on 2026-09-08. The package classifies itself as Development Status :: 3 - Alpha, and the version is derived from git tags at build time via hatch-vcs, so a source checkout without tags will not report a clean version.

The upgrade cost is not the install, which is one script. It is configuration drift. The 2026-06-13 removal of eval/grader.py auto-discovery is the pattern to expect: a task that worked before that date needs grader.entrypoint rewired. The 2026-06-24 isolation change is the other pattern, where the default behaviour of a flag such as agents.isolate_user differs between Docker and host. Pin with CORAL_VERSION=<tag> if you need a run to be reproducible across a paper or a report.

Editorial conclusion

CORAL is for people who already have an objective a script can score: a benchmark harness, a kernel, a simulation, a Kaggle-style task. It is the wrong tool if you cannot write a grader, because the grader is the entire feedback signal and the project removed legacy auto-discovery of eval/grader.py. Before committing compute, verify two things: that your task's grader runs under coral validate, and that your chosen runtime is installed and authenticated, since CORAL does not do that for you. Check whether agents.isolate_user is set if you plan to run on the host rather than in Docker, because the private directory holding the grader venv and answer keys is only shielded from agents by that setting.

Frequently asked questions

What is CORAL and who is it for?

CORAL is open-source infrastructure for running autonomous coding agents against a codebase with a grader, so that experiments are scored automatically and agents share attempts and notes. It is aimed at people with a measurable objective, such as the benchmark-style tasks in the examples directory, rather than at general code assistance.

How do I install CORAL?

The README gives a single command that fetches install.sh from the main branch and pipes it to sh; the script installs the latest coral release globally via uv tool install. You can pin a release with CORAL_VERSION=<tag>, and Python 3.11 or newer is required.

Which coding agents does CORAL support?

The README lists claude_code as the default runtime, plus codex, dsh, cursor, kiro, opencode and pi, selected through agents.runtime. Each agent must be installed and authenticated separately; CORAL does not provide the models.

How do I write a grader for a CORAL task?

Graders are wired through grader.entrypoint pointing at a packaged grader, because legacy eval/grader.py auto-discovery was deprecated and removed as of the 2026-06-13 note. The project also publishes two reusable LLM-judge grader packages, called rubric judges, for open-ended tasks.

Official sources

  1. Human-Agent-Society/CORAL on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/human-agent-society-coral.svg)](https://hysenlabs.com/projects/human-agent-society-coral)