Model or dataset
helixml/helix avatar
helixml/helix

Helix (helixml/helix): a self-hosted fleet of coding agents with a desktop each

♾️ Private Agent Fleet with Spec Coding. Each agent gets their own GPU-accelerated desktop. Run Claude, Codex, Gemini and open models on a full private AI Stack ♾️

814 stars85 forksGoNOASSERTION

At a glance

What is it?
Helix runs Claude Code, Codex, Gemini CLI and other ACP harnesses as parallel agents inside isolated GPU-accelerated desktops, driven from a spec-task Kanban board. It is a Go control plane you deploy yourself, and the trade-off is operational weight.
Who is it for?
Adopt Helix if you already run Kubernetes or Docker internally, want agent desktops on your own GPUs rather than on laptops, and can staff the runner and provider configuration the docs describe. Skip it if a single developer running one agent in one terminal is your whole workflow, or if you cannot give it a machine and a GPU budget.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Helix is aimed at: one agent, one laptop, one person watching

The README makes the target explicit. You already run Claude Code, or Codex, or Gemini, locally: "one agent, one terminal, tied to your machine and your attention." Helix's argument is that this does not scale past one person, because the agent's state, credentials and desktop live on whoever started it. The project's answer is to move the harness onto servers you control and give each agent its own computer.

Who it is for follows from that. Teams that already have a Kubernetes cluster or a Docker host, already pay for hosted LLM APIs or run open models on their own GPUs, and want several agents working at once without each developer babysitting a terminal. It is explicitly self-hostable end to end, including air-gapped, which is the part that matters for anyone who cannot send repository contents to a hosted agent service.

Who it is not for is just as clear. If your workflow is one developer, one repository, one agent in a terminal, Helix replaces a terminal command with a control plane, a Postgres database, runners and a container sandbox per task. That is a real cost, and the README does not pretend otherwise.

Spec tasks and the Kanban board: how a request becomes a pull request

Work is organized into projects. A project connects one or more git repositories and gets a Kanban board where spec tasks move through six columns: Backlog, Planning, Spec Review, In Progress, Pull Request, Merged.

The mechanism is worth stating plainly, because it is the part that differs from a chat window. You write a one-paragraph description of the outcome you want (what should be true when done, not how to do it). Clicking Start Planning sends a planning agent to read the repository and write a spec to a `helix-specs` branch, covering requirements, design and a task breakdown. That spec lands in Spec Review, where you either highlight text to request changes, which sends the agent back to re-plan, or approve it. Only then does an implementation agent code, inside its own isolated desktop sandbox. When it finishes, a pull request is opened in your repository, and the task closes when that PR merges.

The design judgement here is that the PR stays the review gate. Helix does not ask you to trust the agent's summary of its own work; it asks you to review a diff in your existing forge. Per-column WIP limits are also mentioned, which is the only back-pressure the README describes against starting more tasks than you can review.

One desktop per agent, and what the isolation actually consists of

Each agent gets a GPU-accelerated streaming desktop with a browser, terminal, filesystem and GUI apps, not just a shell. You can watch any agent work in real time, zoom into a live screen from a dashboard, and type into the task thread to steer it. The README also states that you can switch to a different agent mid-session and the new one picks up where the last left off, because context carries across the switch.

The density claim is the interesting one: many fully isolated agent desktops on a single machine, with a deduplicated filesystem and per-agent credential and network isolation. The repository layout backs up the desktop part rather than just asserting it. There are separate Dockerfiles named `Dockerfile.sway-helix`, `Dockerfile.ubuntu-helix`, `Dockerfile.zed-build`, `Dockerfile.qwen-build` and `Dockerfile.sandbox`, plus a `qemu-patches/` directory and a `desktop/` tree. The Go module list includes `github.com/go-gst/go-gst`, `github.com/bnema/wayland-virtual-input-go` and `github.com/go-rod/rod`, which is consistent with a Wayland desktop, synthetic input and browser automation rather than a plain container exec.

That is a heavier sandbox than a process in a namespace, and it is the reason the project needs GPUs and a runner pool rather than a single binary. The README does not quantify the per-desktop memory or GPU cost, so sizing a host is something you will have to measure yourself.

Installing Helix on Docker and creating a first spec task

The README gives a quickstart installer for Docker. It downloads a script, makes it executable and runs it with sudo, and the installer prompts before making changes to the system. By default the dashboard is then available on `http://localhost:8080`.

bash
curl -sL -O https://get.helixml.tech/install.sh
chmod +x install.sh
sudo ./install.sh

For a deployment with a DNS name, the README points at `./install.sh --help` or the docs, and says TLS termination is documented. Kubernetes installs go through the Helm charts under `charts/`, which the README describes as the production path; the truncated text stops at that sentence, so read the docs before assuming a chart name or values file.

The compose file shows what the stack expects from the environment. It runs an `api` service from `ghcr.io/helixml/controlplane`, publishes port 8080 and an asset SSH proxy on 2224, and reads a `.env` file. Several keys are worth knowing before first boot:

yaml
environment:
  - SERVER_PORT=8080
  - POSTGRES_HOST=postgres
  - RUNNER_TOKEN=${RUNNER_TOKEN-oh-hallo-insecure-token}
  - ADMIN_USER_IDS=${ADMIN_USER_IDS-all}
  - ADMIN_USER_SOURCE=${ADMIN_USER_SOURCE-env}
  - OPENAI_API_KEY=${OPENAI_API_KEY:-}
  - ANTHROPIC_API_KEY=${ANTHROPIC_API_KEY:-}

Two of those defaults are development defaults, not production ones. `RUNNER_TOKEN` falls back to `oh-hallo-insecure-token`, and `ADMIN_USER_IDS` defaults to `all`, with the compose comment saying it exists to "lock down dashboard in production". Set both deliberately. Dynamic LLM providers can also be declared as `provider1:api_key1:base_url1,provider2:api_key2:base_url2`, and the comment states that only providers which do not already exist are created.

Once the dashboard is up, the first real use is the loop the README describes: create a project, connect a repository, add a task in Backlog with a one-paragraph outcome, click Start Planning, read the spec on the `helix-specs` branch, approve it, and watch the implementation agent in its desktop until a PR appears.

Where Helix is the wrong tool, and what the README leaves open

The clearest limitation is that Helix is infrastructure. A container per task, a desktop streaming stack, a Postgres database and a runner pool are all things that can fail independently of the agent, and none of that exists when you run a harness in your own terminal. For a solo developer or a small team without a cluster, the operational surface is larger than the problem it solves.

The second is that the review gate is a pull request, and the README says so directly. If your code does not live in GitHub, GitLab or Azure DevOps, the workflow described here does not close the loop. Those three are the only source control integrations listed.

The third is what the documentation does not cover. The README does not describe rollback of a spec task, what happens to an in-flight desktop when a runner dies, or how the deduplicated filesystem behaves when two agents write the same path. It also does not state per-desktop resource requirements, so capacity planning is on you. Air-gapped operation is claimed, but the README does not walk through mirroring images or models, which is the part that usually decides whether an air-gapped install is feasible.

How Helix differs from running Claude Code or Codex directly

The honest alternative is the harness itself. Claude Code, OpenAI Codex and Gemini CLI each run locally, in your terminal, against your files, with no control plane. The difference is not capability per agent; it is where the agent lives and who can see it. A local harness has no shared board, no persistent desktop a teammate in another time zone can open, and no fleet dashboard. It also has no isolation boundary, which cuts both ways: nothing to configure, and nothing stopping an agent from touching anything your user account can touch.

Helix's answer is to keep the harness and replace the surroundings. You can swap harnesses per task, including Goose and Zed Agent, or anything that speaks ACP, and point them at hosted providers or at a self-hosted OpenAI-compatible endpoint such as vLLM. The README's cost argument is that context carries across a swap, so a task can start on a cheap local model and escalate only when needed. That is a claim about workflow, not a benchmark, and the README publishes no numbers for it.

The second comparison is with the broader private GenAI stack Helix also ships. Knowledge and RAG with document ingestion, a web scraper, Kodit and LlamaIndex backends with PGVector embeddings, tracing of every agent step and token cost, multi-tenancy with OIDC, billing and metering, cron and webhook automation, and Slack, Discord and email notifications. If you only want the agent fleet, you are adopting that surface too.

Maintenance, releases and the licence question

The repository is not archived, and the last push was on 2026-09-10. Releases are frequent and small: 2.12.11 on 2026-09-07, 2.12.12 on 2026-09-09 and 2.12.13 on 2026-09-10. The compose file pins images through `HELIX_VERSION`, defaulting to `latest`, so an upgrade is a version bump plus a restart rather than a rebuild, and the Dockerfile pins its Go base image by digest with a comment telling maintainers to update the digest when upgrading Go. That is a project that expects operators to track versions.

The upgrade cost that matters is not the control plane, it is the sandbox images and the runners. The Dockerfile pulls a prebuilt embedding model stage from `registry.helixml.tech` by digest and downloads `libtokenizers.a` through `github.com/helixml/kodit/tools/download-ort` for CGo, which is why the build stage is Debian rather than Alpine. Air-gapped operators will need to mirror those artifacts.

The licence is listed as NOASSERTION, and the repository has a `LICENSE.md`. That means GitHub could not classify it automatically, so read the file before you plan to redistribute the images or offer Helix as a service. This is not legal advice; it is a pointer to the one file that decides the question.

Editorial conclusion

Adopt Helix if you already run Kubernetes or Docker internally, want agent desktops on your own GPUs rather than on laptops, and can staff the runner and provider configuration the docs describe. Skip it if a single developer running one agent in one terminal is your whole workflow, or if you cannot give it a machine and a GPU budget. Verify first: which agent harnesses you actually need, whether your git host is GitHub, GitLab or Azure DevOps, and whether the NOASSERTION licence in LICENSE.md matches how you intend to redistribute the images.

Frequently asked questions

What is Helix (helixml/helix) and what does it do?

Helix is a self-hostable control plane that runs coding agents as a fleet, each in its own GPU-accelerated desktop sandbox with a browser, terminal, filesystem and GUI apps. You organize work as spec tasks on a Kanban board, and each task ends in a pull request in your own repository.

How do I install Helix on Docker?

The README quickstart downloads `https://get.helixml.tech/install.sh`, makes it executable and runs it with sudo; the installer prompts before changing the system, and the dashboard is then available on `http://localhost:8080`. For a DNS name it points to `./install.sh --help` or the docs.

Which agent harnesses and LLM providers does Helix work with?

The README lists Claude Code, OpenAI Codex, Gemini CLI, Qwen Code, Goose and Zed Agent, plus anything that speaks ACP, and says you can swap harnesses per task. Providers include hosted APIs, Anthropic through Helix's proxy including Vertex AI and Bedrock, and any OpenAI-compatible endpoint such as a self-hosted vLLM server.

Which git hosts can Helix open pull requests on?

The README lists GitHub, GitLab and Azure DevOps as the source control integrations, and describes the pull request as the real review gate that closes a task when merged.

Official sources

  1. helixml/helix on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/helixml-helix.svg)](https://hysenlabs.com/projects/helixml-helix)