Model or dataset
usehelix/helix avatar
usehelix/helix

Helix ships under four names and protects the recipient, not the amount

Self-healing infrastructure for AI agent payments. 90.3% auto-recovery.

825 stars5 forksTypeScriptMIT

At a glance

What is it?
Helix wraps a function, diagnoses a failure through a six-stage pipeline, and remembers the repair in a SQLite Gene Map so the next occurrence costs nothing. It is built for agent payments on chain, which is why the interesting parts are its identity sprawl and its safety contract: the invariant it refuses to break is the recipient and calldata, while the amount and gas strategy are explicitly negotiable.
Who is it for?
Take Helix if you run an agent that moves money and you want failures diagnosed and repaired automatically, and if you are willing to read the strategy list before switching out of observe mode.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 119 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The GitHub repo, the npm package, the PyPI package and the image are four identities

One product, four distribution identities, and the page mixes them. The repository is `usehelix/helix`. The npm package is `@helix-agent/core`, and it is what the install line tells you to use. The Python package is `helix-agent-sdk`. The container image is `adrianhihi/helix-server`, which is what both the Docker command and the compose file reference.

The release titles show the drift. v2.5.0 is titled with `@vial-agent/core` extracted, v2.6.0 adds Self-Refine and a DSPy prompt optimizer, and v2.7.0 adds agent payment intelligence with VialOS beta. So the published package was renamed at least once, and a reader who searches for the name in an older release note will find a package that no longer exists.

Two details on the page are smaller and still worth fixing. The stargazers badge points at a different GitHub account, `adrianhihi/helix`, not at the repository the page lives in, and one badge link in the badge row is just an empty anchor. The npm badge is also listed twice with identical targets.

And the CLI section carries an explicit hazard warning: use `@helix-agent/core`, because `npx helix` installs a wrong third-party package. That is a real name collision on npm, and it is the kind of thing that ends up in somebody's dependency tree.

A fifth route needs no client library at all, just an HTTP POST to the repair endpoint with the error text and the platform name:

bash
curl -X POST http://localhost:7842/repair \
  -H 'Content-Type: application/json' \
  -d '{"error": "nonce too low", "platform": "coinbase"}'

The safety invariant covers recipient and calldata, and leaves the amount open

This is the part to read twice before enabling anything.

The safety layer is described as seven pre-execution constraints, a four-layer adversarial defence covering reputation, verification, anomaly detection and auto-rollback, cost ceilings, and strategy allowlists. The one invariant stated outright is that it never modifies the recipient or the calldata. That is the correct invariant for an on-chain payment: you cannot redirect funds to another address.

The execution strategy list is where the amount appears. It names `refresh_nonce`, `speed_up`, which multiplies gas by 1.3, `reduce_request`, which divides value by 2, `backoff_retry` from one second up to a sixteen second cap, `renew_session`, `split_transaction`, and `remove_and_resubmit`.

Three of those change what is paid rather than how it is paid: halving the value, splitting one payment into pieces, and cancelling and resubmitting. So the protected part of a transaction is its destination, and the adjustable part is its size, its fee and its shape. The mode table says `observe` is zero risk, `auto` is low risk because it only changes how and never what, and `full` adds fund movement strategies at medium risk. Which strategies land in which mode is not stated, and that mapping is the single most important thing to establish yourself.

The description says 90.3 percent recovery and the benchmarks say 99.9 percent

The repository description claims 90.3 percent auto-recovery. The benchmark section on the same page does not contain that number anywhere.

What the benchmarks do say is specific enough to check. An A/B test over 1,083 Base Mainnet transactions, twelve hours long, reports Helix at 99.9 percent against blind retry at 81.9 percent. Five frontier models, named as GPT-4o-mini, GPT-4o, Claude Opus 4.6, GPT-5.4-mini and GPT-5.4, were each given a bare `execution reverted` error and all five failed, against 100 percent for the pipeline. The warm path for the Gene Map is given as 2,140 milliseconds falling to 1.1 milliseconds, and $0.49 falling to $0.00 per repair.

Those are three different populations measured three different ways: a transaction trace, a synthetic error string, and a cache lookup. The 99.9 percent is a success rate on a set of real transactions, not a repair rate on arbitrary errors, and the second benchmark is a pass or fail on one error message.

The good news is that this is auditable. `benchmark-results.json` and `comparison-log.json` are committed in the repository root, and the evaluation harness is pointed at in `experiments/`. Whether the 90.3 percent in the description comes out of those files is something you can determine rather than assume.

The page counts its own modules three different ways

The module inventory does not hold still. The tree diagram of the underlying runtime lists eight entries: the repair pipeline engine, the Gene Map, Self-Refine, Meta-Learning, the safety verifier, Self-Play, federated learning, and the prompt optimizer. The Architecture section is headed with 15 learning and safety modules. The beta endpoint description says the status route reports 13 modules and 5 adapters.

Eight, thirteen and fifteen, on one page, in three places.

The Architecture section is the most complete account and it groups things three ways. Learning covers ten modules: the Gene Map with reinforcement learning, Meta-Learning for few-shot patterns, a causal graph for prediction, negative knowledge holding anti-patterns, adaptive weights for auto-tuning, Self-Play for exploration, federated learning, Gene Dream for memory consolidation, a prompt optimizer that improves its own classification through an LLM, and auto strategy generation that invents new fixes through an LLM. Safety is the constraint set. Execution is the eight strategies.

The tree diagram itself ends mid-line. After the header for the payment vertical, powered by the runtime, the last visible entry is a Coinbase row that stops after the characters 17 er, so the adapter list is incomplete on the page.

That said, the pieces named do line up with the rest: Meta-Learning is described in the tree as three similar fixes becoming a pattern so the fourth is instant, which matches the warm Gene Map number, and Self-Play is the same idea as the `self-play 10` CLI command.

The Gene Map is a SQLite file, and the warm path skips the model entirely

The repair cycle is a six-stage pipeline drawn as a single line: an error occurs, then Perceive asks what broke, Construct finds fixes, Evaluate scores them, Commit executes, Verify asks whether it worked, and Gene remembers.

The output of that last stage is the Gene Map, described as a SQLite knowledge base scored by reinforcement learning. The architectural claim is that once a fix is stored, the next occurrence of the same error is handled from that store in under a millisecond, with no diagnosis, no LLM call and no cost.

That is the part worth understanding before you deploy it, because it means the runtime has two very different behaviours. On a miss, you pay for diagnosis and construction. On a hit, you pay nothing and change nothing you did not intend, which makes the store the real risk surface: whatever ends up in it will be applied automatically next time, to any agent using the same store.

The storage location is visible in the compose file, a named volume mounted at `/app/data`, so the Gene Map lives with the container rather than in your application. The volume name in that file is `helix-data`, and it is the only persistent state in the default deployment.

The shipped server starts in observe mode while the library example defaults to auto

The quick start wraps a function in auto mode, and the headline code sample is one line: wrap the payment call with `mode: 'auto'` and await it. Every deployment artefact, by contrast, starts in the zero-risk mode.

The Dockerfile ends with a command that serves on port 7842 in observe mode, and sets the mode as an environment variable as well. The compose file passes the same value through `HELIX_MODE=observe` and publishes the port on the host. The image is Alpine based Node 20, built in two stages, and the build stage installs python3, make and g++ because better-sqlite3 compiles natively.

That default is the right choice, and it also means the safe path is the one you get for free while the risky path is the one the documentation models. Note the gap between the two: the library's argument is a one-line change to an existing payment call, and the server's argument is a mode flag.

Two configuration keys are worth reading in that compose file. `HELIX_LLM_API_KEY` and `HELIX_ADMIN_KEY` are both declared with an empty default, using the shell-style fallback of an empty string, and the admin key is what would gate an administrative surface on a port that is published to the host by default. What happens with an empty admin key is not documented on this page, and that is the question to answer before running the container anywhere shared.

Root scripts typecheck and test one workspace out of eight

The repository is a monorepo of eight workspaces: gene-map, vial-core, core, adapter-api, mcp, api, cli, and dashboard. The root `package.json` is private and named helix-monorepo.

Its scripts reveal the shape of the verification. The build script builds `packages/core` and then type-checks the API package with tsc. The test script runs tests in `packages/core` only. The typecheck script is `tsc --noEmit` against `packages/core`. The MCP server has its own build script and is not part of either the build or the test command. So of eight workspaces, one has both tests and a typecheck entry point in the root scripts, and the API, MCP, CLI and dashboard have neither wired up there.

The rest of the root package file is interesting for a different reason. Its dependencies are payment libraries: the Coinbase CDP SDK, the Coinbase SDK, Privy's node client, two x402 packages for core and EVM, and Circle's x402 batching package. Those sit in the root of a private workspace file rather than in the package that uses them. Meanwhile `openai` is a devDependency, which means the two modules that call a language model, the prompt optimizer and auto strategy generation, depend on a package that a production install of the published package would not bring in.

The end-to-end scripts are named for what they exercise: a zero-eth test, a full test, a coinbase test, and a payment agent, each run through tsx, with a bash setup step.

A file named = sits in the repository root

The top level of the tree contains an entry whose name is a single equals sign. That is what a mistyped shell redirect leaves behind, for example an `=` where a filename should have been, and it is sitting next to the Dockerfile and the license.

It is a small thing, but it is the kind of evidence about how a repository is assembled, and the rest of the root tells a consistent story. There are committed result files, `benchmark-results.json`, `comparison-log.json` and `sdk-fixtures.json`, which is what makes the published benchmark numbers checkable rather than decorative. There is `STRATEGIES.md` beside a `helix.config.example.json`, so the strategy list in the README has a home and a configuration template. There is an `install.sh`, a `railway.toml` for one deployment target, a Dockerfile plus compose file, and a `skills/` directory that no section of the page explains.

Two static pages sit alongside the documentation, `research.html` and `vision.html`, neither linked from the text above. And the page itself stops mid-sentence in the self-evolution section, at a line that begins Auto Strategy Generation and stops on the word invent, so the third level of that description, architecture evolution, is announced in the diagram and cut off in the prose.

Editorial conclusion

Take Helix if you run an agent that moves money and you want failures diagnosed and repaired automatically, and if you are willing to read the strategy list before switching out of observe mode. Leave it if you cannot accept a runtime that may halve a payment value or restructure a transaction, because that is a documented repair strategy rather than a bug, or if you need an independently verifiable accuracy figure, because the repository description and the benchmark section quote different headline numbers. Three things to check before enabling auto mode: whether `reduce_request`, `split_transaction` and `remove_and_resubmit` are allowlisted in your configuration, what happens when the admin key is empty on a published port, and whether the failure you are protecting against is actually a rate limit or a nonce problem rather than something the pipeline has never seen.

Frequently asked questions

What is Helix and how does it repair a failure?

It wraps a function and runs a six-stage pipeline when that function fails: Perceive for what broke, Construct to find fixes, Evaluate to score them, Commit to execute, Verify to check the outcome, and Gene to remember it. Successful fixes go into the Gene Map, a SQLite knowledge base scored by reinforcement learning.

What do the observe, auto and full modes do in Helix?

Observe diagnoses only and never touches your call, at zero risk. Auto diagnoses, fixes parameters and retries, described as low risk because it only changes how and never what. Full adds fund movement strategies and is marked medium risk. The README does not say which strategies belong to which mode.

How do I install Helix?

Four documented routes: `npm install @helix-agent/core` for TypeScript, `pip install helix-agent-sdk` for Python, `docker run -d -p 7842:7842 adrianhihi/helix-server` for the server, and a raw REST call posting an error and platform to `http://localhost:7842/repair`.

What benchmarks does Helix publish?

A twelve hour A/B test over 1,083 Base Mainnet transactions reporting 99.9 percent against 81.9 percent for blind retry, five named frontier models all failing on a bare `execution reverted` where the pipeline is reported at 100 percent, and a warm Gene Map path from 2,140 milliseconds to 1.1 milliseconds and $0.49 to $0.00 per repair. The repository description separately claims 90.3 percent auto-recovery.

Can I write my own adapter for Helix?

Yes. Implement the PlatformAdapter interface with a name, a perceive(error) method that returns a code, category and strategy or null, and a getPatterns method, then pass it to wrap(myFunction, { adapter: myAdapter, mode: 'auto' }). The example maps a rate limit message to the backoff_retry strategy.

Is there an npm naming problem to watch out for with Helix?

Yes. The README warns that `npx helix` installs a wrong third-party package and to use `@helix-agent/core` instead. Release titles also show the package was published earlier as `@vial-agent/core`, so older instructions may name a package that no longer exists.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. usehelix/helix on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/usehelix-helix.svg)](https://hysenlabs.com/projects/usehelix-helix)