# AI Market Maker: a hedge fund simulator with a hard veto switch

> LangGraph desks, an arbitrating signal step, a Risk Guard that can refuse a trade outright, and a pinned data directory with sha256 hashes. The name promises market making. The code delivers directional crypto trading with unusually careful plumbing.

**olaxbt/ai-market-maker** — Agentic AI Hedge Fund OS (AIMM)

- Repository: https://github.com/olaxbt/ai-market-maker
- Stars: 2,105 · Forks: 300
- Language: Python
- License: AGPL-3.0
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/olaxbt-ai-market-maker

## What it is called, and what it actually does

Start with the mismatch, because it is the first thing a reader will trip over. The repository is named ai-market-maker and the README title says market maker, but nothing in the system places orders on both sides of a book, quotes a spread, or earns a fee for providing liquidity. Market making is a specific business with specific mechanics: quoting, inventory risk, adverse selection, and fees. This system does directional crypto trading with a hedge fund's organizational shape.

The repository description is more honest about that, calling it an Agentic AI Hedge Fund OS, abbreviated AIMM throughout the configuration. The README describes it as an open-source, hedge-fund-style trading stack for crypto. So the accurate summary is a multi-agent trading framework with a risk governance layer, not a market maker, and anyone evaluating it for spread capture should stop here.

What it does have is an unusual amount of structure around the trading decision. Specialist agents act as desks with named responsibilities, a LangGraph layer orchestrates them, a Risk Guard can veto any trade, and the README claims centralized policy, benchmarking against buy-and-hold, and full traceability.

The Python package metadata backs the naming pattern: the project is `ai-market-maker` at version 0.1.0, described as Crypto Market-Making AI Agent, requiring Python 3.11 or newer. The version number is worth noting against the fact that the repository publishes no tagged releases at all. A 0.1.0 with no releases is a trunk-branch project, which is a fine thing to be and a specific thing to plan around.

Scale indicators: 2,105 stars, 300 forks, and only 3 open issues. That last number is the outlier. A project at this star count usually carries dozens or hundreds of open issues, and three suggests either unusually clean triage or that people are filing elsewhere, most likely on Discord, Telegram or X rather than on GitHub.

## The LangGraph flow: four stages and a veto

The architecture is documented as a rough four-stage flow, and reading it tells you more about the priorities than any feature list would.

Stage one is a market scan plus the Tier-0 desks, drawn from whichever desks the deploy configuration enables, with technical analysis and macro named as examples. Stage two is risk plus desk debate: risk context is assembled, and an optional desk debate driven by an LLM is available but explicitly off in the shipped presets. Stage three is a signal arbitrator, and the ordering inside this stage is the interesting part. Static `agents.*.weight` math runs first. Then, optionally, desk chain-of-thought if LLM mode is enabled. Then, optionally, an arbitrator LLM overlay. The output is BUY, SELL or HOLD.

That ordering is a risk decision expressed as code. The deterministic weighted calculation is authoritative, and the model layers sit on top as refinements rather than as the foundation. Many multi-agent trading systems invert this, letting an LLM arbitrate directly, which produces decisions that vary between identical runs and cannot be regression-tested. Here a model can be switched off and the system still produces a signal.

Stage four is the portfolio path: a proposal becomes subject to the Risk Guard veto, and only then executes. The README is insistent that the Risk Guard holds final veto power rather than merely logging, which is the distinction that matters. A component downstream of execution can only report what went wrong. A component upstream of execution can prevent it.

On execution specifically, the current state is paper trading. The shipped mode executes on Binance Testnet, with a local backtester for offline work and a Hyperliquid adapter that is dry-run only, reached through an OMS layer. Live execution is possible but gated. The environment template sets `AI_MARKET_MAKER_ALLOW_LIVE=0`, `HYPERLIQUID_TESTNET=1` and `HYPERLIQUID_DRY_RUN=1`, so a fresh configuration cannot accidentally trade real funds. That default is the single most responsible line in the repository.

The stated goals section is candid about what is missing: full position lifecycle, multi-asset portfolio management, configurable leverage, and improved long/short handling are all listed as near-term rather than shipped. Longer-term items are deeper agentic capability, better OpenClaw integration, and additional execution venues and data sources.

## Pinned data and benchmarks, or what reproducibility actually means here

The reproducibility claim is the most technically interesting part of the project and also the most limited, so it deserves a precise reading.

The README describes pinned historical data in a `data/` directory containing daily OHLCV, funding rates, FRED and DefiLlama macro series, a Fear and Greed index, and news dailies, together with a sha256 manifest, and states that `--csv-only` reruns match as a result. It also notes the tape is bar-aligned with no live Nexus feed during backtests.

This is a real and correct engineering practice. Version-controlled data with content hashes removes the most common source of irreproducible backtests, which is silently re-downloaded data that has changed underneath a result. If a number in a report came from data whose hash you can verify, you can at least establish that the input was stable.

What the hashes do not fix is the other half of the problem. An agentic backtest requires an LLM, and the project states plainly that backtesting requires an agentic LLM with an OpenAI or Atlas Cloud key. A model call is not deterministic, and no manifest can make it so. So the guarantee is: given the same data, the rerun sees the same inputs. The guarantee that is not available: given the same data, the rerun produces the same decision.

That is a meaningful distinction and it is the honest boundary of what this feature delivers. Treat the pinned tape as eliminating data drift, not as eliminating nondeterminism.

The benchmarking side is better specified. Every backtest includes built-in benchmarks against buy-and-hold, and the README frames the inclusion of benchmarks as a discipline rather than a feature, with the line that there is no hand-waving. For an agentic system, that comparison is the only number that speaks to a question anyone actually has, since the alternative is a return figure with no reference point.

The environment template adds the execution parameters that decide how a backtest terminates, including `BACKTEST_API_MAX_STEPS=5000` and `STRATEGY_INTERVAL_SEC=900`, a fifteen-minute bar cadence. There is also `OPENAI_MODEL=gpt-4o-mini` as the default, which is a cheap model by the standards of trading systems and worth noticing given the model is being asked to arbitrate trade signals.

## Running it: Docker, TA-Lib, and one shared key you should replace

The quick start is honest about the awkward parts, which is a good sign. The sequence clones the repo, installs uv, installs TA-Lib before the Python dependencies, syncs with the dev extra, installs pre-commit hooks, copies the environment template, then brings up the whole stack with Docker Desktop required:

```bash
# 4. Install Python dependencies
uv sync --extra dev
uv run pre-commit install

# 5. Set up environment
cp .env.example .env

# 6. Run the platform stack: DB + migrate + API + worker + web
# Requires Docker Desktop. Futu OpenD optional (`--profile with-futu`).
#
docker compose up --build -d
```

TA-Lib comes first because it is a C library, and the README recommends Conda for it, showing a Miniconda download and a `conda install -y ta-lib -c conda-forge` line. The dependency list confirms the weight of the stack: ccxt for exchange connectivity, yfinance and futu-api for market data, langgraph and langchain for orchestration, statsmodels for the quantitative layer, openai for the model calls, eth-account for signing, fastapi and uvicorn and gunicorn for serving, alembic for migrations, SQLAlchemy and psycopg for storage, and tweepy for the social feed.

The stack is multi-service, which the compose file makes concrete. A `secrets-init` service generates credentials into a `./.secrets/` directory on first boot if the environment variables are left empty, the `db` service runs postgres:16 on port 5433 with a healthcheck using `pg_isready`, and a `migrate` service waits for the database to be healthy and then runs `alembic upgrade head`. The API and worker have their own Dockerfiles, and there are separate compose files for production and for a leaderboard variant. Port 5433 rather than the default is called out as a deliberate choice to avoid colliding with an existing local Postgres.

Two details in the environment template deserve attention. First, live trading defaults to off and Hyperliquid defaults to dry-run, as noted above. Second, the template ships a pre-filled Nexus demo key with the comment that it is shared and rate-limited and should be replaced. That is a courtesy to anyone evaluating the project quickly, but it also means a copied `.env` carries a credential that is not yours. Replace it before anything else runs.

The configuration split is worth noting as a design choice: strategy lives in a deploy JSON at `config/deploy.active.json`, policy and application defaults live in their own JSON files, and the `.env` file holds secrets only. Keeping strategy out of the environment file is what makes a strategy versionable and a run reproducible.

The dashboard is a Next.js app in `web/`, served on port 3000 and bound to localhost by default, with routes for a research console, a leaderboard, and a getting-started view. There is a dedicated documentation set referenced for weights and thresholds and for graph notes, plus a `PRODUCTION.md` and a `SECURITY.md` in the tree.

## OpenClaw packaging, the agent contract, and where the LLM actually sits

Three structural choices tell you what the authors think a multi-agent system is.

The first is a shared contract. Every agent follows the same Input, Process, Output, Feedback shape, and the README lists standardized agents as a reason the project stands out. A uniform contract is what makes desks swappable and what lets the orchestrator treat a macro desk and a technical-analysis desk as the same kind of component. Most multi-agent demos skip this and end up with a bespoke script per agent.

The second is the OpenClaw packaging. The repository ships a `SKILL.md`, a manifest and dedicated runners under an `openclaw/` directory, and calls the system OpenClaw-ready. This is the same skill-bundle pattern that has appeared across the 2026 tooling landscape: a portable folder of instructions and metadata that an assistant can load. The topics list confirms the association alongside agentic-trading, crypto, multiagent-systems and openclaw. The tree shows both `openclaw/` at the top level and `shared/` infrastructure in the other agent-skills project, which suggests a shared convention rather than a one-off.

The third is the governance layer described as a unified agent interface plus governance layer, alongside centralized policy. Combined with the event ledger, full traces and reasoning logs, this is a system designed to be watched. The README lists transparency as a differentiator and points to a web dashboard for telemetry and traces.

Where the model actually sits is the part to keep straight, because it is narrower than the marketing implies. The LLM appears in three optional places: desk debate, which is off in shipped presets; desk chain-of-thought inside the arbitrator, gated on `llm_enabled`; and an arbitrator overlay. The desk debate being off by default is a deliberate choice worth naming, since debate between agents is the most expensive and least predictable part of this architecture.

The README also discloses a commercial relationship. Atlas Cloud is named as the sponsor, described as an inference platform offering access to 300 or more curated models across modalities, and its coding plan promotion is linked from the README. Sponsored inference is ordinary and not a defect, but it does mean the default model path runs through the sponsor's endpoint, which is worth weighing if the data involved is sensitive. Note also that the key name appears spelled both `ATLASCLOUD_API_KEY` in the code blocks and `ALTASCLOUD_API_KEY` in the environment template comment, which is a small documentation inconsistency worth watching for if setup fails.

## Conclusion

The most valuable thing in this repository is a design decision that most trading bots do not make: separating the thing that wants to trade from the thing that is allowed to trade. Every specialist desk produces a proposal, an arbitrator reduces them to a signal, and a Risk Guard component holds a veto that runs after the portfolio step and before execution. When a component that can only say no sits on the execution path, you can inspect its behaviour and argue with it. When the model that generates the idea also decides whether to act, you have nothing to audit. Everything else in the stack, the pinned data directory with a sha256 manifest so a rerun matches, the built-in buy-and-hold benchmark, the event ledger, points the same way: toward a system whose output can be checked rather than trusted. Three practical notes before you evaluate it. An LLM API key is mandatory, so offline reproduction of a backtest is not available and the pinned tape only fixes the data half of reproducibility. The live-trading switch defaults to off and Hyperliquid defaults to dry-run, which is the correct posture for a system with this much autonomy. And the repository publishes no releases while the package version reads 0.1.0, so there is no upgrade path to reason about: you track the trunk branch or you do not use it.

## FAQ

### Is this actually a market maker?

No, and the naming is misleading. Nothing in the system places orders on both sides of an order book, quotes a spread, or earns a liquidity provision fee. The repository description calls it an Agentic AI Hedge Fund OS, and the README describes a hedge-fund-style trading stack where specialist agent desks produce directional signals that are arbitrated and then passed to a Risk Guard veto before execution. Evaluate it as a directional trading framework, not as a market making business.

### Can the Risk Guard stop a trade from executing?

Yes. The Risk Guard runs after the portfolio proposal step and before execution, and the README states it holds final veto power rather than merely logging. That placement is the point: a component downstream of execution can only report a problem, while one upstream of it can refuse the action. The desk debate stage is separately gated and described as off in the shipped presets.

### Can I run a backtest without an LLM API key?

No. The README states that backtesting is agentic and requires an LLM, with an OpenAI or Atlas Cloud key set through OPENAI_API_KEY or ATLASCLOUD_API_KEY. The environment template also defaults OPENAI_MODEL to gpt-4o-mini. The pinned data directory with its sha256 manifest makes the input side of a rerun reproducible, but model calls remain nondeterministic, so identical reruns are not guaranteed to produce identical signals.

### Will it trade real money if I install it?

Not by default. The shipped mode executes on Binance Testnet, the Hyperliquid adapter runs dry-run only, and the environment template sets AI_MARKET_MAKER_ALLOW_LIVE=0 with HYPERLIQUID_TESTNET=1 and HYPERLIQUID_DRY_RUN=1. Note also that the .env.example file ships a pre-filled shared Nexus demo key marked rate-limited, which you should replace with your own before the first run.

## Sources

- [Issues](https://github.com/olaxbt/ai-market-maker/issues)
- [License: AGPL-3.0](https://github.com/olaxbt/ai-market-maker/blob/main/LICENSE)
- [olaxbt/ai-market-maker on GitHub](https://github.com/olaxbt/ai-market-maker)
- [README](https://github.com/olaxbt/ai-market-maker/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/olaxbt-ai-market-maker
