Maze: one run_id across the SDK, LangGraph and a DAG editor
A distributed framework for LLM agents
At a glance
- What is it?
- Maze is a Python framework that turns agent programs into distributed workflows on top of Ray, with separate gpu, cpu and io queues, model wait modelled as its own state, and a single Run object that the SDK, the LangGraph adapter and the Workbench all share. Version 1.0.2 is on PyPI, but the repository publishes no releases.
- Who is it for?
- Use Maze when your agent work is genuinely a DAG spread over more than one machine and you need runs that survive a restart, with separate queues so a model task cannot starve an IO task. It is the wrong amount of machinery for a single process agent, and it asks you to run a Head node, workers and Ray before you write anything.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One run_id is the only identity anything shares
The central claim is convergence. The Python SDK, the LangGraph adapter, the Workbench editor and application specifications all submit through the same `maze.workflow/v1` contract and receive the same Core-owned Run. A Core `run_id` is the public run identity, and static workflows, persisted DynamicRuns and application specs share one set of APIs for snapshots, events, logs, artifacts, cancellation and retry.
That is the architectural point. Rather than maintaining separate execution paths per client, Maze makes the specification the boundary and the Run the unit of state. A DAG written with the `@workflow` decorator, a LangGraph graph, and a workflow typed into the Workbench are the same kind of object once submitted.
The layered diagram underneath runs client to `maze.workflow/v1` to a submit endpoint, into Maze Core for runs, events, logs and artifacts, then out to a scheduler and Ray with head and worker nodes. Everything above Ray, meaning the workflow contracts, resource semantics, durable state, scheduling, model lifecycle, artifacts and operational APIs, is what this project contributes.
gpu, cpu and io queues are separate so nothing blocks anything
Independent `gpu`, `cpu` and `io` queues exist for a specific reason: so one resource class cannot block another. An agent workflow that is mostly waiting on a disk write should not sit behind a queue of tasks waiting on a GPU, and a video job should not delay a text call that needs no accelerator at all.
The scheduler orders tasks that are ready, and the placement strategy separately chooses which registered node executes one. Two plugins ship: `FCFS`, first come first served, which is the default, and `HACS`, the algorithm aligned with the project's research paper. Node placement is chosen independently, so you can run HACS ordering without giving up least-loaded placement, and vice versa.
Warm standby workers are on by default, which means a worker is already waiting when work arrives rather than being created at the moment it is needed. Queue diagnostics are exposed alongside cluster resources, which is how you tell a genuinely idle queue from one whose tasks are parked somewhere else.
Model wait is a state, not a queue item
This is the detail that distinguishes Maze from a plain Ray deployment. When a task needs a local model instance that has not been deployed yet, the task waits for routing. That wait is explicit, and the rule is that a waiting task is not counted as a dispatchable GPU, CPU or I/O queue item. When routing becomes available the task returns to its resource queue rather than sitting outside the accounting.
Without that rule, a workload that spins up local vLLM or Transformers instances would look busy while doing nothing dispatchable. With it, the queue depths mean something, and scheduling a CPU task is not blocked by a model that has not finished loading.
Model execution itself is handled as a lifecycle. Maze discovers local checkpoints, deploys reusable vLLM or Transformers instances on demand, routes model tasks to them, and manages GPU reservations together with scale-in and scale-out. The 2026-08 news entry records that this path was validated with both engines, including automatic deployment, explicit model-wait state, cancellation and deterministic GPU cleanup.
FCFS is the default, and HACS is a flag plus five variables
Turning on the paper's scheduler is one extra flag on the Head, independent of placement:
maze start --head \
--port 8000 \
--strategy least-loaded \
--scheduling-algorithm HACSThe behaviour is then tuned by five environment variables, each with a documented default and a constraint. `MAZE_HACS_ALPHA` defaults to 2 and must be greater than 0. `MAZE_HACS_BETA` defaults to 5 and must be greater than 1. `MAZE_HACS_INITIAL_DCT_SECONDS` defaults to 60 and must be greater than 0. `MAZE_HACS_DCT_EMA_ALPHA` defaults to 0.2 and must be greater than 0 and at most 1. `MAZE_HACS_STARVATION_SECONDS` defaults to 600 and must be greater than 0.
Mechanically, HACS refreshes ready-task priorities at dispatch time and maintains its DCT exponential moving average from the durations of completed workflows. The starvation threshold is the one to watch, because it is what stops a low-priority task from waiting indefinitely behind newer arrivals.
The launcher checks its ports before it starts anything
Starting the whole stack detached is one command:
maze start --head --port 8000 --playground --detachAfterwards the Workbench is on port 5173 and the Core API on port 8000. Maze validates its configured ports before startup rather than failing later, and it prints the path of the detached log, which matters because a detached process that fails silently is the default outcome otherwise.
The port arithmetic has one wrinkle worth knowing. Change the UI port with `--playground-port`, and the Workbench backend defaults to `--playground-port + 1`; it uses 3001 when the UI port is left at its default. `--playground-backend-port` overrides that arithmetic outright when you need it to.
The same subcommands that start it manage it: `maze status`, `maze doctor` and `maze stop`. Warm standby workers can be switched off with `MAZE_STANDBY_WORKERS_ENABLED=0` passed to the same start command.
A Ray node still has to register as a Maze worker
Ray provides the distributed execution underneath, and joining a cluster is a two-step thing rather than one. You start a Maze worker that re-registers periodically, which is what lets it survive a Head or Ray restart:
maze start --worker \
--addr HEAD_IP:8000 \
--agent \
--heartbeat-interval 20The requirement is explicit: a Ray node must also register as a Maze worker before Maze will schedule tasks to it. Being in the Ray cluster is not enough. `maze stop --worker` stops a local worker.
Cluster inspection and repair are four subcommands, all pointed at the Head:
maze cluster resources --server-url http://HEAD_IP:8000
maze cluster queues --server-url http://HEAD_IP:8000
maze cluster join-command --server-url http://HEAD_IP:8000
maze cluster reconcile-workers --server-url http://HEAD_IP:8000`join-command` hands back the exact command to run on a new machine, and `reconcile-workers` is the repair path when the cluster's view and reality disagree.
Runs keep state across process restarts
Durability is the feature the rest of the design serves. A run retains task state, structured errors, events, logs, retries, timeouts, cancellation, placement and content-addressed artifacts, and it keeps them across process restarts.
Content-addressed artifacts are the part with the longest reach. Because an artifact is named by what it contains rather than by when it was produced, a rerun can reference the same output without regenerating it, and the June news entry records content-addressed artifacts and runtime fault-tolerance traces arriving together.
The hardening work of July 2026 is the other half: worker re-registration, run-level deadlines, explicit scheduler-failure states, and restart-safe Run discovery. Each of those closes a case where a dead component previously left a run in an unknown state, which is exactly the failure mode that makes distributed agent frameworks hard to trust in production.
The package pins Ray exactly and ships a placeholder author email
The PyPI name is `maze-agent` and the version in pyproject.toml is 1.0.2, installed with `pip install maze-agent` or from source with a clone and `pip install -e .`. Python support is bounded at both ends, `>=3.10,<3.14`, and the classifier says Beta with an audience of developers and researchers.
The dependency list is where to look before you deploy. Ray is pinned to exactly 2.50.1 rather than a range, networkx to 3.4.2, cloudpickle to 3.1.1, and websocket-client to 1.9.0, while the model stack is looser: transformers from 5.12.1, accelerate from 1.14.0, safetensors from 0.8.0 and openai from 2.15.0. vLLM is not a base dependency at all; it lives in an `inference-vllm` extra pinned to 0.26.0.
Two rough edges sit alongside. The author and maintainer fields both say Maze Development Team with the placeholder address [email protected], and the repository has no GitHub releases despite shipping 1.0.2 to PyPI, so there is no tag to correlate a production install with.
Editorial conclusion
Use Maze when your agent work is genuinely a DAG spread over more than one machine and you need runs that survive a restart, with separate queues so a model task cannot starve an IO task. It is the wrong amount of machinery for a single process agent, and it asks you to run a Head node, workers and Ray before you write anything. Before you tune HACS, read the environment variable table rather than the flag, because the starvation threshold and the DCT smoothing constant change behaviour more than the algorithm switch does.
Frequently asked questions
What is the Maze framework for LLM agents?
A Python framework that turns agent programs into distributed, observable workflows, scheduling task-level work across heterogeneous resources while execution, recovery, model serving and artifacts sit behind one runtime API. It runs on Ray and is MIT licensed.
How do I start Maze and the Workbench?
Run maze start --head --port 8000 --playground --detach. The Workbench is then on http://localhost:5173 and the Core API on http://localhost:8000. Maze validates its ports before startup and prints the detached log path; manage the service with maze status, maze doctor and maze stop.
How do I add a worker node to Maze?
Start a Maze worker with maze start --worker --addr HEAD_IP:8000 --agent --heartbeat-interval 20 so it re-registers after a Head or Ray restart. A Ray node must also register as a Maze worker before Maze schedules tasks to it, and maze stop --worker stops a local one.
What is HACS in Maze and how do I enable it?
A scheduling plugin aligned with the project's SC26 paper, enabled with --scheduling-algorithm HACS alongside a placement strategy such as least-loaded. It refreshes ready-task priorities at dispatch time and maintains a DCT exponential moving average from completed workflow durations, tuned by five MAZE_HACS environment variables.
How does Maze handle a task that is waiting for a local model?
Model wait is explicit. A task waiting for a local model instance is not counted as a dispatchable gpu, cpu or io queue item, and it returns to its resource queue once routing is available, so model loading does not distort queue depths.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/maze-agent-maze)