# LlamaDeploy: what run-llama/llama_deploy does now that it is deprecated

> LlamaDeploy was the deployment layer for LlamaIndex agentic workflows. The README now marks the project deprecated and points to llama-agents, so the real question is not how to adopt it but whether you already depend on it.

**run-llama/llama_deploy** — Deploy your agentic worfklows to production

- Repository: https://github.com/run-llama/llama_deploy
- Website: https://docs.llamaindex.ai/en/stable/module_guides/llama_deploy/
- Stars: 2,068 · Forks: 227
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/run-llama-llama-deploy

## The deprecation notice is the first thing to read

The README opens with a caution block that says, in its own words, that the project is deprecated and that workflows should be served with llama-agents instead. That single paragraph changes how every other fact about the repository should be read. The package still exists on PyPI, the source tree is still there, and the last push to main was on 2026-04-06, which is when v0.9.2 was also released. Nothing about that activity contradicts the deprecation notice. A release can ship after a project is declared deprecated, and this one did.

The practical consequence is that llama_deploy is a migration problem, not an adoption problem. If you are evaluating deployment tooling for a new agentic workflow today, the README itself tells you to look elsewhere. If you already run llama_deploy in production, the parts worth understanding are the ones you would have to replace: the service definitions, the message queue transport, and the control plane that ties them together. The repository layout still shows all of it, including docker/, templates/, examples/, and a system_diagram.png at the top level.

## What LlamaDeploy actually orchestrated

LlamaDeploy sat between a LlamaIndex workflow and the infrastructure that runs it. The pyproject.toml dependency list is the clearest description of the mechanism: fastapi and uvicorn for the HTTP surface, pydantic-settings for configuration, PyYAML for the deployment descriptor, gitpython, and llama-index-workflows. The optional extras name the transports directly, with kafka pulling aiokafka and kafka-python-ng, rabbitmq pulling aio-pika, and redis pulling the redis client. Observability is a separate extra that pulls OpenTelemetry packages including the Jaeger exporter.

That shape implies a control plane plus one or more services that communicate over a message queue, with a YAML file describing the deployment. The README does not spell out the message flow, and the truncated pyproject.toml does not list every observability dependency, so the exact routing between control plane and services is not something the repository front page documents. What can be said from the layout is that examples/ contains quick_start/, python_fullstack/, python_dependencies/, google_cloud_run/, and llamacloud/ directories, which suggests the intended paths were local development, a full-stack app, a Cloud Run deployment, and a LlamaCloud integration.

## Installing llama-deploy and running the quick start example

The package is published as llama-deploy on PyPI, and the README carries the PyPI version badge for that name. The project requires Python 3.10 or newer and below 4.0, according to the requires-python field in pyproject.toml. The README shows a uv badge, and the repository includes a uv.lock, so uv is the tooling the maintainers used.

The README does not publish install commands, so the package name from the PyPI badge is the only install detail the repository front page gives. The extras are named in pyproject.toml under [project.optional-dependencies]: kafka, rabbitmq, redis, and observability. Because the README truncates the observability list, check the installed metadata rather than assuming a particular exporter is included.

For a first run, the repository ships examples/quick_start/. The README does not reproduce its commands, so read the files in that directory before running anything. What the layout tells you is that a deployment involves a YAML descriptor plus at least one service, which is why PyYAML is a base dependency. Expect to start a control plane process and a service process separately, and expect them to need a shared broker if you are not using the default local path.

One caveat worth stating plainly: the README does not document rollback, and it does not document what happens to in-flight workflow runs when a service restarts. For a deployment tool, that silence matters more than the install steps.

## The transports and the operational surface you inherit

Choosing llama_deploy meant choosing a message queue. The extras make that explicit: Kafka, RabbitMQ, and Redis are each a separate dependency group, and the dev group pulls aio-pika, redis, aiokafka, and kafka-python-ng together so the test suite can exercise all three. That is a real constraint, not a configuration detail. A Kafka deployment carries Kafka's operational weight, including broker management and consumer group behavior. A Redis deployment is lighter to start and gives you fewer delivery guarantees to reason about.

The observability extra is where the project's ambition is most visible and least documented on the front page. OpenTelemetry instrumentation for asyncio plus a Jaeger exporter implies distributed tracing across services, which is the right instinct for a multi-service agent system. Whether every transport propagates trace context correctly is not something the README states, and the truncated dependency list leaves the exporter set incomplete. Treat tracing as something to verify against your own broker rather than something the repository promises.

The other inherited surface is the HTTP layer. fastapi and uvicorn mean the control plane exposes an API, and the presence of python-multipart suggests file upload handling somewhere in that API. websockets is a base dependency too, so streaming responses are part of the design. None of that is unusual, but it is all code you would be maintaining after the deprecation notice.

## Where LlamaDeploy is the wrong tool

The clearest case is any new project. The README states the project is deprecated and names llama-agents as the replacement for serving workflows. Starting a new deployment on a deprecated package means starting on a migration path.

The second case is a single-process workflow. If your agent runs inside one Python process and you are comfortable managing that process with systemd, a container runtime, or a platform like Cloud Run, the control plane and message queue in llama_deploy add moving parts without adding capability. The examples/google_cloud_run/ directory suggests the project itself anticipated that environment, but a single container running one workflow does not need a broker.

The third case is a team without queue operations experience. Kafka in particular is not something you adopt casually, and the optional extra makes it a one-line install that hides a substantial operational commitment. If nobody on the team has run a broker in production, the failure modes will arrive before the benefits do.

Finally, there is the documentation gap. The README is short, the deprecation notice is at the top, and the deeper material lives on the docs site under the LlamaIndex module guides. Anything the README does not cover, such as rollback behavior or restart semantics for in-flight runs, is a question you would have to answer by reading the source.

## llama-agents and the difference in approach

The README points to llama-agents, whose repository is run-llama/workflows-py, as the way to serve workflows. That is the only alternative the project itself endorses, and it is worth being precise about what changes. LlamaDeploy was built as a deployment layer around LlamaIndex workflows: a control plane, a YAML descriptor, and a message-queue transport chosen through extras. The replacement lives in the workflows repository itself, which suggests the serving concern moved closer to the workflow runtime rather than staying in a separate deployment package.

That is a meaningful shift for anyone with an existing deployment. The pieces you would have configured separately, the descriptor and the transport wiring, are the pieces most likely to look different in a runtime that owns workflow execution directly. The README does not provide a migration guide, and it does not claim API compatibility. Anyone moving from one to the other should read the replacement's documentation rather than porting a llama_deploy YAML file and expecting it to load.

For teams that want a general-purpose orchestrator instead of a LlamaIndex-specific one, the transports llama_deploy already supported point at the broader ecosystem: running workflows as services behind Kafka, RabbitMQ, or Redis is something a general orchestrator can also do, at the cost of writing the LlamaIndex integration yourself.

## Licence, maintenance and the cost of staying

The licence is MIT, per the LICENSE file at the repository root and the repository metadata. MIT is permissive: it allows commercial use, modification, and redistribution with the copyright notice and permission notice retained. It does not grant patent rights and it provides no warranty, which is standard for the licence and not a statement about this project. None of this is legal advice; check the LICENSE file and your own counsel if the distinction matters to you.

On maintenance, the facts are narrow. The repository is not archived. The last push to main was on 2026-04-06, and v0.9.2 was released the same day. The two prior releases, v0.9.0 and v0.9.1, landed in July 2025. That is the release cadence visible in the repository, and the README's deprecation notice is the more important signal about where effort is going.

Upgrade cost for an existing installation is the cost of pinning and auditing. llama-index-core is constrained to >=0.11.17,<0.14.0, and pydantic is excluded at exactly 2.10, which is an unusual pin that will surface as a resolver conflict if your environment already carries that version. The optional extras each pin their clients to a major range. Before upgrading anything, list which extras your services import, because dropping an extra that a service still needs will fail at import time rather than at startup.

## Conclusion

Read the deprecation notice before anything else: the README states the project is deprecated and directs workflow serving to llama-agents. Teams with an existing llama_deploy deployment should pin the current release, check which optional extras (kafka, rabbitmq, redis, observability) their services actually import, and treat any new workflow as a llama-agents project instead. Teams starting fresh should not install it at all, and anyone who needs a supported deployment path should verify the replacement's own documentation rather than assume the two APIs match.

## FAQ

### What does it mean to deploy an AI?

In this project's terms, it means running an agentic workflow as one or more services behind a control plane rather than inside a single script. The pyproject.toml dependencies show the shape: FastAPI and uvicorn for the HTTP surface, a message queue transport chosen through the kafka, rabbitmq or redis extras, and a YAML descriptor for the deployment.

### What is llama used for?

The repository covers llama_deploy specifically, a deployment layer for LlamaIndex workflows. It is built on llama-index-core and llama-index-workflows, and the README now marks it deprecated in favor of llama-agents for serving workflows.

### What does "deploy" mean?

Here it means taking a workflow defined with llama-index-workflows and running it as a service that a control plane can reach. The repository layout shows the supporting pieces: a docker/ directory, templates/, and five example deployments covering quick start, a full-stack app, Python dependencies, Google Cloud Run, and LlamaCloud.

## Sources

- [License: MIT](https://github.com/run-llama/llama_deploy/blob/main/LICENSE)
- [Project website](https://docs.llamaindex.ai/en/stable/module_guides/llama_deploy/)
- [README](https://github.com/run-llama/llama_deploy/blob/main/README.md)
- [Releases](https://github.com/run-llama/llama_deploy/releases)
- [run-llama/llama_deploy on GitHub](https://github.com/run-llama/llama_deploy)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/run-llama-llama-deploy
