# Onyx (onyx-dot-app/onyx): a self-hosted AI chat and RAG platform you run yourself

> Onyx is the application layer for LLMs: a chat interface with agentic RAG, connectors and agents that you can host on your own hardware. The README documents a one-line installer, but the real decision is between Lite mode and the full stack.

**onyx-dot-app/onyx** — Open Source AI Platform - AI Chat with advanced features that works with every LLM

- Repository: https://github.com/onyx-dot-app/onyx
- Website: https://onyx.app
- Stars: 32,296 · Forks: 4,515
- Language: Python
- License: NOASSERTION
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/onyx-dot-app-onyx

## What Onyx solves, and who it is actually for

Onyx describes itself as "the application layer for LLMs", which is a precise way of drawing its boundary. It is not a model and not a model server. It is the interface and the retrieval machinery around a model you already have, whether that model runs on your own hardware through Ollama, LiteLLM or vLLM, or is reached through Anthropic, OpenAI or Gemini. The README states that Onyx supports all major LLM providers, both self-hosted and proprietary.

The problem it addresses is the gap between a raw model endpoint and something a team can use. A team that wants answers grounded in its own documents needs an index, a connector to populate that index, a chat surface, and some way to control who sees what. Onyx bundles those pieces: over 50 indexing connectors, an agentic RAG pipeline, custom agents, web search, code execution, artifacts and voice mode are all listed as features. The intended audience is stated fairly openly: individual users through to large enterprises, with collaboration, SSO and RBAC called out as enterprise concerns.

The honest reading is that Onyx is for organisations that have decided their data should not leave their infrastructure, or that want to choose the model behind the chat window. If neither of those matters to you, the deployment cost is hard to justify.

## How the indexing and retrieval path is put together

The repository layout tells you more about the architecture than the feature list does. There is a backend/ directory for the Python service, a web/ frontend, a model_server for inference, a cli/ package that ships as onyx-cli, plus desktop/, mobile/ and widget/ targets. The pyproject.toml confirms the split: the shared dependency set includes fastapi, litellm, openai, google-genai, cohere and voyageai, while the backend dependency group adds the connector and job machinery, including celery, dask and distributed, asyncpg for Postgres, and a long list of per-service SDKs such as atlassian-python-api, boxsdk, jira, discord.py and google-api-python-client.

That dependency list is the mechanism in miniature. Connectors pull documents from external systems, a job queue distributes the sync work across workers, and the results land in a vector plus keyword index. The README describes the retrieval side as "hybrid index + AI Agents for information retrieval", meaning the query path mixes vector similarity with keyword matching and then lets an agent decide how to search. The README notes that a benchmark is "to release soon", so the quality claim is currently unverified by any published number.

The README also states that Standard mode adds components that Lite mode leaves out: the vector and keyword index, background containers for job queues and connector syncing, model inference servers for indexing and inference, and Redis plus MinIO for caching and blob storage. Lite mode, by contrast, is described as a lightweight Chat UI that runs under 1GB of memory. In other words, the retrieval pipeline is not a configuration flag. It is a set of services you either run or do not.

## Installing Onyx and getting to a first answer

The README gives a single deployment command. It pipes a script from onyx.app into bash, which means you should read that script before running it, but it is the documented path.

```bash
curl -fsSL https://onyx.app/install_onyx.sh | bash
```

After the script finishes, the README does not spell out the URL or port the UI appears on, so check the output of the installer and the deployment docs at docs.onyx.app/deployment/overview. The README does direct you there for "Docker, Kubernetes, Helm/Terraform" guides and for major cloud providers.

If you would rather work from the repository, the Makefile shows the development path. The craft targets build and load images into a local kind cluster named onyx-dev.

```bash
make craft-up
make craft-backend-image
```

The first target runs deployment/helm/dev/craft-up.sh; the second builds onyxdotapp/onyx-backend:dev, loads it into kind, and restarts the onyx-sandbox-proxy and onyx-api-server deployments in the onyx namespace. There is also make craft-refresh-images, which rebuilds the backend and sandbox images and refreshes the PodTemplate, and make craft-check-images, which runs refresh-images.sh with a --check flag.

Once the UI is up, the first real use is to point a connector at a document source and then ask a question about it. The README does not walk through the connector UI, so the docs are the place to look. Note that this flow only exists in Standard mode: Lite mode has no index to populate.

## The Lite versus Standard split is the real decision

Most self-hosted AI tools make you guess at the resource cost. Onyx states it. Lite mode is "a lightweight Chat UI" that requires under 1GB of memory and runs a less complex stack, and the README recommends it for people who want to test Onyx quickly or who only care about the Chat UI and Agents. Standard mode is "the complete feature set", recommended for serious users and larger teams.

The trade-off is stark and worth stating plainly: if you deploy Lite, you do not get RAG. There is no vector index, no keyword index, no connector sync workers, no inference servers for indexing, and no Redis or MinIO. A Lite deployment is a chat front end with agents bolted on. Anyone who reads "Onyx enables LLMs through advanced capabilities like RAG" and then installs the small version will find the feature missing, and the README's own wording is the only warning.

The other constraint is the stack itself. Standard mode brings Celery workers, Dask and distributed, Postgres via asyncpg, Redis and MinIO, plus model inference servers. That is a multi-service deployment with a job queue, and job queues fail in ways that chat front ends do not: a connector sync can stall, a worker can fall behind, and the index can drift from the source system. The README does not document rollback or index repair, and it does not describe what happens when a sync fails partway. Treat connector health as an operational surface you own.

## Where Onyx is the wrong tool

Onyx is a poor fit if you want a hosted assistant and nothing else. The README points to Onyx Cloud for people who want to try it without deploying, but the project's centre of gravity is self-hosting, and the value only appears when you have data to index and a reason to keep it in-house.

It is also a poor fit for a single developer who wants a local chat window. Lite mode is small, but it still runs a container stack, and the pyproject.toml pins Python 3.13 or newer. If all you need is a local model with a text box, a desktop client talking directly to Ollama is less machinery for the same outcome.

A third case is the team that needs only retrieval over one source. Onyx's connector breadth is its selling point, and that breadth is also surface area: each connector is another SDK pinned in the dependency group, another OAuth flow, another thing that can break on an upstream API change. A single-source retrieval need is often better served by a smaller pipeline you can read end to end.

Finally, the licence boundary matters. The README states there are two editions: Community Edition under MIT, covering core Chat, RAG, Agents and Actions, and an Enterprise Edition with features aimed at larger organisations. The repository's licence field is reported as NOASSERTION even though the README badge says MIT, so if you are evaluating the codebase rather than the marketing page, read LICENSE and the pricing page before you plan around a feature.

## How Onyx differs from wiring your own RAG stack

The obvious alternative is not another product but the do-it-yourself route: a vector database, an embedding model, a retrieval script and a chat UI, assembled from parts. The difference in approach is where the work sits. A hand-built stack gives you full control over chunking, ranking and prompt construction, and you can read every line. It also gives you every maintenance task: connector auth, incremental sync, permission filtering, and the UI.

Onyx takes the opposite position. It pre-commits to a hybrid index and to an agent that decides how to retrieve, and it ships the connectors and the interface. You trade control over the retrieval internals for not having to build them. The README's claim that search and answer quality is "best in class" is not backed by a published benchmark in the repository, so the honest framing is that you are buying integration, not a measured quality advantage.

The second alternative is a managed assistant that indexes your documents for you. That removes the operational load entirely, and for many teams it is the correct answer. The reason to choose Onyx instead is the reason the project exists: the model, the index and the connectors stay on infrastructure you control, and you can point the chat at a self-hosted model through Ollama, LiteLLM or vLLM.

## Maintenance, releases and licence cost

The repository is not archived and the last push was on 2026-09-09, so the project is being worked on. Recent releases show two parallel tracks: v4.7.1 on 2026-09-08 for the main application and cli/v1.4.1 on 2026-09-09 for onyx-cli, with a cli-latest tag from 2026-07-29. The CLI and the server version separately, which is worth knowing before you file a bug against the wrong component.

Upgrade cost depends on which mode you run. A Lite deployment is a small container set and an upgrade is close to a restart. A Standard deployment carries Postgres, Redis, MinIO and a job queue, and the pyproject.toml shows how tightly the backend pins its dependencies: litellm==1.93.0, fastapi==0.133.1, pydantic==2.12.5, celery==5.5.1. Pinned versions make builds reproducible and make upstream security bumps a deliberate act rather than an automatic one. The README does not describe a migration procedure for the database schema, though alembic appears in the backend dependency group, which suggests migrations exist.

On licensing, the README splits the product into Community Edition under MIT and an Enterprise Edition with extra features. MIT is permissive, but the split means some capabilities you may see in the interface are not covered by that licence. This is not legal advice; read LICENSE and the pricing page, and if you are deploying commercially, have someone confirm which edition you are actually running.

## Conclusion

Adopt Onyx if you need a chat interface with RAG over your own connectors and you are willing to operate a container stack; the Community Edition is MIT licensed and the README's install command is the fastest way to see whether the Lite mode fits. Do not adopt it if you only want a hosted chatbot or if you cannot run Docker or Kubernetes. Before committing, verify which features fall under the Enterprise Edition rather than the MIT Community Edition, and confirm the resource profile of Standard mode against the under 1GB figure the README gives for Lite.

## FAQ

### What is the Onyx app?

Onyx describes itself as the open source AI platform and the application layer for LLMs, providing a chat interface with RAG, agents, web search, code execution and artifacts. It supports self-hosted and proprietary model providers and ships over 50 indexing connectors.

### Can I self-host Onyx?

Yes. The README documents a single-command install and states that Onyx supports deployments in Docker, Kubernetes, Helm/Terraform, with guides for major cloud providers. It also offers a hosted option through Onyx Cloud for people who would rather not deploy.

### Is Onyx free to use?

The README states that the Onyx Community Edition is available freely under the MIT license and covers the core features for Chat, RAG, Agents and Actions. An Enterprise Edition exists with extra features aimed at larger organisations, and the pricing page lists the difference.

### Is Onyx AI free?

The Community Edition is MIT licensed and free according to the README, while the Enterprise Edition carries additional features for larger organisations. The repository's licence field is reported as NOASSERTION, so check the LICENSE file and the pricing page rather than the badge alone.

## Sources

- [Official documentation](https://onyx.app)
- [Official README](https://github.com/onyx-dot-app/onyx#readme)
- [Project repository](https://github.com/onyx-dot-app/onyx)
- [Release notes](https://github.com/onyx-dot-app/onyx/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/onyx-dot-app-onyx
