RESTai: a self-hosted AIaaS platform for RAG, agents and REST endpoints
RESTai is an AIaaS (AI as a Service) open-source platform. Supports many public and local LLM suported by Ollama/vLLM/etc. Precise embeddings usage, tuning, analytics etc. Built-in image/audio generation with dynamic loading generators. Live chat deployment. Built-in block based graphical language. Prompt versioning and much more...
At a glance
- What is it?
- RESTai wraps LLM providers, vector stores and a React dashboard into one service you can reach over HTTP. It is aimed at teams that want RAG and agents behind an API without assembling the stack themselves.
- Who is it for?
- Adopt RESTai if you want a self-hosted HTTP layer over LLMs, vector stores and a dashboard, and you are willing to run Python 3.11 to 3.13 plus Postgres or MySQL for anything beyond a demo. Do not adopt it if you need a minimal library you can embed in an existing application, or if you cannot keep the web server and the cron sidecar pointed at the same database and staging volume.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 12 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Who RESTai is for, and the problem it removes
Building a RAG endpoint usually means gluing together a document loader, a chunker, an embedding model, a vector store, a retriever, a prompt template, and a web framework. RESTai ships that assembly as a single service. The README describes it as an AIaaS platform where you "Create AI projects and consume them via a simple REST API", and the repository layout backs that up: main.py, database.py, migrations/, modules/, and a frontend/ directory holding a React app.
The intended user is a team that wants an internal LLM service with a UI, not a library to embed. Projects are the unit of configuration, and each one carries its own LLM, system prompt, tools and settings. The README lists RAG, Agents, Block (visual logic) and Inference as the project types. If your requirement is a Python function that answers questions over a folder of PDFs, this is a heavier answer than you need. If your requirement is a service other teams can call, with token accounting and per-project limits, the scope fits.
How a RESTai project is assembled at runtime
The service is a FastAPI application served by uvicorn, with SQLAlchemy over SQLite, MySQL or PostgreSQL, and Alembic migrations in migrations/ driven by alembic.ini. The dependency list in pyproject.toml shows the retrieval layer: llama-index-core plus the Chroma and Postgres vector store integrations, langchain and langchain-openai, and reranking described in the README as ColBERT or LLM-based.
A RAG project ingests documents into a vector store and answers queries through retrieval plus an LLM call. The optional Knowledge Graph layer adds a second index: the README states that when it is enabled, every ingested document runs through a NER pipeline (dslim/bert-base-NER by default) and extracted entities are persisted in a queryable graph. A custom postprocessor then boosts chunks whose source documents mention entities found in the query. Extraction runs as a background task, which is why the compose file includes a separate cron service reading the same database and the same /app/data volume.
That sidecar arrangement is the main architectural constraint. The compose comments warn that the cron runner "connects to the same DB and uses the same secrets/Fernet key", and that a staging path outside the shared volume "leaves every job stuck at queued". Ingestion is therefore not a purely in-process operation; it depends on a second process being alive and correctly configured.
Installing RESTai from PyPI and serving the first project
The README's shortest path installs the restai-core package, which the README notes includes the pre-built React frontend, so Node.js is not required. The three commands below create the database and admin user, apply migrations, and start the server on port 9000.
pip install restai-core
restai init # Create database + admin user
restai migrate # Run migrations
restai serve # -> http://localhost:9000/admin (admin / admin)To move off the default port and run multiple workers, the README shows an env file plus flags:
restai serve -e .env -p 8080 -w 4The .env.example file in the repository lists the keys you would put in that file. RESTAI_PORT defaults to 9000, RESTAI_DEFAULT_PASSWORD defaults to admin, and EMBEDDINGS_PATH points at ./embeddings/. Database credentials are the MYSQL_* and POSTGRES_* groups; if none are set, the compose file says the platform defaults to SQLite.
RESTAI_PORT=9000
EMBEDDINGS_PATH="./embeddings/"
RESTAI_DEFAULT_PASSWORD="admin"
POSTGRES_HOST="xxxxxxx"
POSTGRES_USER="xxxxxxx"
POSTGRES_PASSWORD="xxxxxxx"
POSTGRES_DB="xxxxxxx"The Docker route avoids Python entirely. The README gives a single command and notes the image is multi-arch for linux/amd64 and linux/arm64, and also published as ghcr.io/apocas/restai:latest:
docker run -p 9000:9000 apocas/restai:latestAfter the server is up, open /admin and log in with admin / admin, then change that password. Note that the docker-compose comments state provider credentials are no longer bootstrapped from environment variables: they live per-LLM in /admin/llms, encrypted at rest and scoped to teams. HF_TOKEN is the exception, because HuggingFace hub libraries read it directly. If you are migrating from an older deployment, that change affects how you configure providers.
The cron sidecar and the staging path are where deployments break
The most concrete failure mode visible in the repository concerns bulk ingestion. The docker-compose.yml comments explain that UPLOAD_STAGING_PATH defaults to a path on the shared /app/data volume because "the API service stages each upload here and the cron service reads it back to index it", and that a path outside the shared volume leaves jobs stuck at queued. The same comments apply to SQLITE_PATH: the cron sidecar and the web server must hit the same database file, and drift between the two environment blocks is described as capable of silently corrupting encrypted columns or seeding a parallel database.
That is a real operational constraint rather than a bug. It means a Kubernetes deployment needs the volume mounted into both workloads, and it means the cron service is not optional if you use bulk ingestion. A single-container deployment with a local SQLite file and no cron process will accept uploads but never index them.
The second limitation is scope. The README positions RESTai as a platform with a dashboard, teams, RBAC, OAuth/LDAP, TOTP 2FA and per-project rate limiting. If your actual need is a retrieval function inside an existing Python service, you are adopting a database schema, a migration chain, a frontend build and a background worker to get it. The pyproject.toml dependency set is correspondingly large, including Selenium, opencv-python-headless, python-pptx and google-cloud-aiplatform. That is the cost of the breadth.
RESTai compared with building on LlamaIndex directly
The natural alternative is to use the retrieval libraries RESTai itself depends on, llama-index-core and langchain, and write your own FastAPI routes. The difference is where the state lives. With LlamaIndex or LangChain alone, projects, prompts, users, token counts and rate limits are your problem; you write the models, the migrations and the endpoints. RESTai already has them, plus an admin UI and analytics for tokens, cost and latency per project.
That trade runs the other way too. A hand-written service has no cron sidecar requirement, no Fernet key that must match between two processes, and no admin password defaulting to admin. It also lets you choose a vector store without checking whether a llama-index-vector-stores-* package exists for it. RESTai's own pyproject.toml contains a comment explaining that the floor for one vector store integration is 1.6.2 rather than 0.6 because earlier releases import a symbol removed by weaviate-client 4.20. Dependency coupling like that is normal for a platform, but it is the kind of thing you inherit rather than manage. If your team already runs LangChain in production and only needs an endpoint, the platform may be more surface area than the problem.
Maintenance, upgrades and what the Apache-2.0 licence leaves open
The repository is not archived, and the last push was on 2026-09-08. Releases are frequent and versioned: v6.4.1 on 2026-08-09, v6.4.0 on 2026-08-04, v6.3.28 on 2026-06-28. The pyproject.toml version matches the latest release at 6.4.1. Python support is pinned to >=3.11,<3.14, so a 3.14 environment will not install the package.
Upgrades differ by install method. For PyPI the README gives two commands, an upgrade followed by a migration against the same env file:
pip install --upgrade restai-core
restai migrate -e .envFrom a source checkout, make update fetches the latest release tag, installs dependencies, runs migrations and rebuilds the frontend, and the README says it auto-detects GPU for GPU-specific dependencies. The Docker path is pinned by tag, and the README explicitly recommends :6.2.13, :6.2 or :6 over :latest. That advice matters more here than for a stateless service, because the compose file ties the database schema, the Fernet key and the staging volume together.
RESTai is Apache-2.0, the same licence as several of its dependencies. The licence permits commercial use and modification, and the README does not describe any separate enterprise edition or licence key. Note that the platform pulls in providers and model weights under their own terms, and the README does not document licensing for the models you attach. That is a question for your own review, not something the repository answers.
Editorial conclusion
Adopt RESTai if you want a self-hosted HTTP layer over LLMs, vector stores and a dashboard, and you are willing to run Python 3.11 to 3.13 plus Postgres or MySQL for anything beyond a demo. Do not adopt it if you need a minimal library you can embed in an existing application, or if you cannot keep the web server and the cron sidecar pointed at the same database and staging volume. Before committing, verify the migration path on a copy of your own data, confirm which vector store you will run in production, and check that your provider credentials survive the move to the per-LLM admin page.
Frequently asked questions
What is RESTai used for?
It is an AIaaS platform for creating AI projects and consuming them through a REST API. The README lists RAG, agents, a block-based visual language and inference as project types, with a React dashboard for analytics and administration.
How do I install RESTai?
The README gives three routes: pip install restai-core followed by restai init, restai migrate and restai serve; a source checkout using make install and make dev; or docker run -p 9000:9000 apocas/restai:latest. The PyPI package includes the pre-built frontend, so Node.js is not required.
Does RESTai need a separate database?
The docker-compose file states the platform defaults to SQLite if no database is configured, and the .env.example lists optional MYSQL_* and POSTGRES_* settings. The compose comments warn that the web server and the cron sidecar must point at the same database, since drift can corrupt encrypted columns or create a parallel database.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apocas-restai)