GitHamza0206/simba: a self-hosted customer service assistant with evaluation built in
OpenSource Production ready Customer service with built in Evals and monitoring
At a glance
- What is it?
- Simba is an Apache-2.0 RAG customer service stack: a FastAPI backend, Celery ingestion workers, Qdrant or FAISS retrieval, and an npm chat widget. The evaluation layer is the reason to look at it, and the operational surface is the reason to hesitate.
- Who is it for?
- Adopt Simba if you are an engineering team that already runs Postgres, Redis and a vector store, and you care about measuring retrieval and generation quality rather than shipping a chat box this week. Do not adopt it if you need multi-tenant isolation, webhook integrations or a fine-tuning pipeline, all of which the README lists as unchecked roadmap items, or if you want a hosted product with no infrastructure to run.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 104 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Simba targets: you cannot tell whether your support bot got worse
Most open source customer service bots ship a retrieval pipeline and a chat UI, then leave quality measurement to the operator. Simba inverts that. The README states the project is "designed from the ground up around evaluation and customization", and the repository layout backs the claim: there is a simba package, a separate simba_sdk package, a frontend, and a packages directory holding the npm widget. The intended user is a team that already has support documentation and wants to know whether a change to chunking, embeddings or the LLM actually improved answers.
The metrics listed are retrieval precision, recall and relevance scores, plus generation faithfulness, answer relevancy and latency, plus conversation analytics such as user satisfaction and resolution rates. The repository does not include a public benchmark, so the value here is the presence of the measurement hooks, not a demonstrated accuracy number. Treat it as instrumentation you would otherwise have to build yourself.
What actually runs: FastAPI, Celery, Redis, Postgres and a vector store
The architecture diagram in the README shows a browser widget calling a FastAPI service, which queries a vector store (Qdrant or FAISS) and an LLM (OpenAI or a local model). Document ingestion is asynchronous: a Celery worker pulls from Redis and writes into the vector store. The Makefile confirms the queue name, running the worker with -Q ingestion, and lists ports for the supporting services: server on 8000, Redis on 6379, Postgres on 5432, Qdrant on 6333, MinIO on 9000 with its console on 9001.
Two details in pyproject.toml matter more than the diagram. First, the project depends on langgraph-checkpoint-postgres and psycopg, which the comment says is "Required for async postgres checkpointer", so Postgres is not optional state storage for the agent graph. Second, Docling is the default document parser, with mistral and unstructured offered as extras and switchable through the PARSER_BACKEND variable in .env.example. That means the ingestion path carries a heavy parsing dependency by default.
Installing Simba with Docker and taking the first real step
The README calls Docker the recommended path. Clone the repository, create a .env file with an OpenAI key, then build and start. The DEVICE variable selects CPU or NVIDIA GPU images.
git clone https://github.com/GitHamza0206/simba.git
cd simbaOPENAI_API_KEY=your_openai_api_keyDEVICE=cpu make build && make upAfter the build completes, the README says the dashboard is at http://localhost:3000. If you prefer to run it without Docker, the package published to PyPI is simba-core, and the README gives two entry points:
pip install simba-coresimba server
simba frontNote the Python constraint in pyproject.toml: requires-python is ">=3.11,<3.13". A 3.13 interpreter will not install the package. The repository also ships a .python-version file, so check it before building a virtualenv by hand. The README's Docker path is the one with the most supporting detail; the manual path assumes you already have Redis, Postgres, Qdrant and MinIO reachable through the variables in .env.example.
Embedding the chat widget in a site, and what the npm package assumes
Website integration is the shortest part of the README. Install the widget and render it with an apiUrl pointing at your Simba instance:
npm install simba-chat-widgetimport { SimbaChat } from 'simba-chat-widget';
function App() {
return (
<SimbaChat
apiUrl="https://your-simba-instance.com"
theme="light"
/>
);
}The README shows only these two props. There is no documented authentication flow for the widget, no per-user identity parameter, and no rate limiting mentioned. For a public marketing site that is probably fine. For a logged-in product where the assistant should answer differently per account, the missing identity parameter is a gap you would have to close in the backend yourself, and the README does not describe how.
Where Simba is the wrong tool: single tenancy and an unbuilt roadmap
The roadmap is the clearest statement of limits. Multi-tenant support, an advanced analytics dashboard, webhook integrations and a fine-tuning pipeline are all unchecked. If you are building a SaaS where each customer needs an isolated knowledge base, the current code does not advertise that capability, and bolting it on means touching the vector store, the Postgres schema and the API surface at once.
The second limitation is operational weight. A working deployment needs Redis, Postgres, a vector store, MinIO for object storage, and a Celery worker process, before you count the frontend. The Makefile has a services target that starts only the infrastructure, which tells you the maintainers expect you to run those pieces locally during development. If your team has no one comfortable owning a Celery queue and a Postgres checkpointer, a managed chatbot service will cost less in engineering time than this stack does.
A third point: the README's customization table lists Cohere, ColBERT and cross-encoder rerankers, and sentence-transformers is a declared dependency for cross-encoder reranking, but the README does not document the configuration keys that select a reranker. The switch exists in the dependency graph; the operator-facing documentation for it is not in the README.
How Simba differs from a hosted assistant and from a bare RAG template
The obvious alternative is a hosted customer service product, where you upload documents and get a widget with no servers. The difference is not features, it is where the pipeline lives. With Simba you hold the vector store, so you can swap Qdrant for FAISS or Chroma, swap OpenAI for Anthropic or a local model, and change the parser through PARSER_BACKEND. You also own the uptime, the Redis queue and the Postgres backups. A hosted product removes that work and removes that control in the same move.
The closer comparison is a generic RAG template built on LangChain. Simba is assembled from the same parts (langchain, langgraph, qdrant-client, fastapi are all declared dependencies), so the code you would write yourself overlaps heavily. What you get by adopting it is the assembled shape: an ingestion queue separated from the request path, a Postgres-backed graph checkpointer, a published npm widget, and evaluation metrics wired in from the start. If you have already written that plumbing, Simba's advantage shrinks to the widget and the metrics, and you should read the simba and simba_sdk packages before deciding.
Licence, maintenance window and upgrade cost
Simba is Apache-2.0, declared in both pyproject.toml and LICENSE.md. That permits commercial use and modification, and it does not oblige you to publish your changes. It also means no vendor can relicense the code out from under you. The usual Apache-2.0 caveats apply: the licence text governs, and this is not legal advice.
The repository is not archived, and the last push was on 2026-06-18. The most recent release in the list is v0.4.0 from 2025-03-11, while pyproject.toml declares version 0.5.0, so the package metadata is ahead of the tagged releases. If you pin to a tag, expect to be behind main.
Upgrade cost is driven by the dependency set rather than by Simba itself. LangChain, LangGraph, FastAPI, Docling and the vector store clients all move quickly, and the project pins only lower bounds (for example langchain>=0.3.14, docling>=2.15.0). There is a uv.lock and a pnpm-lock.yaml, so a lockfile-based install is reproducible, but a fresh resolve can pull a newer LangChain than the code was written against. The classifier in pyproject.toml says "Development Status :: 4 - Beta", which is the honest label for the state of the API.
Editorial conclusion
Adopt Simba if you are an engineering team that already runs Postgres, Redis and a vector store, and you care about measuring retrieval and generation quality rather than shipping a chat box this week. Do not adopt it if you need multi-tenant isolation, webhook integrations or a fine-tuning pipeline, all of which the README lists as unchecked roadmap items, or if you want a hosted product with no infrastructure to run. Verify first that the Python version constraint in pyproject.toml (>=3.11,<3.13) matches your runtime, and confirm the Makefile targets on your machine before trusting the README quick start.
Frequently asked questions
How do I install Simba?
The README recommends Docker: clone the repository, create a .env file with OPENAI_API_KEY, then run DEVICE=cpu make build && make up, or DEVICE=cuda for an NVIDIA GPU. Without Docker, the published package is simba-core and the README gives simba server and simba front as the entry points.
Which Python version does Simba require?
pyproject.toml sets requires-python to ">=3.11,<3.13", so 3.11 and 3.12 are supported and 3.13 is not. The repository also includes a .python-version file.
What services does Simba need to run?
The Makefile lists Redis on 6379, Postgres on 5432, Qdrant on 6333 and MinIO on 9000 with its console on 9001, plus a Celery worker consuming the ingestion queue. The API server runs on 8000 and the README points the dashboard at localhost:3000.
Does Simba support multiple tenants?
No. Multi-tenant support appears in the README roadmap as an unchecked item, alongside an advanced analytics dashboard, webhook integrations and a fine-tuning pipeline.
Can I swap the LLM or vector store in Simba?
The README's customization table lists Qdrant, FAISS and Chroma for the vector store, OpenAI, Anthropic and local models for the LLM, and Cohere, HuggingFace and OpenAI for embeddings. The README does not document the configuration keys for each swap.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/githamza0206-simba)