# GitHamza0206/simba: a self-hosted customer service assistant with evaluation built in

> Simba is an Apache-2.0 RAG customer service stack: a FastAPI backend, Celery ingestion workers, Qdrant or FAISS retrieval, and an npm chat widget. The evaluation layer is the reason to look at it, and the operational surface is the reason to hesitate.

**GitHamza0206/simba** — OpenSource Production ready Customer service with built in Evals and monitoring 

- Repository: https://github.com/GitHamza0206/simba
- Website: https://simba.mintlify.app
- Stars: 1,543 · Forks: 115
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/githamza0206-simba

## The problem Simba targets: you cannot tell whether your support bot got worse

Most open source customer service bots ship a retrieval pipeline and a chat UI, then leave quality measurement to the operator. Simba inverts that. The README states the project is "designed from the ground up around evaluation and customization", and the repository layout backs the claim: there is a simba package, a separate simba_sdk package, a frontend, and a packages directory holding the npm widget. The intended user is a team that already has support documentation and wants to know whether a change to chunking, embeddings or the LLM actually improved answers.

The metrics listed are retrieval precision, recall and relevance scores, plus generation faithfulness, answer relevancy and latency, plus conversation analytics such as user satisfaction and resolution rates. The repository does not include a public benchmark, so the value here is the presence of the measurement hooks, not a demonstrated accuracy number. Treat it as instrumentation you would otherwise have to build yourself.

## What actually runs: FastAPI, Celery, Redis, Postgres and a vector store

The architecture diagram in the README shows a browser widget calling a FastAPI service, which queries a vector store (Qdrant or FAISS) and an LLM (OpenAI or a local model). Document ingestion is asynchronous: a Celery worker pulls from Redis and writes into the vector store. The Makefile confirms the queue name, running the worker with -Q ingestion, and lists ports for the supporting services: server on 8000, Redis on 6379, Postgres on 5432, Qdrant on 6333, MinIO on 9000 with its console on 9001.

Two details in pyproject.toml matter more than the diagram. First, the project depends on langgraph-checkpoint-postgres and psycopg, which the comment says is "Required for async postgres checkpointer", so Postgres is not optional state storage for the agent graph. Second, Docling is the default document parser, with mistral and unstructured offered as extras and switchable through the PARSER_BACKEND variable in .env.example. That means the ingestion path carries a heavy parsing dependency by default.

## Installing Simba with Docker and taking the first real step

The README calls Docker the recommended path. Clone the repository, create a .env file with an OpenAI key, then build and start. The DEVICE variable selects CPU or NVIDIA GPU images.

```bash
git clone https://github.com/GitHamza0206/simba.git
cd simba
```

```bash
OPENAI_API_KEY=your_openai_api_key
```

```bash
DEVICE=cpu make build && make up
```

After the build completes, the README says the dashboard is at http://localhost:3000. If you prefer to run it without Docker, the package published to PyPI is simba-core, and the README gives two entry points:

```bash
pip install simba-core
```

```bash
simba server
simba front
```

Note the Python constraint in pyproject.toml: requires-python is ">=3.11,<3.13". A 3.13 interpreter will not install the package. The repository also ships a .python-version file, so check it before building a virtualenv by hand. The README's Docker path is the one with the most supporting detail; the manual path assumes you already have Redis, Postgres, Qdrant and MinIO reachable through the variables in .env.example.

## Embedding the chat widget in a site, and what the npm package assumes

Website integration is the shortest part of the README. Install the widget and render it with an apiUrl pointing at your Simba instance:

```bash
npm install simba-chat-widget
```

```jsx
import { SimbaChat } from 'simba-chat-widget';

function App() {
  return (
    <SimbaChat
      apiUrl="https://your-simba-instance.com"
      theme="light"
    />
  );
}
```

The README shows only these two props. There is no documented authentication flow for the widget, no per-user identity parameter, and no rate limiting mentioned. For a public marketing site that is probably fine. For a logged-in product where the assistant should answer differently per account, the missing identity parameter is a gap you would have to close in the backend yourself, and the README does not describe how.

## Where Simba is the wrong tool: single tenancy and an unbuilt roadmap

The roadmap is the clearest statement of limits. Multi-tenant support, an advanced analytics dashboard, webhook integrations and a fine-tuning pipeline are all unchecked. If you are building a SaaS where each customer needs an isolated knowledge base, the current code does not advertise that capability, and bolting it on means touching the vector store, the Postgres schema and the API surface at once.

The second limitation is operational weight. A working deployment needs Redis, Postgres, a vector store, MinIO for object storage, and a Celery worker process, before you count the frontend. The Makefile has a services target that starts only the infrastructure, which tells you the maintainers expect you to run those pieces locally during development. If your team has no one comfortable owning a Celery queue and a Postgres checkpointer, a managed chatbot service will cost less in engineering time than this stack does.

A third point: the README's customization table lists Cohere, ColBERT and cross-encoder rerankers, and sentence-transformers is a declared dependency for cross-encoder reranking, but the README does not document the configuration keys that select a reranker. The switch exists in the dependency graph; the operator-facing documentation for it is not in the README.

## How Simba differs from a hosted assistant and from a bare RAG template

The obvious alternative is a hosted customer service product, where you upload documents and get a widget with no servers. The difference is not features, it is where the pipeline lives. With Simba you hold the vector store, so you can swap Qdrant for FAISS or Chroma, swap OpenAI for Anthropic or a local model, and change the parser through PARSER_BACKEND. You also own the uptime, the Redis queue and the Postgres backups. A hosted product removes that work and removes that control in the same move.

The closer comparison is a generic RAG template built on LangChain. Simba is assembled from the same parts (langchain, langgraph, qdrant-client, fastapi are all declared dependencies), so the code you would write yourself overlaps heavily. What you get by adopting it is the assembled shape: an ingestion queue separated from the request path, a Postgres-backed graph checkpointer, a published npm widget, and evaluation metrics wired in from the start. If you have already written that plumbing, Simba's advantage shrinks to the widget and the metrics, and you should read the simba and simba_sdk packages before deciding.

## Licence, maintenance window and upgrade cost

Simba is Apache-2.0, declared in both pyproject.toml and LICENSE.md. That permits commercial use and modification, and it does not oblige you to publish your changes. It also means no vendor can relicense the code out from under you. The usual Apache-2.0 caveats apply: the licence text governs, and this is not legal advice.

The repository is not archived, and the last push was on 2026-06-18. The most recent release in the list is v0.4.0 from 2025-03-11, while pyproject.toml declares version 0.5.0, so the package metadata is ahead of the tagged releases. If you pin to a tag, expect to be behind main.

Upgrade cost is driven by the dependency set rather than by Simba itself. LangChain, LangGraph, FastAPI, Docling and the vector store clients all move quickly, and the project pins only lower bounds (for example langchain>=0.3.14, docling>=2.15.0). There is a uv.lock and a pnpm-lock.yaml, so a lockfile-based install is reproducible, but a fresh resolve can pull a newer LangChain than the code was written against. The classifier in pyproject.toml says "Development Status :: 4 - Beta", which is the honest label for the state of the API.

## Conclusion

Adopt Simba if you are an engineering team that already runs Postgres, Redis and a vector store, and you care about measuring retrieval and generation quality rather than shipping a chat box this week. Do not adopt it if you need multi-tenant isolation, webhook integrations or a fine-tuning pipeline, all of which the README lists as unchecked roadmap items, or if you want a hosted product with no infrastructure to run. Verify first that the Python version constraint in pyproject.toml (>=3.11,<3.13) matches your runtime, and confirm the Makefile targets on your machine before trusting the README quick start.

## FAQ

### How do I install Simba?

The README recommends Docker: clone the repository, create a .env file with OPENAI_API_KEY, then run DEVICE=cpu make build && make up, or DEVICE=cuda for an NVIDIA GPU. Without Docker, the published package is simba-core and the README gives simba server and simba front as the entry points.

### Which Python version does Simba require?

pyproject.toml sets requires-python to ">=3.11,<3.13", so 3.11 and 3.12 are supported and 3.13 is not. The repository also includes a .python-version file.

### What services does Simba need to run?

The Makefile lists Redis on 6379, Postgres on 5432, Qdrant on 6333 and MinIO on 9000 with its console on 9001, plus a Celery worker consuming the ingestion queue. The API server runs on 8000 and the README points the dashboard at localhost:3000.

### Does Simba support multiple tenants?

No. Multi-tenant support appears in the README roadmap as an unchecked item, alongside an advanced analytics dashboard, webhook integrations and a fine-tuning pipeline.

### Can I swap the LLM or vector store in Simba?

The README's customization table lists Qdrant, FAISS and Chroma for the vector store, OpenAI, Anthropic and local models for the LLM, and Cohere, HuggingFace and OpenAI for embeddings. The README does not document the configuration keys for each swap.

## Sources

- [GitHamza0206/simba on GitHub](https://github.com/GitHamza0206/simba)
- [License: Apache-2.0](https://github.com/GitHamza0206/simba/blob/main/LICENSE)
- [Project website](https://simba.mintlify.app)
- [README](https://github.com/GitHamza0206/simba/blob/main/README.md)
- [Releases](https://github.com/GitHamza0206/simba/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/githamza0206-simba
