# Jonex: a self-hosted multimodal parsing and knowledge engine for enterprise agents

> Jonex bundles document and video parsing with an ontology-driven knowledge layer and exposes it through a single gateway. It is worth a look if you want source-grounded retrieval you can run yourself, and it is not a drop-in library.

**yuezhiai/jonex** — All-in-One Multimodal Parsing Engine + Ontology-Powered, LLM Wiki-Driven AI-Ready Knowledge Engine

- Repository: https://github.com/yuezhiai/jonex
- Website: https://jonex.ai
- Stars: 1,228 · Forks: 224
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/yuezhiai-jonex

## The problem Jonex targets: raw content that never becomes queryable

Most teams that want retrieval-augmented generation end up assembling the same pipeline by hand: an OCR or ASR step, a chunker, an embedding service, a vector store, and a prompt wrapper. Each stage has its own config, and the provenance of an answer is usually lost somewhere between the chunker and the vector index. Jonex positions itself as that whole chain in one system. The README describes it as an "end-to-end enterprise AI knowledge platform that turns raw content into reusable knowledge services", covering ingestion, multimodal parsing, domain knowledge compilation, vector and graph indexing, source-grounded retrieval, feedback loops and business applications.

The intended user is not an individual developer prototyping a chatbot. The repository ships multi-tenant support (the badge says "Multi-Tenant Ready"), a demo tenant named tenant_jonex_demo, and a SECURITY.md with a production checklist. That is the shape of a platform an internal team would run for several business units, not a weekend project. The distinguishing claim is the ontology layer: the README states that "Ontology compiles domain reasoning into the knowledge layer before retrieval begins", which is a different bet from the usual approach of embedding everything and hoping similarity finds the right passage.

## How the pipeline is put together

The repository layout is the clearest evidence of the architecture. Top-level entries include api_gateway/, jonex_core/, capabilities/, mcp_server/, frontends/, deploy/ and scripts/. The Python package metadata in pyproject.toml declares three packaged namespaces: api_gateway, capabilities and jonex_core. Capabilities and a capability_runtime.example.yaml suggest a plugin model where parsing and other functions are registered rather than hardcoded.

On the storage side, the dependency list is explicit about what runs underneath: sqlalchemy and asyncpg for PostgreSQL, redis for cache, neo4j as the ontology graph database (the requirements.txt comment labels it as such), and jieba for Chinese word segmentation in ontology full-text search. Object storage is handled through boto3 and the Tencent COS SDK. Milvus appears in the Makefile's dev-infra-up target alongside PostgreSQL, Redis, etcd and MinIO. LightRAG and something called Atomic RAG are started by the dev-deps-up target.

Data flow, as the README's five-minute walkthrough describes it: you create a domain space, create a knowledge base, pick a parser profile for the content type, then upload files or connect a REST API or S3-compatible source. Parsing and knowledge compilation run, and you can then inspect parsing results, the compiled ontology, relationships and the knowledge graph before searching. All external APIs go through one gateway, which is why the login example carries an X-Tenant header.

## Installing Jonex with Docker Compose

Docker Compose is the documented path. The requirements listed are Docker Engine or Docker Desktop, Compose v2 with Buildx, make on macOS or Linux, and enough disk and time for a first build that pulls container images, Python dependencies and RAG models. Clone and initialize first:

```bash
git clone https://github.com/yuezhiai/jonex.git
cd jonex
make init
```

The init target creates deploy/.env for the platform, database, object storage and LLM Gateway, deploy/.env.rag for LightRAG, embeddings and parsing, deploy/.env.mcp for the MCP server, and frontend .env files. Then edit deploy/.env to point at at least one OpenAI-compatible LLM and embedding provider:

```env
LLMGW_UPSTREAM_LLM_HOST=https://your-openai-compatible-host/v1
LLMGW_UPSTREAM_LLM_API_KEY=your_llm_api_key

LLMGW_UPSTREAM_EMBED_HOST=https://your-embedding-host/v1
LLMGW_UPSTREAM_EMBED_API_KEY=your_embedding_api_key
```

Two consistency constraints matter here. EMBEDDING_MODEL must be identical in deploy/.env and deploy/.env.rag because it is used to build the vector index, and LIGHTRAG_API_KEY must match in the same two files. Build and start:

```bash
make build
make up
make ps
```

The first build creates the shared jonex/python-base:local image before Compose builds the platform services in parallel. Then open http://localhost/ and sign in with the local demo credentials admin / admin123 in the tenant_jonex_demo. The README carries an explicit security warning that these are for local evaluation only and must be changed or removed before binding to a non-loopback interface.

## A first knowledge search, and the API entry point

The documented first-use flow is eight steps: sign in, open Core Business and create or select a domain space, create a knowledge base and organize it with folders or tags, select a parser profile for the content you plan to ingest, upload files or configure a REST API or S3-compatible data source, wait for parsing and compilation, inspect the results and graph, then open Knowledge Search, ask a question, verify the references and submit feedback.

Step eight is the part worth pausing on. The README treats reference verification as part of the normal loop rather than an audit feature, which fits the source-grounded framing. External calls go through the gateway, so a programmatic query starts with a login:

```bash
curl -X POST "http://localhost/api/v1/auth/login" \
  -H "Content-Type: application/json" \
  -H "X-Tenan
```

The README excerpt ends mid-header here, so the exact tenant header name and the rest of the login payload are not fully visible in the published README. Treat the gateway login as the confirmed entry point and read the full API section in the repository before scripting against it. For local development outside Docker, the Makefile offers make dev-infra-up (PostgreSQL, Redis, etcd, MinIO, Milvus) or make dev-deps-up (middleware plus LightRAG and Atomic RAG), with make dev-gateway and make dev-frontend in separate terminals and the UI on port 8080. The toolchain there is Python >=3.12.13, Node.js >=20.18.0 with 22 LTS recommended, pnpm >=9.0.0 and uv.

## Where Jonex is heavy, and where it is the wrong choice

The first constraint is operational weight. A working deployment involves PostgreSQL, Redis, etcd, MinIO, Milvus, Neo4j, LightRAG and Atomic RAG, plus a gateway and multiple frontends. That is a lot of moving parts for a team whose actual problem is searching a few hundred PDFs. If your corpus fits in a single vector store and you do not need an ontology graph, the graph half of Jonex is cost without benefit.

The second constraint is the dependency posture. requirements.txt pins fastapi to >=0.109.0,<0.110.0 and pydantic to >=1.10.0,<2.0.0, with a comment explaining that the range is limited to versions compatible with pydantic v1. Pinning to pydantic v1 in a Python 3.12 project is a deliberate compatibility choice, but it also means you cannot freely upgrade FastAPI without checking the whole stack, and any library you add that requires pydantic v2 will conflict. The pyproject.toml build configuration packages only api_gateway, capabilities and jonex_core, so the frontends and deploy directories are not part of the distributable.

The third issue is release hygiene. No releases are listed for the repository, and the version in pyproject.toml is 0.1.0. There is a CHANGELOG.md, but nothing in the README indicates tagged versions you can pin to. If your organisation requires release artifacts for a deployment approval, that is a gap to resolve before you commit.

Finally, the default credentials and the demo tenant are a real hazard, not a formality. The README's own warning says to complete the production checklist in SECURITY.md before deployment. Licence terms are another open question: the repository reports NOASSERTION, so the LICENSE and NOTICE files need a direct read, and THIRD_PARTY_NOTICES.md matters if you redistribute.

## How Jonex differs from a plain RAG stack

The obvious alternative is a general-purpose RAG framework such as LightRAG, which Jonex itself runs as a component rather than competing with. The difference in approach is scope. A RAG framework gives you retrieval over a corpus you have already cleaned; Jonex starts earlier, at ingestion, and owns the parsing step for documents, audio and video, which is why the ASR and VLM connections appear in the configuration and why asr and video are repository topics. If your content is already clean text, that upstream machinery is dead weight.

A second comparison is a managed knowledge platform or a hosted document-AI service. Those remove the operational burden entirely, at the cost of sending your content to a vendor and accepting their retrieval behaviour. Jonex's entire premise is the opposite trade: you run it, you can inspect the compiled ontology and the graph, and the retrieval path is visible in your own infrastructure. For regulated content that cannot leave your network, that difference decides the choice.

The third comparison is a graph database plus a custom ontology, built in-house. That gives you full control over the schema, and it is what Jonex is doing internally with Neo4j and the ontology compilation step. Building it yourself means owning the parser integrations, the multi-tenant model, the feedback loop and the frontends. Jonex's value proposition is that those exist already; its cost is that you inherit its version pins and its deployment topology.

## Conclusion

Adopt Jonex if you need a self-hosted pipeline that turns documents, audio and video into a queryable knowledge base with visible references, and you are willing to run PostgreSQL, Redis, MinIO, Milvus and LightRAG alongside it. Do not adopt it if you want a pip-installable library, a single-binary service, or a project with published releases, because the repository shows no release tags. Before deploying anything, verify the licence file, confirm EMBEDDING_MODEL is identical in deploy/.env and deploy/.env.rag, and replace the default admin account.

## FAQ

### What is Jonex and what does it do?

Jonex is a self-hosted platform that combines an all-in-one multimodal parsing engine with an ontology-driven knowledge engine, so raw documents, audio and video become searchable, source-grounded knowledge services. The README describes it as an end-to-end enterprise AI knowledge platform covering ingestion, parsing, indexing, retrieval and feedback in one governed system.

### How do I install Jonex?

Docker Compose is the documented path: clone the repository, run make init to generate the deploy/.env files, configure at least one OpenAI-compatible LLM and embedding provider, then run make build, make up and make ps. The UI is then available at http://localhost/ with the local demo credentials in the tenant_jonex_demo.

### What are the runtime requirements for Jonex?

For Docker deployment you need Docker Engine or Docker Desktop, Docker Compose v2 with Buildx, make on macOS or Linux, and enough disk space and time for the first build, which downloads container images, Python dependencies and RAG models. For local development the toolchain is Python >=3.12.13, Node.js >=20.18.0, pnpm >=9.0.0 and uv.

### Which databases and services does Jonex depend on?

The dependency list and Makefile targets name PostgreSQL, Redis, etcd, MinIO, Milvus, Neo4j as the ontology graph database, and LightRAG plus Atomic RAG for the RAG stack. The make dev-infra-up target starts PostgreSQL, Redis, etcd, MinIO and Milvus, while make dev-deps-up adds LightRAG and Atomic RAG.

### What should I check before deploying Jonex to production?

The README warns that the default admin / admin123 credentials are for local evaluation only and must be changed or removed before binding Jonex to a non-loopback interface or sharing the deployment, and it points to the production checklist in SECURITY.md. The licence is reported as NOASSERTION, so the LICENSE and NOTICE files also need a direct read.

## Sources

- [Issues](https://github.com/yuezhiai/jonex/issues)
- [Project website](https://jonex.ai)
- [README](https://github.com/yuezhiai/jonex/blob/mian/README.md)
- [yuezhiai/jonex on GitHub](https://github.com/yuezhiai/jonex)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/yuezhiai-jonex
