Model or dataset
skygazer42/MimirQ avatar
skygazer42/MimirQ

MimirQ: a Chinese-first enterprise RAG pipeline where every stage can be inspected

中文优先的企业 RAG 知识库:可控解析、治理、切块、混合检索、重排、引用、图谱、评测与 Dify 接入。

559 stars89 forksPythonApache-2.0

At a glance

What is it?
MimirQ is an Apache-2.0 Python and Next.js knowledge base that splits parsing, governance, chunking, hybrid retrieval, reranking and evaluation into separate, replaceable stages. It is built for teams that need to trace which stage produced a wrong answer.
Who is it for?
Adopt MimirQ if your team already knows that answer quality depends on parsing and chunking choices, and you want those choices visible per document instead of hidden behind an upload button. Do not adopt it if you want a low-code application platform for a simple, stable workflow, or if you cannot run Docker with at least 4 CPU cores, 16 GB RAM and 50 GB of disk.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem MimirQ targets: a wrong answer with no traceable cause

The README states the project's premise directly: the hard part of an enterprise knowledge base is not vectorizing documents, it is making errors locatable, policies replaceable and quality regression-testable. MimirQ grew out of a real government knowledge base delivery, where a bad answer forced the team to guess whether the fault sat in parsing, governance, chunking, recall, reranking or generation drifting away from its citations.

That framing decides the audience. This is not aimed at someone who wants to upload a folder and start chatting. It is aimed at the engineer or delivery lead who has to estimate, accept and govern a knowledge pipeline over months, and who needs to point at a stage and say this is where it broke. The README's own comparison table is honest about the trade-off: for simple business logic, stable flows and low-code priority, application platforms such as Dify or FastGPT are usually faster. MimirQ positions itself as the layer you use when the knowledge chain has to be swapped, audited and regression-tested per business case.

Pipeline architecture: evaluation, parsing, governance, chunking, hybrid recall, reranking, Golden regression

The README lays out the intended flow as a single line: data evaluation, scenario-specific parsing, cleaning and governance, business chunking, vector and full-text indexing, hybrid recall, reranking with citations, then Golden regression. Each arrow is a place where input, output and version can be inspected.

Parsing is the most explicit example. The README recommends sampling and evaluating data before choosing a parser, and names concrete options for different material: MinerU or DeepDoc for complex layouts and scans, Docling when formulas, tables and layout structure are dense, and lighter paths such as MarkItDown for digital-native Office files or plain text. It also states plainly that high-risk material still needs human verification. The repository ships 30 parsing backends, 86 chunking strategies and 13 reranker categories, and the README is careful to call those numbers implementation breadth rather than a quality claim.

Chunking follows the same logic. Instead of one fixed length with a fixed overlap window, the project splits by heading, section, business record or parent-child relationship. The index layer can use a vector store such as Milvus and combine BM25, vector retrieval and reranking, while the application above it can be Dify, LangGraph, PydanticAI or a plain API service. That separation is the actual design argument: the retrieval stack is not welded to the UI.

Installing MimirQ with Docker Compose and checking the API

The README lists prerequisites before anything else: Docker 20.10 or later with Docker Compose 2.0 or later, GNU Make, and Python 3.9 or later solely to generate configuration for the one-command Docker start. Source development additionally needs Python 3.11 or later, Node.js 20 or later and pnpm 10.26. The stated hardware floor is 4 CPU cores, 16 GB RAM and 50 GB of disk.

The first step clones the repository and runs the initializer. According to the README, make init only creates missing .env and web/.env.local files and does not overwrite existing configuration.

bash
git clone --depth 1 --single-branch https://github.com/skygazer42/MimirQ.git
cd MimirQ
make init

After that, edit .env. The README says LLM_API_KEY is required for default model calls, while LLM_API_BASE and LLM_MODEL override the endpoint, EMBEDDING_API_BASE, EMBEDDING_API_KEY and EMBEDDING_MODEL configure a separate embedding service, and ENABLE_RERANKER plus RERANKER_API_BASE, RERANKER_API_KEY and RERANKER_MODEL turn on a reranker. To create the first administrator automatically, set INITIAL_ADMIN_EMAIL, INITIAL_ADMIN_USERNAME and INITIAL_ADMIN_PASSWORD.

bash
make up-web
make api-ping

The Docker path runs the frontend, API, worker and dependency services in containers. The README says to open http://localhost:3000 afterwards, and if no administrator was preconfigured, register the first account in the page. make api-ping is the readiness check for the API. Stopping is make down; make docker-reset clears persisted data and make docker-purge also removes the project's service images. Both are irreversible, and the README points to the Docker Compose deployment guide for the exact deletion scope.

Optional parser containers and the hardware boundaries they impose

MimirQ defaults to a built-in DeepDoc parser and starts other parsers only when the material requires it. The README gives a scenario table with the matching make target for each. Docling Serve runs on CPU with make up-docling, and a CUDA variant runs with make up-docling-gpu. Marker handles PDF to Markdown on CPU via make up-marker. ETL4LLM covers mixed layout, table and image documents with make up-etl4llm. PaddleOCR-VL handles scans and OCR on an NVIDIA GPU, with the README suggesting around 10 GiB reserved, through make up-paddlevl. MinerU pipeline and MinerU VLM both need an NVIDIA GPU and a first-time model download, via make up-mineru and make up-mineru-vlm. olmOCR targets high-precision PDF OCR and the README suggests a 48 GiB class GPU, through make up-olmocr. MagicPDF converts formula and table heavy PDFs to Markdown on GPU with make up-magicpdf. Qianfan-OCR sends PDFs and images to an external vision OCR service using an upstream URL and API key, so no local GPU is needed, via make up-qianfanocr.

The Docling note is the clearest constraint in the README. Its heavy dependencies and models live only in a separate container, the CPU and GPU profiles share port 5001 and therefore cannot run at the same time, and the GPU image is about 11.13 GB. The repository states it has been verified on an RTX 3070 Ti with 8 GiB of VRAM across PDF, DOCX, Markdown tables and the ParserFactory fallback-free path. That is a narrow, named configuration, not a general GPU guarantee. Anyone planning PaddleOCR-VL, MinerU or olmOCR should treat the suggested VRAM figures as the starting point for their own capacity check.

Where MimirQ is the wrong tool

The README answers this itself. If the business logic is simple, the flow is stable and low-code is the priority, Dify or FastGPT are usually faster to ship. If you want DeepDoc and GraphRAG as one integrated package, the README names RAGFlow as the mature choice. MimirQ is not trying to win those cases.

The cost of the design is operational surface. A single-command Docker start exists, but the full picture includes PostgreSQL, Redis, Milvus, MinIO, a Python API, a worker and a Next.js frontend, plus optional parser containers that each carry their own image size and GPU requirement. The .env.example already warns that production must set MIMIRQ_DB_CREATE_ALL_ON_STARTUP to false and MIMIRQ_DB_RUNTIME_MIGRATIONS_ENABLED to false, and run make db-upgrade separately before starting the API and worker. That is a deliberate deployment step, not a default you can ignore.

There is also a data-residency consideration. Parsers such as Qianfan-OCR send documents to an external service, and any configured LLM, embedding or reranker endpoint receives content depending on your setup. Teams with material that cannot leave the network need to keep those paths local, which usually means GPU hardware. The README's own caveat that high-risk documents still require human verification applies here too: no parser selection removes the need for review.

MimirQ compared with Dify, RAGFlow, FastGPT, AnythingLLM and LangChain

The README's comparison table splits the space three ways. Low-code application platforms such as Dify and FastGPT are the faster route when the business process is simple and stable. RAGFlow is the mature integrated option if you want DeepDoc and GraphRAG together out of the box. MimirQ occupies the third slot: the knowledge chain has to be replaceable per business, auditable, and regression-tested.

The practical difference is where the seams are. In an application platform, the parsing and retrieval stages are typically configured through the product's own interface and swapped as a unit. In MimirQ, the README describes parsers, indexers, retrieval and models as switchable per document and per scenario, with the application layer sitting above as a separate concern. That is why the README suggests using MimirQ as Dify's external knowledge layer rather than as a replacement for Dify. It is a component with a narrower responsibility and a correspondingly larger configuration burden. If you never intend to change a parser or inspect a chunk boundary, that burden buys you nothing.

Licence, maintenance and what upgrades actually involve

MimirQ is licensed under Apache-2.0, and the repository carries a NOTICE file alongside the LICENSE, which is the usual pattern for attribution requirements in this licence family. Apache-2.0 permits commercial use and modification and includes a patent grant; it also requires that you preserve licence and notice information in distributions. This is a description of what the licence text does, not legal advice, and any organisation redistributing MimirQ or a modified version should review the LICENSE and NOTICE files with its own counsel.

On maintenance, the last push to the default branch was on 2026-09-09, and the repository is not archived. The release history shows v1.0.0 on 2026-07-28, preceded the same day by v0.8.0 and v0.7.17, while the README badge and text point to v1.0.1 as the current stable version with a release note at docs/releases/v1.0.1.md. The version numbering and the density of releases in July suggest the project is still moving, but the README does not document a rollback procedure, and it does not state a support window or a compatibility policy for the .env keys between versions. Anyone pinning a version should read docs/releases/README.md before upgrading, because the environment file is long and several keys, including the database startup flags, change behaviour between development and production.

Editorial conclusion

Adopt MimirQ if your team already knows that answer quality depends on parsing and chunking choices, and you want those choices visible per document instead of hidden behind an upload button. Do not adopt it if you want a low-code application platform for a simple, stable workflow, or if you cannot run Docker with at least 4 CPU cores, 16 GB RAM and 50 GB of disk. Before committing, verify the parser path for your own document types, check that MIMIRQ_DB_CREATE_ALL_ON_STARTUP and MIMIRQ_DB_RUNTIME_MIGRATIONS_ENABLED are set to false for production, and confirm which optional parser containers your hardware can actually run.

Frequently asked questions

What is MimirQ used for?

It is an enterprise RAG knowledge base that runs a controlled pipeline from data evaluation and scenario-specific parsing through governance, business chunking, vector and full-text indexing, hybrid recall, reranking with citations, and Golden-set regression. The README describes using it as a standalone knowledge layer or as Dify's external knowledge layer.

How do I install MimirQ?

Clone the repository, run make init to create the missing .env and web/.env.local files, fill in LLM_API_KEY and any optional model settings, then run make up-web and check readiness with make api-ping. Docker 20.10 or later, Docker Compose 2.0 or later, GNU Make and Python 3.9 or later are the stated prerequisites.

Does MimirQ need a GPU?

No. The default built-in DeepDoc parser and the CPU options such as Docling Serve, Marker and ETL4LLM run without one. PaddleOCR-VL, MinerU, MinerU VLM, olmOCR and MagicPDF require an NVIDIA GPU, and the README gives suggested VRAM figures for PaddleOCR-VL and olmOCR.

Can MimirQ run alongside Dify on the same machine?

The README states that MimirQ fixes an independent Compose project name, mimirq, so it will not treat a Dify installation on the same host as its own. The README also suggests using MimirQ as Dify's external knowledge layer.

Which vector database does MimirQ use?

The README describes the index layer as able to use a vector store such as Milvus, combined with BM25 and reranking. The .env.example sets MILVUS_IMAGE to milvusdb/milvus:v2.6.11 and defines MILVUS_HOST_DOCKER and MILVUS_PORT_DOCKER for the Compose network, while the requirements file also lists chromadb and faiss-cpu for other vector backends.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. skygazer42/MimirQ on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/skygazer42-mimirq.svg)](https://hysenlabs.com/projects/skygazer42-mimirq)