# RAGFlow runs document parsing and RAG orchestration out of one Go binary

> A Retrieval-Augmented Generation engine written mostly in Go, with DeepDoc doing layout analysis, OCR and table recognition inside the same process through CGO. Local deployment wants 4 cores, 16 GB of RAM and 50 GB of disk, and the 1.0 line is still a release candidate.

**infiniflow/ragflow** — RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses reliable RAG with Agent capabilities to create a superior context layer for LLMs.

- Repository: https://github.com/infiniflow/ragflow
- Website: https://ragflow.io
- Stars: 91,300 · Forks: 10,825
- Language: Go
- License: Apache-2.0
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/infiniflow-ragflow

## The Go service owns the API, the ingestor and the document parser at once

The architecture note in the README is short and unusual: API, Admin, Ingestor and Syncer are provided by a unified Go service, and DeepDoc runs within that Go process handling layout analysis, OCR and table recognition. The Go services call native document parsing libraries and ONNX Runtime through CGO.

That single-process design has a direct consequence for deployment. CGO means the parsing work is not in a sidecar you can scale or replace independently, and it means the parsing capability is bound to whatever native libraries the image ships. The repository reflects that: there are four Dockerfiles, Dockerfile_base, Dockerfile_ci, Dockerfile_tei and Dockerfile.scratch.oc9, each covering a different build target rather than one image with feature flags.

Two capabilities are explicitly optional. MCP and Sandbox Executor can be enabled as needed, which means they are not part of the default deployment and their absence is not a misconfiguration.

## A starting host of 4 cores, 16 GB and 50 GB, and the requirements move from there

The documented starting configuration is 4 CPU cores, 16 GB of RAM and 50 GB of available disk space. The README is careful that this is a starting point, not a ceiling: actual requirements depend on the document engine, the data volume, the parsing tasks and the concurrency, and local models plus other optional components may need more on top.

Docker has its own floor, 24.0.0 or later with Compose v2.26.1 or later. Go does not have to be installed on the host, because the deployment runs in containers.

The consequence for a reader is that the resource number is not the decision, the document engine choice is. A team feeding it scanned PDFs through OCR will parse far more per document than a team feeding it clean text, and the 50 GB figure covers a starting corpus rather than a growing one. Sizing has to be done against the actual ingestion queue rather than against the recommended figures.

## Elasticsearch wants vm.max_map_count at 262144, and Infinity usually does not

The first startup step is conditional on the storage backend. If you are using Elasticsearch, the Docker host needs vm.max_map_count set to at least 262144, and the README notes this step is usually unnecessary with Infinity.

The check and the fix are both given as shell commands:

```bash
sysctl vm.max_map_count
```

```bash
# In this case, we set it to 262144:
sudo sysctl -w vm.max_map_count
```

The failure this prevents is the kind that arrives late and looks unrelated to RAG. A host that starts the stack and then fails under load, or a container that exits on startup, is often the Elasticsearch memory-map limit rather than anything in the retrieval pipeline. The setting also does not survive a reboot as written, since the command shown writes the value for the running system only, so a host that is restarted quietly returns to its old limit.

Choosing Infinity avoids this step and the tuning that comes with it. That is the single largest difference between the two deployment shapes for anyone deciding what to run.

## The Self-Managed container Sandbox needs gVisor, and nothing else does

gVisor appears in the prerequisites list with a precise scope: it is required only when using the Self-Managed container Sandbox. Other Sandbox providers do not require gVisor on the RAGFlow host.

The reason this matters is that the sandbox is where untrusted code runs, and running it under a container runtime that shares the host kernel is a different proposition from running it under gVisor's user-space kernel. The README draws the line at the RAGFlow host rather than at the sandbox provider, so a team running a different Sandbox provider can skip the install entirely and a team running the self-managed one cannot.

What the README does not settle is what the sandbox runs and with what limits. There is no documented CPU, memory or timeout budget for sandboxed execution in the deployment section, so a workload that lets the model execute arbitrary code has no published ceiling to design against. Anyone exposing that path to untrusted input should assume they are sizing it themselves.

## Agentic retrieval has four thinking modes, and the cost scales with the one you pick

Retrieval is no longer a single query. The Agentic Retrieval feature lets the model analyse a question, break it down when needed, retrieve knowledge and verify the evidence, then gather more context across multiple rounds of retrieval and reasoning.

The control is a thinking mode: Low, Medium, High or Ultra, chosen according to question complexity. Ultra is not a quality label, it is a statement about how much work the model does before answering, and the number of retrieval rounds is what a reader is really selecting.

The practical consequence is that the same dataset costs a different amount per question depending on the mode, and the mode is a runtime choice rather than a property of the dataset. A deployment that routes every query to Ultra to maximise answer quality pays for the hardest questions on every question, including the ones a single retrieval pass already answers. There is no documented default mode, so the routing policy is the operator's to write.

## Knowledge Compilation emits seven artifact types, and regenerating them is a separate step

Added on 19 August 2026, Knowledge Compilation turns document and dataset content into structured artifacts. The templates generate Wikis, Graphs, Trees, PageIndex, Mind Maps, Timelines and Skills, and the compilation models and processing rules are configurable.

Two things are worth separating. What compilation produces is a second representation of knowledge that sits alongside the chunks, and the artifacts are meant for organization and reuse rather than for the retrieval path itself. And the artifacts are not static: they can be viewed, updated and regenerated, which means the source content, the compiled artifact and the retrieval index are three things that can drift apart.

For a reader, the consequence is that a change to a source document does not by itself update a compiled Wiki, and a stale artifact reads exactly like a correct one. Nothing in the deployment documentation describes a refresh trigger, so keeping the artifacts honest is a scheduling decision the operator has to make.

## 1.0.0-rc1 shipped on 29 September 2026 while the 0.27 line is what most people run

The newest release is v1.0.0-rc1, published 29 September 2026, one day before the last push to main on 25 September 2026. Before it, v0.27.2 came out on 10 September 2026, and a nightly tag sits between them, dated 1 December 2025.

A release candidate is a specific thing to plan around. The 1.0 line is where the Go-native architecture described above is what ships, and the 0.27 line is what a deployment pinned before September is running. The project also ships a cloud service at cloud.ragflow.io, so the fastest way to see the current behaviour is not to run the release candidate locally.

The other recent changes are dated in the same list, which is useful for judging how fast the surface is moving: sitemap-based website ingestion on 10 September 2026, Knowledge Compilation and Agentic RAG both on 19 August 2026, and Google BigQuery ingestion with incremental synchronization on 2 July 2026. Four features in under three months is a project that will not sit still while you evaluate it.

## Two languages, two lockfiles, and Python pinned to a single minor version

RAGFlow is primarily Go, with go.mod requiring Go 1.27, and it is also a Python project, with pyproject.toml naming the package ragflow at version 1.0.0-rc1 and requiring Python >=3.13,<3.14. That upper bound is a hard one: a host on 3.12 or 3.14 does not satisfy it.

Both toolchains are committed, with go.sum and uv.lock alongside the two manifests. The Python dependency list is explicit about what is not needed, keeping a commented block of modules described as not necessary, including nltk, numpy, openai, openpyxl, pandas, pillow, pymysql, scikit-learn and tiktoken. The active list is short by comparison, with beartype, captcha, elasticsearch-dsl, mammoth, mysql-connector-python, peewee, reportlab, ruamel-yaml and agentrun-sdk among the entries that are live.

For a reader, the consequence is that the two halves are versioned in lockstep and neither can drift. There is no supported configuration where the Go service and the Python side sit on different release trains, which simplifies deployment and removes the option of upgrading one without the other.

## Conclusion

RAGFlow fits a team that has a document-heavy corpus, wants to see the chunking it produced, and is willing to run a stack that starts at 4 cores, 16 GB of RAM and 50 GB of disk. It does not fit a laptop demo or a deployment where the host cannot be tuned, because the Elasticsearch path wants vm.max_map_count at 262144 and the Self-Managed container Sandbox wants gVisor. Before committing, read the release notes for 1.0.0-rc1 rather than the 0.27 line, and decide which storage backend you will run, since the two do not need the same host configuration.

## FAQ

### What is ragflow used for?

RAGFlow is a Retrieval-Augmented Generation engine that combines document parsing with retrieval so answers can carry traceable citations. It extracts knowledge from unstructured documents, chunks them with selectable templates, and supports multi-step agentic retrieval.

### What are the disadvantages of ragflow?

A local deployment starts at 4 CPU cores, 16 GB of RAM and 50 GB of disk, and actual needs grow with the document engine, data volume and concurrency. The Elasticsearch path also needs vm.max_map_count at 262144, and the Self-Managed container Sandbox needs gVisor.

### Is ragflow free to use?

The repository carries an Apache-2.0 LICENSE file and the project describes itself as open source. Running it is a separate question from the licence, since the documented starting hardware is 4 cores, 16 GB of RAM and 50 GB of disk.

### how to install ragflow on mac

Docker deployment is the documented local path and does not require Go on the host. You need Docker 24.0.0 or later with Compose v2.26.1 or later, and you can try the hosted service at cloud.ragflow.io before installing anything.

### is ragflow open source

Yes. The project describes itself as a leading open-source Retrieval-Augmented Generation engine, the repository carries an Apache-2.0 LICENSE, and the source for the Go service, the Python side, the SDK and the web front end is all in the tree.

## Sources

- [Official documentation](https://ragflow.io)
- [Official README](https://github.com/infiniflow/ragflow#readme)
- [Project repository](https://github.com/infiniflow/ragflow)
- [Release notes](https://github.com/infiniflow/ragflow/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/infiniflow-ragflow
