All comparisons
Comparison

dify vs ragflow: a platform for building LLM apps versus a retrieval engine for your own documents

Dify is a TypeScript LLM application platform that bundles a visual workflow canvas, RAG pipeline, agents, model management and an HTTP API behind one Docker Compose stack. RAGFlow is a Go retrieval engine that puts deep document parsing and visible chunking at the centre. Dify is the broader build surface; RAGFlow is the deeper retrieval layer, and they can be combined rather than only compared.

Published September 21, 2026

At a glance

Projectlanggenius/difyinfiniflow/ragflow
LicenceCustom licenceCustom licence: read the LICENSE fileApache-2.0Permissive: commercial use allowed
MaintenanceCommits in the last dayLast push September 29, 2026Commits in the last six monthsLast push September 25, 2026
LanguageTypeScriptGo
GitHub stars157,51291,300
Read moreOur analysisGitHubOur analysisGitHub

Which one to choose

dify

Choose dify if you need to ship a chat app, agent or workflow quickly on a visual canvas, want model management and logging in the same product, and can operate a multi-container deployment. It is the stronger default for teams whose main problem is orchestration, not parsing.

ragflow

Choose ragflow if retrieval quality over your own messy PDFs, slides and scans is the hard part, you want to see and edit chunks before they are indexed, and you can give it at least 16 GB of RAM and 50 GB of disk. It is the weaker choice if you want one tool that also builds the whole application around retrieval.

Two products, two centres of gravity

The repository descriptions already disagree about what the product is for. Dify describes itself as an LLM app development platform whose interface combines AI workflow, RAG pipeline, agent capabilities, model management and observability. Its README lists workflow, prompt IDE, RAG, agents, LLMOps and backend-as-a-service as separate features of one workspace. RAGFlow describes itself as a retrieval-augmented generation engine that fuses RAG with agent capabilities to create a context layer for LLMs. Its README leads with document understanding, template-based chunking, grounded citations and heterogeneous data sources. That difference in framing is the whole comparison. In Dify, retrieval is one node type inside a larger graph you draw. In RAGFlow, retrieval is the product, and the agent features are built around the context it produces. If your hardest problem is deciding what should happen after a user asks a question, Dify starts closer to your problem. If your hardest problem is that your PDFs come back as garbage, RAGFlow starts closer to it.

Architecture: a multi-service canvas versus a Go retrieval service

Dify is written primarily in TypeScript and ships as a Docker Compose stack. The README's quick start is three commands: change into the docker directory, copy .env.example to .env, and run docker compose up -d, after which the dashboard is available at localhost/install. The stated minimum is 2 CPU cores and 4 GiB of RAM, and Docker Compose v2.24.0 or later is required. The stack is not a single binary, and adopting Dify means operating a multi-container deployment; a single static binary is explicitly not what it offers. RAGFlow is written primarily in Go and also installs through Docker Compose, but its prerequisites are steeper: at least 4 CPU cores, 16 GB of RAM, 50 GB of disk, Docker 24.0.0 or later, Compose v2.26.1 or later, and Python 3.13 or later. It also requires vm.max_map_count of at least 262144, which the README documents with a sysctl command and a warning that the change does not survive a reboot unless it is made permanent. The practical reading is that Dify asks less of the host before it will start, while RAGFlow asks for a machine sized for parsing and indexing documents. The README does not document rollback for either project, so treat upgrade paths as something to confirm from the documentation before you commit.

Getting each one running

Dify's path is the shorter one. Install Docker and a recent Compose, copy the environment file, bring the stack up, then open the install URL in a browser to initialise. The README points at a FAQ for setup problems and at a separate guide for deploying from source if you intend to contribute or modify the code. Nothing in the quick start suggests a kernel tuning step. RAGFlow's path has a step that will stop a careless install: the vm.max_map_count check. The README tells you to read the current value with sysctl vm.max_map_count, raise it to at least 262144 with sudo sysctl -w if it is lower, and then notes that this resets after a reboot, so the value has to be added to or updated in the system configuration to persist. That is a host-level change, not a container setting, and it is the kind of thing that works in a test and fails after a restart. RAGFlow also lists gVisor as required only if you intend to use the code executor sandbox feature, which means the agent side of the product can carry an extra dependency that the retrieval side does not. Two more things to check on RAGFlow before committing: the RAGFLOW_IMAGE variable in docker/.env against the release you intend to run, and the deepdoc directory to see which parsers you are actually getting. Neither project publishes a rollback procedure in its README.

Operations, scaling and the resource floor

Dify's floor is low enough that a developer laptop can run it, but the ceiling is a multi-service deployment you have to keep healthy. The README lists observability integrations with Opik, Langfuse and Arize Phoenix, and LLMOps features for monitoring application logs and performance over time, which is the part of operations Dify actually gives you tooling for. It does not give you a single process to supervise. RAGFlow's floor is the constraint most teams underestimate: 16 GB of RAM and 50 GB of disk is the number people miss. That floor is not arbitrary. Deep document understanding, template-based chunking and multi-recall with fused re-ranking all consume memory and storage during ingestion. If you plan to run RAGFlow alongside other services on one host, add its floor to theirs rather than assuming it shares nicely. RAGFlow's README also lists data synchronisation from Confluence, S3, Notion, Discord and Google Drive as a feature added in November 2025, which matters operationally because it turns ingestion into a recurring job with credentials to manage rather than a one-off upload. Neither project's README documents a horizontal scaling story, so plan capacity per host and verify before you promise throughput.

Where each one falls short

Dify's weakness is depth at the retrieval layer and the shape of its deployment. Do not adopt it if you need a single static binary, if your workflow logic is mostly custom code, or if you cannot accept the licence terms as they stand. A visual canvas is a benefit until your logic outgrows it, at which point you are maintaining both a graph and the code around it. The README documents text extraction from PDFs, PPTs and other common formats, but it does not describe the chunking review workflow or citation tracing that RAGFlow puts at the front of its feature list. If your documents are the problem, Dify gives you a pipeline without giving you RAGFlow's visibility into it. RAGFlow's weaknesses are the mirror image. It is a retrieval engine with agent components, not a general application platform; the README's feature list is about ingestion, chunking, citations and data sources, and its agent features arrive as templates and components rather than as a workspace for building arbitrary applications. RAGFlow does not fit anyone on ARM64 without building their own image, anyone with less than 16 GB of RAM, or anyone who wants a managed vector store they never touch. The kernel tuning requirement and the gVisor dependency for the code executor add two more places an install can go wrong. Neither project is a drop-in replacement for the other.

Licence and maintenance implications

RAGFlow is Apache-2.0, which is a standard permissive licence and the simpler of the two to reason about inside a company. Dify is classified by GitHub as Other, a custom licence GitHub cannot classify, and the README points to a Dify edition overview page alongside the repository. Read LICENSE at the repository root, and do not adopt Dify if you cannot accept the licence terms as they stand. That is not a statement that the licence is bad; it is a statement that you have to read it, because a custom licence can carry conditions a permissive one does not, and the classification alone tells you nothing about which conditions apply. On maintenance, both repositories are unarchived and both were pushed recently: Dify on 2026-09-15 and RAGFlow on 2026-09-17, both within days of the most recent release in each list. Dify's latest release is 1.17.0 from 2026-08-25, following 1.16.1 and 1.16.0 in July. RAGFlow's latest is v0.27.1 from 2026-08-28, following v0.27.0 on 2026-08-19, with a nightly build from 2025-12-01 also listed. The release cadence on both sides is regular, and neither shows the dormancy that would make a fork or a pinned version the safer bet. What the release lists do not tell you is how disruptive an upgrade is; the README does not document rollback for either project.

Choosing for a concrete scenario

A two-person team building an internal assistant over a handful of well-formatted documents should take Dify. The canvas, the model management and the logging cover most of what they need, the 2-core, 4 GiB floor runs on modest hardware, and the HTTP API means the assistant can be called from an existing tool without rewriting it. A team whose corpus is scanned contracts, slide decks and mixed spreadsheets should take RAGFlow, because template-based chunking and visible chunk editing are the difference between an answer you can defend and one you cannot. A team that has both problems should consider running them together: RAGFlow as the retrieval service behind an API, Dify as the application layer that calls it. That combination is not documented in either README as an integrated setup, so it is a design decision to validate rather than a supported path to assume. A team with an ARM64 fleet, or one that cannot spare 16 GB of RAM on a single host, should start with Dify and treat RAGFlow as a later addition on dedicated hardware. A team that needs a single static binary, or whose logic is mostly custom code, should not adopt Dify; the Compose stack is the reason.

Bottom line

Pick Dify when the application is the hard part and the documents are manageable; pick RAGFlow when the documents are the hard part and you can give it a properly sized host. If both are true, run RAGFlow as the retrieval layer and Dify as the build layer, but verify the integration yourself because neither README describes it. Before committing to Dify, read LICENSE at the repository root, run the Compose stack in a staging host and check that the containers come up clean, and confirm from the documentation how you would upgrade and roll back across a version boundary. Before committing to RAGFlow, check the RAGFLOW_IMAGE variable in docker/.env against the release you intend to run, confirm vm.max_map_count is permanently set, and read the deepdoc directory to see which parsers you are actually getting.

Sources

  1. langgenius/dify repository
  2. langgenius/dify README
  3. infiniflow/ragflow repository
  4. infiniflow/ragflow README