Model or dataset
AdamPlatin123/Open-Deep-Research-workflow-on-Dify avatar
AdamPlatin123/Open-Deep-Research-workflow-on-Dify

Deep Researcher On Dify: a multi-source research workflow you import as a DSL file

Deep Research workflow on Dify: cascaded multi-source search → outline → cited long-form report. Credited in Awesome-Dify-Workflow.

323 stars61 forksUnknownLicense varies

At a glance

What is it?
Open-Deep-Research-workflow-on-Dify reproduces Deep Research inside Dify as a cascaded search and report pipeline. It ships as a single YAML export, and the README documents the model choices, the rate-limit trap and the timeout error you will meet.
Who is it for?
Adopt it if you already run Dify and want a working reference for cascaded retrieval plus cited report generation without building the orchestration yourself. Do not adopt it if you need a maintained product: the last push was on 2026-08-18, the To Do List still lists subtitle duplication and a large-scale rewrite as open items, and there is no release beyond the 2025-02-16 initial one.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 44 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Deep Researcher On Dify fills between a chat model and a research report

A single prompt to a chat model returns a single pass of text. It does not search your own documents, it does not iterate over sub-questions, and it does not attach citations. Deep Researcher On Dify is built for the case where you want all three, but you want them assembled inside Dify rather than in a separate research product. The README frames it as a reproduction of Deep Research's core function, combining multi-source retrieval (a local knowledge base plus web search) with multi-model collaboration, and it states the system can generate a structured research report of ten-thousand-character scale within five minutes. That five-minute figure is the author's description of the workflow's behaviour, not a result measured here.

The intended user is someone who already has Dify running and has material to search: an internal knowledge base, a document set, or a topic that needs web sources on top. The repository is a workflow export, not a library. Nothing installs into your application code. You import a YAML file into Dify and the pipeline appears as a graph of nodes you can inspect, edit and rewire. That shape is the point: the README describes the design as modular, with the underlying models and data sources meant to be swapped.

How the cascade works: topic parsing, question generation, hybrid retrieval, multi-model writing

The Mermaid diagram in the README is the clearest statement of the data flow. A user question enters a topic-parsing node. That node branches: one path generates questions that the user answers, and another performs topic analysis. Both paths converge on secondary-topic extraction, which feeds a hybrid iterative retrieval engine, which feeds multi-model collaborative generation, which produces the structured report.

The parsing stage uses Gemini 2.0 Flash, and the README says it supports analysis across four dimensions. The retrieval stage is described as multi-channel: a local knowledge base plus Wikipedia, Google and Bing APIs. The writing stage is described as paragraph-level generation using models including deepseek-r1-distill, with Markdown as the output format.

The part worth reading twice is the rhythm control. The README calls it a 2>1 model cascade, implemented through conditional branches and conversation-turn markers. In practice that means the workflow does not fire every model at every step; it uses branch conditions and turn state to decide what runs when. This is the mechanism that keeps a long pipeline from spending its whole budget on the first retrieval pass. It is also the part most likely to need adjustment, because the README's own To Do List names balancing RPM against processing time as an open problem. The conditional structure is visible in the exported YAML, so you can read the branch conditions directly rather than inferring them.

Installing the workflow into Dify and running a first research request

There is no package manager step. The repository contains three top-level entries: the README, a PDF, and the workflow file named `Deep Researcher On Dify .yml`. Installation is an import into Dify. The README does not spell out the menu path, so treat the import screen of your Dify build as the entry point and select the YAML file from the repository.

Before importing, confirm you have the credentials the nodes reference. The README names Gemini 2.0 Flash for topic parsing, deepseek-r1-distill among the generation models, and Wikipedia, Google and Bing as search channels. A node whose provider is not configured in your Dify instance will fail at run time, not at import time.

Once the workflow is in place, open it and check the model bindings before the first run:

yaml
# after import, inspect the model nodes in the Dify editor
# topic parsing      -> Gemini 2.0 Flash
# generation         -> deepseek-r1-distill (paragraph-level)
# retrieval channels -> local knowledge base, Wikipedia, Google, Bing

Then start a run from the Dify app interface with a question broad enough to need sub-topics. The README's diagram shows the workflow asking you questions before retrieval begins, so expect an interactive turn rather than a single submit-and-wait. Answer those questions; the answers feed the secondary-topic extraction that drives retrieval.

If you are on a Google free API key, the README gives a specific remedy rather than a general warning. It states the Google default limit is 15 RPM and that too many requests in a short window will error. The suggested fix is to insert a local model node to throttle RPM. That is a structural edit to the graph, not a setting:

bash
# README guidance, not a command:
# Google free API default limit is 15 RPM
# insert a local model node to throttle RPM before the search nodes

The other documented failure is the LLM node's `Timepouterror` (the README's spelling). It appears when local models are under heavy request pressure. The README offers two options: switch to an online API service, or raise the timeout value in Dify's configuration file. It does not give the exact key name or a recommended value.

Where the workflow breaks: timeouts, rate limits and duplicate subtitles

The README is unusually direct about its own weak points, and they are worth taking at face value. The first is the `Timepouterror` on LLM nodes. Running the retrieval and generation stages against local models puts sustained pressure on the inference endpoint, and the node gives up before the model responds. Switching to a hosted API moves the pressure elsewhere but introduces cost and a network dependency. Raising the timeout in Dify's config file trades a hard failure for a longer wait, which matters when the pipeline already has several sequential model calls.

The second is the Google free-tier ceiling. Fifteen requests per minute is low for a workflow that queries multiple channels iteratively. The README's own answer, inserting a local model node as a throttle, is a workaround that adds a node to the graph and therefore adds latency to every run. The To Do List acknowledges the tension directly: balancing RPM against processing time is listed as unfinished work.

The third is output quality. The To Do List includes fixing the occasional appearance of multiple subtitles in an answer. If you are generating a report for publication rather than for your own reading, that defect lands in the final artifact, and the README does not describe a post-processing step that removes it.

Finally, the To Do List includes a large-scale refactor to make the workflow adapt its questioning to the complexity of the user's question. Until that exists, a simple question and a complex one follow the same path. If your use case is mostly simple lookups, this pipeline is heavier than the task requires.

Open-Deep-Research-workflow-on-Dify against nickscamara's open deep research

The related searches around this project include nickscamara open deep research, and the comparison is a real fork in the road. The difference is where the orchestration lives. This project puts it in Dify: the pipeline is a graph of nodes in a visual editor, the state is conversation turns and branch conditions, and the artifact you version is a YAML export. nickscamara's open deep research is a code-first project, where the loop, the search calls and the report assembly are expressed in source you run yourself.

That distinction decides most adoption questions. If your team already operates Dify and wants a research app that non-engineers can open and edit, the Dify workflow avoids building a service around a research loop. If you need to embed research as a function inside an existing Python or TypeScript codebase, call it from tests, or control retry and concurrency in your own runtime, a code-first implementation fits better, because a Dify workflow is configured rather than imported as a library.

The trade-off runs the other way too. A Dify workflow is inspectable by anyone who can read a graph, which is a genuine advantage over a research script only one person understands. But you inherit Dify's execution model, including the timeout behaviour described above, and you cannot fix a node's failure by patching a function. You change the graph or you change Dify's configuration.

Licence and the cost of keeping a workflow export current

The README states LGPL3.0. That is a copyleft licence, and the repository is a workflow definition rather than a linked library, so the usual linking question does not map cleanly onto it. If you plan to redistribute a modified version, or ship it inside a product, read the licence text and get your own advice rather than relying on a summary here. The README also notes the workflow was first published in this repository on 2025-02-13 and was later included in Awesome-Dify-Workflow, where the README says the upstream file is byte-identical to this repository's version and the upstream README credits @AdamPlatin123 with a link back here. That provenance is useful if you need to know which copy you are importing.

Upgrade cost is the part to weigh before committing. The repository has one release, published 2025-02-16. The last push was on 2026-08-18. There is no changelog and no migration note, so an upgrade is a re-import of the YAML, and any edits you made in the Dify editor have to be reapplied or diffed by hand. The To Do List is the closest thing to a roadmap, and it names a large-scale refactor. If that refactor ships, expect the node structure to change rather than the behaviour to be patched in place, which makes local modifications expensive to carry forward.

There is also a version-coupling cost that the README does not address. The workflow depends on Dify's node types and DSL format, and on external model and search providers. A change on either side can break an import. Nothing in the repository pins those versions.

Editorial conclusion

Adopt it if you already run Dify and want a working reference for cascaded retrieval plus cited report generation without building the orchestration yourself. Do not adopt it if you need a maintained product: the last push was on 2026-08-18, the To Do List still lists subtitle duplication and a large-scale rewrite as open items, and there is no release beyond the 2025-02-16 initial one. Verify two things first: that your Dify instance imports the YAML without missing nodes, and that your Google API quota tolerates the request pattern, because the README states the default limit is 15 RPM and that excess requests error out.

Frequently asked questions

Can Dify.AI be self-hosted, and does that matter for Deep Researcher On Dify?

The README does not describe Dify's deployment options. What it does say is that the workflow is built on the Dify platform and imported as a YAML file, so wherever you run Dify, the import step and the model credentials are yours to configure.

What are the stages of the Deep Researcher On Dify workflow?

The README's diagram shows topic parsing, question generation with user answers, topic analysis, secondary-topic extraction, a hybrid iterative retrieval engine, multi-model collaborative generation and a structured report. The README describes the design as a 2>1 model cascade controlled by conditional branches and conversation-turn markers.

What do I need before importing Open-Deep-Research-workflow-on-Dify into Dify?

You need a Dify instance and credentials for the providers the nodes reference: Gemini 2.0 Flash for topic parsing, deepseek-r1-distill among the generation models, and Wikipedia, Google and Bing as search channels, plus a local knowledge base if you want that retrieval path. A node whose provider is not configured will fail at run time rather than at import.

Why does the workflow throw a Timepouterror on LLM nodes?

The README attributes it to request pressure when using local models. It suggests switching to an online API service or raising the timeout value in Dify's configuration file, and does not give the exact key or a recommended value.

Is there a newer release of Deep Researcher On Dify than the initial one?

The repository lists a single release, Publish (Initial Release), dated 2025-02-16. The last push to the default branch was on 2026-08-18, and the To Do List still lists subtitle duplication and a large-scale refactor as open items.

Official sources

  1. AdamPlatin123/Open-Deep-Research-workflow-on-Dify on GitHub
  2. Issues
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/adamplatin123-open-deep-research-workflow-on-dify.svg)](https://hysenlabs.com/projects/adamplatin123-open-deep-research-workflow-on-dify)