# dzhng/deep-research: a 500-line deep research agent you can read in one sitting

> The repository is a small TypeScript implementation of an iterative research loop built on Firecrawl and OpenAI-compatible models. Its value is legibility, not features, and that trade-off shapes how you run it.

**dzhng/deep-research** — An AI-powered research assistant that performs iterative, deep research on any topic by combining search engines, web scraping, and large language models.  The goal of this repo is to provide the simplest implementation of a deep research agent - e.g. an agent that can refine its research direction overtime and deep dive into a topic.

- Repository: https://github.com/dzhng/deep-research
- Stars: 19,738 · Forks: 2,000
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/dzhng-deep-research

## The problem: research loops are easy to describe and hard to read

Most agent frameworks ask you to learn their abstractions before you can change one line of behaviour. This repository takes the opposite position. The README states the goal directly: keep the repo under 500 lines of code so it is easy to understand and build on top of. That constraint is the product. What you get is a single research loop you can hold in your head, not a platform.

The intended user is a developer who wants to see how iterative research actually works, or who wants a base to fork. It is not aimed at someone who wants a research tool with a login page. There is no hosted service, no dashboard, and no published release. The README points to the author's other work at Duet but the repository itself is a standalone script.

The loop it implements is the now-familiar pattern: ask follow-up questions, generate search queries, read results, extract learnings, generate new directions, repeat. The difference from commercial deep research products is that here you can see every step and change any of them.

## How the recursive loop works, from the flowchart to the file layout

The README includes a mermaid flowchart that describes the control flow precisely enough to reason about. Input is a user query plus two parameters, breadth and depth. The agent generates SERP queries, processes the results, and splits the output into two buckets: learnings and directions. A decision node checks whether depth is greater than zero. If it is, the next direction is assembled from prior goals, new questions, and accumulated learnings, and that context feeds back into the research step. If it is not, the loop terminates and emits a markdown report.

Breadth and depth do different jobs. Breadth controls how many queries are generated at each level, and the README recommends 3 to 10 with a default of 4. Depth controls how many times the loop recurs, recommended 1 to 5 with a default of 2. Raising breadth widens a single pass; raising depth compounds passes, and each pass carries the previous learnings forward as context.

The repository layout matches this description. Source lives under src/, with run.ts as the interactive entry point and api.ts as an Express server entry. The presence of express, cors, and uuid in package.json confirms the API mode is a real second interface, not a stub. Concurrency is handled with p-limit, and js-tiktoken is present for token counting. The whole thing is TypeScript run through tsx rather than compiled ahead of time.

## Installing dzhng/deep-research and running your first query

The README gives a Node.js path and a Docker path. For the Node path, clone the repository and install dependencies with npm.

```bash
npm install
```

You need two credentials before anything will run: a Firecrawl API key for search and content extraction, and an OpenAI API key for the o3-mini model. Create a .env.local file in the repository root. The .env.example file shows the expected shape, including the optional self-hosted Firecrawl base URL and the context size setting.

```bash
FIRECRAWL_KEY="your_firecrawl_key"
# FIRECRAWL_BASE_URL="http://localhost:3002"
OPENAI_KEY="your_openai_key"
CONTEXT_SIZE="128000"
```

Start the assistant with npm start. The package.json script is tsx --env-file=.env.local src/run.ts, which is why the file must be named .env.local specifically.

```bash
npm start
```

You will be prompted for a research query, then breadth, then depth, then follow-up questions that refine the direction. The README states the final output is written as report.md or answer.md in your working directory, depending on which modes you selected. A report.md file sits at the top level of the repository, which is consistent with that behaviour.

If you would rather run it in a container, rename .env.example to .env.local, set your keys, build the image, then bring the compose service up.

```bash
docker compose up -d
docker exec -it deep-research npm run docker
```

The compose file mounts the repository into /app and sets tty and stdin_open to true, which is what makes the interactive prompt usable inside the container. Note that the Dockerfile is based on node:18-alpine while package.json declares engines.node as 22.x. That mismatch is visible in the files and worth resolving before you build.

## Pointing the agent at local models, OpenRouter, or DeepSeek R1

The model layer is more flexible than the default configuration suggests. To use a local LLM, the README says to comment out OPENAI_KEY and instead uncomment OPENAI_ENDPOINT and OPENAI_MODEL, setting the endpoint to a local server address such as http://localhost:1234/v1 and the model name to whatever is loaded there. The .env.example file names the variable CUSTOM_MODEL rather than OPENAI_MODEL in its example block, so read both files before assuming the key name.

```bash
# OPENAI_ENDPOINT="http://localhost:11434/v1"
# CUSTOM_MODEL="llama3.1"
```

There is also a dedicated path for DeepSeek R1. Setting a Fireworks API key causes the system to switch over to R1 instead of o3-mini automatically, according to the README. The dependency list includes @ai-sdk/fireworks, which is consistent with that claim.

```bash
FIREWORKS_KEY="api_key"
```

This is the most interesting design decision in the repository. Model choice is inferred from which environment variables are present rather than passed as an explicit argument. It makes the common case frictionless and the debugging case harder, because a stale key in your environment silently changes which model runs.

## Concurrency limits are where a free Firecrawl key breaks

The README is unusually direct about the failure mode here. Concurrent processing is a listed feature, and the CONCURRENCY_LIMIT environment variable controls it. The guidance is that a paid or self-hosted Firecrawl instance can raise the limit for speed, while a free instance may hit rate limit errors and should be reduced to 1, at the cost of running much slower.

That is a real constraint, not a footnote. The agent's speed is bounded by an external service you do not control, and the default behaviour is not documented in the README beyond the existence of the variable. The .env.example shows FIRECRAWL_CONCURRENCY="2" in a commented line, which suggests two is a reasonable starting point, but the README's own text refers to CONCURRENCY_LIMIT. Two different names appear across the two files, so verify which one your version of the source actually reads before tuning.

There is a second, quieter cost. Each recursion level regenerates queries and reprocesses results, so raising depth multiplies both API calls and wall-clock time. Depth 5 is not five times the work of depth 1 in any predictable way, because the number of directions generated at each level depends on what the model finds.

## What it does not give you: tests, releases, and a stable interface

The package.json test script is echo "Error: no test specified" && exit 1. There is no test suite. For a project whose entire output is a model-generated report, that means correctness is something you assess by reading the markdown it produces, not by running a command.

There are no published releases, so there is no version to pin and no changelog to read. The package version is 0.0.1 and the license field in package.json says ISC while the repository LICENSE file and the README both say MIT. That inconsistency is minor but it is the kind of thing you want to settle before vendoring the code into a commercial product.

The last push to the default branch was on 2026-04-11. The repository is not archived, but five months without a push means you should treat the current state as the state, and plan to maintain your fork yourself.

Finally, the README does not document rollback, retry behaviour, or what happens when Firecrawl returns partial results. If your use case depends on knowing that a failed scrape was retried, this repository does not answer the question.

## How it compares to GPT Researcher and to hosted deep research modes

The closest thing to a stated alternative in the README is the community Python port linked from it, deep-research-python, which reimplements the same loop for a Python environment. If your stack is Python, that is the direct swap, and the difference is language and ecosystem rather than approach.

The more meaningful comparison is with GPT Researcher. That project is also an open source research agent, but it is structured as a fuller application with a configurable report generation pipeline and its own retrieval abstractions, and it is substantially larger. The difference in approach is exactly the one this repository's README advertises: here the loop is small enough to read end to end, and there is no abstraction layer between you and the search calls.

Hosted deep research modes in ChatGPT, Gemini, and Perplexity are a different category again. They handle retrieval, ranking, and citation for you and give you a polished report. They also give you no way to change how queries are generated or how learnings are carried between iterations. Choosing this repository means choosing control over convenience, and accepting that you own the retrieval quality.

## Licence and the cost of keeping a fork alive

The LICENSE file and the README both state MIT, which permits commercial use and modification with attribution. The package.json license field says ISC, which is functionally similar in permissiveness but is a different identifier. This is not legal advice; if the distinction matters to your organisation, resolve it with the copyright holder rather than assuming either field governs.

The upgrade cost is low in one sense and high in another. Low, because there are no releases to track and no dependency upgrade treadmill imposed by a maintainer. High, because the runtime cost is per query and scales with breadth times depth, and the model choice determines that cost. Running o3-mini through OpenAI, R1 through Fireworks, or a local model through OPENAI_ENDPOINT are three very different bills for the same loop. The .env.example CONTEXT_SIZE value of 128000 hints at how much context the loop is designed to carry, and larger context means more tokens per iteration.

## Conclusion

Adopt this if you want to read and modify the whole research loop yourself, and you already have a Firecrawl key and an OpenAI key or a local OpenAI-compatible endpoint. Skip it if you need a hosted product, a stable API surface, or a test suite, since package.json defines the test script as an error stub and no releases are published. Before relying on it, verify two things: that your Firecrawl plan tolerates the CONCURRENCY_LIMIT you set, and that the model you point OPENAI_ENDPOINT at actually returns the structured output the loop expects.

## FAQ

### What is dzhng/deep-research?

It is a TypeScript implementation of an iterative research agent that combines search engines, web scraping, and large language models. The README describes the goal as the simplest implementation of a deep research agent, kept under 500 lines of code.

### Is dzhng/deep-research free?

The code is MIT licensed and free to use and modify, but running it requires a Firecrawl API key and an OpenAI API key, or a self-hosted Firecrawl instance and a local OpenAI-compatible endpoint. The README notes that a free Firecrawl tier may hit rate limits and should be run with a concurrency limit of 1.

### How do I use dzhng/deep-research?

Install dependencies with npm install, put FIRECRAWL_KEY and OPENAI_KEY in a .env.local file, then run npm start. You are prompted for a query, a breadth value, a depth value, and follow-up questions, and the result is written as report.md or answer.md in your working directory.

### Does dzhng/deep-research have deep research mode like ChatGPT?

The repository implements its own iterative research loop rather than integrating with ChatGPT's deep research mode. It generates SERP queries, extracts learnings, and recurses based on a depth parameter you set at the prompt.

### How do I add deep research to my own project with dzhng/deep-research?

The README states the goal is to keep the repository under 500 lines so it is easy to build on top of. The source lives under src/, with run.ts as the interactive entry and api.ts as an Express server entry, so you can call the loop from your own code rather than the CLI.

## Sources

- [dzhng/deep-research on GitHub](https://github.com/dzhng/deep-research)
- [Issues](https://github.com/dzhng/deep-research/issues)
- [License: MIT](https://github.com/dzhng/deep-research/blob/main/LICENSE)
- [README](https://github.com/dzhng/deep-research/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/dzhng-deep-research
