GPT Researcher: a planner plus crawler agent that writes cited reports
An autonomous agent that conducts deep research on any data using any LLM providers.
At a glance
- What is it?
- GPT Researcher splits research into a planner that writes questions and crawler agents that answer them, then publishes a cited report. It is a Python 3.11 package with a FastAPI server, a CLI, and an MCP client, and its weakest point is that retrieval quality depends entirely on the search backend you configure.
- Who is it for?
- Adopt GPT Researcher if you need a report generator you can point at your own search backend, your own LLM endpoint, or your own PDFs in my-docs, and you are willing to read the source when a run drifts. Do not adopt it if you need a fixed, auditable pipeline: the planner decides its own questions, so two runs on the same query will not produce the same report.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What GPT Researcher is for, and who actually needs it
The README states the problem plainly: manual research can take weeks, LLMs trained on stale data hallucinate on current topics, and a single model call cannot hold enough context to write a long report. GPT Researcher is built for the case where you want a written, cited document rather than a chat answer. The README describes it as an autonomous agent that "conducts deep research on any data using any LLM providers", producing reports that it says exceed 2,000 words and aggregate over 20 sources.
The audience is narrower than the tagline suggests. If you are a developer wiring an internal research step into a pipeline, the pip package gives you a class you can await. If you want a browser UI, the repository ships two frontends: a lightweight HTML/CSS/JS one and a Next.js plus Tailwind one. If you are an analyst who just wants a report, the Docker Compose file is the shortest path. What GPT Researcher is not is a search engine. It does not index anything itself. Every fact in its output comes from whatever retriever you point it at, which is why the Tavily key in the setup instructions is not optional in practice.
The planner, the crawler agents, and the publisher
The architecture section describes three roles. A planner generates research questions from your query. Execution agents gather information for each question using a crawler agent. A publisher aggregates the findings into the final report. The README frames this as parallelized agent work, and says the design is inspired by the Plan-and-Solve and RAG papers.
The data flow is linear in shape but concurrent in execution. Each generated question becomes its own retrieval task, and each retrieved resource is summarized and source-tracked before aggregation. That source tracking is what produces the citations in the output. The repository also exposes a multi_agents/ directory and a deep_agents/ directory alongside the core gpt_researcher/ package, so the single-planner flow described in the README is not the only agent topology in the tree.
Two design consequences are worth naming. First, question generation is a model call, so the quality of the whole run is capped by how well your model decomposes the topic. A vague query produces vague sub-questions and a shallow report. Second, because summarization happens per resource before aggregation, the final report is a synthesis of summaries, not of raw pages. If a summarization step drops a number, no later stage can recover it.
Installing GPT Researcher and running a first research job
The README requires Python 3.11 or later. Clone the repository, then export your keys. The README shows both an OpenAI key and a Tavily key, and notes that OPENAI_BASE_URL can point at any OpenAI-compatible endpoint.
git clone https://github.com/assafelovic/gpt-researcher.git
cd gpt-researcher
export OPENAI_API_KEY={Your OpenAI API Key here}
export TAVILY_API_KEY={Your Tavily API Key here}Install the dependencies and start the FastAPI server. The README gives this exact pair of commands and says to visit http://localhost:8000 afterwards.
pip install -r requirements.txt
python -m uvicorn main:app --reloadIf you would rather embed the agent than run the server, the README shows the pip package. Note that conduct_research and write_report are both awaited, so this belongs inside an async function.
from gpt_researcher import GPTResearcher
query = "why is Nvidia stock going up?"
researcher = GPTResearcher(query=query)
research_result = await researcher.conduct_research()
report = await researcher.write_report()For a containerised setup, docker-compose.yml defines a gpt-researcher service on port 8000 and a gptr-nextjs service on port 3000, and mounts ${PWD}/my-docs into the container read-write. That mount is the local-document path: files placed in my-docs become available to the agent alongside web search. The same file also defines a gpt-researcher-tests service behind the test profile.
Pointing it at Ollama, MCP servers, or your own documents
The requirements file pulls in langchain-ollama and the ollama package, so a local model is a supported configuration rather than an afterthought. The README does not walk through the Ollama environment variables, so you are relying on the docs site for the exact keys. What the README does document is the custom endpoint escape hatch: setting OPENAI_BASE_URL redirects the OpenAI-compatible calls elsewhere, which is how you front a local or third-party model that speaks the same API.
MCP support is the more interesting extension. Setting RETRIEVER to a comma-separated list enables hybrid retrieval, and the README gives this example.
export RETRIEVER=tavily,mcp # Enable hybrid web + MCP researchThe corresponding Python example passes an mcp_configs list containing a server name, a command, args, and an env block. In the README's sample the server is github, launched with npx and the @modelcontextprotocol/server-github package, with a GITHUB_TOKEN in the environment. The practical effect is that the same run can cite a GitHub repository and a web page. That is a real capability, not a marketing line, but it also means your report's provenance now spans two systems with different freshness guarantees.
The IMAGE_GENERATION_* variables in docker-compose.yml control inline illustrations: IMAGE_GENERATION_ENABLED defaults to false, the model defaults to gemini-2.0-flash-preview-image-generation, and IMAGE_GENERATION_MAX_IMAGES defaults to 3. A report with generated images is harder to audit than one without, which is a reason to leave that flag off for anything factual.
Where GPT Researcher breaks down
The retrieval layer is the failure point. GPT Researcher does not fetch the web itself in any meaningful sense; it delegates to Tavily, to DuckDuckGo via the ddgs dependency, or to whatever RETRIEVER you name. If your retriever returns thin snippets, the crawler has thin material, the summarizer produces thin summaries, and the publisher writes a confident report over nothing. The README's claim of aggregating over 20 sources is a target, not a guarantee, and nothing in the documentation describes what happens when a sub-question returns zero usable results.
Cost and latency scale with the question count, and the question count is decided by the model at runtime. There is no documented cap on how many questions the planner generates, so a broad query can fan out into many retrieval and summarization calls. The README does not document a budget ceiling, a maximum step count, or a rollback mechanism for a run that goes wrong. You stop it and start over.
Determinism is another gap. The README lists determinism among the problems the project addresses, but a planner that generates its own questions is inherently non-deterministic across runs. If your use case requires the same input to produce the same output for audit purposes, this is the wrong tool. A fixed retrieval pipeline with a templated summarization prompt would serve you better, even though it produces duller reports.
GPT Researcher compared with STORM and Perplexity
The closest open source comparison is STORM, which people search for alongside this project. Both generate research questions and both produce cited long-form output. The difference is in the control surface. STORM is a research and writing system you run to produce an article; GPT Researcher is packaged as a service, with a FastAPI app, a Docker Compose stack, a Next.js frontend, and a pip-installable class. If you want to embed research as a step inside a larger application, GPT Researcher's Python API and MCP client are the reason to pick it. If you want to study or reproduce a specific question-generation method, STORM's narrower scope is easier to reason about.
Perplexity is a different category entirely. It is a hosted product with its own index and its own answer format. GPT Researcher has no index. That is the trade: you supply the retrieval, and in exchange you can point the agent at a private MCP server, a local model through Ollama, or a folder of documents mounted at my-docs. For anything involving internal data, that difference decides the choice. For a quick public-web answer, Perplexity is faster and cheaper than standing up a FastAPI service and paying for two API keys.
Licence, maintenance, and the cost of upgrading
The repository's LICENSE file is Apache-2.0, and the GitHub metadata lists the project as Apache-2.0. The packaging files disagree: pyproject.toml declares license = "MIT", and setup.py declares license="MIT" with an MIT classifier. Both files also carry version 0.14.7 while the releases page shows v3.6.1 as the most recent tag. Those two version numbers do not track each other, and anyone pinning a dependency should check which one their tooling reads. This is a factual discrepancy in the tree, not a legal opinion, and if the licence matters to your organisation you should read the LICENSE file yourself.
The last push to the default branch was on 2026-08-24, the same day v3.6.1 was tagged as a major-fix release. That is recent enough that the project is not dormant, and there is no archive flag on the repository. The upgrade cost is the part to plan for. requirements.txt pins langchain>=1.0.0 and langgraph>=0.2.76, alongside numpy>=2.0.0,<2.3.0, so a LangChain major-version bump ripples through the whole dependency set. The repository also carries an evals/ directory and a tests/ directory with pytest configured in pyproject.toml, which gives you a way to check a version bump against the project's own tests before you ship it. Note that numpy is capped below 2.3.0, so a transitive dependency demanding a newer numpy will conflict.
Editorial conclusion
Adopt GPT Researcher if you need a report generator you can point at your own search backend, your own LLM endpoint, or your own PDFs in my-docs, and you are willing to read the source when a run drifts. Do not adopt it if you need a fixed, auditable pipeline: the planner decides its own questions, so two runs on the same query will not produce the same report. Before committing, verify three things on your own key: whether your retriever returns enough usable snippets, whether your model follows the structured output format the publisher expects, and whether the report's citations actually resolve to pages that support the sentences they are attached to.
Frequently asked questions
What is GPT Researcher?
It is an autonomous agent that conducts research on a query and produces a cited report, using a planner to generate questions, crawler agents to gather sources, and a publisher to aggregate the findings. The README describes it as designed for both web and local research, and it ships as a pip package, a FastAPI server, and a Docker Compose stack.
How do I use GPT Researcher?
The README gives two paths. Install the pip package and await conduct_research() followed by write_report() on a GPTResearcher instance, or clone the repository, set OPENAI_API_KEY and TAVILY_API_KEY, run pip install -r requirements.txt, and start the server with python -m uvicorn main:app --reload on port 8000.
Is GPT Researcher free?
The source is licensed Apache-2.0 per the LICENSE file and repository metadata, so there is no licence fee. Running it still costs money: the README's setup requires an OpenAI API key and a Tavily API key, and a local model through Ollama avoids the first of those but not the retrieval service.
How does GPT Researcher compare with Perplexity?
Perplexity is a hosted product with its own index and answer format, while GPT Researcher has no index and relies on the retriever you configure. The trade is that GPT Researcher can be pointed at a private MCP server, a local model via Ollama, or documents mounted at my-docs, which a hosted product cannot do.
How does GPT Researcher compare with STORM?
Both generate research questions and produce cited long-form output. GPT Researcher is packaged as a service with a FastAPI app, a Docker Compose stack, two frontends, and a pip-installable class, while STORM is narrower in scope, which makes its question-generation method easier to study and reproduce.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/assafelovic-gpt-researcher)
Community notes