node-DeepResearch: the agent that loops until it knows
Keep searching, reading webpages, reasoning until it finds the answer (or exceeding the token budget)
At a glance
- What is it?
- node-DeepResearch is Jina AI's open-source TypeScript agent that keeps searching, reading webpages and reasoning until it finds an answer or exhausts its token budget. It deliberately targets concise, deeply searched answers over long-form generated reports, runs on Gemini, OpenAI or local models, and exposes an OpenAI-compatible server.
- Who is it for?
- Use node-DeepResearch when your questions need iterative web investigation and you want the answering loop under your own control, on your own model provider, with citations rather than a polished essay as the output. Use the commercial Deep Research products when the deliverable is a long-form report and you accept their closed loops.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 151 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Search, read, reason, repeat until the budget dies
The entire architecture of node-DeepResearch is one loop, drawn in the README as a flowchart: a query enters, and the agent searches, reads, reasons, then searches again, continuing until an answer emerges or the token budget is exceeded. That second exit condition is the honest one, because iterative agents have no natural stopping point and a budget is how you guarantee termination. The README then draws a boundary that most projects in this space avoid drawing: unlike the Deep Research features of OpenAI, Gemini and Perplexity, this project focuses solely on finding the right answers through the iterative process and explicitly does not optimize for long-form articles, which it calls a completely different problem. If you need quick, concise answers from deep search, the framing says, you are in the right place, and if you want generated reports, you are not. That clarity about what a tool refuses to do is rarer and more useful than most feature lists. The loop is also the answer to a question every search user has asked, when is the first result not enough: here the agent itself makes that judgment, and the token budget is the only arbiter of how long the judgment may take.
One loop, three reasoning providers
The reasoning inside the loop is pluggable across Gemini, OpenAI and local models, while the searching and reading of webpages runs through Jina Reader, with a free API key carrying one million tokens available from jina.ai. The documented setup is environment variables and a single command:
export GEMINI_API_KEY=... # for gemini
# export OPENAI_API_KEY=... # for openai
# export LLM_PROVIDER=openai # for openai
export JINA_API_KEY=jina_... # free jina api key, get from https://jina.ai/reader
npm run dev $QUERYThe default reasoning model is the latest gemini-2.0-flash, with the OpenAI path one commented pair away. Installation is a git clone, a cd and an npm install, and while the package exists on npm, the README recommends against using it because the code remains under active development, advice that doubles as a statement about where the project sees its maturity. A YouTube video walks the installation for those who prefer it demonstrated.
Demos that count their steps honestly
The demonstration section is unusually candid about quality gradients. A question about the latest Jina AI blog post resolves correctly in three steps, and one about a model's context length in two. Listing all findable Jina employees takes eleven steps and comes out only partially correct, with the author noting their own absence from the result. A future-prediction question, who will be Jina AI's biggest competitor, runs forty-two steps and is called arguably correct, with the author reserving the right to an I-told-you-so moment about Weaviate. The example commands extend the range from no-tool arithmetic, one-step trivia, a thirteen-step ambiguous comparison, to open research questions like election predictions and corporate strategy. Reporting steps alongside correctness, and partial correctness alongside full, sets expectations better than any benchmark table would. Step counts also double as a cost proxy, since every iteration spends tokens on both the searching and the reasoning sides, so the demonstrations quietly teach the economics of the design along with its capabilities.
An OpenAI-compatible server, think tokens and footnotes
Beyond the CLI, the project ships as a server speaking the OpenAI chat completions schema, so any compatible client can drive it, with CherryStudio and Chatbox named as GUI examples. Starting it is one command, optionally with a secret that clients must present as a bearer token:
# Without authentication
npm run serve
# With authentication (clients must provide this secret as Bearer token)
npm run serve --secret=your_secret_tokenThe server listens on localhost port 3000 at POST /v1/chat/completions. Integration guidance for client builders covers the details that matter: the model is a reasoning-plus-search grounded LLM best used for questions needing both, the response may contain think tokens that renderers should handle deliberately, and citations arrive as GitHub-flavored markdown footnotes. The same design backs Jina's hosted API at deepsearch.jina.ai with model name jina-deepsearch-v1, free of the self-hosting entirely, with rate limits between 10 and 30 requests per minute depending on key tier.
Local LLMs, and the structured-output requirement
Local models are supported through Ollama or LM Studio, with a configuration that reuses the OpenAI client pointed at a local endpoint:
export LLM_PROVIDER=openai # yes, that's right - for local llm we still use openai client
export OPENAI_BASE_URL=http://127.0.0.1:1234/v1 # your local llm endpoint
export OPENAI_API_KEY=whatever # random string would do, as we don't use it (unless your local LLM has authentication)
export DEFAULT_MODEL_NAME=qwen2.5-7b # your local llm model nameThe caveat is technical and important: not every LLM works with the reasoning flow, because the agent needs models that support structured output, sometimes called JSON Schema or object output, well. The documented example model is qwen2.5-7b, and the project invites pull requests adding more open-source models to the working list, which frames local support as an actively curated compatibility matrix rather than a blanket claim. The comment noting that a random string suffices for OPENAI_API_KEY, unless the local server has authentication of its own, is a small courtesy that spares newcomers an avoidable failure mode. For privacy-sensitive or cost-sensitive deployments, this path plus the Docker image is the complete self-hosted story.
jsonrepair, duck-duck-scrape and the dependency tells
The package.json reads like a mechanical description of the agent. The Vercel AI SDK, the ai package with @ai-sdk/google and @ai-sdk/openai, drives model interactions across providers. zod and zod-to-json-schema generate the structured output contracts the reasoning loop depends on, and jsonrepair exists to fix the malformed JSON models occasionally emit, a dependency that tells you the authors met reality. jsdom parses fetched pages, sharp processes images from them, and duck-duck-scrape provides a DuckDuckGo search path, exercised by a dedicated search script. The scripts directory names the auxiliary machinery: a query-rewriter for reformulating failed searches, an ngram utility, batch evaluation runners, and a docker test with a five-minute timeout. Express and commander supply the server and CLI surfaces. Nothing in the list is decorative, and each maps to a step in the loop the README drew.
A reference implementation behind a productized API
The repository's own framing positions it as the exact codebase behind Jina's public deployment, with a hosted UI at search.jina.ai, a separate deepsearch-ui repository for the interface, and a stable API at deepsearch.jina.ai, so the open source project functions as the reference implementation of a commercial endpoint. Docker support is complete, a multi-stage node:20-slim build producing a server image whose compose file threads GEMINI, OPENAI, JINA and BRAVE API keys through the environment and exposes port 3000, the Brave key hinting at an additional search provider path. The implementation guide exists as a two-part blog series in English, with Chinese and Japanese editions, and for teams weighing self-hosting against the hosted endpoint the trade is purely operational, the same loop either way, but one side owns the keys, the rate limits and the upgrades while the other reads release notes. Release history is a burst, v1.2.0, v1.3.0 and v1.4.0 all within one February 2025 week, with the last push on 2026-05-01, and the license is Apache-2.0, permissive for building products on the loop.
Editorial conclusion
Use node-DeepResearch when your questions need iterative web investigation and you want the answering loop under your own control, on your own model provider, with citations rather than a polished essay as the output. Use the commercial Deep Research products when the deliverable is a long-form report and you accept their closed loops. Verify first that your chosen model supports structured output, since the reasoning flow depends on it, that your Jina Reader key covers the reading volume you plan, and remember the npm package is explicitly not recommended while development stays active: deploy from the repository or the Docker image instead.
Frequently asked questions
What is node-DeepResearch?
node-DeepResearch is Jina AI's open-source TypeScript agent that keeps searching, reading webpages and reasoning until it finds an answer or exceeds its token budget. It deliberately targets concise, deeply searched answers rather than long-form generated reports.
Which LLMs can node-DeepResearch use?
Gemini, OpenAI models, or a local LLM through Ollama or LM Studio, selected via environment variables. Local models must support structured output, JSON Schema output, to work with the reasoning flow.
Can node-DeepResearch be exposed as an API?
Yes. npm run serve starts an OpenAI-compatible server on localhost port 3000 with a /v1/chat/completions endpoint, optionally protected with a secret bearer token, usable from GUI clients like CherryStudio or Chatbox. Jina also hosts the same capability as its stable DeepSearch API.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jina-ai-node-deepresearch)