Open-source project
ryoiki-tokuiten/Deepthink avatar
ryoiki-tokuiten/Deepthink

Deepthink branches one challenge into parallel strategies and feeds the correction step a pool of deliberately wrong artifacts

Using LLMs for iteratively exploring the solution search space at scale.

738 stars69 forksTypeScriptMIT

At a glance

What is it?
A multi-agent pipeline that explores a solution space with roles for strategies, hypotheses, execution, critique and correction, wired through a sandbox with per-branch directories. The README calls Deepthink mode a huge refactoring in progress, and the repository name and the package name are not the same string.
Who is it for?
Worth reading if you are building an agent loop of your own, because the parts that transfer are structural rather than model specific: information packets resolved per branch, a noise pool that exists to break a critique loop, and a filter that can archive a stuck branch instead of feeding it more iterations. Before depending on it, check three things.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 30 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The repository is Deepthink and the package manifest says Iterative Studio

The two names diverge from the first line. The repository is named Deepthink, the heading on the document is Iterative Studio, and the manifest repeats that as the package name with version 0.0.0 and private set to true. There is no version to install from and no tagged release in the repository, so the manifest is the only version statement anywhere. There is also no engines field, which matters because the test scripts execute TypeScript source directly through Node's type stripping rather than through a build step, so the required Node version is implied by the test command rather than declared.

The scripts block is short enough to state in full:

json
"typecheck": "tsc --noEmit",
"test:unit": "node --experimental-strip-types --test Deepthink/DeepthinkContext.test.ts",
"test:sandbox": "node --experimental-strip-types --test Backend/Sandbox_Environment/repositorySnapshot.test.ts",
"test": "npm run test:unit && npm run test:sandbox"

So the checks are two specific test files rather than a glob, one for context handling and one for the sandbox repository snapshot, and there is no install or setup command anywhere in the document, which sends you to the manifest instead.

Deepthink mode is labelled as mid refactoring by its own author

The first operational mode carries the parenthetical huge refactoring in progress in its heading, and the sandbox section later describes its own directory scheme as largely a scaffolding with a cleaner solution being worked on. Both statements come from the same document, so the state described is a working system that its author does not consider settled.

The design underneath is stated as eight ideas: strategies proximity, hypothesis proximity, parallel exploration, iterative corrections and refinements, cross-strategy learning through curated context, independent hypothesis generation and testing, random structured noise injection, and a meta strategies evolving loop that prunes. Each agent role can be steered individually by the user. The intent is to scale inference time compute on any problem or benchmark, and the document claims support for all major providers plus local models. The top level tree shows the same split as the two modes: a Deepthink directory, an AdaptiveDeepthink directory, and shared pieces in Backend/, Contextual/, Core/, Routing/, Styles/ and UI/, with three design notes sitting loose at the root as markdown files.

Each branch is told which hypotheses it is allowed to see

The interesting routing decision is that hypotheses are not shared with every branch. Strategies are generated first, as distinct high level approaches to the challenge, and in a global context they behave like parallel branches each attempting the challenge its own way. Hypotheses are then generated in parallel about the same challenge and each is tested by an independent agent. The stated rationale is token economics, since something tested once with full attention does not need to be tested again by the agent that inherits the result.

What each branch receives is filtered. A hypothesis usually knows which strategies will execute later, so the hypothesis is generated so the executing agent can benefit, and it is resolved in advance down to exactly what a given branch should see in order to keep its context focused. The output of hypothesis work is called an information packet, or a sub-packet when it has been resolved for one strategy. The execution agent then receives the core challenge, its assigned strategy and the information packet resolved for that strategy, produces its work, and has that work critiqued, in every branch in parallel.

The solution pool exists to hand the correction step wrong artifacts

Before a correction is produced, the default path inserts one more agent whose stated purpose is to add random structured noise so a branch is not stuck in a local minimum. The pool holds artifacts, independent helpful blocks, correction approaches, logic fragments, alternative improvements, or complete alternative solutions that the correction agent may benefit from. The document is explicit that these are not necessarily correct, rigorous or complete, and that wrong artifacts belong in there on purpose. That is the noise.

The stated reason for it is blunt: a typical execution, then critique, then correction sequence never works with these models, and the loop gets stuck. Showing the correction agent wrong artifacts or approaches that were explicitly executed removes its restraint, so it starts considering implementation paths it might otherwise think about privately and never put into the delivered work. The analogy given is random structured noise inside a sanity boundary. Once the pool exists, the correction agent receives the previous work, the critique, the information packet, that branch's solution pool, and curated cross-strategy context, and produces an improved result.

A quality filter can archive a branch and hand the slot a new strategy

On top of the loop sit three refresh points. Hypothesis and information packets can be refreshed every k iterations. History can be distilled into memory banks recording what worked, what improved, and which critique patterns persisted. And the strategies themselves can be rewritten after a post-quality filter pass. The trigger for that last one is stated as the degree to which critique and correction loop without a big delta, which is the question of whether the branch is stuck or the strategy is flawed. The filter then decides whether a branch continues, gets refined, or is replaced with an updated strategy.

Replacement is destructive to the branch directory and the layout makes that visible. A replaced branch is archived whole under Pruned_Strategies with a name carrying the first post-quality-filter pass, and later replacements take ordinal successors, before fresh active slot directories are recreated. After a set number of iterations the final corrections are collected and sent to a final judge, which selects the best execution. So the pipeline has two stopping conditions in different places: a per-branch filter that can retire a branch mid-flight, and one final judge at the end.

Sandbox roles submit artifacts, not transcripts

With the sandbox terminal environment enabled, every Deepthink role receives two tools, sandbox_exec and final_output. Active branches use a Strategy-N directory holding Critique and SolutionPool subdirectories, where execution and correction write branch files directly, critique owns its own directory and the pool owns the other. Hypothesis tests are organized as Hypothesis-vN, but only the currently routed tests are mounted to branch workers.

The submission contract is the part with the most teeth. A role that produces JSON submits its existing role-specific object straight through final_output, and the environment validates that contract inside the tool loop, returning a correction error without throwing away the agent's research. Everything downstream, including the central system, receives only the submitted artifact: intermediate command transcripts and scratchpad data are filtered out rather than forwarded. The stated motivation is that downstream agents should spend their tokens on results. The full contracts, repository schemas, mode behaviour, iteration synchronization and failure policy live in a separate document under Deepthink, with the previous architecture diagram kept alongside it.

Adaptive Deepthink removes the final judge and bounds its own revisions

The second mode is described as an orchestrator-directed, pass-based workflow for divergent strategic search, and its distinguishing feature is the absence of a separate final judge. The worker topology pairs a strategy generator with a strategies proximity worker, a hypothesis generator with a hypothesis proximity worker, a hypothesis test role, and then an execution, critique and correction chain per selected strategy. Arrows run both ways between each generator and its proximity partner, so a pair is a loop rather than a hand-off.

Each generation and proximity pair runs a bounded three-round internal revision loop, which is the main contrast with the first mode. Nothing there is described as unbounded: the loop has a round count, and the mode is pass-based rather than continuous. The description stops mid sentence after the bounded revision loop is introduced, so what the orchestrator does between passes, how the pass boundary is chosen, and how a strategy gets selected out of the set are not written down in the document. The directory pair AdaptiveDeepthink/ alongside Deepthink/ in the tree suggests the two modes are separate implementations rather than one configured two ways.

Dependency list spans four model providers and two of them at once

The dependency set is where the provider claim becomes concrete, and it is wider than one SDK per provider. There is a direct openai package, an @anthropic-ai/sdk, a @google/genai, and an @ai-sdk/openai-compatible entry, which is how any endpoint speaking the OpenAI wire format is reached. On top of that sit three LangChain integration packages plus @langchain/core and @langchain/langgraph, so the orchestration layer is available in two different styles within the same project. Whether any given run uses one or both is not stated.

The rest of the list is ordinary application surface. CodeMirror language modes for eleven languages including C++, Go, Python, Rust, SQL and YAML behind a react-codemirror wrapper, a diff and diff2html pair for rendering changes, js-tiktoken for token counting, katex with react-markdown and remark-gfm for rendering model output that contains math and tables, msgpack for serialization, nanoid for identifiers, axios, lucide-react icons and React 19 on both react and react-dom. The user interface is served by vite with a preview script, so the pipeline runs as a local web app rather than a command line tool.

Editorial conclusion

Worth reading if you are building an agent loop of your own, because the parts that transfer are structural rather than model specific: information packets resolved per branch, a noise pool that exists to break a critique loop, and a filter that can archive a stuck branch instead of feeding it more iterations. Before depending on it, check three things. Deepthink mode is described by its own author as mid refactoring, the manifest declares version 0.0.0 with private set and the repository has no tagged release, so there is no version to pin. And the sandbox directory layout is called scaffolding by the author with a cleaner solution in progress, so scripts written against Pruned_Strategies and Hypothesis-vN naming should expect that to move.

Frequently asked questions

What is the solution pool agent in Deepthink for?

It adds random structured noise, including artifacts that are deliberately wrong or incomplete, so the correction step leaves a local minimum instead of looping through execution, critique and correction without progress.

What are information packets and sub-packets in Deepthink?

They are the output of independently generated and tested hypotheses, with a sub-packet resolved in advance for one specific strategy so a branch only receives the context relevant to it rather than every hypothesis result.

How does Deepthink decide to abandon a stuck branch?

A post-quality filter pass looks at whether critique and correction are looping without a big delta, and it decides whether a branch continues, is refined, or is replaced, archiving the old one under Pruned_Strategies before recreating its slot directories.

What does final_output do differently from sandbox_exec in Deepthink?

sandbox_exec is for iterative exploration and testing, while final_output submits completed work, validating a role specific JSON contract in the tool loop and returning a correction error without discarding the research, so downstream agents see only the artifact.

How does Adaptive Deepthink differ from Deepthink mode?

Adaptive Deepthink is orchestrator directed and pass based for divergent strategic search, runs a bounded three-round revision loop inside each generator and proximity pair, and has no separate final judge.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. ryoiki-tokuiten/Deepthink on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ryoiki-tokuiten-deepthink.svg)](https://hysenlabs.com/projects/ryoiki-tokuiten-deepthink)