Model or dataset
OpenRaiser/NanoResearch avatar
OpenRaiser/NanoResearch

NanoResearch ships two dependency lists and the second one is short

🦞+🔬 NanoResearch: The Autonomous AI Research Assistant

1,337 stars90 forksPythonMIT

At a glance

What is it?
A nine-stage pipeline that takes a research topic to a LaTeX paper, and the stage that makes it different from a writing tool is the one that submits the generated training code to a local GPU or a SLURM cluster, reads the real logs back, and builds the figures from them. It also retries failed runs twice after editing its own code. The packaging is shakier than the pipeline, and the section that promises measured figures contains none.
Who is it for?
NanoResearch is worth reading for the execution stage, because a pipeline that actually runs the experiment and refuses to fabricate the numbers is rarer than the marketing around it suggests. Two things to weigh before you point it at a cluster.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 39 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Two dependency lists exist and one of them is missing a declared package

The project manifest and a plain requirements file both enumerate the runtime dependencies, and they are not the same list. The manifest declares thirteen. The requirements file lists twelve. The package that is in the first and absent from the second is the YAML parser, and the manifest pins it with a lower bound like everything else. There is nothing subtle about the omission; it is one line. What makes it more than cosmetic is which file the instructions tell you to use. The quick start clones the repository and installs the project in editable mode with its development extra, which resolves from the manifest, not from the requirements file. So the file nobody is told to use is the file that is wrong, and it will stay wrong until someone runs a tool that diffs the two, because nothing in the build reads it. It is also the kind of drift that gets caught the moment someone installs for a different reason than the documented one.

The section that proves the figures are real contains no figures

The page opens with a claim worth believing if it is true: that the tool actually runs computational experiments, submits the generated code to a GPU cluster, collects real results, generates the figures, and outputs a complete LaTeX paper whose every number, table and chart comes from an experiment that ran rather than from a model's invention. It makes that claim three times in three places, including in the headline callout and in a comparison table. Then it presents the evidence, and the evidence is four tables. Each table has two or three cells. Each cell contains a caption and a line break. None of them contains an image. So the section whose stated purpose is 论文实测展示, the display of figures measured from real runs, renders as a grid of empty labelled boxes, and the closing note asserting that the figures were generated by the pipeline and sourced from real training logs is attached to nothing. The claim could still be true. On this page, it is unevidenced.

The execution stage edits its own code and resubmits, twice

This is the part that separates the project from a writing tool, and it is one bullet in a list. When a training run fails, the pipeline analyses the error log, fixes the code and re-executes. The configuration sets the retry budget to two by default, and the stage also submits batch scripts to a cluster, monitors job status, detects local GPUs, and streams logs and metrics as the run progresses. It can switch between local and cluster execution automatically depending on task complexity, and the default profile is a quick local one. Put the three facts together and you have a system that writes training code, runs it on real hardware, reads the failure, rewrites the code, and spends more compute, up to a configurable limit. That is the honest version of autonomous research, and it is also the part with the largest blast radius: the cost of a bad suggestion is measured in GPU hours rather than in a wrong sentence.

The resume command ships an unsubstituted placeholder

Resumability is a headline feature: any stage can be recovered, and there is a dedicated command for it.

bash
nanoresearch resume --workspace ~/.nanoresearch/workspace/research/{session_id} --verbose

The workspace argument ends in a brace expression that nothing on the page tells you how to fill. The export command beside it has the same placeholder in the same position, and so does the launch example in the verification step. So the three commands that constitute the practical workflow of a long run all carry an unresolved template variable, and a reader has to infer both what the session identifier is and where it is printed. Given that the export step is the one that produces the paper, this is an unfortunate place for the documentation to be eliding its own argument. A dry-run command is provided for verification, which is the right mitigation, though the page does not point out that it is also the way to discover the workspace path.

The only release in the repository is a bundle of static assets

The release history has one entry and its name is assets, with a description of static assets. It is dated in March and it is not a version. Meanwhile the project manifest reads zero-point-one-point-zero, and there is no tag anywhere in the list that corresponds to it. So the releases page is being used as a deployment bucket for the web front end rather than as a version history, which means there is no tag to pin, no changelog to diff between versions, and no way to ask the repository which commit you have. For a tool whose selling point is reproducibility, and whose stated use case includes generating reproducible experiment results in batch, that is a gap worth noticing. The last commit to the default branch came at the end of August, so the branch is moving while the release surface is not.

Per-stage temperature is the real routing mechanism

The comparison table claims multi-model collaboration where other tools have a single model, and the mechanism it describes is routing by stage. The configuration makes that concrete, because each stage carries its own model, temperature, token ceiling and timeout. Ideation runs at a temperature of one half. Planning drops to two tenths. Code generation drops again to one tenth. All three carry the same token ceiling of sixteen thousand and the same ten-minute timeout. So the multi-model claim is really two claims: you may route stages to different models, and the project lowers temperature monotonically as it moves from ideas to code, which is a sensible default for a pipeline whose output is executable. What is not documented is which model names are valid, what the temperature knob does at each stage in practice, or whether the same model can be used throughout, since the example simply repeats your own model in all three places.

Four front ends over one engine, including an MCP server at the top level

The pipeline is reachable four ways, and the repository layout shows it. There is a command line, which itself has two presentations: a full-screen panel interface and a plain streaming log, the latter chosen for redirecting and for script integration, with the stated reason that the plain form suits automation. There is a mode for working inside a coding agent. There is a messaging bot for a Chinese workplace chat platform. And there is a server directory at the top level for the Model Context Protocol, which means the whole research pipeline can be driven by an external tool rather than only by its own CLI. That is a broader surface than the quick start suggests, and it is the reason the repository has both a CLI entry point declared in the manifest and a top-level server package. It also means there are four places where a bug in the same stage logic would show up.

An unexplained training extra sits outside the described scope

The manifest declares three optional dependency groups and two of them are obvious. The development group brings a test framework, its async plugin and an HTTP mocking library. The mathematics group brings a computer algebra library. The third does not match anything in the visible description: it pulls a deep learning framework, a transformers library, a parameter-efficient fine-tuning library, an accelerator library and a safetensors reader, under the name sdpo. Nothing on the page explains what sdpo stands for or what it enables, and the dependency list is large enough that installing it is not a small decision. The core dependency set is otherwise modest and readable, thirteen packages covering schemas, one provider client, HTTP, parsing, templating, plotting, a CLI framework, terminal output, the protocol server, imaging and PDF handling. The mismatch between that tidy core and one unexplained training stack is the thing to ask about.

Editorial conclusion

NanoResearch is worth reading for the execution stage, because a pipeline that actually runs the experiment and refuses to fabricate the numbers is rarer than the marketing around it suggests. Two things to weigh before you point it at a cluster. It edits its own code and re-submits when a training run fails, with a configurable retry count, so the blast radius of a bad suggestion is a compute budget rather than a wrong paragraph. And its own claims are unevenly supported: the assertion that every figure comes from a real run is repeated three times, while the section meant to demonstrate it contains only labels.

Frequently asked questions

What does NanoResearch actually do that a writing tool does not?

It runs the experiments. The pipeline generates a training project, submits it to a local GPU or to a SLURM cluster, monitors the job, parses the real training logs into structured evidence, and generates the paper's figures from that data before writing the LaTeX from it.

How does NanoResearch handle a failed training run?

The execution stage analyses the error log, fixes the generated code and re-executes, with the retry count set in configuration and defaulting to two. Execution can also switch automatically between local and cluster depending on task complexity, and the default profile is a quick local one.

How do I install and configure NanoResearch?

Clone the repository and install it in editable mode with its development extra, which needs Python 3.10 or newer. Then write a configuration file in your home directory pointing at your own OpenAI-compatible endpoint and key, with a template format, execution profile, writing mode, retry count and per-stage model settings; three environment variables can override the base URL, key and timeout.

Can NanoResearch resume a run that failed partway through?

Yes, by design: any stage can be recovered. The resume command takes a workspace path, and the documented examples for both resume and export carry an unsubstituted session identifier placeholder, so you will need to fill that in from your own run. A dry-run flag is available for verifying a configuration first.

What are the stages in the NanoResearch pipeline?

Nine, in order: ideation with literature search and hypothesis generation, planning into an experiment blueprint, environment setup, code generation, execution on local hardware or a cluster, analysis of logs and metrics, figure generation, LaTeX writing from the collected evidence, and a review stage that checks and revises the sections. A separate self-evolving variant adds skill and memory evolution across multiple rounds.

Official sources

  1. Issues
  2. License: MIT
  3. OpenRaiser/NanoResearch on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/openraiser-nanoresearch.svg)](https://hysenlabs.com/projects/openraiser-nanoresearch)