AutoResearchClaw: A 23-Stage Autonomous Research Pipeline
Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper.
At a glance
- What is it?
- AutoResearchClaw is an MIT-licensed Python tool that takes a research idea as input and runs it through 23 automated stages to produce a complete academic paper. It supports human-in-the-loop supervision and routes experiments to domain-specialist agents for fields including high-energy physics, biology, and statistics.
- Who is it for?
- AutoResearchClaw is suited for researchers who want to explore whether a hypothesis can survive automated experiment design, execution, and peer-style review without manually running each stage. It is the wrong tool for producing papers that require proprietary lab equipment, wet-lab results, or experimental data that cannot be generated computationally.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 42 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What AutoResearchClaw Does and Who Needs It
AutoResearchClaw is a Python research automation framework that takes a text description of a research idea and runs it through 23 sequential stages to produce a structured academic paper. It targets researchers and AI engineers who want to evaluate whether a hypothesis can be turned into a complete paper without manually driving each stage of the research process.
The project's arXiv paper is at arxiv.org/abs/2606.23050, titled 'AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration'. The repository description is 'Fully autonomous and self-evolving research from idea to paper'.
The tool is not a writing assistant. It performs literature search, experiment design, code generation, execution, result analysis, and paper drafting in sequence. This scope makes it useful for generating exploratory results quickly and costly to use for research requiring specialized lab environments or datasets that cannot be downloaded or simulated.
The 23-Stage Pipeline: Mechanism and Data Flow
The pipeline splits a research process into 23 stages. The early stages handle idea refinement and hypothesis formation. Middle stages (10 through 13, per the v0.5.0 release notes) execute experiments using domain-specific agents or code generation. Later stages handle quality auditing, writing, and revision.
The v0.2.0 release introduced three multi-agent subsystems: CodeAgent for code generation, BenchmarkAgent for evaluation, and FigureAgent for producing figures. The v0.2.0 notes also describe a 4-round paper quality audit that includes AI-slop detection and a 7-dimension review scoring system with a NeurIPS checklist.
The v0.3.0 release added MetaClaw cross-run learning. Pipeline failures generate structured lessons that are reinjected into all 23 stages. The README states this is opt-in, controlled by setting metaclaw_bridge.enabled: true in configuration, and the release notes report a measured improvement in robustness across controlled experiments when it is active.
The pipeline supports a --resume flag with automatic detection. Runs interrupted partway through can be continued from the last completed stage rather than restarting.
Installing the CLI and Loading Skills
The package is named researchclaw and its CLI entry point is registered in pyproject.toml:
[project.scripts]
researchclaw = "researchclaw.cli:main"The project requires Python 3.11 or later. Dependencies are declared in pyproject.toml under the [project] section; the core set includes pyyaml, rich, arxiv, and numpy. Optional dependency groups add web search (scholarly, crawl4ai, tavily-python), PDF support (PyMuPDF), and a Hugging Face hub client.
The v0.3.2 release notes document a skills system. Pre-loaded skills covering scientific writing, experiment design, chemistry, biology, and more are included. Custom skills can be added with:
researchclaw skills installAlternatively, a SKILL.md file dropped into .claude/skills/ is picked up automatically. The release notes state that 20 pre-loaded skills are included as ready-to-use references.
Domain-Specialist Agents and ARC-Bench
The v0.5.0 release (2026-05-20) introduced two significant additions. First, the experiment execution stages now route to specialist agents based on the research domain. High-energy physics tasks go to a ColliderAgent that runs a Lagrangian-to-simulation chain using FeynRules, MadGraph5, and Delphes via the Magnus cloud. Biology tasks use COBRApy for genome-scale metabolic modelling. Statistics tasks use a simulation-study agent. A generic Docker executor covers chemistry, materials science, and other fields.
Second, the release introduced ARC-Bench: a 55-topic open-ended benchmark spanning ML (25 topics), HEP (10), quantum (10), biology (7), and statistics (3). Each topic ships a manifest containing a research question, conditions, metrics, datasets, and a rubric for graded scoring. The benchmark is available in the repository under experiments/arc_bench/ and also on Hugging Face as AIMING-Lab-UNC/ARC-Bench.
The pipeline auto-selects the appropriate executor from the research domain. The README does not describe how the domain classification decision is made beyond the statement that the pipeline uses automatic routing.
Human-in-the-Loop Co-Pilot Modes
The v0.4.0 release (2026-04-01) added a Human-in-the-Loop system with six intervention modes: full-auto, gate-only, checkpoint, step-by-step, co-pilot, and custom. Per-stage policies allow fine-grained control over which stages pause for human input.
The release also introduced several specific collaboration features: an Idea Workshop for hypothesis co-creation before the pipeline starts, a Baseline Navigator for reviewing experiment design, a Paper Co-Writer for collaborative drafting, and a SmartPause feature that triggers intervention dynamically based on the pipeline's own confidence score.
CLI commands for interactive control include attach, status, approve, reject, and guide. Anti-hallucination claim verification and cost budget guardrails are also part of v0.4.0. Pipeline branching allows parallel hypothesis exploration from a single starting idea.
The full documentation for the co-pilot system is in docs/HITL_GUIDE.md in the repository.
Where the Pipeline Breaks Down
The README acknowledges directly that the system can produce hallucinated claims. The anti-fabrication system in v0.3.2 uses a VerifiedRegistry and an experiment diagnosis-and-repair loop to address this, but claim verification is documented as one of the improvements added in v0.4.0, not a guarantee of correctness in all runs.
The domain coverage is finite. Fields that do not map to any of the specialist agents (ColliderAgent, COBRApy, statistics agent, generic Docker executor) must run under the Docker executor, which is less specialized. The README does not describe what happens when the Docker executor cannot reproduce a required simulation environment.
The pipeline generates computational research only. Physical experiments, proprietary datasets, wet-lab procedures, and any result that cannot be produced by running software are outside its scope. The README does not document what stage count or execution time typical runs require, so cost estimation before running is not straightforward from the documentation alone.
AutoResearchClaw versus Using a General-Purpose Agent Directly
A researcher can use a general-purpose coding agent like Claude Code or Codex CLI directly for research assistance. The fundamental difference is structural: AutoResearchClaw imposes a 23-stage sequence with explicit checkpoints, quality audits, and domain routing, whereas a direct agent session is conversational and unstructured.
This matters for reproducibility. AutoResearchClaw logs each stage and supports --resume, so a run that fails partway through can be continued. A direct agent session typically produces results that are hard to replay exactly. The ARC-Bench benchmark exists specifically to measure reproducibility and quality across the full pipeline, a level of evaluation infrastructure that direct agent use does not provide.
AI Scientist, published by Sakana AI, is another autonomous research system targeting ML research specifically. The key difference in scope is that AutoResearchClaw's v0.5.0 specialist agents extend beyond ML to include high-energy physics, biology, and statistics, and its HITL modes allow human supervision at specific stages. Claims about AI Scientist's current capabilities come from its own documentation rather than this repository.
For teams already using a CLI agent backend, the v0.3.2 release notes state that AutoResearchClaw can run on top of Claude Code, Codex CLI, Copilot CLI, Gemini CLI, or Kimi CLI as agent backends, so it does not require switching infrastructure.
Editorial conclusion
AutoResearchClaw is suited for researchers who want to explore whether a hypothesis can survive automated experiment design, execution, and peer-style review without manually running each stage. It is the wrong tool for producing papers that require proprietary lab equipment, wet-lab results, or experimental data that cannot be generated computationally. Before adoption, verify that the target domain has a compatible specialist agent or can run under the generic Docker executor, and confirm that the LLM backend costs are within the budget controls available in v0.4.0. The last push to the public repository was on 2026-08-19.
Frequently asked questions
How does AutoResearchClaw work?
AutoResearchClaw runs a research idea through 23 sequential stages covering hypothesis formation, literature search, experiment design, code generation and execution, result analysis, paper writing, and quality auditing. Domain-specialist agents handle experiment execution for fields such as high-energy physics, biology, and statistics.
How do you use AutoResearchClaw?
The CLI entry point is the researchclaw command. Skills can be added with researchclaw skills install or by placing a SKILL.md file in .claude/skills/. The Human-in-the-Loop mode documentation is in docs/HITL_GUIDE.md.
What research domains does AutoResearchClaw support?
Version 0.5.0 added specialist agents for high-energy physics (using FeynRules, MadGraph5, and Delphes), biology (using COBRApy for metabolic modelling), and statistics. A generic Docker executor covers chemistry, materials science, and other fields. ML research is handled by the default sandbox.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/aiming-lab-autoresearchclaw)