Open-source project
EvoMap/AutoResearch avatar
EvoMap/AutoResearch

AutoResearch: A Stateful Agent Workflow from Research Idea to Paper-Ready Evidence

AI/ML research agents from idea to paper-ready evidence. An EvoMap open-source project.

2,805 stars249 forksPythonApache-2.0

At a glance

What is it?
AutoResearch is an open-source Python workflow for AI and ML research that takes a research idea through signal discovery, experiment planning, implementation, execution, and independent review to produce a traceable evidence package. It stores all intermediate state to disk so researchers can inspect, pause, or resume any stage.
Who is it for?
AutoResearch is the right tool for ML researchers who want a structured, auditable pipeline from research signal collection to experiment results, and who have access to multiple LLM API endpoints. It is not a tool for general web research or for producing research papers automatically without researcher oversight: the README is explicit that the system cannot guarantee every conclusion is correct and is designed to preserve evidence for researcher review.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Problem AutoResearch Addresses in AI Research

Large language models can generate plausible-sounding research ideas and experimental claims, but they tend to validate their own output, cite sources they invented, and present negative results as successes. A single LLM acting as both idea generator and evaluator is prone to circular validation.

AutoResearch addresses this by grounding idea generation in real external signals (recent ArXiv papers, community discussions, open-source trends), adding domain knowledge from a local knowledge base maintained by the researcher, and running important claims through three or more distinct LLM models in a cross-model review stage. The README describes the goal as keeping hallucination out of the process: real signals in, unsupported generation minimized.

The system also handles negative results explicitly. Instead of forcing every experiment toward a success story, AutoResearch preserves evidence and stops when a hypothesis fails. The README describes this as supporting negative results by recording failure causes in the logs.

The last push was on 2026-09-16. There are no GitHub releases.

Three Starting Points: Idea Discovery, Direct Execution, and Complete Workflow

The README describes three paths into the workflow. When a researcher does not have a specific idea, they run idea_generation.py to discover candidate directions from recent external signals and cross-review them. When a researcher already has an idea, they go directly to the experiment execution stage. The complete workflow runs idea generation first and then moves into execution for a selected plan.

These paths are reflected in the repository structure. idea_generation.py is the entrypoint for the discovery path. run_pending_forge.py resumes idea seeds already present in data/pending_forge_seeds.json without re-collecting online sources. A check_idea.py script is available for evaluating an existing idea against the pipeline.

The workflow is stateful. Research plans, code, run logs, metrics, failure causes, critic reports, and blind reviews are written to disk. This means a long-running experiment can be interrupted and resumed. The README describes this as the workflow being stateful and recoverable.

Setting Up AutoResearch: Clone, Install, and Configure

AutoResearch requires a Linux or SSH machine with Git, Python 3.10 or later, and python3-venv. Clone the repository and run the bringup script:

bash
git clone https://github.com/EvoMap/AutoResearch.git
cd AutoResearch
bash scripts/bringup.sh

The bringup.sh script creates a virtual environment in .venv, installs Python dependencies, runs baseline tests and a secret scan, and checks the current model configuration. It does not contact any model API on the first run. A BLOCKED result at the end of the first run is expected until API credentials are configured.

Create local configuration files without overwriting any existing ones:

bash
test -f .env || cp .env.example .env
test -f config/providers.local.json || \
  cp config/providers.example.json config/providers.local.json

Edit .env with real API credentials. The .env.example shows the expected environment variables: ANTHROPIC_API_KEY, GEMINI_API_KEY, OPENAI_API_KEY, and optionally TAVILY_API_KEY for web search. Edit config/providers.local.json to declare which models handle which roles (screener, judge, ideator, planner, agent, critic, and so on).

After configuring credentials, run the preflight check:

bash
set -a
. ./.env
set +a
.venv/bin/python scripts/preflight.py --live

This sends a small number of real API requests. Exit code 0 means all normal roles have models assigned and that multi-model stages meet the independence requirement. AutoResearch does not require a fixed combination of Anthropic, Google, or OpenAI endpoints. You can use any compatible LLM APIs, as long as the cross-review stages have access to at least three distinct underlying model identities.

Idea Generation: How Cross-Domain Discovery Works

The idea generation pipeline runs in five stages. Stage 1 collects recent research signals from several public channels including ArXiv, community discussions, and open-source trends. Stage 2 aggregates, deduplicates, screens, and deeply assesses candidate directions. Stage 3 intersects each accepted signal with selected directions from the local knowledge base in knowledge_base/. Stage 4 asks three or more distinct models to develop ideas independently and cross-review the candidates. Stage 5 checks freshness and agreement, then produces experiment plans for accepted ideas.

The local knowledge base is a directory of Markdown files where the researcher records domain knowledge, constraints, and common failure patterns. AutoResearch reads from it but never rewrites it automatically.

The cross-review stage is the mechanism that reduces circular validation. The README specifies that at least three distinct models must participate, meaning a single endpoint that happens to expose multiple model names counts only once if the underlying models are the same. The preflight check verifies this independence requirement.

Experiment Execution and What Gets Written to Disk

When running in execution mode, AutoResearch takes a selected experiment plan and drives an agent through implementation, pilot testing, full-scale experiment execution, result analysis, critic review, and blind review. The README describes this as a complete record from research signals to paper-ready evidence.

The pilot-before-scaling feature tests feasibility at lower cost before committing to full experiments. If the pilot fails, the workflow records the failure and stops rather than continuing into full-scale runs.

All outputs are written to the outputs/ directory (configurable via AUTORESEARCH_OUTPUT_DIR in .env). Logs go to logs/. Data artifacts go to data/. The ARCHITECTURE.md file in the repository describes the relationship between these directories and the pipeline stages.

The requirement for negative results means that when an experiment does not support the hypothesis, AutoResearch writes the failure evidence to disk and records the stop reason. The README explicitly states that the system will not force every experiment into a success story.

Limitations and What AutoResearch Cannot Do

AutoResearch produces evidence packages, not finished papers. A researcher is expected to review the outputs, evaluate the methodology, and write the paper from the collected evidence. The README is direct about this: the system cannot guarantee that every conclusion is correct, and the design goal is to preserve evidence for researcher review, not to replace it.

The system requires multiple LLM API endpoints with different underlying model identities for the cross-review stages. A researcher with access to only one provider or one model cannot meet the independence requirement for Idea Forge and critic stages.

The requirements.txt lists Python dependencies including httpx, beautifulsoup4, feedparser, arxiv, and openreview-py. Optional GPU and Claude Code dependencies are managed separately, per the comment in requirements.txt. The AUTORESEARCH_PROXY_URL environment variable supports routing API calls through a proxy, which may be necessary in institutional network environments.

The workflow is designed for AI and ML research. It is not a general-purpose web research agent and does not document use for other research domains such as biology, economics, or social science.

How AutoResearch Compares to GPT-Researcher

GPT-Researcher is a Python-based open-source agent that conducts web research on a given question and produces a structured report. It is a general-purpose research tool oriented toward gathering and summarizing information from the web.

AutoResearch targets a different task. It is designed for ML research where the output is not a summary of existing information but an evidence package from original experiments. The key differences are: AutoResearch runs actual ML experiments, not just web searches; it requires multiple distinct LLM models for cross-model review of ideas; and it produces structured logs, critic reports, and blind reviews alongside code and metrics, not just text summaries.

For a researcher who needs to survey the literature on a topic, GPT-Researcher is the appropriate tool. For a researcher who wants to go from a research hypothesis to an experimentally supported evidence package, AutoResearch addresses that end of the pipeline.

Editorial conclusion

AutoResearch is the right tool for ML researchers who want a structured, auditable pipeline from research signal collection to experiment results, and who have access to multiple LLM API endpoints. It is not a tool for general web research or for producing research papers automatically without researcher oversight: the README is explicit that the system cannot guarantee every conclusion is correct and is designed to preserve evidence for researcher review. Before running the full workflow, run scripts/preflight.py --live to confirm that your configured models meet the independence requirements for multi-model review stages.

Frequently asked questions

What is AutoResearch?

AutoResearch is an open-source Python agent workflow for AI and ML research. It takes a research idea through signal discovery, multi-model idea review, experiment planning, code implementation, execution, and independent blind review to produce a traceable evidence package ready for paper writing.

How do I install AutoResearch?

Clone the repository, then run bash scripts/bringup.sh. This creates a virtual environment, installs dependencies, and runs baseline tests. Then copy .env.example to .env and config/providers.example.json to config/providers.local.json and fill in your LLM API credentials.

How do I use AutoResearch to generate research ideas?

Run .venv/bin/python idea_generation.py. This collects recent signals from ArXiv and other public channels, intersects them with your local knowledge base, runs cross-model idea review with at least three distinct models, and produces experiment plans for accepted ideas.

Official sources

  1. EvoMap/AutoResearch on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/evomap-autoresearch.svg)](https://hysenlabs.com/projects/evomap-autoresearch)