# ARIS (Auto-claude-code-research-in-sleep): Markdown Skills for Overnight ML Research

> ARIS is a set of Markdown-only skills that run an autonomous research loop inside Claude Code, Codex CLI, OpenClaw or any other LLM agent. The pitch is cross-model adversarial review; the catch is that it is a methodology, not a platform, and you supply the models.

**wanshuiyin/Auto-claude-code-research-in-sleep** — ARIS (Auto-Research-In-Sleep), Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in, works with Claude Code, Codex, OpenClaw, or any LLM agent.

- Repository: https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep
- Stars: 16,681 · Forks: 1,417
- Language: Python
- License: MIT
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/wanshuiyin-auto-claude-code-research-in-sleep

## The problem ARIS targets: research runs that stop when you close the laptop

Most LLM-assisted research workflows are interactive. You ask a question, the model answers, you read the answer, you ask again. The agent is idle whenever you are not typing. ARIS is built around the opposite assumption: the loop should keep running while you sleep, and the human should review a finished artefact rather than supervise each step.

The repository describes itself as "lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation." That sentence is the whole scope. There is no runtime to install beyond the agent you already use, no database, and no server process that owns your experiment state. The skills are instruction files that an agent reads and follows.

Who it is for: people who already run Claude Code, Codex CLI, Cursor, Trae, Antigravity, GitHub Copilot CLI, OpenClaw or DeepSeek Harness from a terminal, and who treat the agent as a collaborator on a research task rather than a chat window. If your work is a single question with a single answer, the loop machinery adds overhead you will not recover.

## How the ARIS loop works: author model, reviewer model, research wiki

The mechanism visible in the repository is a division of labour between models. One model authors the work; a different model reviews it. The README states plainly that on DeepSeek Harness "Codex still the independent reviewer," which is the clearest statement of the design intent: the reviewer is deliberately not the author.

The .env.example file shows how many reviewer backends exist. There are blocks for Gemini review (GEMINI_REVIEW_*), Claude review (CLAUDE_REVIEW_*), a generic LLM chat endpoint (LLM_*), and a manual review mode (MANUAL_REVIEW_*) that the file annotates as "human-in-the-loop, zero API cost." Each backend gets a server name, a binary or API key, a timeout, and a debug log path. The timeouts are generous by default: 600 seconds for Gemini and Claude, 86400 seconds for manual review.

Memory is handled by what the README calls a research-wiki. The stated reason is concrete: long stories break when the model forgets earlier details or judges its own work. The wiki is the persistent record the loop reads back, and the separate reviewer is the check on self-assessment. Those two pieces, external memory plus a non-author reviewer, are the load-bearing parts of the architecture. Everything else is prompt text.

The manual review mode is worth noting because it inverts the usual assumption. With MANUAL_REVIEW_MODE=browser, MANUAL_REVIEW_AUTO_OPEN=true and MANUAL_REVIEW_PORT=17900, a run can pause and wait for a person instead of an API. The default timeout of 86400 seconds means it will wait a full day.

## Installing ARIS and running a first review loop

There is no package to pip install for the core skill set. The README directs you to use ARIS "as a skill-based workflow" inside an agent you already have, and points at per-agent adaptation documents: docs/CURSOR_ADAPTATION.md, docs/TRAE_ARIS_RUNBOOK_EN.md, docs/ANTIGRAVITY_ADAPTATION.md, docs/COPILOT_CLI_ADAPTATION.md and docs/OPENCLAW_ADAPTATION.md. Codex users are pointed at skills/skills-codex/.

The one install command the README gives verbatim is for DeepSeek Harness, where ARIS ships as a plugin. It fetches from npm itself, so there is no separate install step, but the README warns that pnpm must be on PATH:

```bash
# DeepSeek Harness only: installs all 82 skills as one plugin
dsh plugin --profile web add dsh-aris
```

Configuration is driven by environment variables. The repository ships .env.example with every key the reviewer backends read. Copy it and fill in the backends you intend to use. The Gemini reviewer block, for example, expects an API key and a model name:

```bash
cp .env.example .env
# then edit .env, for example:
GEMINI_API_KEY=your-key-here
GEMINI_REVIEW_BACKEND=api
GEMINI_REVIEW_API_MODEL=gemini-2.5-flash
```

A first real use does not require any of the API backends. The manual review mode is the cheapest way to see the loop work end to end, because it substitutes a browser page for a paid model:

```bash
MANUAL_REVIEW_MODE=browser
MANUAL_REVIEW_AUTO_OPEN=true
MANUAL_REVIEW_PORT=17900
MANUAL_REVIEW_PENDING_DIR=.aris/pending_review
```

With those set, a run that reaches a review step opens a browser page and waits. What you should see is a pending review item under .aris/pending_review and a process that has not exited. That is the loop pausing for you, not failing.

There is also a built-in monitor for macOS. The README says it needs no clone, no pip and no browser, and gives this:

```bash
cd aris-monitor && ./run.sh
# a borderless panel floats top-right; click a row to jump to that terminal
```

The panel lights up when a session is waiting for your approval. On any other platform, the README points at the third-party Claude Fleet dashboard instead.

## Where ARIS stops being the right tool

ARIS is explicitly a methodology rather than a platform, and the README says so: "ARIS is a methodology, not a platform. What matters is the research workflow." Take that seriously, because it means several things you might expect are simply not there.

There is no scheduler. Nothing in the repository decides when a run starts or how many runs to launch. The "in sleep" part of the name describes when you are, not a cron facility. If you want jobs to fire at 02:00, you build that yourself.

There is no execution sandbox for experiments. The description mentions experiment automation, but the repository layout shows skills, templates, tools, mcp-servers and tests. Nothing there constrains what a generated experiment script does to your machine or your cluster. The reviewer model checks reasoning and claims; it does not check that a shell command is safe.

Cross-model review costs money and needs keys. Every reviewer backend except manual review reads an API key from the environment. If you only have one provider, the adversarial property collapses: the same model family authors and reviews, and the README's own framing about models that "judge their own work" applies to you.

The open question the repository does not answer is reproducibility of the loop itself. Two runs of the same skill with the same models may take different paths, and the README does not document a seed, a pinned model snapshot, or a way to replay a run. If you need an auditable research pipeline, that gap matters more than any feature.

## ARIS against a plain Claude Code session

The honest alternative is the thing you already have: a single Claude Code or Codex CLI session with a well-written CLAUDE.md or AGENTS.md and a human reading the output. That setup has one model, one context, and a person in the loop by default.

The difference is structural, not cosmetic. A plain session has no second model to disagree with the first, and no external wiki to survive context loss. ARIS adds both, at the cost of configuration. You must decide which model reviews which, supply keys for each, and accept that a review step can block for up to 600 seconds by default, or 86400 in manual mode.

The repository's own ecosystem shows the trade-off. The HERO project is described as a roughly 550-token block for CLAUDE.md or AGENTS.md that bounds what an agent proposes, and the README notes it exists because "ARIS's reviewer is good, and it also proposed hashes nobody reads." That is a candid admission from the same author: an independent reviewer can over-engineer, and a short constraint file may be a better fix than another review round. If your problem is an agent that proposes too much, HERO is the lighter option. If your problem is an agent that cannot tell whether its own result is real, the cross-model loop is the relevant one.

## Licence, maintenance and the cost of upgrading

ARIS is MIT licensed. In practice that means you can copy the Markdown skills into your own repository, modify them, and ship them inside a commercial product without asking. The licence file is at the repository root. This is not legal advice; read LICENSE yourself if the distinction matters to your organisation.

The MIT grant interacts well with the Markdown-only design. Because the skills are text, forking them is a copy operation rather than a build. It also means there is no dependency graph to audit and no transitive package risk in the skill set itself. The Python in the repository sits in tools, tests and mcp-servers, and those are where a supply-chain review would actually focus.

The maintenance picture is active. The last push was on 2026-08-21, and the release history shows v0.4.23 on 2026-08-03, v0.4.24 on 2026-08-09, and dsh-aris-v0.1.0 on 2026-08-21. The cadence suggests the project is still moving, and the newest release is a separate distribution channel for DeepSeek Harness rather than a core change.

Upgrade cost is the part to think about before you fork. The README references 82 skills on the DeepSeek Harness path. If you edit those skills locally, you own the merge on every release. Because the artefacts are Markdown, a git-based fork with periodic rebases is workable, but there is no documented migration guide and no versioned skill schema. Pin to a tag if you depend on specific skill behaviour.

## Conclusion

Adopt ARIS if you already drive an LLM agent from a terminal and want a repeatable structure for idea discovery, experiment runs and cross-model review, and if you are willing to wire up the API keys yourself. Do not adopt it if you want a packaged research platform with a job scheduler, a database and a web UI, or if you need a single model to be both author and reviewer. Before committing, verify three things: which skills exist under skills/ and skills/skills-codex/, whether your chosen agent can load them, and which reviewer backend your .env.example entries actually point at.

## FAQ

### Does ARIS let Claude Code do research while I sleep?

ARIS is designed around that use case: it sets up a loop where one model authors research output and a different model reviews it, and the manual review mode can wait up to 86400 seconds for a human. The repository does not include a scheduler, so something outside ARIS has to start the run.

### What does auto mode mean in the context of ARIS and Claude Code?

In ARIS, auto refers to the autonomous research loop rather than a Claude Code setting. The loop chains idea discovery, experiment automation and cross-model review, with a research-wiki holding memory across steps so the model does not lose earlier details.

### Can Claude Code be used for research with ARIS?

Yes. The README lists Claude Code as one of the supported hosts, alongside Codex CLI, Cursor, Trae, Antigravity, GitHub Copilot CLI, OpenClaw and DeepSeek Harness. ARIS is used there as a skill-based workflow rather than as an installed framework.

## Sources

- [Official README](https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep#readme)
- [Project repository](https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep)
- [Release notes](https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/wanshuiyin-auto-claude-code-research-in-sleep
