# darwin-skill: A Ratchet-Based Optimizer for Claude Code Agent Skills

> darwin-skill is a Claude Code skill that applies a ratchet mechanism to improve other skills: score, improve one dimension at a time, and keep only gains. Inspired by Karpathy's autoresearch and incorporating frameworks from two Microsoft Research papers, it adds mandatory human checkpoints that distinguish it from fully automated skill optimization.

**alchaincyf/darwin-skill** — 达尔文.skill —— 一个让你的Skill无限进化的系统：评估→改进→测试→保留或回滚 | Autoresearch-inspired autonomous skill optimization for Claude Code. Evaluate, improve, test, keep or revert.

- Repository: https://github.com/alchaincyf/darwin-skill
- Stars: 6,129 · Forks: 642
- Language: HTML
- License: MIT
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/alchaincyf-darwin-skill

## Skill Optimization as a Systematic Problem

The README describes the problem directly: when you have ten skills, you can review them manually. When you have sixty or more, you need a system. darwin-skill addresses this by treating SKILL.md files the way Andrej Karpathy's autoresearch treats model training code: define a measurable objective, make one controlled change, test it, and keep only the changes that move the objective in the right direction.

The project targets Claude Code and other agent tools that support the SKILL.md format. The README names Claude Code, Codex, OpenClaw, Trae, and CodeBuddy as tools in this category. Any SKILL.md file in those environments is a valid target for darwin-skill.

The key design premise is that a skill with perfect formatting can still perform poorly at runtime. darwin-skill evaluates both structural quality (does the skill follow the required format?) and effect quality (does the skill actually produce better agent outputs?). A skill that passes a structural review but fails effect tests scores low on the dimensions that matter most. The README states that effect testing carries the highest weight in the scoring system.

## The Ratchet: Forward Only, Never Backward

The core mechanism is a ratchet. Each optimization round can only result in two outcomes: the skill improves and the change is committed, or the skill does not improve and the change is reverted. Scores never decrease over time because every non-improving change is rolled back before the next round starts.

The README describes the ratchet explicitly: a round that produces a score of 75 when the current best is 78 is automatically reverted. The effective baseline remains at 78, and subsequent improvements start from there. This prevents the gradual accumulation of low-quality changes that would otherwise degrade skill quality over many edits.

The rollback mechanism uses `git revert`, not `git reset --hard`. This is listed as a hard rule in the anti-pattern blacklist: the README explicitly prohibits using `git reset --hard` as the rollback tool. The reason is that `git revert` creates a new commit that documents the rollback, preserving the history of what was tried and rejected. Using `git reset --hard` would erase that history.

Early stopping prevents diminishing returns: if a single round produces a gain of less than one point, the system stops automatically. This avoids accumulating small redundant edits that add length without adding substance.

## Installing darwin-skill and Running an Optimization

darwin-skill installs via the skills CLI.

```bash
npx skills add alchaincyf/darwin-skill
```

After installation, you trigger an optimization by telling the agent to optimize a skill by name, or to optimize all skills. The agent then runs through the evaluation phases described below. For environments where npm or GitHub access is unavailable, the README describes an alternative: download the zip archive from the project's distribution URL, extract it, and place the SKILL.md file at `~/.claude/skills/darwin-skill/`.

The README includes two security notes worth following before starting. First, run the optimization inside a git repository. darwin-skill commits changes and uses git revert for rollbacks; without a git repo, the mechanism cannot function. Second, commit or stash any local changes to the skill before starting, so that darwin-skill can cleanly keep or revert its experimental edits without mixing them with your own uncommitted work.

The repository also contains a test-prompts.json file at the top level. The optimization loop uses this file to run effect tests that verify whether a proposed change actually produces better agent outputs. Preparing test prompts specific to the skill being optimized is necessary for the effect validation gate to produce meaningful results.

## The 9-Dimension Evaluation Framework

darwin-skill v2.0 scores each skill on nine dimensions with a total of 100 points. The first six dimensions come from structural analysis. The final three come from the SkillLens paper published by Microsoft Research (arXiv:2605.23899), which the README describes as providing an empirically validated rubric where that specific set of dimensions predicted improvement accuracy 73.8% of the time.

The three new dimensions added in v2.0 are Failure Mechanism Encoding, Actionable Specificity, and High-Risk Action Blacklist. Failure Mechanism Encoding requires that known failure paths are explicitly coded into the skill rather than covered by a general instruction to avoid mistakes. Actionable Specificity requires that the skill uses no vague phrasing like 'consider', 'may', 'depending on the situation', or 'as appropriate'. High-Risk Action Blacklist requires that destructive operations such as rm, git reset --hard, and force push are explicitly listed as prohibited in the skill.

The README notes that the same AI cannot both edit and evaluate: SkillLens found that LLM self-evaluation accuracy is only 46.4%. To address this, each evaluation round uses two independent evaluator sub-agents. These evaluators are not reused between rounds to avoid anchoring effects, where an evaluator that approved a previous round is more likely to approve the next one regardless of actual improvement.

## Three Human-in-the-Loop Checkpoints

The README describes darwin-skill's human-in-the-loop approach as the core distinction from SkillOpt's fully automated design. Three mandatory pause points require human confirmation before the system continues.

Phase 1 is the baseline evaluation. The system generates an evaluation report and pauses for the human to review it and decide which dimensions to address. This gives the operator a chance to reject the evaluation before any edits happen.

Phase 2 is single-dimension optimization. After each edit, the system reaches a CHECKPOINT marked in red and displays the diff alongside the score change. The human must confirm before the next dimension is targeted. The README prohibits improving more than one dimension per round, which means Phase 2 can repeat several times before moving on.

Phase 2.5 is an optional test prompt run, where the operator can run the skill against specific prompts to validate effect before committing.

Phase 3 is the regression test. If the cumulative improvement from the optimization session falls below a defined threshold, the system stops entirely. This prevents the operator from accepting marginal gains that do not represent a real improvement in skill performance.

The dry-run ratio is also monitored. If more than 30% of the optimization runs are dry runs (where the skill is evaluated but not edited), the system issues a warning.

## Where darwin-skill Is the Wrong Tool

darwin-skill cannot improve a skill if you cannot describe what good output looks like. The effect validation gate in Phase 3 runs test prompts and scores the outputs. If test-prompts.json is empty or contains prompts that do not exercise the skill's core behavior, the effect gate will not catch regressions or confirm improvements. The improvement loop can then proceed on structural scores alone, which the README explicitly warns is insufficient.

The tool also cannot operate outside a git repository. The ratchet mechanism depends on git commit and git revert. Running it on a SKILL.md file that is not tracked in a git repository is not supported.

Skills that involve shell commands, git operations, credentials, local file access, or publishing workflows carry higher risk during optimization. The README specifically recommends reviewing every checkpoint diff carefully for these skills, since an automated edit could introduce a destructive pattern that looks structurally valid but behaves incorrectly at runtime.

The README also notes that the same optimization framework should not be used if the skill's purpose is itself unclear. Optimization requires a measurable target. A skill written for a vaguely defined task will receive structural scores but the effect dimension scores will not converge.

## darwin-skill vs. SkillOpt: Human Control vs. Full Automation

SkillOpt is a Microsoft Research project (also at arXiv:2605.23904) that provides a validation-gated editing framework for agent skills. It is available as `pip install skillopt` and is designed to run fully autonomously: it generates edits, runs validation, and commits or reverts without waiting for a human at each step.

darwin-skill draws on SkillOpt's validation-gated design but adds the three mandatory human confirmation points described above. The README states the rationale directly: skill quality is more subjective than a training loss metric, and automated validation alone is not sufficient for deciding whether to accept an edit.

The relationship between the two projects is documented in both directions. The darwin-skill README credits SkillOpt for the validation-gated framework. The SkillOpt repository listed darwin-skill as an integrated project on 2026-06-03, with the note 'gbrain, gbrain-evals, and darwin-skill have all integrated SkillOpt.' Teams who want full automation should look at SkillOpt directly. Teams who want the same core mechanism with manual review at each phase are the target audience for darwin-skill.

## Conclusion

darwin-skill is the right choice for Claude Code users who maintain a significant library of SKILL.md files and want a repeatable improvement process with human oversight at each phase. It is the wrong choice for teams who need fully automated optimization with no manual confirmation steps, or for skills that have no articulable test cases to put in test-prompts.json. Before running an optimization, make sure the target skill lives in a git repository with no uncommitted changes, and prepare a test-prompts.json with prompts that exercise the skill's intended behavior, since the validation gate depends on those test runs to confirm that a proposed edit actually improves real outputs.

## FAQ

### What file format does darwin-skill optimize?

darwin-skill optimizes SKILL.md files, which are the skill definition format used by Claude Code, Codex, Trae, and other agent tools that support the skills ecosystem. The skill being optimized is the SKILL.md file; darwin-skill itself is also delivered as a SKILL.md.

### How does darwin-skill avoid the problem of an AI evaluating its own edits?

Each evaluation round launches two independent evaluator sub-agents that are separate from the agent making the edits. A new set of evaluators is started for each round rather than reusing the previous round's evaluators, which the README describes as preventing anchoring effects. The README cites SkillLens research showing LLM self-evaluation accuracy is only 46.4% as the justification for this design.

### Is darwin-skill compatible with agent tools other than Claude Code?

The README lists Claude Code, Codex, OpenClaw, Trae, and CodeBuddy as tools that support the SKILL.md format that darwin-skill targets. Any tool that uses that format is a candidate for darwin-skill optimization, though the README is written with Claude Code as the primary context.

## Sources

- [alchaincyf/darwin-skill on GitHub](https://github.com/alchaincyf/darwin-skill)
- [Issues](https://github.com/alchaincyf/darwin-skill/issues)
- [License: MIT](https://github.com/alchaincyf/darwin-skill/blob/master/LICENSE)
- [README](https://github.com/alchaincyf/darwin-skill/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/alchaincyf-darwin-skill
