# ClawBio: a local-first bioinformatics skill library for AI agents

> ClawBio packages genomics analyses as Agent Skills folders that run on your own machine, from pharmacogenomics demo cards to GWAS lookups across nine databases. It is a Python 3.11+ project built on OpenClaw, and its MCP server is already scheduled for removal.

**ClawBio/ClawBio** — 🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free.

- Repository: https://github.com/ClawBio/ClawBio
- Website: https://clawbio.github.io/ClawBio/
- Stars: 1,150 · Forks: 279
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/clawbio-clawbio

## The gap ClawBio fills between genomics tools and chat agents

Most genomics work happens in two disconnected places. Command-line tools and Galaxy wrappers do the computation but expect you to know which tool, which reference build, and which flags. General AI assistants are good at conversation and bad at reproducible computation, because they have no fixed skill boundary and no demo data to check against. ClawBio's claim is narrower than "AI for bioinformatics": it packages each analysis as a skill with runnable demo data, so the same unit can be executed from a terminal, imported as a Python function, or read by an editor that understands Agent Skills folders. The audience is therefore people who already work with variants and want a local, inspectable layer between their data and an assistant. The README is explicit that most skills are pure Python and need no LLM at all; only the Bio Orchestrator routing path requires an API key. That is a meaningful design choice, because it means the pharmacogenomics or GWAS skills keep working offline and deterministically, and the AI part is optional rather than load-bearing.

## How a skill runs: folders, demo data and reproducibility bundles

The architecture is deliberately flat. Each skill lives as a plain Agent Skills folder under skills/, and the entry point is clawbio.py at the repository root or the installed clawbio package. The README gives the library form directly: run_skill and list_skills are importable from clawbio, so a skill is a function call with a demo flag rather than a service. Demo data is built on the Corpasome, a single CC0-licensed individual whose 23andMe SNP chip data has been available since launch, with 30x Illumina whole-genome subsets (GRCh37) covering roughly 4M SNPs, about 600K indels, and structural variants. That reference choice matters: because the demo genome is public and fixed, a demo run is comparable between machines. Many analyses write a reproducibility/ bundle containing replay commands, environment metadata, and output checksums, though the README warns the exact files vary by skill and some replays still need the original external inputs present. The v0.5.0 release added a validation layer around this: an Alzheimer's disease ground truth set of 34 genes across three evidence tiers with 20 negative controls, a mock API server for Ensembl, GWAS Catalog and ClinPGx so CI runs offline, and a scorer reporting gene recovery rate, FDR, precision, recall, F1 and a tier-weighted composite. A swappable fine-mapping interface pits ABF against SuSiE on the same data through a method registry. This is the most interesting part of the project, because it treats agent skills as something you can score, not just run.

## Installing ClawBio and running your first pharmacogenomics card

The README's quick start is two commands. The first installs the package from PyPI and requires Python 3.11 or newer; the second runs the pharmacogenomics skill against bundled demo data, so no genome file and no API key are needed.

```bash
pip install clawbio
clawbio run pharmgx --demo
```

A conda route exists as well, and the README names the channel: conda install -c bioconda clawbio. Expect the demo to print a pharmacogenomic result for the built-in sample rather than for your own genome. If you would rather call it from Python, the README shows the same skill through the library API, which returns a result object instead of printing to the terminal.

```python
from clawbio import run_skill, list_skills
result = run_skill("pharmgx", demo=True)
```

For editor integration, the README says every skill is a plain Agent Skills folder and that Cursor, VS Code, Codex and Zed all read ~/.agents/skills/. Copy or symlink the folders you need from skills/ into that directory, or into a project-local .agents/skills/. No server is involved in this path, which is the direction the project is pushing. Claude Code users get a plugin route instead: /plugin marketplace add ClawBio/ClawBio. If you want all skills with full demo data, the README recommends a source checkout with uv, running uv sync and then uv run python clawbio.py run pharmgx --demo. The repository also ships a Makefile with demo, list, lint, test and demo-all targets, so make demo and make list are the shortest way to see what is installed.

## The MCP deprecation is the biggest migration risk

Version 0.7.0 deprecated the MCP server and the release title says so plainly: "Genomics skills, in your editor. MCP server deprecated." The README repeats that it will be removed in 0.8.0, that existing configurations keep working until then, and that they print a notice on start. The optional dependency extra is kept only so that uvx --from 'clawbio[mcp]' clawbio mcp keeps functioning, and pyproject.toml pins the mcp dependency below 2.0 with a truncated comment explaining why. Anyone who built a workflow around the MCP server in 0.6.1 has a bounded window to move to the folder-based Agent Skills approach. The migration documentation is at docs.clawbio.ai/reference/mcp, not in the README itself, so the README alone does not tell you what the replacement looks like for your editor. That is a real cost, and it is worth weighing before adopting the MCP path at all: starting there in September 2026 means starting on a surface the maintainers have already announced they are removing. The folder approach is the safer bet, but it assumes your editor reads ~/.agents/skills/, which the README asserts for Cursor, VS Code, Codex and Zed without documenting per-editor caveats.

## Where ClawBio is the wrong tool

Several limits are visible in the project's own files. It classifies itself as Development Status 4 - Beta in pyproject.toml, and the README states that the Corpasome reference data is provided for research and educational purposes only. Nothing here is a clinical pipeline, and the warfarin example in the README, with its AVOID and DO NOT USE wording, is a demonstration of the skill's output format rather than a dosing authority. Reproducibility is partial: the README says many analyses write a reproducibility/ bundle, that the files vary by skill, and that some replays still require the original external inputs. A checksum proves an output was not altered; it does not make a replay self-contained. The skill inventory is also uneven. The README counts 97 skills but only 91 with runnable demo data, so six skills cannot be exercised from the bundled data, and you only find out which by trying or by reading the skill folders. Several skills depend on hosted inference: the gi-promoter, gi-splice, gi-enhancer, gi-chromatin, gi-expression and gi-annotation skills need a Genomic Intelligence API key, and the .env.example ships a shared hackathon key whose rate limits are set server-side and readable only from the RateLimit-Limit and RateLimit-Remaining response headers. If your work cannot leave your machine, those six skills are out regardless of the local-first framing. Finally, the Bio Orchestrator's LLM routing needs a FLock API key, so the orchestration layer is not offline either.

## How ClawBio differs from Galaxy and from generic agent frameworks

The README positions ClawBio next to Galaxy rather than against it: the project claims 8,182 Galaxy tools alongside its own skills. The difference in approach is where the interface lives. Galaxy wraps tools in a web platform with a server, a database and a job history, which makes sharing and provenance strong but ties you to that deployment. ClawBio puts each analysis in a folder that a local agent or editor reads, with no server in the folder path, and leans on a fixed public demo genome so a demo run is comparable across machines. The trade-off runs the other way too: Galaxy's tool catalogue is a long-standing community effort with its own wrappers, while ClawBio's skills are the project's own and its validation infrastructure only dates to v0.5.0. Against generic agent skill collections, the difference is the ground truth. A general skill library tests that a tool ran; ClawBio's benchmark suite scores gene lists against a curated Alzheimer's disease truth set with negative controls and reports FDR alongside recall. That is a bioinformatics-specific notion of correctness, and it is the reason the project's benchmark numbers are more meaningful than a passing test count. The 4,723 tests and 74 benchmark tests at the v0.5.0 baseline are worth reading as engineering discipline, not as a quality guarantee.

## Licence, maintenance and upgrade cost

The licence signal is inconsistent and you should resolve it before depending on the code. The README badge says MIT, and pyproject.toml declares license = "MIT" with license-files = ["LICENSE"]. The repository metadata reports NOASSERTION, which usually means a licence file that GitHub's classifier could not match to a known template, but it can also indicate additional terms. The repository contains a LICENSE file; read it rather than the badge. If the terms matter to your organisation, that is a five-minute check that removes the ambiguity.

Maintenance is current: the last push was on 2026-09-10, and v0.7.0 landed on 2026-09-03, a week earlier. The release cadence visible in the repository runs from v0.5.2 in June 2026 through v0.6.1 in July to v0.7.0 in September, so roughly monthly. The upgrade cost is concentrated in the MCP removal. If you are on 0.6.1 with an MCP configuration, 0.8.0 is the breaking release, and the migration guide lives at docs.clawbio.ai/reference/mcp. If you use the folder-based skills, upgrades look additive: new folders under skills/, new demo data, and an expanding benchmark suite. The dependency list is another cost to note. pyproject.toml pulls in pandas, numpy, scipy, scikit-learn, biopython, matplotlib, pydeseq2, rocrate, google-cloud-bigquery and openai, so the installed footprint is not small even though most skills are pure Python. The requirements pin minimum versions rather than exact ones, which means a fresh install can resolve to newer libraries than the maintainers tested against.

## Conclusion

Adopt ClawBio if you want genomics analyses that run locally, carry demo data, and can be dropped into an editor that reads Agent Skills folders; the pip install plus clawbio run pharmgx --demo path is short enough to evaluate in one sitting. Do not adopt it if your workflow depends on the MCP server, which the README says is deprecated as of 0.7.0 and will be removed in 0.8.0, or if you need clinical-grade interpretation rather than research and educational output. Verify first that every skill you plan to use has runnable demo data, since the README puts that at 91 of 97 skills, and read docs/reproducibility.md before assuming a replay bundle reproduces without the original external inputs. The licence metadata is the other thing to settle early: the README badge and pyproject.toml both say MIT, while the repository reports NOASSERTION.

## FAQ

### What is ClawBio?

ClawBio is a bioinformatics-native AI agent skill library built on OpenClaw, described as local-first and reproducible. It ships 97 skills, 91 of them with runnable demo data, and installs with pip install clawbio on Python 3.11 or newer.

### Which AI is best for bioinformatics?

The project does not rank AI systems for bioinformatics. ClawBio's own position is that most of its skills are pure Python and need no LLM, and that only the Bio Orchestrator routing path requires a FLock API key.

### Will bioinformatics be taken by AI?

The README does not address this question. ClawBio frames its skills as reproducible, inspectable units that run locally, with benchmark scoring against curated ground truth rather than autonomous interpretation.

### Who is Claw AI?

The README does not describe a product or person called Claw AI. ClawBio is the project name, and it is built on OpenClaw, which the README cites as a separate upstream repository with 180k+ GitHub stars.

### What are 10 applications of bioinformatics?

The README does not list general bioinformatics applications. It names the analyses ClawBio itself covers: pharmacogenomics, GWAS lookup across nine genomic databases, polygenic risk scores, UK Biobank field search, and the six Genomic Intelligence skills for promoter, splice, enhancer, chromatin, expression and annotation work.

## Sources

- [ClawBio/ClawBio on GitHub](https://github.com/ClawBio/ClawBio)
- [Issues](https://github.com/ClawBio/ClawBio/issues)
- [Project website](https://clawbio.github.io/ClawBio/)
- [README](https://github.com/ClawBio/ClawBio/blob/main/README.md)
- [Releases](https://github.com/ClawBio/ClawBio/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/clawbio-clawbio
