BioMate-KB turns 200 Bioconductor packages into vignette-grounded Claude skills
BioMate-KB Bioconductor Skills — 200 packages (top 100 by downloads + 100 rising stars) as vignette-grounded Claude/agent skills, with per-package workflow recipes
At a glance
- What is it?
- The top 100 most-downloaded and top 100 rising-star Bioconductor packages, formatted as agent skills with per-package workflow recipes and verified function names.
- Who is it for?
- BioMate-KB is a quietly rigorous piece of agent infrastructure. It picks packages by download data, grounds every skill in the package vignette with a tracked verification score, splits multi-analysis packages into explicit recipes, pins its Bioconductor version, and admits where its taxonomy is messy.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 103 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the knowledge base contains
BioMate-KB is a skill bundle in Claude Code Skills format covering 200 Bioconductor packages: the 100 most-downloaded packages plus the top 100 rising stars. Each skill teaches an agent when to choose a package, which workflows it supports, what parameters to set, how to interpret results, and which pitfalls to avoid. The selection is quantitative, not aesthetic: the top 100 cover about 56 percent of Bioconductor analysis-package download volume, and combined with the rising stars the bundle reaches roughly 57 percent.
Beyond this 200-package sample, the full BioMate KB claims runnable workflows for 1,818 Bioconductor packages, around 87 percent of analysis-package download volume, available through BioMate Cloud. The GitHub repository is the open, inspectable slice of that larger catalog.
How the 200 were picked
The two sets of 100 come from different rules. The top 100 is ranked purely by the official Bioconductor download score and includes both analysis tools like DESeq2, edgeR, limma, and fgsea, and the foundational data-structure, I/O, and annotation packages nearly every analysis imports, such as GenomicRanges, Biostrings, SingleCellExperiment, and AnnotationHub.
The rising-star 100 is restricted to analysis packages, excluding infrastructure, data-container, and GUI packages. Eligibility requires a first release no earlier than 2021 and fast growth in 2025, measured as year-over-year download growth plus at least 3,000 distinct download IPs, ranked by 2025 downloads. The point of this second set is timing: it surfaces newly important methods before they reach the top by raw volume, which is where a skills library can actually change what an agent recommends.
Recipes, not descriptions
The distinguishing format choice is the workflow recipe. A package that supports multiple analyses lists each as its own recipe subsection under a Workflows heading. The README gives concrete examples: DESeq2 gets separate recipes for standard, multi-factor, and likelihood-ratio-test analyses; crisprScore contributes six scoring recipes covering on-target, off-target, and indel scoring.
Across the 200 packages there are 390 workflow recipes, and 89 packages are multi-workflow. This is the difference between a skill that says what DESeq2 is and one that says how to run it, which is what an agent needs when a user asks a question in plain language and expects executed analysis rather than a lecture. The same recipe structure carries into the directory layout: twelve domain folders from transcriptomics (97 packages) down to metabolomics (1 package).
Vignette grounding and verification
The project stakes its credibility on a verification claim: every R function named in a skill is verified to appear in that package's own Bioconductor vignette, with a mean verification score of 0.91. The skills are grounded against primary sources rather than generated as free-form LLM prose.
Anyone who has watched an agent confidently call a nonexistent R argument knows why this matters. Hallucinated function names are the dominant failure mode for bioinformatics agents, because the training data mixes Bioconductor versions and packages with similar APIs. Anchoring each skill to its package's current vignette moves the failure mode from plausible invention to checked reference, and the 0.91 mean means the exceptions are tracked rather than hidden.
Version pinning as a feature
The skills are grounded against Bioconductor 3.21, pinned explicitly. The README explains why in one sentence: release is a moving pointer that drops packages as it advances, and several rising stars had already dropped out by 3.23. The pinned version is recorded in MANIFEST.json under bioconductor_version, and you can re-fetch a different snapshot by setting the BIOC_VERSION environment variable and running the extraction script.
This is the kind of detail that separates a maintained knowledge base from a frozen dump. R ecosystems move slowly enough that people forget to version them, then an agent trained across versions produces code mixing incompatible idioms. A declared, re-fetchable pin gives you a reproducible target and an honest answer to the question of which Bioconductor the agent actually knows.
Domains and coverage in practice
The domain breakdown shows where Bioconductor usage concentrates, and the repository lays it out as its top-level map:
skills/ (200 packages · 12 domains) ⭐ = rising star
├── transcriptomics/ (97 — DESeq2, edgeR, limma · ⭐ standR, Voyager, sechm, crisprScore)
├── genomics/ (33 — GenomicRanges, Biostrings · ⭐ rBLAST, syntenet, ggmanh)
├── general/ (19 — Biobase, DOSE · ⭐ immunotation, faers, mosbi)
├── proteomics/ (16 — MSnbase, mixOmics, mzR · ⭐ MatrixQCvis, MsDataHub, TargetDecoy)
├── epigenomics/ (10 — ChIPseeker, minfi · ⭐ HiCExperiment, HiContacts, epigraHMM)
├── single-cell/ (8 — SingleCellExperiment · ⭐ demuxmix, hoodscanR, MuData)
├── variant-calling/ (4 — VariantAnnotation, snpStats, vsn)
├── metagenomics/ (4 — phyloseq, microbiome, DirichletMultinomial)
├── imaging/ (4 — flowCore, EBImage · ⭐ lisaClust, cytoviewer)
├── annotation/ (2 — biomaRt, KEGGgraph)
├── enrichment/ (2 — enrichplot, ReactomePA)
└── metabolomics/ (1 — ⭐ rgoslin)Transcriptomics dominates with 97 packages including DESeq2, edgeR, limma, and rising stars standR, Voyager, sechm, and crisprScore. Genomics holds 33, general infrastructure 19, proteomics 16, epigenomics 10, single-cell 8, and the remaining domains shrink from variant-calling's 4 down to metabolomics' single rgoslin.
The project is candid that these labels are coarse. Transcriptomics is a catch-all absorbing most single-cell, spatial, and gene-set tools, so scater and scran sit as single-cell while fgsea and GSVA are filed as enrichment, and the small single-cell and annotation folders are remnants of the same imperfect classifier. The labels are kept for traceability, with the advice to treat domain folders as a rough guide rather than a strict ontology. That honesty is preferable to a tidy fake taxonomy.
A worked example and the wider KB
The README includes a tutorial that shows the intended end state: a plain-English bulk RNA-seq request routed through DESeq2 and clusterProfiler, differential expression on 120 human GRCh38 samples comparing control against treated, GO enrichment via enrichGO identifying apoptotic regulation as the top pathway, and an interactive STRINGdb protein-protein interaction panel over the top 2,841 differentially expressed genes, with nodes colored by log2 fold change and edges weighted by STRING confidence. The demo closes with an AI findings summary, a citable methods section, and citation export.
The selection methodology also documents its own competition: other public skill libraries are task-oriented or Python-centric, and the closest R-focused one has 13 skills about R packaging rather than Bioconductor analysis. That gap is the project's actual contribution, and the examples directory (deseq2, edger, limma) plus PACKAGES.md with all 200 entries make it verifiable before you trust it.
Editorial conclusion
BioMate-KB is a quietly rigorous piece of agent infrastructure. It picks packages by download data, grounds every skill in the package vignette with a tracked verification score, splits multi-analysis packages into explicit recipes, pins its Bioconductor version, and admits where its taxonomy is messy. For labs running bioinformatics through Claude Code or similar agents, it replaces hallucinated R with checked references, and the 200-package open slice is enough to evaluate whether the full 1,818-package catalog is worth adopting.
Frequently asked questions
Which Bioconductor version are the skills grounded against?
Bioconductor 3.21, pinned explicitly and recorded in MANIFEST.json as bioconductor_version. The pin exists because the release pointer drops packages as it advances, and you can re-fetch a different snapshot with the BIOC_VERSION environment variable and the extraction script.
How were the 200 packages selected?
Two ranked sets of 100. The top 100 come from the official Bioconductor download score and cover about 56 percent of analysis-package download volume. The 100 rising stars are analysis packages first released in 2021 or later with fast 2025 growth and at least 3,000 distinct download IPs, ranked by 2025 downloads.
What makes these skills trustworthy for R code generation?
Every R function named in a skill is verified to appear in that package’s own Bioconductor vignette, with a mean verification score of 0.91, so the skills are grounded in primary sources rather than free-form generated prose. Multi-analysis packages carry one recipe per workflow, 390 recipes across the 200 packages.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/biomate-ai-biomate-bioconductor-kb)