Model or dataset
exon-research/genomi avatar
exon-research/genomi

Genomi: a local-first agent runtime for reading your own DNA

Local-first, open-source Claude Science alternative, before Claude Science is a thing. Turn your AI agent into personal DNA expert.

483 stars62 forksPythonApache-2.0

At a glance

What is it?
Genomi gives an MCP-capable coding agent a private workspace over your variants, public genetics evidence and report tools. It is a developer preview at v0.1.0, and its limits are worth knowing before you hand it a whole-genome file.
Who is it for?
Adopt Genomi if you already run an MCP host such as Claude Code or Codex, you have a VCF or gVCF you are allowed to process locally, and you want an agent to query your own variants against public evidence instead of uploading them to a service. Do not adopt it if you need a clinical interpretation, a validated diagnostic pipeline, or a stable 1.0 interface: the release is v0.1.0 and GenomiLab is labelled a developer preview.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 31 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Genomi is aimed at: three billion base pairs and no reader

A consumer genotyping file is already awkward to work with. A whole-genome VCF is worse: millions of observed variants, tens of thousands of genes, and almost no way to ask a question in plain language. The README puts the scale directly, describing roughly 3 billion base pairs, 20,000+ genes and millions of observed variants per person, and notes that no clinician, lab or individual holds that in their head.

The target user is not a bioinformatician with a Galaxy instance. It is someone who already runs an AI coding agent, has a raw DNA file, and wants the agent to answer questions like why a drug does nothing for them or what their variants say about a specific risk. Genomi's pitch is that the agent runtime, not a web service, is the right place for that work, because the genome never leaves the machine. That is a defensible position. It is also a position that only makes sense if you are comfortable running a Python package and registering an MCP server, which is the real entry barrier here.

How the Active Genome Index and MCP tooling fit together

The core object is the Active Genome Index, or AGI. According to the pyproject comments, parse_source writes a bgzip-compressed canonical VCF per Active Genome Index, and downstream capabilities read it through BGZFile virtual-offset random access. Crucially, that comment states the intake file is never reopened: once the AGI exists, the agent works from the canonical copy, not from your original export.

That design explains several dependencies. pysam provides the htslib bindings for the compressed VCF. pyliftover handles GRCh37 to GRCh38 coordinate conversion via UCSC chain files, which matters because public resources and personal files are not always on the same build. duckdb is the query engine over the index, and isal is listed as optional gzip acceleration for ClinVar and population VCF imports, with a documented fallback to subprocess gzip or stdlib gzip when unavailable.

Around that core sits the agent integration. The README describes Genomi as an open-source AI agent runtime that works with Claude Code, Codex, OpenClaw, Hermes and any MCP-capable host, giving the agent a private workspace with variants in a local index, public genetics evidence ready to query, memory of what was explored, and report tools. GenomiLab is the layer above it: a patient-facing disease-investigation application where the host agent chairs a board of 2 to 5 specialist subagents, and the GenomiLab service records milestone states for a loopback portal rather than raw agent messages or chain of thought. That last detail is a deliberate privacy boundary, and it is the most interesting architectural choice in the README.

Installing Genomi through your agent, and a first query

The project does not lead with a pip command. It leads with an instruction you paste into your agent, which then reads a remote install guide covering dependency checks, library selection, MCP registration, optional genome-source import and verification. That guide lives at INSTALL_FOR_AGENTS.md in the repository.

text
Install and configure Genomi by following the instructions here:
https://raw.githubusercontent.com/exon-research/genomi/master/INSTALL_FOR_AGENTS.md

If Genomi is already packaged or present on the machine, the README says the canonical install or update path is the `genomi install` command or the MCP operation `genomi.install`. The source bootstrap is only for hosts that do not have Genomi yet. Note the ordering: the agent-facing guide is the primary path, and the command is the fallback for an existing install, not the other way round.

For a first real use, the README's own TL;DR is the shortest route to a working mental model. It asks the agent to read a bundled plain-text file and explain what makes Genomi different from other agent harnesses and why it is useful for private DNA analysis.

text
Hey please read this and tell me why Genomi is different from other AI
agent harnesses. Why is this actually useful for understanding my DNA privately?
https://raw.githubusercontent.com/exon-research/genomi/master/llms-full.txt

What you should see is the agent fetching that file and summarising the runtime rather than answering from general knowledge. If it starts speculating about genomics without reading the URL, the MCP wiring has not taken effect. Genome intake comes next: the README states you can give the host agent the local VCF, gVCF or other supported genome-source path, and the agent prepares and selects the AGI. On a fresh Genomi home, that path alone creates a usable local placeholder profile that you can rename later; only an existing ambiguous multi-user home needs a follow-up choice about which profile owns the genome.

Where Genomi stops: no clinical answers, no rollback, no second agent

The README is unusually direct about refusal. One of the linked demonstration videos is titled around Genomi knowing when to say no and I don't know. That is a feature, not a gap, but it also sets expectations: this is an evidence-grounded research tool, not a diagnostic one. Nothing in the README claims clinical validity, and the description of GenomiLab as a developer preview reinforces that the disease-investigation workflow is not finished.

There are harder constraints. The AGI is built once from your intake file and the intake is never reopened, so if the index is wrong the agent will keep reasoning over the wrong canonical VCF until you rebuild it. The README does not document rollback. GenomiLab requires a current Genomi user with a query-ready AGI already selected before the Research Desk opens, and it explicitly does not discover or launch a second embedded agent server: the host owns the conversation, task lifecycle, planning, subagents, streaming, resume and cancellation. If your host cannot do those things, GenomiLab has nothing to attach to.

The dependency list is also a real adoption cost. The Paperclip client is pinned to a specific git commit, and the comment around it says the Proto prerequisite checks use an isolated fixed-endpoint child with explicit user-owned Modal credentials, while stating that no Proto tool is deployed or run. A pinned git dependency plus a credential-bearing prerequisite check is the kind of thing that fails quietly on a locked-down machine. Test the install in the environment you actually intend to use.

Genomi against a general agent harness or a hosted interpretation service

The obvious alternative is the one you already have: a general MCP-capable agent with no genomics runtime, pointed at a public variant database through some other server. The difference is what the agent can reach. A general harness has no Active Genome Index, no coordinate liftover between GRCh37 and GRCh38, and no canonical compressed VCF that downstream tools read by virtual offset. Every query would mean re-parsing your file, and the agent would have no memory of what was explored.

The second alternative is a hosted interpretation service that takes your file and returns a report. The difference there is direction of data flow. Genomi's stated position is that your genome stays on your machine and the agent does the work. A hosted service inverts that: you get a curated pipeline, but the raw data leaves your control. Neither is strictly better. If you want a fixed, validated report, a hosted clinical-grade service is the right tool and Genomi is not. If you want to iterate on questions against your own file, the local index is the point.

A third comparison is against doing the work yourself with pysam and duckdb in a notebook. That is entirely possible, and Genomi is essentially a packaged version of it with an agent interface and a report layer. The value it adds is the MCP surface, the skill definitions under skills/, and the evidence and hypothesis bookkeeping in GenomiLab. If you would rather write that glue yourself, the Apache-2.0 licence lets you read how they did it.

Licence, maintenance and what an upgrade actually costs

Genomi is Apache-2.0, which permits commercial and private use, modification and redistribution provided you keep the licence and notices, and it includes a patent grant. That matters for a tool that touches personal health data: you are not locked into a vendor. It does not make the output regulated or clinically valid, and nothing here should be read as legal advice about handling health data in your jurisdiction.

On maintenance, the last push to the default branch was on 2026-08-31, and the only listed release is v0.1.0 from 2026-07-01. The repository is not archived. That is a young project with a single tagged release, so treat the API as moving. The upgrade cost is concentrated in two places: the pinned Paperclip git dependency, which will need attention whenever that commit is superseded, and the Active Genome Index format, since a change there would force a rebuild from your original intake file. Because the README does not document rollback, keep that original file in a place the agent cannot overwrite.

Editorial conclusion

Adopt Genomi if you already run an MCP host such as Claude Code or Codex, you have a VCF or gVCF you are allowed to process locally, and you want an agent to query your own variants against public evidence instead of uploading them to a service. Do not adopt it if you need a clinical interpretation, a validated diagnostic pipeline, or a stable 1.0 interface: the release is v0.1.0 and GenomiLab is labelled a developer preview. Before trusting anything it outputs, verify three things on your own machine: that the Active Genome Index was built from the file you intended, that the MCP server is registered with the host you actually use, and that the report distinguishes evidence from hypotheses. If the index build fails on your VCF, there is no documented rollback path in the README, so keep the original file untouched.

Frequently asked questions

What is Genomi and who is it for?

Genomi is an open-source AI agent runtime that turns an MCP-capable agent into a personal DNA expert, giving it a private workspace with your variants in a local Active Genome Index, public genetics evidence and report tools. It is aimed at people who already run a host like Claude Code or Codex and want to query their own genome without sending it to a service.

How do I install Genomi?

The README's primary path is to paste an instruction into your agent telling it to follow the install guide at INSTALL_FOR_AGENTS.md, which covers dependency checks, library selection, MCP registration, optional genome-source import and verification. If Genomi is already present, the canonical install or update path is the genomi install command or the genomi.install MCP operation.

Does Genomi keep my genome on my machine?

The README describes Genomi as local-first and states that your genome stays on your machine while the agent does the work. The Active Genome Index is written as a bgzip-compressed canonical VCF that downstream capabilities read locally, and the pyproject comments state the user's intake file is never reopened.

What file formats can I give Genomi?

The README says you can give the host agent the local VCF, gVCF or other supported genome-source path, and the agent prepares and selects the Active Genome Index from it. On a fresh Genomi home that path alone creates a usable local placeholder profile you can rename later.

Is Genomi a medical or diagnostic tool?

The README does not claim clinical validity. One of the linked demonstrations is specifically about Genomi knowing when to say no and I don't know, which matches its framing as an evidence-grounded research tool rather than a diagnostic one.

Official sources

  1. exon-research/genomi on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/exon-research-genomi.svg)](https://hysenlabs.com/projects/exon-research-genomi)