Biomni: A General-Purpose Biomedical AI Agent That Runs Code Against an 11GB Data Lake
Biomni: a general-purpose biomedical AI agent
At a glance
- What is it?
- Biomni is a Stanford research agent that turns natural-language biomedical requests into Python execution over a bundled data lake. Here is how the install works, what the datalake download costs you, and where the agent model falls short.
- Who is it for?
- Adopt Biomni if you are a computational biologist who already works in conda, has roughly 11GB of disk and bandwidth to spare, and wants an LLM agent that writes and runs Python against a bundled biomedical data lake rather than answering from model memory.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 9 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Biomni is aimed at: research questions that need code, not prose
Ask a general chatbot to plan a CRISPR screen and you get a plausible paragraph. Ask it to actually run the analysis and it cannot, because the answer lives in datasets and libraries it has no access to. Biomni is built for that second case. The README describes it as a general-purpose biomedical AI agent that "autonomously execute[s] a wide range of research tasks across diverse biomedical subfields," combining LLM reasoning with retrieval-augmented planning and code-based execution.
The intended user is a scientist who can read Python but does not want to hand-write every glue script. The README's own examples are the clearest statement of scope: planning a CRISPR screen and generating 32 candidate genes, running scRNA-seq annotation on a path you supply, and predicting ADMET properties for a SMILES string. Those are three different subfields handled by one entry point.
That breadth is the selling point and also the risk. A single agent that spans genomics, cheminformatics and single-cell work has to route each request to the right tool, and the README does not describe how that routing is validated.
How A1 turns a sentence into executed code
The architecture visible in the repository is a Python package with an agent class, A1, exported from biomni.agent. You construct it with a data path and a model name, then call .go() with a natural-language task. The README shows the constructor signature directly: A1(path='./data', llm='claude-sonnet-4-20250514').
Two mechanisms sit behind that call. The first is retrieval-augmented planning, named in the overview. The second is the data lake, a local corpus of biomedical files that the agent reads from the path you pass. The README states the lake is roughly 11GB and is downloaded automatically on first run. So the agent is not answering from model weights alone; it is grounding plans in files that live on your disk.
There is also a centralized configuration layer. biomni.config exposes default_config, and the README shows mutating it globally (default_config.llm = "gpt-4", default_config.timeout_seconds = 1200) so that every agent created afterward inherits the setting. That is convenient for a notebook session and awkward for anything with per-user isolation, because the default is process-wide state.
Model selection is string-based. The .env.example lists LLM_SOURCE with accepted values "OpenAI", "AzureOpenAI", "Anthropic", "Ollama", "Gemini", "Bedrock", "Groq", "Custom". The README adds one sharp edge: if you use Azure, prefix the model name with azure-, as in llm='azure-gpt-4o'.
Installing Biomni: setup.sh, conda, pip, and the keys you actually need
The README is explicit that the environment is large and that you should not assemble it by hand. It points to biomni_env/README.md and a single setup.sh script, then tells you to activate the environment named biomni_e1. Run those two steps first:
conda activate biomni_e1With the environment active, install the published package. The README offers a source install as the route to the latest update:
pip install biomni --upgrade
pip install git+https://github.com/snap-stanford/Biomni.git@mainNext, credentials. Copy the example file and fill it in:
cp .env.example .envThe example file marks only one key as required: ANTHROPIC_API_KEY. OpenAI, Bedrock, Gemini and Groq keys are all labelled optional, and the custom serving block (CUSTOM_MODEL_BASE_URL, CUSTOM_MODEL_API_KEY) is commented out. A minimal working .env therefore needs one line.
Now the first real task. Constructing the agent triggers the datalake download, so expect a long first run:
from biomni.agent import A1
agent = A1(path='./data', llm='claude-sonnet-4-20250514')
agent.go("Predict ADMET properties for this compound: CC(C)CC1=CC=C(C=C1)C(C)C(=O)O")If you want to skip the download, the README documents passing an empty list: A1(path='./data', llm='claude-sonnet-4-20250514', expected_data_lake_files = []). It recommends that for faster testing, limited storage, and cases where your tools do not need lake files. For a browser-based view instead of a script, agent.launch_gradio_demo() serves at http://localhost:7860, requires gradio>=5.0,<6.0, and defaults to requiring the verification code "Biomni2025".
The 11GB first run and the packages that are deliberately missing
Two constraints show up before you write a single analysis. The first is the data lake. A default A1 construction pulls about 11GB. On a laptop with a metered connection, or in a CI job, that is the difference between a quick experiment and an aborted one. The escape hatch exists (expected_data_lake_files = []), but the README frames it as a way to skip the download, not as a supported mode for lake-dependent tools. If your task needs those files, the download is not optional.
The second is dependency conflicts. The README carries a warning that some Python packages are not installed by default "due to dependency conflicts," that you must install them manually, and that you "may need to uncomment relevant code in the codebase." The current list lives in docs/known_conflicts.md, which is not reproduced in the README. Editing library source to enable a feature is a real maintenance cost: your patch is local, and an upgrade can move the code you uncommented.
A third limitation is structural rather than technical. The README documents no evaluation harness, no confidence reporting, and no rollback when a generated analysis is wrong. The package metadata is honest about maturity in its own way: pyproject.toml pins requires-python to ">=3.11" and the release history is still in 0.0.x, with v0.0.8 published on 2025-10-27. Treat generated code as a draft that a domain expert reviews, not as a result.
Biomni compared with wiring your own agent on LangChain
The obvious alternative is not another biomedical agent. It is building the same loop yourself on a general agent framework. Biomni's own dependency list makes the comparison concrete: pyproject.toml declares pydantic, langchain and python-dotenv, plus an optional gradio extra. If you already run LangChain, you can construct a tool-calling agent with your own tool set and your own data directory.
The difference is what you inherit. A hand-built agent gives you control over every tool, every prompt and every file it can touch, and no 11GB download. Biomni gives you the biomedical tooling and the data lake already assembled, at the cost of accepting its environment, its setup.sh and its dependency-conflict list. There is no documented way to swap in your own data lake structure, and the README does not describe a plugin or registration API for adding tools to A1.
So the choice is between assembly work and environment work. If your tasks are narrow and your data is already local, a small LangChain agent is less machinery. If you want broad biomedical coverage without curating dozens of tools, Biomni's bundle is the reason to use it.
Licence, maintenance and what an upgrade costs you
Biomni is Apache-2.0. The LICENSE file is included, pyproject.toml declares license = "Apache-2.0" with license-files = ["LICENSE"], and the repository also carries a license_info.md. Apache-2.0 is a permissive licence with an explicit patent grant and a requirement to preserve notices; it is not a copyleft licence. That matters here because the package bundles access to datasets whose own terms are separate from the code licence, and the README does not discuss data licensing. If you plan to redistribute anything derived from the data lake, check the upstream sources for those files rather than assuming the Apache-2.0 grant covers them. This is a description of what the repository states, not legal advice.
On maintenance: the repository is not archived, and the last push was on 2026-09-07, which is recent. The release cadence visible in the repository is roughly monthly around v0.0.6 through v0.0.8. Upgrading is not free, though. The known-conflicts workflow means a version bump can invalidate local edits, and the Gradio interface is pinned to 5.x because of API changes in 6.0, so a Gradio upgrade in your environment can break the demo launch. Pin your versions and re-read docs/known_conflicts.md after each upgrade.
Editorial conclusion
Adopt Biomni if you are a computational biologist who already works in conda, has roughly 11GB of disk and bandwidth to spare, and wants an LLM agent that writes and runs Python against a bundled biomedical data lake rather than answering from model memory. Do not adopt it if you need a hosted product with no API key management, if you cannot run the provided setup.sh environment, or if you expect the agent to be correct without review: the README does not document a validation step for generated analyses. Before committing, verify three things yourself: that biomni_env/README.md installs cleanly on your platform, that your Anthropic key is set because ANTHROPIC_API_KEY is the only key marked required, and that the tools you need are not on the known_conflicts.md list, since those packages are deliberately absent from the default environment.
Frequently asked questions
Is Biomni free?
The code is Apache-2.0 licensed and installs from PyPI or GitHub at no cost. Running it is not free in practice, because it calls a hosted LLM and the .env.example marks ANTHROPIC_API_KEY as required, so you pay your model provider.
What model does Biomni use?
The README's examples use llm='claude-sonnet-4-20250514'. The .env.example lists LLM_SOURCE values of OpenAI, AzureOpenAI, Anthropic, Ollama, Gemini, Bedrock, Groq and Custom, and the README notes that Azure models must be prefixed with azure-.
Is Biomni open source?
Yes. The repository is public under Apache-2.0, with LICENSE and license_info.md at the top level, and pyproject.toml declares the same licence.
How do you use Biomni?
Install the environment from biomni_env/README.md, activate biomni_e1, install the biomni package, set your API keys, then create an A1 agent and call agent.go() with a natural-language task. The first construction downloads a datalake of about 11GB unless you pass expected_data_lake_files = [].
What does Biomni do?
It is a general-purpose biomedical AI agent that combines LLM reasoning with retrieval-augmented planning and code-based execution. The README's examples cover CRISPR screen planning, scRNA-seq annotation and ADMET property prediction.
Is there an alternative to Biomni?
You can build the same loop on a general agent framework. Biomni itself depends on langchain and pydantic, so a hand-built LangChain tool-calling agent with your own tools and data directory is the closest alternative, trading Biomni's bundled biomedical tooling and data lake for full control over what the agent can touch.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/snap-stanford-biomni)