Evo 2: running the Arc Institute DNA language model locally
Genome modeling and design across all domains of life
At a glance
- What is it?
- Evo 2 is a DNA language model that scores and generates sequences at up to 1M base pair context. The install is the hard part: the 40B, 20B and 1B checkpoints need FP8 through Transformer Engine and Hopper-class GPUs, while the 7B checkpoints run in bfloat16 on a lighter stack.
- Who is it for?
- Adopt Evo 2 if you have NVIDIA hardware and a Linux host, and start with evo2_7b, which the README says runs in bfloat16 without Transformer Engine. Do not adopt it if you need training or finetuning in this repository, since the README points elsewhere for that, or if you are on macOS.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 104 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Evo 2 does that a general sequence tool does not
Evo 2 is a DNA language model trained autoregressively on OpenGenome2, a dataset the README describes as 8.8 trillion tokens from all domains of life. It models sequences at single-nucleotide resolution with a context window of up to 1 million base pairs on the long-context checkpoints. That combination is the point: a variant effect, a regulatory element or a gene cluster can be scored in the context of a very long stretch of surrounding sequence rather than a short window.
The audience is computational biologists and ML engineers who want to run inference on their own hardware. The README is explicit that the repository is for local inference or generation through the Vortex inference code, and it separates that from training and finetuning, which it points to another section for. If your work is scoring likelihoods across sequences, extracting embeddings for a downstream classifier, or generating DNA, this is the intended path. The published description of the method is in Nature, and the architecture is StripedHyena 2, pretrained with Savanna.
How Evo 2 loads a checkpoint and returns logits
The public surface is small. You construct an Evo2 object with a checkpoint name, tokenize a sequence, wrap the token ids in a tensor, and call the model. The forward pass returns a tuple whose first element is the output; indexing it gives logits shaped (batch, length, vocab). The tokenizer is reachable as evo2_model.tokenizer, so you do not need to bring your own nucleotide encoding.
Device placement is handled by Vortex rather than by you. The README states that for the 40B model, Vortex automatically splits the model across available CUDA devices. That is a real convenience, but it also means the memory arithmetic is not something you tune directly; you supply the GPUs and the loader decides how to shard. The 40B checkpoint is documented as requiring multiple H100 GPUs.
Embeddings are a separate return value from the same call. The README notes that intermediate embeddings work better than final embeddings and shows a layer name of the form blocks.28.mlp.l3. That is a design detail worth taking seriously: if you were planning to use the last hidden state as a sequence representation, the project's own guidance points you at an inner layer instead.
Installing Evo 2 and running a first forward pass
The full install expects you to put Transformer Engine and Flash Attention in place before Evo 2 itself. The README recommends conda for Transformer Engine. Run these in order, then install the package:
conda install -c nvidia cuda-nvcc cuda-cudart-dev
conda install -c conda-forge transformer-engine-torch=2.3.0
pip install flash-attn==2.8.0.post2 --no-build-isolation
pip install evo2If you only intend to run 7B checkpoints, there is a lighter path that skips Transformer Engine and FP8 hardware. The README notes a compatible PyTorch must be installed before Flash Attention, and gives an example index URL for a CUDA 12.8 build:
pip install torch==2.7.1 --index-url https://download.pytorch.org/whl/cu128
pip install flash-attn==2.8.0.post2 --no-build-isolation
pip install evo2There is also a source route, which is what you want if you plan to read or modify the code:
git clone https://github.com/arcinstitute/evo2
cd evo2
pip install -e .The README's verification step is a generation test rather than a toy import. Pass the checkpoint you actually installed:
python -m evo2.test.test_evo2_generation --model_name evo2_7bThe README says to run this after configuration changes or on different hardware, which tells you the project treats silent numerical drift as a real risk rather than a theoretical one.
For a first real use, the forward example is the smallest thing that produces output. It scores a four-base sequence and prints the logits and their shape:
import torch
from evo2 import Evo2
evo2_model = Evo2('evo2_7b')
sequence = 'ACGT'
input_ids = torch.tensor(
evo2_model.tokenizer.tokenize(sequence),
dtype=torch.int,
).unsqueeze(0).to('cuda:0')
outputs, _ = evo2_model(input_ids)
logits = outputs[0]Expect a tensor of shape (batch, length, vocab). The first call also downloads the checkpoint, because the Dockerfile notes that models are fetched by the library on first use and are not baked into the image. Mounting the Hugging Face cache is the README's suggested way to avoid re-downloading, and the Docker example does exactly that:
docker build -t evo2 .
docker run -it --rm --gpus '"device=0"' -v ./huggingface:/root/.cache/huggingface evo2 bashThe Dockerfile is built from nvcr.io/nvidia/pytorch:25.04-py3 and installs evo2 directly. Its comments state that the container assumes a host with NVIDIA GPUs of compute capability 8.9 or higher for full FP8 support, and that you must pass --gpus at runtime.
The FP8 split is the constraint that decides your hardware
The checkpoint table is not a menu of sizes. It is a hardware filter. The README states that the 40B, 20B and 1B models require FP8 via Transformer Engine for numerical accuracy and need an NVIDIA Hopper GPU. Only the 7B family (evo2_7b, evo2_7b_262k, evo2_7b_base) can run in bfloat16 without Transformer Engine on any supported GPU.
So the light install is not a convenience for small experiments. It is the only configuration that works on non-Hopper hardware, and it caps you at the 7B checkpoints. If your question needs the 1M context window, note that evo2_7b and evo2_7b_262k carry long context while evo2_7b_base is listed at 8K, so the base checkpoints are not interchangeable with the long-context ones.
Two more boundaries are worth stating plainly. The official OS is Linux, with WSL2 described as limited support, so macOS is out. And CUDA 12.1 or newer with cuDNN 9.3 or newer is required, with Python 3.11 or 3.12 and a recommended Torch 2.6.x or 2.7.x. The pyproject file pins requires-python to >=3.11,<3.13, so a 3.13 environment will not install the package.
The optional Triton kernels are a separate axis. They come from a Vortex pull request and require vtx>=1.1.0, which the pyproject dependency list also carries. You opt in per model load, and the test scripts accept a matching flag:
from evo2 import Evo2
evo2_model = Evo2('evo2_7b', use_kernels=True)python -m evo2.test.test_evo2_generation --model_name evo2_7b --use_kernelsTreat this as a tuning knob to enable after the default path works, not before. The README does not document rollback or a fallback if a kernel produces different results on your GPU.
When Evo 2 is the wrong tool
If you need to train or finetune a DNA model, this repository is not the place to start. The README frames the repo as being for running Evo 2 locally for inference or generation, and directs training and finetuning to a separate section rather than presenting it as a supported workflow here. A team expecting an end-to-end train-and-deploy pipeline will spend its first week discovering that boundary.
The second case is hardware. Without a Hopper-class GPU, you are limited to the 7B checkpoints, and the 40B model is documented as needing multiple H100s. There is no CPU path in the README and no quantized variant listed among the checkpoints.
The third case is reproducibility on unfamiliar hardware. The README's own instruction to validate outputs after configuration changes or on different hardware is an admission that results are not guaranteed to be identical across setups. If your work requires bit-identical scores across machines, that is a risk you have to measure, not assume away.
Finally, if you only need a hosted model and never intend to manage a GPU host, the README points to an NVIDIA Hosted API and to self-hosting through NVIDIA NIM. Running this repository is the option you choose when you want the weights local.
How Evo 2 compares with a protein or general genomic foundation model
The nearest alternative in kind is a protein language model such as ESM, and the difference is the alphabet and the pretraining corpus, not the interface. A protein model takes amino acid sequences and has nothing to say about non-coding DNA, regulatory regions or the intergenic sequence that makes up most of a genome. Evo 2 is trained on nucleotides from all domains of life, which is what lets it score a stretch of DNA directly and generate new nucleotide sequence.
Compared with a task-specific genomic model, the difference is context length and generality. A variant effect predictor trained on a fixed window cannot reason about a regulatory element a hundred kilobases away. Evo 2's long-context checkpoints are built for exactly that, at the cost of the hardware described above. The trade is real in both directions: Evo 2 gives you a general model you adapt, not a tuned predictor you call.
The hosted NVIDIA API is the third option. It removes the install entirely, but it removes local weights with it. If your sequences cannot leave your infrastructure, that option is closed and the local path is the only one.
Licence, maintenance and upgrade cost
The repository is licensed Apache-2.0, and the LICENSE and NOTICE files sit at the top level alongside AUTHORS. Apache-2.0 permits commercial use and modification and includes a patent grant, but it also carries notice and attribution obligations, so redistributing a modified copy means keeping the notices intact. That is a description of the licence text, not legal advice; have counsel read it against your distribution model.
The last push to the default branch was on 2026-06-19, and the most recent release is v0.5.0 from 2026-02-28, which the release notes describe as the Evo 2 20B release. The pyproject version is 0.6.0, which is ahead of the latest tagged release, so the package on PyPI and the tagged release do not line up one to one. Check which version you actually installed before filing a bug.
The upgrade cost is dominated by the dependency chain rather than by Evo 2 itself. Transformer Engine is pinned to 2.3.0 in the README's conda command, Flash Attention to 2.8.0.post2, and vtx to >=1.1.0 in the dependency list. Moving any one of those means rebuilding Flash Attention from source with --no-build-isolation, which is the step most likely to fail on a new machine. Budget for that, and keep the verification command in your CI rather than running it once by hand.
Editorial conclusion
Adopt Evo 2 if you have NVIDIA hardware and a Linux host, and start with evo2_7b, which the README says runs in bfloat16 without Transformer Engine. Do not adopt it if you need training or finetuning in this repository, since the README points elsewhere for that, or if you are on macOS. Before committing, verify two things on your own machine: that flash-attn==2.8.0.post2 builds against your PyTorch, and that python -m evo2.test.test_evo2_generation --model_name evo2_7b passes, which the README presents as the check to run after any configuration change or hardware switch.
Frequently asked questions
What was Evo 2 trained on?
Evo 2 was trained autoregressively on OpenGenome2, which the README describes as a dataset containing 8.8 trillion tokens from all domains of life.
What makes the Evo 2 model special?
It models DNA at single-nucleotide resolution with up to 1 million base pair context, using the StripedHyena 2 architecture and pretrained with Savanna. The README also notes that intermediate embeddings work better than final embeddings for downstream use.
Which Evo 2 checkpoints need FP8 hardware?
The README states that evo2_20b, evo2_40b, evo2_40b_base and evo2_1b_base require FP8 via Transformer Engine and an NVIDIA Hopper GPU. The 7B checkpoints can run in bfloat16 without Transformer Engine.
Can I run Evo 2 on Windows or macOS?
The README lists Linux as the official OS and WSL2 as limited support, with no macOS option. CUDA 12.1 or newer and cuDNN 9.3 or newer are also required.
How do I check that my Evo 2 install works?
The README's verification step is python -m evo2.test.test_evo2_generation with a checkpoint name such as evo2_7b. It says to run this after configuration changes or on different hardware.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/arcinstitute-evo2)