DeepVariant: CNN-based Germline Variant Calling from Sequencing Reads
DeepVariant is an analysis pipeline that uses a deep neural network to call genetic variants from next-generation DNA sequencing data.
At a glance
- What is it?
- DeepVariant is an open-source pipeline from Google that identifies genetic variants in aligned sequencing reads by converting pileups into images and classifying them with a convolutional neural network. It supports Illumina WGS and WES, PacBio HiFi, Oxford Nanopore, and several other sequencing technologies, and is deployed primarily through Docker.
- Who is it for?
- DeepVariant is a well-maintained option for human germline variant calling when you work with Illumina WGS, WES, PacBio HiFi, Oxford Nanopore R10.4.1, or hybrid data. The models are trained only on human data; non-human organisms require extra caution as documented in the repository.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Problem DeepVariant Solves and Who It Is For
Variant calling is the process of identifying positions in a sequenced genome that differ from a reference. Traditional callers such as GATK HaplotypeCaller use local assembly and probabilistic genotyping graphs. DeepVariant reframes the problem: it converts candidate variant sites into pileup image tensors, one per site, and classifies each image using a convolutional neural network to determine whether the site is a variant and what its genotype is.
The target users are computational biologists and bioinformatics engineers running germline variant-calling workflows on human sequencing data. Clinicians and researchers who need results in standard VCF or gVCF format compatible with downstream tools such as GATK pipelines, Cromwell, and cohort-merge utilities will find DeepVariant fits into existing workflows because its output conforms to the same format conventions.
DeepVariant supports germline calling on diploid organisms only. The README explicitly states that the genotypes it supports are hom-alt, het, and hom-ref, which corresponds to ploidy two. Organisms with other copy numbers are outside the documented scope.
Architecture: Pileup Images, CNNs, and Three-Stage Processing
DeepVariant processes data in three sequential stages that correspond to three internal tools.
The first stage, make_examples, reads an aligned BAM or CRAM file along with a reference FASTA and generates candidate variant sites as pileup image tensors. Each tensor encodes the sequencing reads at that site as an image where pixel values represent base calls, base quality, mapping quality, strand orientation, and related signals.
The second stage, call_variants, loads the saved tensors and passes them through the CNN model. The model outputs a genotype probability distribution for each candidate site.
The third stage, postprocess_variants, reads the probability distributions and writes the final VCF or gVCF file.
The Docker image bundles all three stages along with their dependencies. The recommended command runs all three stages in sequence:
BIN_VERSION="1.10.0"
docker run \
-v "YOUR_INPUT_DIR":"/input" \
-v "YOUR_OUTPUT_DIR:/output" \
google/deepvariant:"${BIN_VERSION}" \
/opt/deepvariant/bin/run_deepvariant \
--model_type=WGS \
--ref=/input/YOUR_REF \
--reads=/input/YOUR_BAM \
--output_vcf=/output/YOUR_OUTPUT_VCF \
--output_gvcf=/output/YOUR_OUTPUT_GVCF \
--num_shards=$(nproc)The `--model_type` flag accepts exactly one of WGS, WES, PACBIO, ONT_R104, or HYBRID_PACBIO_ILLUMINA. Selecting the wrong model type for the input data is a common source of accuracy degradation.
Supported Sequencing Technologies and Case Studies
DeepVariant ships distinct trained models for each major sequencing technology. The models included in version 1.10.0 cover WGS (Illumina or Element), WES (Illumina or Element), PacBio HiFi, Oxford Nanopore R10.4.1 Simplex, Complete Genomics T7 and G400, Roche SBX-D and SBX-Fast, Illumina RNA-seq, PacBio Iso-Seq/MAS-Seq, and a hybrid mode combining PacBio HiFi with Illumina WGS.
The repository's docs/ directory contains a separate case study for each technology. These case studies use publicly available test data and walk through the exact command line for that data type. They are the authoritative reference for per-technology configuration.
A separate project called DeepTrio extends DeepVariant for parent-child trio and duo calling. DeepTrio uses the same pileup-to-image approach but processes child and parent reads together to exploit inheritance constraints. It currently supports Illumina WGS, Illumina WES, and PacBio HiFi. An external tool, GLnexus, is used to merge the VCF outputs from multi-sample DeepTrio runs.
A pangenome-aware variant of DeepVariant is also documented. It supports Illumina and Element WGS and WES mapped with either BWA or the graph mapper vg.
GPU Acceleration and the Fast Pipeline Mode
DeepVariant can run on GPUs. The Dockerfile in the repository includes build arguments for an NVIDIA CUDA base image and a GPU build flag:
sudo docker build --build-arg=FROM_IMAGE=nvidia/cuda:12.3.2-cudnn9-devel-ubuntu22.04 --build-arg=DV_GPU_BUILD=1 -t deepvariant_gpu .An experimental mode documented in the Fast Pipeline case study overlaps the make_examples stage (which runs on CPU) with the call_variants stage (which runs on GPU). This overlap reduces wall-clock time on machines where the CPU and GPU can work simultaneously. The README describes this as an experimental mode, not the default configuration.
GPU builds require NVIDIA drivers and CUDA runtime compatibility with the base image version. The Singularity container format is also supported as an alternative to Docker for HPC environments where Docker is not available; the Quick Start documentation covers this path.
For large cohorts, the `--num_shards=$(nproc)` flag in the standard command distributes make_examples work across all available CPU cores. This is the most direct lever for scaling throughput on CPU-only systems.
DeepVariant vs GATK HaplotypeCaller: Differences in Approach
GATK HaplotypeCaller, developed by the Broad Institute, is the most widely used alternative for germline variant calling. The fundamental difference in approach is that GATK performs local de novo assembly at candidate sites using a read threading graph, then evaluates genotype likelihoods from the assembled haplotypes. DeepVariant skips assembly and instead encodes the raw pileup as a multi-channel image and relies on the CNN to learn variant patterns from training data.
The practical consequence is that DeepVariant's accuracy on a given technology depends on the quality and diversity of its training data, while GATK's accuracy depends more directly on the local assembly quality and the accuracy of the statistical model. DeepVariant won the PrecisionFDA Truth Challenge V2 in 2020 for the ONT, PacBio, and Multiple Technologies categories and the 2016 PrecisionFDA Truth Challenge for best SNP performance, which are reported as accuracy benchmarks in the README.
DeepVariant is BSD-3-Clause licensed, which permits commercial use. It runs exclusively through Docker or a custom build from source. GATK is available with its own license and has broader integration into existing bioinformatics workflow managers.
Limitations: Ploidy, Non-Human Organisms, and Model Scope
The ploidy constraint is hard: DeepVariant supports only copy-number two. Organisms with other ploidy levels, or analyses of cancer genomes where copy number is variable (somatic calling), are outside the scope of the germline models. The README points to a separate repository, google/deepsomatic, for somatic use cases.
All bundled models are trained on human data. The README notes directly that using the models on non-human organisms introduces pitfalls and links to a blog post on non-human variant calling for guidance. This is a meaningful limitation for labs working on plant, insect, or bacterial genomes.
While DeepVariant produces standard VCF output, it does not include a joint genotyping step for cohort analysis. Multi-sample cohort calling requires an external merge step using GLnexus, which is documented in the trio-merge case study but is a separate dependency.
The last push to the repository was on 2026-09-24 and version 1.10.0 was released on 2026-03-05, indicating ongoing active development.
Editorial conclusion
DeepVariant is a well-maintained option for human germline variant calling when you work with Illumina WGS, WES, PacBio HiFi, Oxford Nanopore R10.4.1, or hybrid data. The models are trained only on human data; non-human organisms require extra caution as documented in the repository. The hard constraint is ploidy: only diploid organisms with copy-number two are supported. Teams using Docker should pin to a versioned image such as `google/deepvariant:1.10.0` and read the per-technology case study before running production samples, because model_type selection directly determines accuracy.
Frequently asked questions
What is DeepVariant used for?
DeepVariant is used for calling germline genetic variants, specifically single nucleotide polymorphisms and small insertions and deletions, from aligned sequencing reads in BAM or CRAM format. It outputs standard VCF or gVCF files for use in downstream analysis.
How do I install DeepVariant?
The recommended installation is via the official Docker image. Pull and run it with `docker run google/deepvariant:1.10.0`, mounting your input and output directories. A GPU build is available by passing build arguments to the Dockerfile; Singularity is supported for HPC environments.
What is the difference between DeepVariant and GATK HaplotypeCaller?
GATK HaplotypeCaller performs local de novo assembly at candidate sites and applies probabilistic genotyping, while DeepVariant converts pileup data into images and classifies each image with a convolutional neural network. DeepVariant's accuracy depends on its training data; GATK's depends on local assembly quality and the statistical model.
Does DeepVariant support non-human organisms?
The bundled models are trained exclusively on human data. The README notes that using them on non-human organisms introduces pitfalls and links to a dedicated blog post for guidance on non-human variant calling with DeepVariant.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/google-deepvariant)