DeepChem: Deep Learning for Drug Discovery and Materials Science
Democratizing Deep-Learning for Drug Discovery, Quantum Chemistry, Materials Science and Biology
At a glance
- What is it?
- A Python library that brings deep learning to molecular and materials research. DeepChem handles the data preparation and model pipeline that researchers would otherwise build from scratch, with prebuilt support for TensorFlow, PyTorch, and JAX.
- Who is it for?
- DeepChem is for research teams working on molecular problems where deep learning might help: drug discovery, materials screening, protein folding, or quantum chemistry. Skip it if your use case fits narrow sklearn models or if you need bleeding-edge transformers before they land in DeepChem.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 43 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Why molecular deep learning needs more than TensorFlow
Applying deep learning to drug discovery or materials science means solving three problems at once: converting molecules to tensors, building or adapting neural network architectures, and handling tabular data alongside molecular structures. A researcher starting fresh would write a data loader, a tokenizer for SMILES strings or molecular graphs, and a training loop. DeepChem provides the data loaders for standard benchmarks (BACE, HIV, Tox21), the featurizers that turn molecules into numerical representations, and the model classes that know how to consume them. That saves weeks of integration work and means a chemist can focus on the problem instead of the infrastructure.
How DeepChem connects RDKit chemistry to neural networks
DeepChem sits between RDKit (a chemistry toolkit) and TensorFlow, PyTorch, or JAX (the deep learning backend). You define a dataset using a DeepChem Dataset object, which pairs molecules (SMILES strings or 3D coordinates) with target values (activity, solubility, toxicity). A Featurizer converts each molecule into a numerical representation: a fixed-size vector, an adjacency matrix, or a graph structure. A Model takes the features, trains a neural network, and returns predictions. The project includes prebuilt models (graph convolutional networks, message-passing networks) and a train-test-validate loop. The same dataset and model interfaces work whether you swap the backend to PyTorch or JAX, which matters if a team wants to try different frameworks without rewriting the whole pipeline.
Installing and running a first model
DeepChem installs as a Python package with its core dependencies (NumPy, pandas, scikit-learn, RDKit). The stable version is available from PyPI:
pip install deepchemOr via conda:
conda install -c conda-forge deepchemIf you want to use neural networks, install one of the deep learning frameworks as well:
pip install deepchem[torch]Or for TensorFlow:
pip install deepchem[tensorflow]The GitHub repository includes tutorials in the examples/tutorials directory, written as Jupyter notebooks and runnable in Google Colab. A beginner tutorial loads the BACE dataset (molecules with activity measurements), featurizes them, trains a simple neural network, and evaluates it. That cycle takes a few minutes once the libraries are installed, which tells you quickly whether deep learning is worth the effort for your molecules.
Soft dependencies and framework integration for neural networks
DeepChem's core requires six mandatory packages: joblib, NumPy, pandas, scikit-learn, SciPy, and RDKit. But it also lists soft dependencies for optional functionality: if your task uses graph neural networks, you need Deep Graph Library (DGL); for transformers and modern language models, you need Hugging Face transformers; for quantum simulations, you need PySCF. The documentation states that an ImportError at runtime tells you which one to install. This design means a fresh pip install is lightweight, but a team building serious neural network models ends up installing TensorFlow or PyTorch anyway, adding 500 MB or more. DeepChem supports three deep learning backends: TensorFlow via `pip install deepchem[tensorflow]`, PyTorch via `pip install deepchem[torch]` with additional dependencies including pytorch-lightning and DGL, and JAX via `pip install deepchem[jax]` with dm-haiku and optax. Each framework requires CUDA separately for GPU support before installing DeepChem. In zsh, square brackets are special characters, so you must escape the installation commands with quotes: `pip install 'deepchem[torch]'`. Docker images are available from DockerHub at `deepchemio/deepchem:x.x.x` for stable versions or `deepchemio/deepchem:latest` for development builds.
Standard benchmarks and custom data pipelines
DeepChem ships datasets for common problems: BACE (Pfizer blood-brain barrier prediction), HIV (activity screening), Tox21 (toxicity prediction with 12K compounds), and Clintox (clinical trial outcomes). These let you test the pipeline immediately and compare models against published baselines in the examples directory. The examples/ directory includes domain-specific tutorials: examples/adme/ for absorption-distribution-metabolism-excretion, examples/kinase/ for kinase inhibition prediction, examples/hiv/ for activity screening, examples/hopv/, examples/delaney/, and others. They also represent a limitation: if your molecules are from a private compound collection, you need to build your own featurizer or data loader, which puts you back in the infrastructure business. The repository provides examples for loading CSV and SDF (structure data format) files and integrating custom Python featurizers, but the README does not document custom featurizers in detail beyond directing you to readthedocs. The Dataset API expects pairs of molecules (SMILES strings or 3D coordinates) and target values, and you can pass a custom featurizer function to convert molecules into numerical representations. The data loader and model interface stay consistent; only the featurization layer changes.
Modular pipeline for domain experts and model interpretability
DeepChem exposes each layer: you can inspect the featurized representations, swap models, and combine predictions from multiple networks. This matters for teams where a domain expert (a chemist or biologist) needs to understand what the model learned. If a graph convolutional network learns that ring count is predictive, you can extract that fact. You can also compose ensemble models that combine multiple architectures' predictions, letting different models specialize on different failure modes. The modular design allows you to replace the featurizer without changing the training loop, or replace the model without reloading the dataset. This composability is critical in molecular research where interpretability and transparency matter for publication. The Discord community and discussion forum host scientists and developers, not just the core team, which means you can ask whether a design choice is typical for molecular learning or a bug in your setup. Weights & Biases integration allows tracking of model training and evaluation metrics across experiments.
Active development; release lag behind bleeding-edge techniques
DeepChem's last push to GitHub was 2026-08-20, about one month ago. The project is not archived and continues receiving updates. The most recent release was 2.8.0 in April 2024 (version 2.8.0.pre came out on 2024-04-02). Updates come steadily but lag behind the latest transformers and models: if a new technique becomes mainstream in molecular machine learning, you may wait weeks or months for it to land in DeepChem. The project is MIT licensed and managed by open-source contributors, not a commercial company, which means no vendor lock-in but also no service level agreement. The setup.py file in the repository defines environment-specific dependencies for JAX, PyTorch, and TensorFlow frameworks, allowing you to build development environments with all dependencies or minimal ones. Contributing is documented in CONTRIBUTING.md. The discord community and discussion forum at forum.deepchem.io host scientists and developers asking questions and proposing features. The community maintains Model and Tutorial wishlists as GitHub issues, guiding the project's roadmap.
Editorial conclusion
DeepChem is for research teams working on molecular problems where deep learning might help: drug discovery, materials screening, protein folding, or quantum chemistry. Skip it if your use case fits narrow sklearn models or if you need bleeding-edge transformers before they land in DeepChem. Start by reading the tutorials for your domain on the GitHub repository, then check which deep learning framework you have time to learn.
Frequently asked questions
What is deepchem used for?
DeepChem applies deep learning to molecular discovery problems: drug screening, toxicity prediction, protein properties, and quantum chemistry simulations. The library handles the data loading and featurization that would otherwise require manual work.
How do you install DeepChem?
Install the base library with pip install deepchem, then add the deep learning framework you want: pip install deepchem[torch] for PyTorch, pip install deepchem[tensorflow] for TensorFlow, or pip install deepchem[jax] for JAX.
What is DeepChem?
DeepChem is a Python library that brings deep learning to molecular and materials research. It provides data loaders, featurizers that convert molecules to numbers, and prebuilt neural network models for drug discovery and related fields.
Is DeepChem free to use?
Yes, DeepChem is open source under the MIT license. There are no usage fees or commercial restrictions.
How does DeepChem compare to Chemprop?
DeepChem is a general-purpose library for molecular machine learning with support for multiple backends (TensorFlow, PyTorch, JAX) and many model types. Chemprop is a specialized tool for property prediction using message-passing neural networks. DeepChem gives more flexibility; Chemprop focuses on one architecture.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/deepchem-deepchem)