Self-hosted service
deepmodeling/deepmd-kit avatar
deepmodeling/deepmd-kit

DeePMD-kit: fine-tune a pretrained Deep Potential model and run it under LAMMPS

A deep learning package for many-body potential energy representation and molecular dynamics

2,050 stars654 forksPythonLGPL-3.0

At a glance

What is it?
DeePMD-kit turns quantum-mechanical reference data into machine-learning interatomic potentials. Version 3.2.0 pushes a pretrained-first workflow: download a DPA4 checkpoint, fine-tune it, freeze it, and hand it to a molecular dynamics engine.
Who is it for?
Adopt DeePMD-kit if you already have quantum reference data and a simulation engine that needs a fast surrogate potential, and you are comfortable with a Python 3.10+ stack plus a compiled extension. Do not adopt it if you want a pure-Python library with no build step, or if your problem is a single small molecule where a classical force field is accurate enough.
Can I use it commercially?
Yes, with conditions. LGPL-3.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What DeePMD-kit actually solves, and who it is for

Density functional theory gives you forces and energies you can trust, at a cost that limits you to hundreds or thousands of atoms for picoseconds. Classical force fields run for microseconds but encode assumptions about bonding that break down for reactive or novel chemistry. DeePMD-kit occupies the space between: it trains a neural network on quantum-mechanical reference data so that the resulting potential reproduces that data at a fraction of the cost, then exports the trained model to a molecular dynamics engine. The README describes the package as turning "quantum-mechanical reference data into fast, scalable interatomic potentials," and lists the intended range as finite molecules, covalent systems, periodic solids and metals.

The audience is computational chemists, materials scientists and physics groups who already generate DFT or other ab initio data and want to run larger or longer simulations with it. It is not a tool for someone who wants a ready-made potential without training data. The README's framing makes the intended entry point explicit: "a pretrained model can be your starting point, not just your end result." That sentence is the whole design thesis of the 3.x line. You download a DPA4 checkpoint, fine-tune the full model on your system, then test, export and deploy through the same workflow. Training from scratch is still supported, and the README presents it as the second of two starting points.

The model portfolio splits by intent. DPA4 is positioned for accuracy, DPA4C for simulation throughput and scale. Beyond energy and force, the README lists virials, Hessians, spin and magnetic forces, dipoles, polarizabilities, electronic density of states and atomic populations as supported targets. That breadth is unusual for this class of package and is worth checking against your actual observable before you invest in data generation.

How a training run becomes a simulation: the data flow

The pipeline has five stages, and the README draws them as a flowchart. First you choose a starting point, either a pretrained DPA4 checkpoint or a model configuration for training from scratch. Second you prepare target reference data in DeePMD's NumPy format, or convert structures and trajectories with the companion project dpdata. Third you fine-tune the pretrained model or optimize a new DPA4 or DPA4C model, using single-task, multi-task or distributed training. Fourth you validate and export, through `dp test`, `dp freeze`, backend conversion, embedding extraction and supported compression paths. Fifth you run the simulation through Python or native APIs, or load the model into a supported engine.

The training configuration is a JSON file. The README's fine-tuning example downloads a checkpoint and a matching released training configuration, which tells you the project expects you to start from the hyperparameters the checkpoint was released with rather than inventing your own. That is a meaningful constraint: the released JSON is the documented baseline, and deviating from it moves you outside what the maintainers have validated.

Export produces a frozen model, and the README lists backend-aware model formats with conversion paths between compatible architectures. Compression is a separate step with a stated payoff: the README says model compression can deliver more than 10x inference speedup and reduce memory usage by as much as 20x on supported descriptors and workloads, with the caveat that actual gains depend on the model, system and hardware. Treat those two numbers as an upper bound advertised for the best case, not a prediction for your system.

Deployment targets are broad. The README names LAMMPS, i-PI, ASE, GROMACS, JAX MD, nvalchemi, OpenMM, Amber, CP2K and ABACUS, and lists CLI, Python, C, C++ and Node.js as interfaces. External GNNs such as MACE and NequIP can be connected through plugins, and hybrid potentials can be composed with analytical ZBL or long-range corrections.

Installing DeePMD-kit and fine-tuning a DPA4 checkpoint

The README states the requirement plainly: Python 3.10 or later. The fastest documented path is a shell script that installs the package and puts a `dp` executable on your path. Run it, then confirm the CLI is present.

bash
curl -fsSL https://dp1s.deepmodeling.com | bash
dp --version
dp -h

You should see a version string and the command list. If `dp` is not found, the script's install location is not on your PATH; the installation guide covers pip, conda-forge, containers, offline packages, GPU builds, LAMMPS, i-PI and source installation as the alternative routes.

The fine-tuning walkthrough in the README has two steps: fetch a built-in checkpoint, then fetch the training configuration released with it. The example uses DPA4-Neo, described as one of the recommended general-purpose sizes.

bash
dp pretrained download DPA4-Neo-OMat24-v20260805
curl -fsSL \
    https://huggingface.co/deepmodelingcommunity/DPA4-OMat24/resolve/main/DPA4-Neo-OMat24-v20260805.json \
    -o input_finetune.json

After this you have a checkpoint directory and an `input_finetune.json` that matches it. Point the training data section of that JSON at your own system's data in DeePMD's NumPy format, or convert your structures and trajectories with dpdata first. Then start training against that configuration file.

The README's own sequence after training is `dp test` to validate, `dp freeze` to export, and then backend conversion or compression if you need them. It does not give the exact `dp test` and `dp freeze` argument lists in the excerpt shown here, so read the model guide and the testing and freeze pages in the documentation before running them. The 1.x-era habit of hand-writing a full input.json and training from random initialization still exists, but the README now treats it as the second path, not the default one.

Where DeePMD-kit is the wrong tool

The support matrix is the first trap. The README states directly that backend and interface support varies by model and feature, and that the web documentation marks compatibility and limitations on each feature page. That sentence exists because the combinations are not uniform. If you pick DPA4C for throughput and then discover your target engine's interface only handles a different model family, you have made an expensive choice at the wrong end of the pipeline. Check the feature page for your model and your engine before generating data, not after.

The second trap is the build. The project declares itself Production/Stable and classifies as C, C++ and Python, with scikit-build-core and a custom backend in `backend.dp_backend` driving the build. A pip install from source is not a pure-Python affair. On clusters where you cannot pull a container and cannot reach conda-forge, expect to spend time on the toolchain, particularly for GPU builds. The README points to offline packages and source installation as the escape hatches, which is an acknowledgment that the easy path does not cover every environment.

The third is scope. DeePMD-kit is a potential-fitting framework, not a workflow manager and not a data-generation engine. If your bottleneck is producing the reference data, this package does not help you. If you need a potential for a system where no reasonable pretrained checkpoint exists and you cannot afford the reference calculations, fine-tuning will not rescue you. And if your system is small enough that a well-parameterized classical force field reproduces the observables you care about, the training and validation effort here is wasted.

Finally, the pretrained-first workflow has a subtler cost. Fine-tuning a released checkpoint means inheriting its training distribution. The README's own caveat about compression gains depending on model, system and hardware applies in spirit to fine-tuning as well: the released configuration is a starting point the maintainers validated, not a guarantee about your chemistry.

How DeePMD-kit differs from MACE and NequIP

MACE and NequIP are the obvious comparison points, and the README makes an unusual move by naming them as connectable through plugins rather than treating them as pure rivals. That tells you something about the architecture: DeePMD-kit is partly a training and deployment framework that can host external GNN architectures, not only an implementation of one model family.

The practical difference is in the surrounding machinery. MACE and NequIP are research codebases that ship a model architecture and a training loop. DeePMD-kit ships a model portfolio (DPA4 for accuracy, DPA4C for throughput), four supported training backends (TensorFlow, PyTorch, JAX, Paddle), a freeze-and-export step, compression, an AOTInductor `.pt2` export path, and a long list of engine integrations including LAMMPS, i-PI, GROMACS, OpenMM, Amber, CP2K and ABACUS. If your deliverable is a paper about an architecture, the smaller projects are a shorter route. If your deliverable is a production simulation running under LAMMPS on a cluster, the integration surface is what you are buying.

The cost of that surface is the compatibility matrix described above. A focused research codebase has fewer combinations to keep working. DeePMD-kit has more, and the README's warning about varying support is the price of it. Choose accordingly: breadth of deployment targets against the probability that your particular model-plus-engine pair is the one with a documented limitation.

Licence, maintenance and what an upgrade costs

DeePMD-kit is licensed LGPL-3.0-or-later, and `pyproject.toml` declares the SPDX identifier as `LGPL-3.0-or-later` with `LICENSE` listed in `license-files`. The LGPL matters if you link the compiled library into a proprietary application: the licence is weaker than GPL but still imposes obligations around the library itself, including relinking and notice requirements. If you only call the Python API or the `dp` CLI from your own scripts, you are in more familiar territory. This is not legal advice; the `LICENSE` file and your own counsel are the sources that count.

The repository is not archived, and the last push was on 2026-09-07, which is recent. Releases are frequent: v3.2.0 on 2026-08-19, a v3.2.0b0 beta on 2026-06-06, and v3.1.3 on 2026-03-19. That cadence is a real upgrade cost. The README's own history shows the framework moving from hand-written configurations toward pretrained checkpoints and released training JSONs, which means an input file written for an earlier minor version may not be the recommended shape today.

Two upgrade hazards are visible in the README and the repository files. First, model formats are backend-aware, and the README mentions conversion paths only for compatible architectures, so a frozen model produced under one backend is not automatically usable under another. Second, the pretrained checkpoints are versioned by name, for example `DPA4-Neo-OMat24-v20260805`, and the training configuration is fetched from a URL that embeds that same version string. Pin both. A checkpoint and a configuration from different releases are not documented as interchangeable.

Contributor friction is low: the repository carries `CONTRIBUTING.md`, `AGENTS.md`, a pre-commit configuration, clang-format settings and a devcontainer definition. The `skills/` and `dpa_adapt/` top-level directories suggest the project is actively growing new surfaces beyond the core training path.

Editorial conclusion

Adopt DeePMD-kit if you already have quantum reference data and a simulation engine that needs a fast surrogate potential, and you are comfortable with a Python 3.10+ stack plus a compiled extension. Do not adopt it if you want a pure-Python library with no build step, or if your problem is a single small molecule where a classical force field is accurate enough. Before committing, verify three things: that a pretrained DPA4 checkpoint covers chemistry close to yours, that the backend you intend to train with (TensorFlow, PyTorch, JAX or Paddle) supports the model family you picked, and that the frozen model loads in your target engine version. The README's own warning is the place to start: backend and interface support varies by model and feature, and the documentation marks compatibility on each feature page.

Frequently asked questions

How do I install DeePMD-kit?

The README gives a one-line install script, `curl -fsSL https://dp1s.deepmodeling.com | bash`, followed by `dp --version` to confirm it worked. It requires Python 3.10 or later, and the installation guide also covers pip, conda-forge, containers, offline packages, GPU builds, LAMMPS, i-PI and source installation.

What is DeePMD-kit?

It is a deep learning package that turns quantum-mechanical reference data into fast, scalable interatomic potentials for molecular and materials simulation. The README describes its scope as finite molecules, covalent systems, periodic solids and metals, with deployment into engines such as LAMMPS, i-PI and ASE.

Is molecular dynamics hard to learn?

The README does not discuss the difficulty of learning molecular dynamics as a field. DeePMD-kit's documentation assumes you already have reference data and a target simulation engine, so it is not positioned as an introduction to the subject.

What are molecular dynamics models?

The README frames the relevant distinction as quantum-mechanical reference data versus fast interatomic potentials trained on it. DeePMD-kit produces the second kind, and lists energy, force, virials, Hessians, spin and magnetic forces, dipoles, polarizabilities and electronic density of states among the properties its models can target.

Official sources

  1. deepmodeling/deepmd-kit on GitHub
  2. License: LGPL-3.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/deepmodeling-deepmd-kit.svg)](https://hysenlabs.com/projects/deepmodeling-deepmd-kit)