DVC: Git for Data and Lightweight ML Pipelines
Data Versioning and ML Experiments. When you make changes, only run the steps impacted by those changes.
At a glance
- What is it?
- DVC is a Python command line tool that versions data and models alongside Git and reruns only the pipeline steps affected by a change. It fits teams already on Git, and it is the wrong tool when you need a server-side data catalogue.
- Who is it for?
- Adopt DVC if your team already works in Git and wants data, models and pipeline stages tracked in the same repository, with the cache pushed to S3, Azure, Google Cloud or SSH storage. Skip it if you need a hosted data catalogue with row-level lineage, or if your data cannot leave a single machine.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 9 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem DVC solves: data files that do not belong in Git
Git tracks code well and large binary files badly. A dataset or model checkpoint committed to Git bloats the repository, and every clone pays for it. DVC's answer is to keep the version information in Git and the bytes somewhere else. The README describes this directly: store data and models in your cloud storage but keep their version info in your Git repo. The repository calls this "Git for data" and compares it to Git-LFS, with the difference that DVC needs no dedicated server.
The audience is machine learning teams that already use Git and do not want to adopt a separate platform to version their inputs. The pyproject.toml lists keywords including data-version-control, machine-learning and reproducibility, and requires Python 3.9 or newer. The project classifies itself as Development Status 4 - Beta even at version 3.67.1, which is worth reading as a statement about API expectations rather than a warning about stability.
How DVC separates Git metadata from the data cache
The mechanism has three parts. First, dvc add moves a file or directory into a local cache outside the repository and writes a small .dvc file in its place. That .dvc file is what you commit. Second, a remote is where the cache is copied for sharing and backup, and the README lists any cloud (S3, Azure, Google Cloud) or on-premise network storage via SSH. Third, dvc push and dvc pull move content between the local cache and that remote.
Pipelines are the second layer. DVC pipelines are computational graphs that connect code and data, specifying input dependencies, the command to run, and the outputs to save. Because each stage declares its dependencies and outputs, DVC can determine which stages a change affects. That is the claim in the project description: when you make changes, only run the steps impacted by those changes. The README's own analogy is Makefiles for ML.
Experiments are the third layer. dvc exp run executes a pipeline under a named experiment, results live in the local Git repo with no server, and dvc exp show compares them. Note the boundary: the cache is local until you push it, so a colleague who clones your repository gets the .dvc pointers and the pipeline definitions but not the data itself.
Install DVC with pip and run a first tracked pipeline
The README lists several installation routes: snap, choco, brew, conda, pip, or an OS-specific package, and it points to https://dvc.org/doc for full instructions. The simplest route for a Python project is pip. The pyproject.toml requires Python 3.9 or newer.
pip install dvcAfter installation, the README's quick start begins by adding your code to Git and then adding a data directory to DVC. The dvc add command prints the size of the data it moved and creates a .dvc file next to the directory. Commit that file, not the images.
git add train.py params.yaml
dvc add images/Next, define stages. The README gives these two commands, where -n names the stage, -d declares an input dependency, -o declares a regular output, -M declares a metrics file, and the trailing text is the command to execute.
dvc stage add -n featurize -d images/ -o features/ python featurize.py
dvc stage add -n train -d features/ -d train.py -o model.p -M metrics.json python train.pyWith the graph defined, run an experiment, edit train.py, and run a second one. The second run should only execute the stages whose dependencies changed.
dvc exp run -n exp-baseline
vi train.py
dvc exp run -n exp-code-change
dvc exp showTo share the data rather than just the pointers, configure a remote and push. The README uses an S3 example.
dvc remote add myremote -d s3://mybucket/image_cnn
dvc pushWhere DVC stops being the right tool
The first limitation is stated by the README itself: the cache is what gets pushed, so a repository without a configured remote is a repository whose data exists on one machine. Cloning gives you the .dvc files and the pipeline definitions, and dvc pull has nothing to fetch. There is no documented rollback path for a cache that was deleted before it was pushed.
The second is scope. DVC versions artifacts and pipelines. It does not store row-level lineage, it does not serve a query interface over your tables, and it does not manage access control on the underlying storage. If your team's question is "which rows fed this model, and who can read them", DVC's .dvc files will not answer it.
The third is the Beta classifier. The project ships frequent releases (3.67.1 on 2026-03-31, 3.67.0 on 2026-03-14, 3.66.1 on 2026-01-08) and the last push to main was on 2026-03-31, but the package metadata still declares Development Status 4 - Beta. Treat CLI behaviour as stable in practice and internal APIs as not.
Finally, a pipeline stage is only as reproducible as its command. DVC reruns what you declared as a dependency; if a stage reads a file you never listed with -d, DVC has no reason to rerun it and no way to know it went stale.
DVC against lakeFS and MLflow: different layers, not substitutes
People searching for DVC alternatives usually mean one of two things. Against lakeFS, the difference is where versioning happens. lakeFS versions the object store itself, giving you branches and commits over the data at rest, so any tool reading that storage sees a consistent snapshot. DVC versions artifacts from the client side and records pointers in Git, so versioning is a property of your repository rather than of the storage layer. If several teams write to the same bucket with different tools, lakeFS's placement is the stronger fit; if the unit of work is a Git repository and its pipeline, DVC's is.
Against MLflow, the difference is the centre of gravity. MLflow is oriented around a tracking server, runs, and a model registry. DVC's README is explicit that experiment tracking happens in your local Git repo with no servers needed, and that comparison covers data, code, parameters, models and plots. The trade-off is real: no server means no central dashboard, and collaboration goes through Git hosting plus whatever remote you configure. A team that wants a queryable run history across many people will find MLflow's server model more natural. A team that wants everything to survive a git clone will prefer DVC.
Licence, upgrade cost and the release cadence
DVC is Apache-2.0, declared in pyproject.toml with license-files pointing at LICENSE. That permits commercial use and modification, and it includes a patent grant. It does not give legal advice; if you redistribute DVC inside a product, read the NOTICE and attribution requirements yourself.
The dependency list in pyproject.toml is long and pinned in places: dvc-data is constrained to >=3.18.2,<3.19.0, dvc-render to >=1.0.1,<2, and dvc-studio-client to >=0.21,<1. Those upper bounds mean a DVC upgrade can require coordinated upgrades of the dvc-data and dvc-render packages, which is the practical cost of staying current. The cadence is fast: three releases between 2026-01-08 and 2026-03-31. Nothing in the README documents a long-term support branch, so plan on tracking recent versions rather than pinning to an old one.
The VS Code extension is a separate install from the Marketplace and, per the README, requires core DVC on your system separately. Two upgrade paths to keep in sync.
Editorial conclusion
Adopt DVC if your team already works in Git and wants data, models and pipeline stages tracked in the same repository, with the cache pushed to S3, Azure, Google Cloud or SSH storage. Skip it if you need a hosted data catalogue with row-level lineage, or if your data cannot leave a single machine. Before rolling it out, verify one thing: run dvc add on a real dataset, confirm the .dvc file is small and committed, and confirm dvc push actually writes to your configured remote, because the README documents no rollback or recovery path for a cache that was never pushed.
Frequently asked questions
What is DVC used for?
DVC versions data and models, stores their version info in a Git repository while the bytes live in cloud or on-premise storage, and defines pipelines that rerun only the steps affected by a change. It also tracks experiments locally in Git, with no server required.
What are the key differences between lakeFS and DVC?
lakeFS versions data at the object-store layer, so every tool reading that storage sees a consistent snapshot. DVC versions artifacts from the client side and keeps pointers in Git, which ties versioning to a repository and its pipeline rather than to the storage itself.
What is Data Version Control (DVC) and how does it work?
DVC is a command line tool and VS Code extension. dvc add moves a file or directory into a local cache and leaves a small .dvc file to commit; a remote such as S3, Azure, Google Cloud or SSH storage holds the cache for sharing and backup; and dvc push and dvc pull move content between the two.
How does data version control work?
Git stores and versions code, including DVC meta-files that act as placeholders for data, while DVC keeps the actual data and model files in a cache outside Git. Pipelines record each stage's dependencies, command and outputs so DVC can determine which stages a change affects.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/treeverse-dvc)