# spacy-models: the release repository behind every spaCy pipeline

> explosion/spacy-models is not a library you import but a release archive of trained pipelines distributed as .whl and .tar.gz files. Here is how installation, version compatibility and the sm/md/lg trade-off actually work, and where the model data falls short.

**explosion/spacy-models** — 💫  Models for the spaCy Natural Language Processing (NLP) library

- Repository: https://github.com/explosion/spacy-models
- Website: https://spacy.io
- Stars: 1,908 · Forks: 322
- Language: Python
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/explosion-spacy-models

## What spacy-models actually is, and who needs it

The repository holds releases of models for the spaCy NLP library. It is not a Python package you install to get functionality; it is the distribution point for trained pipelines that spaCy loads at runtime. The README is explicit about why: the models can be very large and consist mostly of binary data, so they cannot live as files in a GitHub repository. Instead they are attached to releases as .whl and .tar.gz files, which keeps a public release history without bloating the git tree.

That design decision shapes who the project is for. If you are writing an application that needs tokenization, part-of-speech tagging, dependency parsing, lemmatization or named entity recognition in a supported language, you are the target user. You install a pipeline, call spacy.load, and get a callable object. If you are training your own model, this repository is mostly irrelevant to you except as a source of baseline pipelines and naming conventions. The top-level entries confirm the scope: compatibility.json, shortcuts.json, shortcuts-v2.json, shortcuts-nightly.json, a meta/ directory and a tests/ directory. There is no training code here.

## How a model name encodes its components, genre and size

spaCy expects model packages to follow the naming convention [lang]_[name]. For the provided pipelines, the name breaks into three parts. The type describes capabilities: core gives tagging, parsing, lemmatization and named entity recognition; dep gives only tagging, parsing and lemmatization; ent gives only named entity recognition; sent gives only sentence segmentation. The genre describes the training text, for example web for web text or news for news text. The size indicator describes the vector table: sm has no word vectors, md has a reduced table with 20k unique vectors covering roughly 500k words, and lg has a large table with around 500k entries.

The README's own worked example is en_core_web_md: a medium English model trained on written web text (blogs, news, comments) with a tagger, parser, lemmatizer, entity recognizer and a 20k-vector table. Reading the name before installing saves real time. If your pipeline only needs entities, pulling a core model downloads a parser and lemmatizer you will never call. If you need word vectors for similarity, sm is the wrong choice by definition, because it has none.

## Version numbers are a compatibility contract, not a changelog

A model version a.b.c maps to three things. The a digit is the spaCy major version, so 2 means spaCy v2.x. The b digit is the spaCy minor version, so 3 in a 2.3.x context means spaCy v2.3.x. The c digit is the model version itself, and it changes when the config changes: different training data, different parameters, different iteration counts, different vectors.

This is the part that trips people up. The model version is not a semantic version of quality; it is a compatibility marker. The README points to compatibility.json as the detailed overview and notes that this file is also the source of spaCy's internal compatibility check, performed when you run the download command. That means python -m spacy download is doing more than fetching a file: it consults compatibility data to pick a matching version. If you install a .whl by hand with pip, you bypass that check, and a mismatch between the model's spaCy minor version and your installed spaCy is your problem to catch.

## Installing a model and running your first document

The quickstart is a single command with the model name substituted in. The README uses en_core_web_sm as the example.

```bash
python -m spacy download en_core_web_sm
```

This downloads the best-matching version of that model for your spaCy installation, using the compatibility data described above. Once it finishes, the model is importable and loadable.

```python
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp(u"This is a sentence.")
```

The README also documents a second loading path: import the model by its full name and call its load() method with no arguments. The README states this should also work for older models in previous versions of spaCy.

```python
import spacy
import en_core_web_sm

nlp = en_core_web_sm.load()
doc = nlp(u"This is a sentence.")
```

If you would rather control the file yourself, the README gives four pip forms: a local .tar.gz, a local .whl, and the same two fetched from a release URL.

```bash
pip install /Users/you/en_core_web_sm-3.0.0.tar.gz
pip install /Users/you/en_core_web_sm-3.0.0-py3-none-any.whl
pip install https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.0.0/en_core_web_sm-3.0.0.tar.gz
pip install https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.0.0/en_core_web_sm-3.0.0-py3-none-any.whl
```

For a manual install into a custom directory, the README shows the archive layout: the downloaded .tar.gz contains setup.py, a meta.json, and a pipeline package directory holding __init__.py plus a versioned data directory with config.cfg, meta.json and component data. Downloading the archive by browser and unpacking it is the documented route when you do not want the package installed into site-packages.

## The size and language limits you should plan around

The most concrete limitation is stated in the README itself: the models can be very large. That is the reason they are release artifacts rather than repository files, and it is the reason a container image that bundles an lg pipeline grows accordingly. There is no documented way in this repository to fetch only the components you need from a core model; the size tiers are the granularity you get.

The second limit is coverage. This repository publishes pipelines for a set of languages and genres, and the README's index points to the models directory for the current list rather than enumerating it. If your language or domain is not represented, the correct move is training your own pipeline in spaCy, not stretching a nearby model. The type prefixes help here: a dep model has no entity recognizer, so using one for NER will not work regardless of how good the parser is.

A third constraint is version drift. Because the model version's first two digits track spaCy, upgrading spaCy usually means reinstalling models. Nothing in the README describes a rollback procedure for a model that turns out to regress on your data, so pinning the exact version string in your dependency file is the practical safeguard.

## How it compares to NLTK's data and to Hugging Face model hubs

The natural comparison is NLTK, which also ships pretrained data. The difference in approach is packaging. NLTK distributes corpora and models through its own downloader into a data directory that the library searches at runtime. spacy-models distributes each pipeline as an installable Python package with a version number that encodes spaCy compatibility, so your existing dependency tooling can pin it and your deployment can install it from a URL. The trade-off is that NLTK's downloader covers a much broader set of corpora and tasks, while spacy-models covers pipelines for a single library's runtime.

The other comparison is a general model hub. A hub is oriented around sharing checkpoints from arbitrary training runs; spacy-models is oriented around one library's release process, with compatibility.json driving an automatic version check during download. That is narrower and more predictable, but it means you cannot bring an arbitrary checkpoint and expect spacy.load to accept it without the packaging work the README describes. If your requirement is a transformer checkpoint you fine-tuned yourself, this repository is the wrong distribution channel.

## Maintenance cadence, licensing and upgrade cost

The last push to the repository was on 2026-03-20, and the most recent release listed is en_core_web_hftrf-3.8.1 on the same date. The two zh_core_web releases listed above it date to 2024-09-30. That pattern is worth reading carefully: releases arrive when a model is retrained or a compatibility line moves, not on a schedule. The repository is not archived, but the cadence is event-driven, and the README does not promise otherwise.

Upgrade cost is dominated by the versioning rule. A spaCy minor bump can require a matching model bump, and because the third digit changes with training config, the same model name can behave differently across c versions. Pinning exact strings is the only way to make a deployment reproducible.

Licensing is the part to check per model, not per repository. The repository's own license field is not stated in the available repository information, and the README's v1.x table shows individual models carrying different license identifiers, including CC BY-SA and CC BY-NC. A non-commercial identifier on a model has obvious implications for commercial deployment. Treat the licence of the specific model release you install as the governing one, and read it before shipping. This is a factual observation about the table, not legal advice.

## Conclusion

Adopt spacy-models if you need a pretrained pipeline for tagging, parsing, lemmatization or NER and you want it installed as a normal Python package. Do not adopt it if you need a model for a language or domain that has no published release, or if you cannot accept the download size of an md or lg pipeline. Before committing, verify three things: that the model version's first two digits match your spaCy minor version, that the type prefix (core, dep, ent, sent) covers the components you need, and that compatibility.json lists the model you intend to pin. The repository is a distribution channel, not a training toolkit, so plan separately for anything the published pipelines do not cover.

## FAQ

### What are spaCy models?

They are trained pipelines for the spaCy NLP library, distributed from this repository as .whl and .tar.gz release files because they are large and mostly binary data. A name like en_core_web_md tells you the language, the components, the training genre and the vector table size.

### How do I install spaCy models?

Run python -m spacy download followed by the model name, for example en_core_web_sm. This fetches the best-matching version for your spaCy installation using the compatibility data. Alternatively, pip install a .whl or .tar.gz archive from a local path or a release URL.

### What is spaCy used for?

spaCy is an NLP library, and the models here supply its runtime capabilities: tagging, parsing, lemmatization and named entity recognition, depending on the model type. The core type includes all four; dep, ent and sent each cover a narrower subset.

### Is NLTK or spaCy better?

The two distribute pretrained data differently. NLTK ships corpora and models through its own downloader into a data directory; spacy-models ships each pipeline as an installable Python package whose version encodes spaCy compatibility. Which is better depends on whether you need NLTK's broader corpus coverage or spaCy's pinned, package-managed pipelines.

### Is spaCy still relevant?

The repository is not archived, its last push was on 2026-03-20, and releases such as en_core_web_hftrf-3.8.1 shipped on that date. Beyond that, the README does not discuss the project's long-term roadmap.

## Sources

- [explosion/spacy-models on GitHub](https://github.com/explosion/spacy-models)
- [Issues](https://github.com/explosion/spacy-models/issues)
- [Project website](https://spacy.io)
- [README](https://github.com/explosion/spacy-models/blob/master/README.md)
- [Releases](https://github.com/explosion/spacy-models/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/explosion-spacy-models
