# spaCy: what the Python NLP library does, how to install it, and where it stops

> spaCy is an MIT-licensed NLP library for Python and Cython with pretrained pipelines for 70+ languages. This covers the pipeline mechanism, the install path, the cases it does not fit, and how it differs from NLTK.

**explosion/spaCy** — Industrial-strength Natural Language Processing (NLP) in Python. spaCy: Industrial-strength NLP spaCy is a library for **advanced Natural Language Processing** in Python and Cython.

- Repository: https://github.com/explosion/spaCy
- Website: https://spacy.io
- Stars: 33,923 · Forks: 4,727
- Language: Python
- License: MIT
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/explosion-spacy

## The gap spaCy fills: NLP that ships inside a product

The README states the library was "designed from day one to be used in real products." That sentence is the whole positioning. Most Python NLP code in the 2010s was written for research: load a corpus, try a tokenizer, print statistics. spaCy was built the other way around, as an importable component that takes a string and returns annotated objects you can loop over in a request handler.

The target user is a Python engineer who needs named entity recognition, part-of-speech tagging, dependency parsing or text classification in a running service, and who does not want to assemble a tokenizer, a tagger and a parser from separate projects. The README lists tokenization and training for 70+ languages, pretrained pipelines, neural network models for tagging, parsing, named entity recognition and text classification, multi-task learning with pretrained transformers such as BERT, and a training system with model packaging and deployment. The library is written in Python and Cython, and the compiled modules listed in setup.py show where the speed comes from: the tokenizer, the parser internals, the NER transition system, the morphology tables and the Doc, Span and Token containers are all Cython extensions rather than pure Python.

That design choice has a cost. A pure Python library installs anywhere. A Cython library needs a compiled wheel for your platform and Python version, which is why the project maintains a cibuildwheel configuration covering manylinux, musllinux and macOS images. If your target platform is unusual, you are in source-build territory, and the build requirements in pyproject.toml include Cython, cymem, preshed, murmurhash, thinc and numpy.

## How the pipeline actually works: Doc objects, components and a shared Vocab

A spaCy pipeline is a sequence of components. You call it with text, and you get back a Doc. The Doc owns the tokens, and components read from and write to it: the tagger assigns part-of-speech tags, the parser assigns dependency relations and sentence boundaries, the entity recognizer assigns entity spans, and a text classifier assigns document-level categories. Because everything lands on the same object, a downstream component sees what an earlier one produced without any glue code between them.

Underneath, tokens are not Python objects in the usual sense. The compiled modules in setup.py include spacy.tokens.doc, spacy.tokens.span and spacy.tokens.token, and spacy.lexeme and spacy.vocab. Lexemes hold the lexical type, and the Vocab holds the string store and shared lexeme table, so a token is a lightweight view into shared state rather than a standalone object. This is the mechanism behind spaCy's memory behaviour on long documents, and it is also why the API feels different from libraries where a token is a dictionary-like object you can freely copy.

Training is the second half. The README points to a training system with model packaging and deployment, and the repository carries a training config system: requirements.txt pins confection, and the release notes for 3.8.15 mention fixing a click requirement, which points at the command-line layer built on typer and click. Multi-task learning with pretrained transformers is supported, so a pipeline can combine a transformer with task-specific heads. The examples/ directory holds training material, and examples/README.md is the entry point for it.

## Installing spaCy and running a first pipeline

The README points to PyPI and conda-forge, and the package name is spacy. The Makefile in the repository shows the project's own install path for development: it creates a virtual environment, upgrades pip, setuptools and wheel, and installs numpy before building wheels into a wheelhouse.

```bash
python$(PYVER) -m venv $(VENV)
$(VENV)/bin/pip install -U pip setuptools pex wheel
$(VENV)/bin/pip install numpy
```

For normal use, install the published package instead of building it, and remember that the library ships without models. The README links to spacy.io/models for pretrained pipelines, which are separate downloads. Release 3.8.14 is described as a bug fix for model downloading in environments without pip on PATH, so if your deployment shells out to download models from a minimal container, check that version or later.

The repository's own test path shows how the project runs spaCy from a packaged binary, using a pex environment with PEX_PATH pointing at the built spacy pex and pytest invoked with --pyargs spacy -x:

```bash
( . $(VENV)/bin/activate ; \
PEX_PATH=dist/spacy-$(version).pex ./dist/pytest.pex --pyargs spacy -x ; )
```

That is the maintainer workflow, not the end-user one, but it makes the packaging boundary explicit: the library and its dependencies are resolved into a single artefact, and the models are not part of it. In your own code, load a pipeline once at process start rather than per request, then call it on a string to get a Doc and read entities, tags and dependency relations from that one object.

## Where spaCy is the wrong tool

The first limitation is model coverage versus language coverage. The README claims tokenization and training for 70+ languages, but tokenization and training are not the same as a good pretrained pipeline. A language can be tokenizable out of the box and still have no accurate NER model, and the README does not claim otherwise. If your language or your domain is not covered by a released pipeline, you are in the training path, which means annotated data and a training config, not a pip install.

The second limitation is the object model itself. Because tokens are views into shared state, patterns that work on plain Python objects do not always transfer. Copying a token, storing it across requests, or pickling parts of a Doc are not the same operations they would be with a list of strings. Code that treats a token as a durable record will break in ways that are hard to debug.

The third is build friction. The cibuildwheel configuration skips cp39, win32, i686 and cp310-win_arm64 targets, and the comments in pyproject.toml explain that numpy 2.3 or later only ships manylinux_2_28 wheels, which is why the x86_64 and aarch64 build images moved to manylinux_2_28 while ppc64le and s390x stayed on manylinux2014. Practical consequence: on an older glibc distribution or on 32-bit Windows, you may have no wheel and will be compiling from source. The pyproject.toml also shows a Rust toolchain being installed before the build, so a source build is not a small undertaking.

Finally, if your task is corpus statistics, concordances, or comparing five tokenizers on the same text, spaCy's opinionated pipeline is overhead. It wants to be the pipeline, and it is not a toolkit of interchangeable parts.

## spaCy versus NLTK: a pipeline against a toolkit

The comparison people search for is spaCy versus NLTK, and the difference is architectural rather than a matter of accuracy. NLTK is a collection of algorithms and corpora with a common interface style but no shared document object. You choose a tokenizer, you choose a tagger, you call them in sequence, and you pass plain Python structures between them. That is flexible and it is why NLTK remains the standard teaching tool: you can swap a component and see what changes.

spaCy inverts that. One object, the Doc, is annotated in place by components that are registered in a pipeline. You can add, remove and reorder components, and the training system lets you train them jointly, but you are working inside the framework rather than assembling parts. The README's phrase "multi-task learning with pretrained transformers" only makes sense in that framing: shared representations across tasks require a shared pipeline.

The practical split: reach for NLTK when the task is exploratory, when you need a specific algorithm or corpus it ships, or when you want to compare approaches. Reach for spaCy when the task is fixed, the output shape is known, and the code has to run in production with a model you can version and package. Neither is a superset of the other, and the choice is usually made by what the surrounding service needs, not by benchmark tables.

## Maintenance, releases and the upgrade cost you are signing up for

The repository is not archived, and the last push was on 2026-08-24. The most recent release is v3.8.16 on the same date, following v3.8.15 on 2026-08-07 and v3.8.14 on 2026-03-29. The 3.8 line is receiving patch releases, and the README announces version 3.8 as current.

Upgrading is where the cost sits. spaCy pins its own stack tightly: thinc is constrained to >=8.3.12,<8.4.0, numpy to >=2.0.0,<3.0.0, pydantic to >=2.0.0,<3.0.0, and click to >=8.2.1,<9.0.0. If another library in your environment needs a different thinc or pydantic major version, resolution fails and you are choosing between projects. The 3.8.15 release being described as a fix to the click requirement is a small illustration of how a transitive pin can break an install.

Models are versioned separately from the library, and the release notes record a fix for model downloading in environments without pip on PATH. That matters for container images built without pip on the path or with restricted network access at runtime. Pin both the spaCy version and the model version in your lockfile, and install models at image build time rather than at process start.

The licence is MIT, and the repository carries a licenses/ directory alongside LICENSE, which suggests bundled third-party components with their own terms. MIT is permissive, but if you redistribute a built artefact, read the licenses/ contents rather than assuming a single licence covers everything. That is a factual note about the repository layout, not legal advice.

## Conclusion

Adopt spaCy when you need named entity recognition, tagging or parsing inside a Python service and are willing to track model and library versions together. Do not adopt it for raw corpus statistics or experimentation with many tokenizers, where NLTK's separate objects fit better. Before committing, verify that a pretrained pipeline exists for your language and domain, and check how your deployment downloads and stores model packages, since the release notes record a fix for model downloading in environments without pip on PATH.

## FAQ

### What is spaCy used for?

It is used for natural language processing in Python, specifically tokenization, part-of-speech tagging, dependency parsing, named entity recognition and text classification. The README describes it as designed to be used in real products, with pretrained pipelines and a training system for custom models.

### Is spaCy better than NLTK?

They solve different problems. spaCy is a pipeline that annotates a single Doc object with components you can train jointly, while NLTK is a collection of algorithms and corpora you assemble yourself. The right choice depends on whether your task is fixed and production-bound or exploratory.

### How to install spaCy in Python?

Install the spacy package from PyPI or conda-forge into a virtual environment, then download a pretrained pipeline separately, since the library ships without models. Wheels are published for the platforms covered by the project's cibuildwheel configuration.

### How to install the spaCy en_core_web_sm model?

Pretrained pipelines are downloaded separately from the library, and the README links to spacy.io/models for the full list. Release 3.8.14 fixed model downloading in environments without pip on PATH.

### How to use spaCy in Python code?

Load a pipeline once and call it on a string to get a Doc. The Doc exposes tokens, entities, tags and dependency relations, so downstream code reads from one annotated object instead of passing structures between separate tools.

## Sources

- [Official documentation](https://spacy.io)
- [Official README](https://github.com/explosion/spaCy#readme)
- [Project repository](https://github.com/explosion/spaCy)
- [Release notes](https://github.com/explosion/spaCy/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/explosion-spacy
