# spacy-transformers: running BERT, RoBERTa and XLNet inside a spaCy v3 pipeline

> spacy-transformers is the bridge between Hugging Face transformer weights and spaCy v3 components. It is a training and feature-extraction layer, not a drop-in classifier, and the pip extra you install determines whether it works at all.

**explosion/spacy-transformers** — 🛸 Use pretrained transformers like BERT, XLNet and GPT-2 in spaCy

- Repository: https://github.com/explosion/spacy-transformers
- Website: https://spacy.io/usage/embeddings-transformers
- Stars: 1,409 · Forks: 179
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/explosion-spacy-transformers

## What spacy-transformers actually adds to a spaCy pipeline

The package supplies spaCy components and architectures that run transformer models through Hugging Face's transformers library. The README frames the payoff as access to architectures such as BERT, GPT-2 and XLNet from inside spaCy, with automatic alignment of transformer output to spaCy's tokenization. That alignment is the part that matters. Transformer tokenizers split text differently from spaCy, and the package handles the mapping rather than leaving it to you.

The audience is narrow and specific. You need to be building on spaCy v3, because the README states this release requires spaCy v3 and points anyone on the older version at the v0.6.x branch. If you are on spaCy v2, this package in its current form is not for you. The other stated use case is multi-task learning: backpropagating to one transformer model from several pipeline components. That is a training-time concern, not an inference convenience, and it tells you the project is aimed at people training their own pipelines rather than consuming someone else's.

## The transformer component is a feature source, not a task head

This is the limitation most people hit after installing. The README is explicit that the transformer component does not support task-specific heads like token or text classification. You cannot point it at a fine-tuned sentiment model and get labels back. What you get is features that a spaCy component such as ner or textcat can be trained on top of.

The README offers a workaround for the prediction case: if you only want the predictions from an existing Hugging Face text or token classification model, use the wrappers from spacy-huggingface-pipelines to bring that model into a spaCy pipeline. Those are two different jobs, and the README keeps them separate. Reading the feature list as if it covered both is the most common misreading of this project.

Two other design points are worth stating plainly. The transformer data saved in the Doc object is customizable, and how long documents are processed is customizable. Long-document handling is a real constraint with transformer models, so exposing that as a knob rather than a fixed window is a deliberate choice. Serialization and model packaging are listed as working out of the box, which matters because a spaCy pipeline that cannot be saved and reloaded is not deployable.

## Installing spacy-transformers with pip and the CUDA extra

The README gives one install command. Installing from pip pulls in the dependencies, including PyTorch and spaCy, and the README adds an ordering instruction: install this package before you install the models.

```bash
pip install 'spacy[transformers]'
```

For GPU use, the README says to find your CUDA version with nvcc --version and add the version in brackets, giving spacy[transformers,cuda92] for CUDA9.2 and spacy[transformers,cuda100] for CUDA10.0 as its examples. The README also notes that if PyTorch installation itself is the problem, you should follow the instructions on the official PyTorch site for your operating system.

The stated requirements are Python 3.6+, PyTorch v1.5+ and spaCy v3.0+. The repository's requirements.txt is tighter in places: spacy>=3.5.0,<4.1.0, transformers[sentencepiece]>=3.4.0,<4.53.3, torch>=1.8.0, numpy>=1.15.0, srsly>=2.4.0,<3.0.0 and spacy-alignments>=0.7.2,<1.0.0. Trust requirements.txt over the README prose when you are pinning a build, since it is the file the package actually ships.

A first real use is a spaCy v3 training config that names the transformer component and points at a Hugging Face model. The repository ships example configs under examples/configs/, and the spaCy documentation linked from the README, Embeddings, Transformers and Transfer Learning, is where the full config shape lives. Read the shipped example before writing your own; the config system is where most of the setup effort goes, not the install.

## Version pinning is the main maintenance cost

The dependency ranges tell the story. transformers is capped below 4.53.3 and spaCy below 4.1.0, and the most recent release listed, v1.4.0 from 2026-03-17, is described only as an update to the transformers pin. That is a maintenance pattern, not a criticism: the package tracks two fast-moving upstreams, so releases are often pin adjustments rather than new capability. The two releases before it, v1.3.8 and v1.3.9, deal with wheel publishing.

When you plan an upgrade, expect to move spaCy and transformers together. If your environment needs a transformers version above the cap, this package is not the tool for that environment until the pin moves. The build itself compiles a Cython extension, spacy_transformers.align, which is why the project publishes wheels across manylinux and musllinux images and why a source install needs a working compiler and numpy headers. The last push to the repository was on 2026-03-27.

The licence is MIT. That is permissive and imposes no copyleft obligation on your own code, but it says nothing about the licences of the pretrained transformer weights you download through Hugging Face. Those carry their own terms, and the README does not address them.

## Where spacy-transformers is the wrong tool

If your goal is to run an existing fine-tuned classification model and read its labels, this package does not do that, and the README says so directly and points you elsewhere. If you are on spaCy v2, the current release does not apply to you.

There is a third case that the README implies rather than states. If your task does not need transformer features at all, adding this package buys you a PyTorch dependency, a compiled alignment extension and a version-pinning problem in exchange for nothing. spaCy's own pipeline components exist for exactly that situation, and the spaCy documentation distinguishes between them. The install command pulls in PyTorch whether or not your pipeline uses a GPU, so the cost is paid up front.

A fourth boundary is documentation drift. The README warns that the package was extensively refactored for spaCy v3.0 and that previous versions worked considerably differently, directing readers to older tagged versions of the README. Anything you find describing the v2-era API is describing a different library.

## Alternatives and how they differ in approach

The README itself names spacy-huggingface-pipelines as the route for using the predictions of an existing Hugging Face text or token classification model. The difference is structural. spacy-transformers gives you a transformer as a feature source inside a spaCy pipeline you train, with alignment to spaCy's tokenization and the option to backprop from several components into one transformer. spacy-huggingface-pipelines wraps an already-trained task model so its predictions appear in your spaCy documents. One is a training architecture, the other is an inference adapter.

Going the other way, you can skip spaCy and use transformers directly. That removes the alignment layer and the spaCy config system, and it also removes spaCy's Doc container, serialization and model packaging. The README lists serialization and model packaging as out-of-the-box features here, so that is what you give up. The choice comes down to whether you want spaCy's pipeline structure around the transformer or not.

## Conclusion

Adopt spacy-transformers if you already build spaCy v3 pipelines and want transformer features feeding your own ner or textcat components, with multi-task backprop into one shared transformer. Do not adopt it if you want to call an existing Hugging Face text or token classification model for its predictions; the README points that use case at spacy-huggingface-pipelines instead. Before committing, verify that your installed spaCy and transformers versions fall inside requirements.txt (spacy>=3.5.0,<4.1.0 and transformers[sentencepiece]>=3.4.0,<4.53.3), because the v1.4.0 release is described only as an update to the transformers pin.

## FAQ

### Which is better, NLTK or spaCy?

The README does not compare the two. It describes spacy-transformers as providing spaCy components and architectures that use transformer models via Hugging Face's transformers library, and states that this release requires spaCy v3.

### Is spaCy still relevant?

The repository's most recent release listed is v1.4.0 from 2026-03-17, described as an update to the transformers pin, and the last push to the repository was on 2026-03-27. The README documents the package against spaCy v3.

### What is spaCy used for?

The README describes spaCy as the pipeline this package plugs into, with components such as ner and textcat that can be trained on transformer features. It also links to spaCy's documentation on embeddings, transformers and transfer learning.

### What are the Transformers in NLP?

The README names BERT, RoBERTa, XLNet and GPT-2 as the pretrained transformer architectures this package can use, and notes that spaCy aligns transformer output to its own tokenization.

## Sources

- [explosion/spacy-transformers on GitHub](https://github.com/explosion/spacy-transformers)
- [License: MIT](https://github.com/explosion/spacy-transformers/blob/master/LICENSE)
- [Project website](https://spacy.io/usage/embeddings-transformers)
- [README](https://github.com/explosion/spacy-transformers/blob/master/README.md)
- [Releases](https://github.com/explosion/spacy-transformers/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/explosion-spacy-transformers
