# sense2vec: contextually-keyed word vectors for Python and spaCy

> sense2vec is an MIT-licensed Python library from Explosion for loading, querying and training word vectors that are keyed by part-of-speech tag and entity label rather than by surface string alone. It is a small, stable package, and the interesting part is not the API but the decision to key vectors on sense.

**explosion/sense2vec** — 🦆 Contextually-keyed word vectors

- Repository: https://github.com/explosion/sense2vec
- Website: https://explosion.ai/blog/sense2vec-reloaded
- Stars: 1,680 · Forks: 236
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/explosion-sense2vec

## The problem sense2vec solves: one word, several senses

Plain word2vec gives you one vector per token string. That means the string "duck" has a single position in the vector space, whether it means the bird, the verb for avoiding something, or the fabric. The README points at the original Trask et al. 2015 paper and describes the library as a twist on word2vec that produces more detailed word vectors. The twist is in the key: a sense2vec entry is not "natural_language_processing" but "natural_language_processing|NOUN", and not "duck" but "duck|VERB" or "duck|NOUN".

The audience follows from that. This is for people building search, recommendation, tagging or annotation tooling over text where the part of speech changes the meaning, and who want to query nearest neighbours for a phrase rather than for a single token. It is not a general text embedding service. The README positions it as a library for loading, querying and training models, with the most concrete payoff being multi-word phrase vectors keyed by POS tag and entity label.

## How the sense2vec mechanism works: spaCy parses, the key carries the sense

The README shows two entry points. Standalone, you construct a Sense2Vec object and read it from disk, then index into it with a string key that includes the tag. The README example uses the query "natural_language_processing|NOUN" and asserts that it is present in the object before retrieving its vector.

The pipeline path is where the design becomes clear. You load a spaCy model, add a pipe named "sense2vec", and call from_disk on the component with the vector directory. From then on, spaCy's own tokenizer and tagger decide what the key is. In the README example, a document containing "A sentence about natural language processing." exposes the span doc[3:6] whose text is "natural language processing", and that span carries extension attributes: _.s2v_freq for frequency, _.s2v_vec for the vector, and _.s2v_most_similar(3) for neighbours. The neighbours come back as tuples of the phrase and its tag, for example (('machine learning', 'NOUN'), 0.8986967).

So the data flow is: raw text goes through the spaCy pipeline, the sense2vec component maps the resulting tokens or spans to keys derived from their tags, and those keys resolve into vectors loaded from disk. The README also notes optional caching of nearest neighbours for faster most-similar queries, which is a deliberate memory-for-latency trade.

## Installing sense2vec and running a first query

The package ships on pip. The README gives a single install command, and requirements.txt pins the runtime dependencies to spacy>=3.0.0,<4.0.0, wasabi>=0.8.1,<1.2.0, srsly>=2.4.0,<3.0.0, catalogue>=2.0.1,<2.1.0 and numpy>=1.15.0. Note the upper bound on spaCy: this is a v3-era package, and the README explicitly tells spaCy v2 users to install sense2vec==1.0.3 and check out the v1.x branch instead.

```bash
pip install sense2vec
```

Vectors are separate. The README says they are attached to the GitHub release and that large files are split into multi-part downloads. The 2015 archive is a single 573 MB file; the 2019 archive is 4 GB across three parts. Multi-part archives are merged with cat before extraction, exactly as the README shows:

```bash
cat s2v_reddit_2019_lg.tar.gz.* > s2v_reddit_2019_lg.tar.gz
```

Once the archive is extracted, point from_disk at the extracted directory. The README's standalone example is the shortest path to a real answer:

```python
from sense2vec import Sense2Vec

s2v = Sense2Vec().from_disk("/path/to/s2v_reddit_2015_md")
query = "natural_language_processing|NOUN"
assert query in s2v
vector = s2v[query]
freq = s2v.get_freq(query)
most_similar = s2v.most_similar(query, n=3)
```

The README shows the expected output of that last call as machine_learning|NOUN at 0.8986967, computer_vision|NOUN at 0.8636297 and deep_learning|NOUN at 0.8573361. If your run returns neighbours in that neighbourhood, the vectors loaded correctly. If the assert fails, the key is wrong, most likely because the tag does not match the one in the vector file.

There is also a bundled Streamlit script for browsing vectors interactively. The README says to install streamlit and pass one or more paths to pretrained vectors as positional arguments:

```bash
pip install streamlit
streamlit run https://raw.githubusercontent.com/explosion/sense2vec/master/scripts/streamlit_sense2vec.py /path/to/vectors
```

## The spaCy component and where it stops being the right tool

Inside a pipeline, the component exposes the same information through extension attributes rather than through direct dictionary access. You add it by name and load vectors into it:

```python
import spacy

nlp = spacy.load("en_core_web_sm")
s2v = nlp.add_pipe("sense2vec")
s2v.from_disk("/path/to/s2v_reddit_2015_md")

doc = nlp("A sentence about natural language processing.")
freq = doc[3:6]._.s2v_freq
vector = doc[3:6]._.s2v_vec
```

The practical limitation is that the quality of every lookup depends on the tagger that produced the key. The README's own example uses en_core_web_sm, a small model. If your tagger disagrees with the tagger used to build the vectors, the key will not be in the table and the lookup fails rather than degrading gracefully. That is a hard failure mode, not a soft one, and the README does not document any fallback for a missing key.

The second limitation is size. The 2019 vectors are 4 GB compressed, split across three downloads, and the README does not state the extracted size. Whatever it is, it has to be resident or memory-mapped next to your process. For a container that already carries a spaCy model, that is a real deployment constraint. The README also does not document rollback or versioning of vector archives, so if you pin one, you are pinning a release asset URL.

Third, this is not a fine-tuning framework. The README describes training your own vectors from raw text using a pretrained spaCy model plus GloVe or fastText, and it points to a separate section for details. If your goal is to adapt an embedding model to a downstream classification task, sense2vec is the wrong layer; it gives you vectors and similarity queries, not a training loop for a task head.

## sense2vec compared with plain word2vec and with transformer embeddings

The honest comparison is with the thing sense2vec is built on. Plain word2vec, whether via gensim or via fastText, gives one vector per token. sense2vec gives one vector per token-plus-tag, which is why the README's query strings carry a pipe and a POS label. The cost of that extra resolution is that every lookup now requires a tagger, and you cannot look up a key you have not parsed.

Against transformer contextual embeddings the difference is architectural rather than a matter of quality. A transformer produces a vector per token per context, computed at inference time, so the vector for "duck" in a sentence is different from the vector for "duck" in another sentence. sense2vec produces a static vector per sense, looked up from a table. That makes sense2vec cheap at query time and fully serializable, which the README lists as a feature, and it makes it blind to context beyond the tag. Two sentences where "duck" is a verb will get the same vector even if one is about avoiding a question and the other about lowering your head.

The topics list on the repository places sense2vec alongside gensim, gensim-word2vec and spacy, which reflects that lineage. If you want per-sentence contextual vectors, this is not that tool. If you want fast, shippable, phrase-level similarity over a fixed vocabulary, the static table is the point.

## Maintenance status, licensing and what an upgrade actually costs

The repository is not archived, and the last push was on 2026-03-27. The most recent tagged release is v2.0.2 from 2023-04-17, preceded by v2.0.1 in December 2022 and v2.0.0 in February 2021. So the release cadence is slow and the code has been quiet at the tag level for years, even though the default branch has seen activity more recently than the last release. If you need a library with a frequent release train, this is not it, and the version numbers alone tell you that.

The licence is MIT, which is permissive and places few obligations on how you redistribute the code. That covers the library. The pretrained vector archives are a separate question: they are hosted as release assets, and the README does not state a licence for the vector data itself. If you intend to ship the vectors inside a product, check the terms attached to the release assets rather than assuming the MIT licence on the code extends to them. This is not legal advice; it is a pointer to the gap in the documentation.

Upgrade cost is dominated by the spaCy pin. requirements.txt caps spaCy below 4.0.0, so a spaCy 4 migration is not something this package currently supports. Within v3, the README notes that the v1.x branch exists for spaCy v2 users, which means a v2-to-v3 move is a package-version change, not a flag. The vector archives themselves are versioned by release asset, and the README does not describe a migration path between vector versions.

## Conclusion

Adopt sense2vec if you need vectors for multi-word phrases and you want to distinguish the same word used as different parts of speech, and you are already inside the spaCy v3 ecosystem. Do not adopt it if you want a general-purpose embedding model you can fine-tune on your own task, or if you cannot host a 4 GB vector archive next to your application. Before committing, verify three things: that the pretrained archive you need is still attached to the v1.0.0 release, that the spaCy version in your environment falls inside the spacy>=3.0.0,<4.0.0 range in requirements.txt, and that disk and memory can absorb the 573 MB or 4 GB of vectors once extracted.

## FAQ

### What is sense2vec used for?

It is used to load, query and train word vectors that are keyed by part-of-speech tag and entity label, so a phrase like natural_language_processing|NOUN can be looked up and compared with its nearest neighbours. The README also describes using it as a spaCy pipeline component that exposes frequency, vector and most-similar lookups on spans.

### How do I install sense2vec in Python?

The README gives a single pip command, pip install sense2vec, and requirements.txt pins spaCy to the range spacy>=3.0.0,<4.0.0. Vectors are downloaded separately from the GitHub release and passed to from_disk.

### Where do I download the sense2vec vectors?

The README states that the vector files are attached to the GitHub release, and that large files are split into multi-part downloads. The 2015 archive is a single 573 MB file and the 2019 archive is 4 GB across three parts, merged with cat before extraction.

### Is there a sense2vec demo I can try without installing anything?

The README links to an interactive demo at demos.explosion.ai/sense2vec for exploring semantic similarities across Reddit comments from 2015 and 2019. The repository also includes a Streamlit script for browsing vectors locally, which is run with streamlit run and one or more paths to pretrained vectors.

## Sources

- [explosion/sense2vec on GitHub](https://github.com/explosion/sense2vec)
- [License: MIT](https://github.com/explosion/sense2vec/blob/master/LICENSE)
- [Project website](https://explosion.ai/blog/sense2vec-reloaded)
- [README](https://github.com/explosion/sense2vec/blob/master/README.md)
- [Releases](https://github.com/explosion/sense2vec/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/explosion-sense2vec
