KeyBERT: keyword extraction with BERT embeddings and cosine similarity
Minimal keyword extraction with BERT
At a glance
- What is it?
- KeyBERT is a Python package that extracts keywords and keyphrases by comparing BERT embeddings of n-grams against the embedding of the whole document. It installs with pip and runs in a few lines, but the quality of the output follows the quality of the embedding model you point it at.
- Who is it for?
- Adopt KeyBERT when you want a pip-installable keyword extractor that needs no training and fits in a few lines of Python, and when you can accept cosine similarity over sentence-transformer embeddings as the ranking rule. Do not adopt it if you need a trained sequence labeler, a fixed taxonomy of entities, or deterministic output across model versions.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 37 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap KeyBERT fills between TF-IDF and a trained keyphrase model
Statistical extractors such as Rake, YAKE! and TF-IDF rank terms by frequency and co-occurrence. They need no model weights and no GPU, but they cannot tell that two documents use different words for the same subject. Trained keyphrase models handle that, at the cost of a training pipeline and labeled data. The README states the author's own motivation plainly: he could not find a BERT-based solution that did not have to be trained from scratch and could be used by beginners, so the goal became "a `pip install keybert` and at most 3 lines of code in usage."
That is the audience. Developers who need keywords for search indexing, tagging, clustering or topic labeling, and who want semantic matching without owning a training set. The project does not claim novelty in the technique. The README says KeyBERT "is by no means unique" and points to other BERT-based keyphrase extractors, including a paper on the approach. It is a packaging decision as much as an algorithmic one: sentence-transformers for embeddings, scikit-learn for the similarity math, a small API on top.
How the embedding and cosine similarity pipeline actually runs
The mechanism is short enough to describe in full. First, the document is embedded with a sentence-transformer model to produce one document-level vector. Then candidate words and phrases are generated as n-grams from the document, and each candidate is embedded with the same model. Finally, cosine similarity between each candidate vector and the document vector ranks the candidates, and the highest-scoring ones are returned as keywords.
The README describes exactly this sequence: "First, document embeddings are extracted with BERT to get a document-level representation. Then, word embeddings are extracted for N-gram words/phrases. Finally, we use cosine similarity to find the words/phrases that are the most similar to the document."
Two details matter in practice. The candidate set is bounded by the n-gram range you pass, so a range of (1, 1) can only ever return single words and (1, 2) can return two-word phrases. And the default model is a sentence-transformer, not a raw BERT checkpoint, which is why the same `KeyBERT()` call works on a laptop CPU. The README advises `"all-MiniLM-L6-v2"` for English documents and `"paraphrase-multilingual-MiniLM-L12-v2"` for multilingual or non-English text, and links to the sentence-transformers pretrained model list for the full set of options.
Installing KeyBERT and extracting your first keyphrases
The base install pulls numpy, rich, scikit-learn and sentence-transformers, which means PyTorch comes along with it. Python 3.10 or newer is required according to `pyproject.toml`.
pip install keybertThe README also lists optional extras for other backends: `keybert[flair]`, `keybert[gensim]`, `keybert[spacy]` and `keybert[use]`. If you want to avoid PyTorch entirely, the README gives a light-weight path using Model2Vec:
pip install keybert --no-deps scikit-learn model2vecA first run needs a document and one call. The README's minimal example constructs the model with no arguments and extracts keywords from a block of text about supervised learning.
from keybert import KeyBERT
doc = "Supervised learning is the machine learning task of learning a function that maps an input to an output based on example input-output pairs."
kw_model = KeyBERT()
keywords = kw_model.extract_keywords(doc)To get phrases rather than single tokens, set `keyphrase_ngram_range`. With `(1, 2)` and `stop_words=None`, the README shows results such as `('learning algorithm', 0.6978)` and `('machine learning', 0.6305)`. Each tuple carries the phrase and its cosine similarity to the document. Setting `highlight=True` returns the keywords with the matching spans marked in the text.
Max Sum Distance and MMR: the two diversification options, and their cost
Similarity ranking has an obvious failure mode: the top results are near-duplicates. If the document says "supervised learning" ten times, the top five keywords may all be variants of that phrase. KeyBERT exposes two diversification strategies.
Max Sum Distance takes the 2 x top_n candidates most similar to the document, evaluates all possible `top_n`-combinations among them, and picks the combination whose members are least similar to each other. The README is direct about the cost: "This has super-exponential time complexity and therefore not advised if you use a large `top_n`." That is a real constraint, not a footnote. A `top_n` of 5 with `nr_candidates=20` is fine; scaling `top_n` upward is not.
Maximal Marginal Relevance, the second option, also works from cosine similarity but selects candidates by balancing relevance to the document against redundancy with already-selected keywords. MMR is the better default when you want diversity at a larger `top_n`, because its cost grows far more gently than the combinatorial search. The trade-off is that MMR is a greedy selection, so it does not explore the full combination space the way Max Sum Distance does. Neither method fixes a bad candidate set: if your n-gram range excludes the phrase you wanted, no diversification strategy will surface it.
Where KeyBERT returns weak or misleading keywords
The method is unsupervised in the sense that no labeled keywords are used, but it is entirely dependent on the embedding model. Change the model and the ranking changes. The README does not document a reproducibility guarantee across model versions, so a pipeline that pins nothing can produce different keywords after a dependency upgrade. Pin the model name explicitly.
Second, the candidate generation is n-gram based over the raw document. That means the extractor sees whatever tokens the tokenizer produces, including artifacts from your preprocessing. If you strip stop words, you lose phrases whose meaning depends on them; if you keep them, single-word results fill with "the" and "of" unless `stop_words` is set. The README's own examples pass `stop_words=None` to show raw behaviour, which is a demonstration choice, not a production default.
Third, Max Sum Distance is unusable at scale for the reason quoted above. And fourth, the project is not an entity recognizer. It returns phrases that resemble the document, not phrases that belong to a known type. If you need people, organizations or dates with types attached, this is the wrong tool. The README's own framing supports that reading: it presents KeyBERT as "a very basic, but powerful method" rather than a complete information extraction system.
KeyBERT against YAKE! and against training your own extractor
The README names Rake and YAKE! as prior art. The difference is mechanical. YAKE! scores terms from statistical features of the text itself: term frequency, position, casing, co-occurrence with other terms. It runs in milliseconds, needs no model download, and behaves identically on every machine. KeyBERT instead asks a neural encoder whether a phrase means something close to the whole document, which is why it can surface a phrase that appears once while a frequency-based method ignores it.
The cost of that is the model download, the PyTorch dependency, and CPU or GPU time per document. On a large batch of short texts, a statistical extractor will often be the better engineering choice, and KeyBERT's advantage shrinks as documents get shorter, because a document embedding of a two-sentence text carries less signal for the similarity comparison to work with.
The other alternative is to train a keyphrase model on labeled data. That buys you domain-specific accuracy and a fixed output vocabulary, and it costs you annotation, training infrastructure and retraining when the domain drifts. KeyBERT's position is the middle: semantic matching without training, and without a guarantee that the output vocabulary stays stable across model swaps.
Licence, release cadence and what an upgrade costs you
KeyBERT is MIT licensed, with the licence file at the repository root and the classifier `License :: OSI Approved :: MIT License` in `pyproject.toml`. MIT permits commercial use and modification with attribution and the licence text retained. That covers KeyBERT's own code. It does not cover the embedding model you load, and the README points to the sentence-transformers pretrained model list without discussing model licences. Check the licence of whichever checkpoint you select; this is a separate question from the package licence and the README is silent on it.
The version in `pyproject.toml` is 0.9.0, matching the v0.9.0 release. Release spacing has been uneven: v0.7.0 in November 2022, v0.8.0 in September 2023, v0.9.0 in February 2025. The last push to the default branch was on 2026-08-25. There is no documented deprecation policy and no changelog in the repository listing, so an upgrade means reading the release notes and the diff yourself. The dependency floor on sentence-transformers (`>=0.3.8`) is loose, which means a fresh install can pull a much newer version than the one the release was built against. Pinning that dependency is the cheapest insurance against a keyword ranking that shifts without any change to your code.
The repository also carries a Makefile with `test`, `install`, `install-test`, `pypi`, `clean` and `check` targets, and a `tests/` directory, so running the suite locally after an upgrade is a one-command check.
Editorial conclusion
Adopt KeyBERT when you want a pip-installable keyword extractor that needs no training and fits in a few lines of Python, and when you can accept cosine similarity over sentence-transformer embeddings as the ranking rule. Do not adopt it if you need a trained sequence labeler, a fixed taxonomy of entities, or deterministic output across model versions. Before committing, verify two things on your own corpus: which embedding model you will pin, and whether the n-gram range you choose produces phrases that survive stop-word filtering.
Frequently asked questions
How do I install KeyBERT?
Run `pip install keybert`. The README also lists optional extras for other backends (`keybert[flair]`, `keybert[gensim]`, `keybert[spacy]`, `keybert[use]`) and a light-weight path without PyTorch using `pip install keybert --no-deps scikit-learn model2vec`.
What is KeyBERT?
It is a Python package that extracts keywords and keyphrases by embedding the document and its n-gram candidates with a sentence-transformer model, then ranking candidates by cosine similarity to the document. The README describes it as a minimal method that needs no training.
How can I extract keywords from a text with KeyBERT?
Construct a model with `KeyBERT()` and call `extract_keywords(doc)`. Set `keyphrase_ngram_range` to control phrase length, for example `(1, 2)` for two-word phrases, and pass `stop_words` to filter common words from single-token results.
What are the alternatives to KeyBERT?
The README names Rake, YAKE! and TF-IDF, which score terms from statistical features of the text rather than from neural embeddings. It also links several other BERT-based keyphrase extraction projects, noting that KeyBERT is not unique in its approach.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/maartengr-keybert)