BERTopic: BERT embeddings plus c-TF-IDF for interpretable topics
Leveraging BERT and c-TF-IDF to create easily interpretable topics.
At a glance
- What is it?
- BERTopic clusters documents with sentence embeddings and HDBSCAN, then labels each cluster with c-TF-IDF. It is a Python library for teams that need readable topic labels, not a fixed taxonomy.
- Who is it for?
- Adopt BERTopic when you have a corpus of a few thousand or more documents, no fixed label set, and a need to read the topics yourself. Do not adopt it when you need a deterministic classifier, when your corpus is a few hundred rows, or when you cannot afford the embedding step.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What BERTopic does that keyword counting cannot
Classic topic models such as LDA treat a document as a bag of words and infer a distribution over a fixed number of topics. The output is a list of weighted terms per topic, and the reader has to guess what the topic is about. BERTopic takes a different route: it embeds documents with a transformer model, reduces the embedding space, clusters the result, and then describes each cluster by comparing term frequencies inside the cluster against term frequencies in the corpus as a whole. That comparison is c-TF-IDF, a class-based variant of TF-IDF where the class is the topic.
The library is aimed at people who have text and want to see what is in it. The README frames the goal as creating dense clusters that allow for easily interpretable topics while keeping important words in the topic descriptions. The README also lists a wide set of modes beyond the default: guided, supervised, semi-supervised, manual, hierarchical, class-based, dynamic, online or incremental, multimodal, multi-aspect, zero-shot, merge models and seed words. That breadth is the real selling point. You are not choosing one algorithm; you are choosing from a family of configurations around the same embedding and clustering core.
It is not a classifier. Nothing in the repository describes a stable label set that survives new data. If your problem is routing support tickets into five known queues, a supervised model is the right tool. BERTopic is for the case where you do not yet know the queues.
The pipeline: embeddings, UMAP, HDBSCAN, c-TF-IDF
The default path has four stages, and the dependency list in pyproject.toml maps onto them. sentence-transformers produces the document embeddings. umap-learn reduces dimensionality, which matters because HDBSCAN degrades in high-dimensional space. hdbscan does the clustering. Then c-TF-IDF generates the topic representations.
HDBSCAN is density-based and does not require you to specify the number of topics in advance. The trade-off is that it assigns low-density documents to a noise class, conventionally topic -1. On messy real-world text, that noise share can be large, and the documentation does not promise a particular bound. You should measure it on your own corpus rather than assume it will be small.
The modularity is deliberate. Because each stage is a separate component, you can swap the embedding backend, replace UMAP with another dimensionality reduction step, or use a different clustering algorithm. The README points to a lightweight installation path for training with Model2Vec or for inference without transformers, UMAP or HDBSCAN at all. That is the escape hatch when you only need to apply an already-fitted model.
Topic representations are also pluggable. The README links to a text generation page covering LLM-based representation, which replaces the c-TF-IDF keyword list with generated labels. That is a meaningful difference in output: keyword lists are reproducible and cheap, generated labels read better but depend on a model you have to host or call.
Installing BERTopic and fitting a first model
The README gives two install commands. With uv, the Python package manager, the command is uv add bertopic. With pip it is pip install bertopic. Both bring in sentence-transformers by default, which means the first run will download an embedding model.
pip install bertopicIf you want a different embedding backend, the README lists extras: flair, gensim, spacy and use. There is also a vision extra for topic modeling with images.
pip install bertopic[flair,gensim,spacy,use]Once installed, the standard entry point is the BERTopic class. The README's quick start begins by importing BERTopic and a loader for the 20 newsgroups dataset, then fitting on the documents. The README excerpt cuts off before the fitting call, so check the documentation site for the exact method names rather than copying a pattern from a tutorial.
The README links a Start Here notebook on best practices and several Colab notebooks covering custom embedding models, advanced customization and (semi-)supervised modeling. Those are the places to look for a complete, runnable first example. The dependency list includes plotly, so interactive visualization is part of the default install, and a datamap extra adds datamapplot and matplotlib for a different chart style.
Where BERTopic is the wrong tool
The first failure mode is corpus size. Density-based clustering needs enough points to find density. On a few hundred documents you will get either one large topic or a large noise class, and neither is useful. The library will run; the output just will not carry information.
The second is determinism. Embedding models, UMAP and HDBSCAN all involve stochastic or approximate steps. Two runs on the same data can produce different topic numbering and slightly different boundaries. If you need a pipeline that produces identical output on every run and can be diffed across releases, this is a poor fit. The README does not document a reproducibility guarantee.
The third is the cost profile. Every document passes through a transformer. The README's large-data notebook mentions GPU acceleration, which tells you the default CPU path is not intended for very large corpora. If you have millions of short documents and no GPU, the embedding step will dominate your runtime budget.
The fourth is drift. Topics are numbered, not named. Topic 7 in January is not guaranteed to mean the same thing in June. The dynamic and online modes address change over time, but they do not turn the output into a stable schema. If downstream systems key off topic IDs, you own the mapping problem.
BERTopic compared with LDA and with an LLM labeling pass
Against LDA, the difference is where the semantics come from. LDA infers topics from co-occurrence counts of words. BERTopic infers them from positions in an embedding space produced by a transformer trained on large text corpora. That means BERTopic can group documents that share meaning but not vocabulary, which LDA cannot do without heavy preprocessing. The cost is a model download and a much slower fit. LDA runs on a laptop CPU in seconds on modest corpora; BERTopic does not.
Against using a large language model directly, the difference is supervision and cost. You can prompt a model to label each document, but that is a per-document call and the labels will not be consistent across a corpus unless you constrain them. BERTopic does the grouping first and only optionally uses an LLM to name the groups. The README's text generation page describes that second pattern. If your corpus is small and your labels are few, prompting may be simpler. If it is large and open-ended, clustering first and labeling second is cheaper and gives you a structure you can inspect.
One more comparison worth naming: the README's zero-shot mode. That lets you steer topics toward descriptions you supply rather than discovering them purely from data. It sits between unsupervised discovery and supervised classification, and it is the mode to consider if you have a rough idea of the categories but not labeled training data.
Maintenance, licence and upgrade cost
The repository is not archived, and the last push was on 2026-09-09. The most recent release listed is v0.17.4 from 2025-12-03, following v0.17.3 and v0.17.1 in July 2025. The version in pyproject.toml is 0.17.4, consistent with the release tag.
The project is MIT licensed. That is permissive: you can use it commercially, modify it and redistribute it, provided the copyright notice and licence text are preserved. The pyproject.toml declares the licence as a file reference to LICENSE. Note that MIT covers BERTopic itself, not the embedding models you download. Those carry their own licences, and sentence-transformers will pull a model from a hub on first use. Check the model's licence separately before shipping. This is not legal advice.
The upgrade cost is dominated by the dependency chain rather than by BERTopic's own API. The declared dependencies include hdbscan, umap-learn, numpy, pandas, plotly, scikit-learn, sentence-transformers, tqdm and llvmlite, with a comment noting that llvmlite is pinned above 0.36.0 to prevent conflicts with uv. That is a sign the maintainer has already hit resolver problems in this stack. Python support starts at 3.10, and the classifiers list 3.10 through 3.14. If you are on 3.9, you cannot install the current release.
The Makefile shows the project's own workflow: pytest for tests, ruff for formatting and linting, mkdocs serve for docs. There is a test-no-plotly target that uninstalls plotly and runs a subset of tests, which suggests plotly is optional at runtime in some paths even though it is a declared dependency.
Editorial conclusion
Adopt BERTopic when you have a corpus of a few thousand or more documents, no fixed label set, and a need to read the topics yourself. Do not adopt it when you need a deterministic classifier, when your corpus is a few hundred rows, or when you cannot afford the embedding step. Before committing, verify the sentence-transformers model download works on your network, check that HDBSCAN and UMAP install cleanly on your platform, and confirm the HDBSCAN noise share on a sample of your own data is acceptable.
Frequently asked questions
What is BERTopic?
BERTopic is a Python topic modeling technique that combines transformer embeddings with c-TF-IDF. It clusters documents and then describes each cluster with terms that are distinctive to that cluster rather than just frequent in it.
How do I install BERTopic?
The README gives pip install bertopic or uv add bertopic, both of which include sentence-transformers. Extras such as flair, gensim, spacy and use are available if you want a different embedding backend.
What is BERTopic used for?
It is used to discover and describe topics in a text corpus without specifying the number of topics in advance. The README lists guided, supervised, dynamic, online, multimodal and zero-shot variants for different analysis needs.
Is BERTopic better than LDA?
They differ in where semantics come from. LDA infers topics from word co-occurrence counts, while BERTopic clusters transformer embeddings, so it can group documents that share meaning without sharing vocabulary. BERTopic is slower and requires a model download.
Is BERTopic supervised or unsupervised?
The default mode is unsupervised: HDBSCAN finds clusters without a preset topic count. The README also documents supervised and semi-supervised modes, which use labels to guide the topics.
What does BERTopic stand for?
The name combines BERT, the transformer architecture family it builds on, with topic modeling. The README describes it as a topic modeling technique that leverages transformers and c-TF-IDF.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/maartengr-bertopic)