Open-source project
JasonKessler/scattertext avatar
JasonKessler/scattertext

Scattertext has three version numbers, a Python 2.7 note, and numpy pinned below 2.0

Beautiful visualizations of how language differs among document types.

2,344 stars284 forksPythonApache-2.0

At a glance

What is it?
A Python package for finding terms that distinguish one part of a corpus from another and plotting them in an interactive HTML chart, published as an ACL system demonstration in 2017. The tool's ideas are good and still worth reading; the packaging is a nine-year-old artifact with dependency caps from 2023 and a test suite wired to a framework that no longer runs.
Who is it for?
Use this if you want to see how association statistics differ from one another on the same corpus, because that is the thing this package does better than anything else in its niche: it holds the plotting constant and swaps the score, so fifteen statistics can be compared in the same frame. Read the code rather than installing it as a dependency.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 93 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Three version numbers for one package

The repository has exactly one release, and it is called 0.0.2.4.4, dated 2017-03-13. A four-segment version number is not a typo you can resolve by rounding; it is a scheme somebody used once. The packaging file disagrees with it, declaring version 0.2.3, and the README's own heading agrees with the packaging file rather than the tag. So there are three strings in play and no statement about which is authoritative. The commit record does not settle it either: the last push is 2026-07-04, which is three months ago, so the tree is being touched while the only published artefact is from 2017. None of this stops the package from working, and a reader who installs it will get whatever the index holds. It does mean that a bug report citing a version is citing a number that three places in the repository disagree about, which is worth knowing before you file one.

The page tells you to install Python 3.11 and then that it may work on 2.7

The installation section opens by asking for Python 3.11 or higher, which matches the packaging file's declared floor, and then gives the install as one command:

bash
pip install scattertext

Four lines later, in the same section, it says the package should mostly work with Python 2.7, but that it may not. Both sentences are still on the page, and the second one is a fossil: a Python 2 statement that would need a floor declaration of 2.7 to be true, while the manifest declares a floor of 3.11 and a bare floor with no upper bound. Nothing tells you which one the code reflects. There is a related caution that has aged better, that the HTML output looks best in Chrome and Safari, which is a narrow statement about a browser family rather than a version, and one that has quietly stopped being true for anyone who only tests in Firefox.

The dependency set is frozen at a 2023 numeric stack

Look at the pins rather than the names. The numeric libraries are capped: a floor of 1.2.6 and a ceiling below 2.0 for the array library, a floor of 1.7.0 and a ceiling below 1.14 for the scientific one, then an uncapped floor of 1.4 for the machine learning library and 2.0 for the data frames, with a floor of 0.14.1 for the statistics package. Those two ceilings are the whole story. They are the reason an install on a current environment resolves to the previous generation of two core numeric libraries, and they are why a colleague on a different machine may get a resolution error that says nothing about Scattertext. The capping was presumably deliberate to keep a 2017 tool working, and the result is a package that can only be installed into an environment assembled around it rather than one it can join.

spaCy is required and the page also documents running without it

The parser library is a hard requirement in the dependency list with a floor of 3.2. The installation section then says that if you cannot, or do not want to, install it, you can substitute a whitespace tokeniser for the lines that load an English model. Those two statements cannot both be operationally true for a fresh install, since the package will not install without the parser. What the substitution actually changes is runtime behaviour rather than installation: the page names the consequences precisely, saying the fallback is not compatible with the word similarity explorer, and that tokenisation and sentence boundary detection become low-performance regular expressions. That is the more important half of the sentence for anyone using it as a library, because every association score in the package is computed over those token boundaries.

The page recommends four packages that are commented out of the manifest

The installation section recommends installing a Chinese segmenter, the parser, a topic and category lexicon package, an astronomy library, the flash-text pattern matcher, a topic modelling library and a dimensionality reduction library. Of those seven, the parser, the pattern matcher and the topic modelling library are genuine dependencies. The other four are not, and the reason is visible in the manifest, where they sit commented out next to a short list of things once considered and dropped: a text ranking package, a second Japanese segmenter, the lexicon package, the dimensionality reduction library, a plotting library, a statistics plotting library, a notebook library and an annotator. So the page tells you to install four packages the project chose not to depend on, and the manifest keeps the archaeology. Each optional library corresponds to a demo script at the root, which is the pattern that runs through the whole repository.

The package installs a console script the page never documents

The packaging file declares a console entry point named after the package, pointing at a command-line module in the library. Nothing in the README's table of contents mentions a command line, and the installation section only covers the import path. So every install puts an executable on your path that the documentation does not describe, which is the sort of gap that either means the interface is unfinished or that the documentation predates it, and the page offers no way to tell. The same file also wires the test suite to a collector from a framework that has been unmaintained for years, and lists that framework in the test requirements. There is no testing section in the table of contents either. The one continuous integration configuration named anywhere is a Travis file, matching the single badge at the top of the page.

About a hundred demo scripts sit at the repository root

The root is a wall of numbered examples, one per feature, named for the technique they demonstrate rather than for a task: scripts for two different effect sizes, a third non-parametric effect size, a separation score, a log odds ratio with a prior, log relative risk, a zeta statistic, a divergence measure, correlations, a f-score variant, a matrix factorisation baseline, two term weighting schemes, category frequencies, dispersion, dense rank in two variants, emoji, a Chinese example, a Japanese example, and a pair of scripts about plotting films. Each corresponds to a heading in the README's own table of contents, so the documentation and the file list are the same list twice. That is a defensible way to write a manual for a plotting tool, where the reader wants to run one script and see one shape. It also means there is no example directory, and the table of contents has grown to around thirty-five entries.

The axes are ranks, and the label placement is the actual invention

Two details make the chart readable and both are easy to miss. First, the axes are not frequencies. In the worked example the axes are the dense ranks of each term's usage within each of two categories, so a term near the top is simply one of the most frequent terms in that category, and the plot's message comes from the pairing rather than from either axis alone. Second, the package's own description of what it does is about labels: points are selectively labelled so that labels do not overlap other labels or points. That is the contribution over a scatter plot you could draw yourself, and it is why the selective-labeling step has its own section in the table of contents. Everything else in the library is a swappable score feeding the same picture, which is why there are fifteen scoring sections. The worked example on the front page is cut off mid-argument in its final parameter, so read it as a shape rather than as code you can paste.

Editorial conclusion

Use this if you want to see how association statistics differ from one another on the same corpus, because that is the thing this package does better than anything else in its niche: it holds the plotting constant and swaps the score, so fifteen statistics can be compared in the same frame. Read the code rather than installing it as a dependency. The install path is a Python 2.7-era setup file with a numpy cap below the current major, a test suite wired to a collector that no longer runs on modern interpreters, and three different version strings between the tag, the manifest and the page. If you do install it, the two decisions that matter are the Python version, which the page and the manifest disagree about, and the spaCy model, because the documented fallback changes the tokenisation quality for every score you then compute. The chart itself is browser-local and needs no server, which is the part that has aged best.

Frequently asked questions

What is Scattertext?

It is an Apache-2.0 licensed Python package for finding terms and phrases that are more characteristic of one category in a corpus than another, and for displaying them as an interactive HTML scatter plot. It was published as an ACL System Demonstrations paper in 2017 with a title about a browser-based tool for visualising how corpora differ, and the page asks that the paper be cited when the tool is used.

How do I install Scattertext?

Install Python 3.11 or higher, then install the package from the index. The page also recommends installing a Chinese segmenter, the parser, a topic and category lexicon package, an astronomy library, a pattern matcher, a topic modelling library and a dimensionality reduction library to get the most out of it, four of which are commented out of the dependency list rather than required. It adds that the HTML output looks best in Chrome and Safari.

What do the axes in a Scattertext plot mean?

Each axis is the dense rank of a term's usage within one of the two categories, not its raw frequency, so a point near the top of one axis is among that category's most frequent terms. The plot's reading comes from the pairing of the two ranks. Labels are then placed selectively so they do not overlap other labels or points, which is the step the package treats as its own contribution.

Can Scattertext run without spaCy?

The page says to substitute a whitespace tokeniser for the lines that load an English model if you cannot install the parser, and warns that this fallback is not compatible with the word similarity explorer and reduces tokenisation and sentence boundary detection to low-performance regular expressions. The parser is nonetheless a hard requirement in the dependency list, so the substitution is about runtime behaviour rather than about installing the package.

Which term scores does Scattertext support?

The table of contents names a long list, one section each: two standard effect sizes and a non-parametric third, a bi-normal separation score, a log odds ratio with a prior, log relative risk, a divergence measure, a zeta statistic, correlations, a scaled f-score, two term weighting schemes, dense rank in two variants, dispersion, expected versus actual frequencies, and several coordinate and background-frequency options.

Is Scattertext still maintained?

The repository has a single release, tagged 0.0.2.4.4 and dated 2017-03-13, while the packaging file and the README heading both say 0.2.3, so three version strings disagree. The last commit is dated 2026-07-04. The test suite is wired to a long-unmaintained collector, the only continuous integration configuration named is a Travis file, and the numerical dependencies carry upper bounds from 2023.

Official sources

  1. Issues
  2. JasonKessler/scattertext on GitHub
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jasonkessler-scattertext.svg)](https://hysenlabs.com/projects/jasonkessler-scattertext)