Library / SDK
stanfordnlp/stanza avatar
stanfordnlp/stanza

Stanza: the Stanford NLP Python library for 60+ languages

Stanford NLP Python library for tokenization, sentence segmentation, NER, and parsing of many human languages

7,878 stars959 forksPythonNOASSERTION

At a glance

What is it?
Stanza wraps tokenization, sentence segmentation, NER and dependency parsing for 60+ languages in a Python API, with an optional client for Java CoreNLP. It is a good fit when you need linguistic structure, and the wrong tool when you need speed or a pure-Python dependency tree.
Who is it for?
Adopt Stanza if your work needs dependency parses, lemmas or NER across many languages and you can afford a PyTorch model download per language. Do not adopt it if you need a pure-Python dependency tree, or if you cannot ship model weights with your application.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 19 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What Stanza solves, and who ends up using it

Most NLP libraries make you choose between coverage and structure. A tokenizer handles one language well, a parser handles another, and the glue between them is yours to write. Stanza is the Stanford NLP Group's official Python library, and its pitch is that one Pipeline object gives you tokenization, sentence segmentation, named entity recognition and parsing for more than 60 languages, with the same call shape in each. The README describes it as containing "support for running various accurate natural language processing tools on 60+ languages and for accessing the Java Stanford CoreNLP software from Python".

The people who reach for it are usually doing linguistic work rather than text classification. If you need a dependency tree to extract subject-verb-object relations, or lemmas to normalize a corpus before indexing, Stanza returns those as structured objects rather than as a flat token list. The repository also ships biomedical and clinical English model packages, described in the README as covering "syntactic analysis and named entity recognition (NER) from biomedical literature text and clinical notes", which points at a second audience: researchers working with clinical text who need entity types that general-purpose models do not produce.

It is a weaker fit when the task is fuzzy. Stanza does not do sentiment, topic modeling or embeddings. It gives you grammar and entities, and you build on top.

The Pipeline object and what happens on the first call

The mechanism is a chain of neural processors, each loaded from downloaded model weights and run in sequence over a Document. You construct a Pipeline for a language code, and the pipeline decides which processors to load. Calling the pipeline on a string returns a Document whose sentences and tokens carry annotations: token boundaries, sentence boundaries, lemmas, part-of-speech tags, morphological features, dependency heads and relations, and named entities where a model exists.

The data flow is document-in, document-out. There is no separate tokenizer object to thread through; the Document holds the whole annotation state, and the processors mutate it in order. That design is why the README example can be three lines long. It is also why the ordering matters: sentence segmentation runs before parsing, and the parser expects token boundaries to already exist.

Stanza also includes a CoreNLP client. The README states the library supports "accessing the Java Stanford CoreNLP software from Python", and the demo directory contains a CoreNLP interface notebook and a corenlp.py script. That path exists because some CoreNLP functionality, including Semgrex and Ssurgeon for searching and rewriting dependency graphs, is not reimplemented in the Python pipeline. If you need graph pattern matching, you are choosing the client path, not the neural pipeline.

Installing Stanza and running a first pipeline

Stanza supports Python 3.6 or later according to the README, and the recommended install is pip. The package pulls in PyTorch 1.3.0 or above as a dependency, so the install is not small.

bash
pip install stanza

If a previous version is already present, the README gives the upgrade form explicitly:

bash
pip install stanza -U

There is also a conda channel, with one documented caveat: the README notes that installing via Anaconda does not work for Python 3.10, and says to use pip on that version.

bash
conda install -c stanfordnlp stanza

Once installed, the README's getting-started sequence downloads English models, builds a pipeline, and annotates two sentences. The download step is marked optional in the README because the pipeline can fetch models on demand.

python
import stanza
stanza.download('en')
nlp = stanza.Pipeline('en')
doc = nlp("Barack Obama was born in Hawaii. He was elected president in 2008.")

What you should see is a Document object, not a string. The two sentences are separated, and each token carries annotations. Printing the document gives a readable token-and-annotation dump; iterating doc.sentences and then sentence.words is how you get at the fields programmatically. The first run is slow because the models are being fetched and loaded, and the models are cached locally after that.

For development work, the README offers a source install that keeps the checkout editable:

bash
git clone https://github.com/stanfordnlp/stanza.git
cd stanza
pip install -e .

Where Stanza gets in the way: models, memory and scope

The first real constraint is that models are downloaded per language, not bundled. A fresh install has no weights, and the first pipeline call for a new language reaches the network. Recent release notes point at movement in this area: v1.13.0 is titled "Use huggingface_hub for downloads, conparser efficiency gains", so the download path has changed across versions. If you are running in an offline environment, plan for pre-fetching models during image build rather than at request time.

The second constraint is PyTorch. Stanza is a neural library, and the pipeline runs models on CPU or GPU. That is fine for batch annotation and painful for per-request use inside a latency-sensitive service. There is no documented lightweight mode that skips the neural stack while keeping the Python API.

The third is coverage asymmetry. The claim is 60+ languages, but processors are not uniformly available across all of them. Named entity recognition in particular is not guaranteed for every language code, and the README does not publish a per-language processor matrix. You find out by attempting to load the pipeline and seeing which processors come up, or by reading the model download page on the project site.

A fourth point worth stating plainly: the package metadata in setup.py classifies the project as "Development Status :: 4 - Beta". That is the project's own label, not a third-party assessment, and it sits oddly next to a library that research groups cite in published work. Treat it as a signal that interfaces can move between minor versions, which the release history supports.

Stanza compared with spaCy

The obvious alternative is spaCy, and the difference is not one of quality but of design centre. spaCy is built as a production NLP framework: a single pipeline object, fast tokenization, a model ecosystem you install as packages, and an emphasis on throughput. Stanza is built as an interface to Stanford's neural models, with the CoreNLP bridge as a first-class feature.

The practical differences follow from that. spaCy models are distributed as installable packages, so pinning a model version is a dependency-management problem you already know how to solve. Stanza downloads weights at runtime through its own mechanism, which the v1.13.0 notes show moving to huggingface_hub. If your deployment pipeline needs reproducible artifacts, that distinction matters more than any accuracy comparison.

On functionality, the split is cleaner. If you need Semgrex pattern matching over dependency graphs, Stanza's CoreNLP client exposes it and spaCy does not have an equivalent in the same form. If you need a small, fast pipeline with no Java in the picture and no per-language download step, spaCy is the more direct answer. Stanza's own demo directory, with its Semgrex and Ssurgeon scripts, is a fair summary of where its distinctive value sits.

Maintenance, licensing and the cost of upgrading

The repository is not archived. The last push was on 2026-09-08, which is recent, and the release cadence is steady: v1.12.2 on 2026-06-03, v1.13.0 on 2026-06-18, v1.14.0 on 2026-07-15. The README states that maintenance of the repo is currently led by John Bauer. Two of those three releases carry security framing in their titles, "weights_only security fix" and "Security fixes and Lemmatizer efficiency updates", which is worth noting if you are pinned to an older version: the upgrade path has had security content in it.

The licence picture needs care. The setup.py file declares license='Apache License 2.0', while the repository's licence field is reported as NOASSERTION. That mismatch is exactly the kind of thing to resolve before shipping, because it is the model weights, not just the code, that you redistribute if you bundle them. The README asks that you cite the ACL2020 system demo paper if you use the library in research, and cites a separate paper for the biomedical and clinical models. Citation is not a licence obligation, but it tells you the models were released with academic use in mind. This is not legal advice; read the LICENSE file and the model pages yourself.

Upgrade cost is mostly model re-download and interface drift. Because the Beta classifier is the project's own, minor versions can change behaviour, and the v1.13.0 download change is a concrete example of infrastructure moving under the API.

Editorial conclusion

Adopt Stanza if your work needs dependency parses, lemmas or NER across many languages and you can afford a PyTorch model download per language. Do not adopt it if you need a pure-Python dependency tree, or if you cannot ship model weights with your application. Before committing, verify the licence of the models you intend to use, and check whether your target language has a package on the model download page.

Frequently asked questions

How do I use Stanza?

Import stanza, call stanza.download for your language code, then build a pipeline with stanza.Pipeline and call it on a string. The README's example downloads English, builds an English pipeline, and annotates two sentences about Barack Obama.

How do I use Stanza in a sentence?

The README's getting-started example passes a string to the pipeline: doc = nlp("Barack Obama was born in Hawaii. He was elected president in 2008."). The call returns a Document whose sentences and tokens carry the annotations, rather than a plain string.

What is Stanza?

Stanza is the Stanford NLP Group's official Python NLP library, described in the README as containing support for running natural language processing tools on 60+ languages and for accessing the Java Stanford CoreNLP software from Python.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. stanfordnlp/stanza on GitHub
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/stanfordnlp-stanza.svg)](https://hysenlabs.com/projects/stanfordnlp-stanza)
Community notes

Community notes