Open-source project
stanfordnlp/CoreNLP avatar
stanfordnlp/CoreNLP

Stanford CoreNLP: the Java NLP suite you may not be allowed to ship

CoreNLP: A Java suite of core NLP tools for tokenization, sentence segmentation, NER, parsing, coreference, sentiment analysis, etc.

10,121 stars2,719 forksJavaGPL-3.0

At a glance

What is it?
stanfordnlp/CoreNLP is an integrated Java suite covering tokenization, parsing, NER, coreference and sentiment, with varying support for seven languages beyond English. The licence is the full GPL, which the README itself warns rules out use in proprietary software you distribute.
Who is it for?
Adopt Stanford CoreNLP if you need coreference resolution, constituency parsing and Semgrex pattern extraction in one Java pipeline, and if the full GPL is compatible with how you distribute your software, since the README states plainly that it rules out proprietary distribution. Do not adopt it if you need those annotations inside a commercial product you ship, in which case Stanza or spaCy give you a permissive licence and a Python-first workflow.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What the suite covers

CoreNLP is a set of natural language analysis tools written in Java. The README describes the range in one long sentence: base forms of words, parts of speech, whether tokens are names of companies or people, normalised dates times and numeric quantities, sentence structure as syntactic phrases or dependencies, and which noun phrases refer to the same entities.

That last item is the one that keeps people coming back. Coreference resolution, deciding that "the company" and "Acme" and "it" are the same entity, is still not something most NLP libraries hand you, and building it yourself is a research project.

The design is an integrated framework rather than a bag of tools. The README says that starting from plain text you can run all the tools with two lines of code, and that the tools variously use rule-based, probabilistic machine learning and deep learning components. That mix explains the behaviour: the deterministic parts are stable and predictable, and the learned parts carry the accuracy.

The audience is anyone doing linguistic analysis who needs the full stack on one pass, which historically means academic work, and the README says the tools are used by groups in academia, industry and government.

Languages are not equal here

CoreNLP was originally developed for English, and the README is careful about the rest. It now provides, in its words, varying levels of support for Modern Standard Arabic, mainland Chinese, French, German, Hungarian, Italian and Spanish.

"Varying levels" is doing real work in that sentence. English gets the full annotation pipeline. Other languages may have tokenization, part of speech tagging and NER without the same depth of parsing or coreference, and the README does not break down which language gets which annotator.

Anyone planning on a non-English pipeline should check the model jar table in the README, which lists each language's jar and the version it was last updated for, rather than assume parity with English.

The models are distributed separately from the code, which is the next thing to understand before you start.

Building it and getting the models

The README gives three build paths. With Ant, compiling is one command in the project root:

bash
cd CoreNLP ; ant

With Maven, the package target runs the tests and produces the jar:

bash
mvn package

The README's Maven example shows the output named stanford-corenlp-4.5.4.jar, so the version in the documentation trails the current release.

Models are the real setup work. They are not in the repository, and the README says the best way to get them is git-lfs cloning from Hugging Face Hub. The example it gives is French:

bash
git lfs install
git clone https://huggingface.co/stanfordnlp/corenlp-french

For Maven projects the models have to be installed into your local repository, since they are jars that are not published to Maven Central. The README gives a worked command for Spanish, with the note that other languages need the language name changed and that the main models jar needs -Dclassifier=models:

bash
mvn install:install-file -Dfile=/location/of/stanford-spanish-corenlp-models-current.jar -DgroupId=edu.stanford.nlp -DartifactId=stanford-corenlp -Dversion=4.5.4 -Dclassifier=models-spanish -Dpackaging=jar

English additionally has extra and KBP jars for the larger models such as the shift-reduce parser and WikiDict, which the README says are not in the default models jar.

The licence is the constraint

CoreNLP is licensed under the GNU General Public License v2 or later, and the README spells out what that means in practice. It states that this is the full GPL, which allows many free uses but not use in proprietary software that you distribute to others.

That is an unusual thing for a project README to say outright, and it should be read as the deciding factor rather than a footnote. If you build a product that links CoreNLP and ship that product to customers, the GPL obliges you to release your product's source under a compatible licence. Internal use and academic use are fine. Embedding it in a commercial binary is not.

The repository also carries a RESOURCE-LICENSES file and a licenses/ directory, because the models and bundled resources have their own terms separate from the code. The code being GPL does not tell you what the model jars permit, and those are what you actually ship in volume.

This is a description of what the README states, not legal advice, and the right move before adopting it in a company is to have someone who is qualified read both the GPL obligation and RESOURCE-LICENSES.

Recent releases remove things for security reasons

The two most recent releases are both partly subtractive, which is worth knowing before you upgrade.

Version v4.5.10, published on 2025-06-07, removed Lucene and the Patterns project. The release notes explain that older Lucene versions have a security issue, that Lucene 9.12 is not compatible with Java 8, and that the only CoreNLP code using it was the patterns directory. The notes invite anyone using patterns to file an issue so it can return in a future Java 11 compatible release.

Version v4.5.9, published on 2025-04-07, removed the ability to specify an external library for deserialization of annotations in the server, reported as a potential security vulnerability and removed because the protobuf format is complete enough not to need it. It also removed the naturalli demo.

The same release notes add features to Semgrex and Ssurgeon, the graph search and graph editing languages over dependency trees: a uniq operator, an undirected connection operator, and several Ssurgeon edit improvements. Those are the tools people use for pattern-based extraction from parse trees, and they are a genuine strength of this project.

The pattern across both releases is that Java 8 compatibility and security advisories are driving removals. If you depend on removed functionality, you are pinning an older version with a known advisory.

Stanza and spaCy as alternatives

The obvious alternative is Stanza, from the same Stanford NLP group. It is a Python library built on neural networks that covers many of the same annotations, and it is the natural choice if your stack is Python and you want the Stanford approach without a JVM.

The difference is access and licence. CoreNLP is Java, runnable as a server or embedded, and GPL licensed. Stanza is Python, calls into neural models directly, and the experience is a pip install rather than a Maven build plus a model jar hunt.

spaCy is the other option and differs in philosophy. It is built for production pipelines: fast, MIT licensed, with an opinionated pipeline configuration and a strong ecosystem of trained pipelines. It covers tokenization, tagging, parsing and NER well. Where CoreNLP still stands apart is coreference resolution and the constituency parsing plus Semgrex tooling, which is how you do structured extraction from parse trees.

Choose CoreNLP when you need those specific capabilities and GPL is acceptable. Choose Stanza or spaCy when you want a Python-first pipeline and a permissive licence.

Maintenance reality

The last push to the repository was on 2026-09-15. The newest release is v4.5.10 from 2025-06-07, so roughly fifteen months of commits sit outside a tag, and the README's own build instructions still reference version 4.5.4 output. Users of the latest code are expected to build from HEAD.

The repository layout reflects a long history and several build systems at once: pom.xml alongside pom-java-11.xml and pom-java-17.xml, build.gradle, build.xml, plus lib/, liblocal/ and libsrc/ directories. The separate Java 11 and Java 17 POMs tell you which runtime versions are actually supported in practice.

There is an itest/ directory for integration tests, a SECURITY.md, and GitHub Actions running tests. The project is not dormant, it is simply on an academic release cadence, described in the README as several times a year corresponding to a stable commit.

For planning, that means tracking main and building from source is the realistic mode of use, and pinning a released jar means accepting that it will be a year old.

Editorial conclusion

Adopt Stanford CoreNLP if you need coreference resolution, constituency parsing and Semgrex pattern extraction in one Java pipeline, and if the full GPL is compatible with how you distribute your software, since the README states plainly that it rules out proprietary distribution. Do not adopt it if you need those annotations inside a commercial product you ship, in which case Stanza or spaCy give you a permissive licence and a Python-first workflow. If you proceed, plan for model setup as the real work: clone the models with git-lfs from Hugging Face, or install the model jars into your Maven repository with install:install-file, and check the per-language jar table before assuming a non-English pipeline matches English coverage.

Frequently asked questions

Can I use Stanford CoreNLP in commercial software?

The README says the licence is the full GPL, which allows many free uses but not use in proprietary software that you distribute to others. Internal and academic use are fine.

How do I build Stanford CoreNLP?

The README gives an Ant path with cd CoreNLP ; ant and a Maven path with mvn package, which runs the tests and builds the jar.

How do I get the Stanford CoreNLP models?

Models are separate from the code. The README says to use git-lfs and clone from Hugging Face Hub, giving git clone https://huggingface.co/stanfordnlp/corenlp-french as the example, or to download the language jars listed in its table.

Which languages does Stanford CoreNLP support?

It was originally developed for English and the README says it provides varying levels of support for Modern Standard Arabic, mainland Chinese, French, German, Hungarian, Italian and Spanish.

Is Stanford CoreNLP still releasing?

The newest release is v4.5.10 from 2025-06-07 and the repository was last pushed to on 2026-09-15, so recent work is only on main. The README describes releases as arriving several times a year.

Official sources

  1. License: GPL-3.0
  2. Project website
  3. README
  4. Releases
  5. stanfordnlp/CoreNLP on GitHub
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/stanfordnlp-corenlp.svg)](https://hysenlabs.com/projects/stanfordnlp-corenlp)
Community notes

Community notes