# Spark NLP: running transformer pipelines inside Apache Spark

> Spark NLP wraps tokenization, NER, translation and LLM inference as Spark ML stages, so the same pipeline runs on a laptop and on a cluster. Here is how it installs, what the annotate flow returns, and where it stops being the right tool.

**JohnSnowLabs/spark-nlp** — State of the Art Natural Language Processing

- Repository: https://github.com/JohnSnowLabs/spark-nlp
- Website: https://sparknlp.org/
- Stars: 4,158 · Forks: 743
- Language: Scala
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/johnsnowlabs-spark-nlp

## What Spark NLP actually solves

The problem is not "I need to tag entities in a sentence." Dozens of libraries do that. The problem is that the sentences are already sitting in a distributed table, and the NLP library lives in a different process, a different language runtime and a different scaling model. Moving a billion rows out of Spark into a Python service and back is the expensive part, not the inference.

Spark NLP's answer is to make the annotators Spark ML transformers. A DocumentAssembler, a Tokenizer, a WordEmbeddingsModel and a NerDLModel are stages in a Pipeline, the same abstraction Spark uses for feature engineering and classifiers. That means text processing inherits Spark's partitioning, its lazy execution and its scheduling. The README frames the library as providing "simple, performant & accurate NLP annotations for machine learning pipelines that scale easily in a distributed environment," and the annotation types are the mechanism that makes the claim concrete: results are columns, not Python objects.

The audience is therefore narrower than "anyone doing NLP." It is teams with a Spark footprint, a JVM stack, or a Databricks workspace. If you are writing a Flask endpoint that classifies support tickets one at a time, the Spark session is pure overhead.

## How annotations flow through a Spark ML pipeline

Spark NLP does not hand you strings. It hands you annotator types. A DocumentAssembler takes a raw text column and produces a DOCUMENT annotation. Downstream annotators consume that column and emit their own: TOKEN, SENTENCE, POS, LEMMA, EMBEDDINGS, NAMED_ENTITY, and so on. Each is a column in a DataFrame, and each carries metadata such as the character offsets of the original span.

The offsets are the design decision worth noticing. Because every token and entity knows where it came from in the source document, you can map annotations back to the original text without re-tokenizing, and you can chain stages that need alignment, such as a spell checker sitting between a tokenizer and a lemmatizer. An annotator that silently shifts offsets breaks everything after it, which is why the library ships a `checked` annotation in its default pipeline output.

The pretrained path collapses this into one object. `PretrainedPipeline` downloads a named pipeline and its models, wires the stages, and exposes `annotate()`. That is convenient, and it is also the point where you lose visibility: the README's example prints the keys of the result rather than the pipeline definition, so the stages inside `explain_document_dl` are not obvious from the quick start. For production work you will want the explicit annotator chain, because that is where you control memory, batching and which models are loaded.

## Installing Spark NLP and running a first pipeline

The README's quick start pins Java first. Spark NLP expects Java 8 or 11, Oracle or OpenJDK, and the example uses conda for the Python environment. Note the version pairing in the example: `spark-nlp==6.4.2` with `pyspark==3.3.1`. The cheatsheet in the README maps Spark 3.0 through 3.5 to the plain `spark-nlp` package, so the pip install and the PySpark version are separate decisions you have to keep consistent.

```bash
$ java -version
# should be Java 8 or 11 (Oracle or OpenJDK)
$ conda create -n sparknlp python=3.7 -y
$ conda activate sparknlp
# spark-nlp by default is based on pyspark 3.x
$ pip install spark-nlp==6.4.2 pyspark==3.3.1
```

Once installed, the session is started through the library rather than through a plain SparkSession builder. The `start()` function takes `gpu`, `apple_silicon` and `memory` parameters, so accelerator selection and driver memory are set at session creation. The README warns that M1/M2 and AArch64 are under experimental support, which matters if your developers are on Apple laptops and your cluster is x86.

```python
from sparknlp.base import *
from sparknlp.annotator import *
from sparknlp.pretrained import PretrainedPipeline
import sparknlp

spark = sparknlp.start()
pipeline = PretrainedPipeline('explain_document_dl', lang='en')
result = pipeline.annotate(text)
```

The first run downloads the pipeline and its models, so expect network access and a wait. The README's expected output for `result['entities']` on the Mona Lisa sample is `['Mona Lisa', 'Leonardo', 'Louvre', 'Paris']`. If you get an empty list, the usual cause is a version mismatch between the installed `spark-nlp` package and the Spark runtime, not a missing model.

## The model download is the operational cost, not the library

Pretrained pipelines fetch weights at first use. The README advertises 100000+ pretrained pipelines and models across 200+ languages, and that catalogue size is exactly why the download step deserves planning. On a single machine it is a one-time wait. On a cluster where executors start fresh, an unmanaged download path means every executor pulls the same weights, or fails to, depending on network egress and where the cache lives.

The library does not hide this, but the README does not document a cache-location configuration in the quick start either. That is a real gap for anyone deploying to ephemeral containers. The practical workaround is to pre-stage the model directories on shared storage and point the annotators at local paths rather than model names, which means abandoning `PretrainedPipeline` for the explicit annotator chain. That trade is worth naming: the convenient API and the cluster-friendly API are not the same API.

Model importing support widens the surface further. TensorFlow, ONNX, OpenVINO and Llama.cpp GGUF models can all be brought in, and the repository topics list onnx and llamacpp alongside bert and transformers. Each backend brings its own runtime dependency and its own memory profile, so a pipeline that mixes a TensorFlow-based embeddings model with a GGUF generative model has two sets of native libraries to keep compatible across driver and executors.

## Where Spark NLP is the wrong choice

The clearest failure mode is small data. Spark's scheduler, serialization and JVM startup dominate when your corpus is a few thousand documents. A single-process library will finish before `sparknlp.start()` returns.

The second is latency-sensitive serving. Spark NLP is built for batch and streaming jobs over DataFrames. If you need sub-100ms per-request inference behind an HTTP endpoint, the Spark session is a liability, and the model download behaviour is worse: you do not want an executor fetching weights during a request. The library's own positioning is production pipelines at scale, and that is a batch-shaped claim.

The third is when the model you need simply is not in the catalogue. The README lists a long set of architectures and tasks, but a specific fine-tune for your domain is not guaranteed to exist. At that point you are either training your own annotator or importing an ONNX or TensorFlow model, and the import path is more work than the quick start suggests.

Finally, the README's own note that M1/M2 and AArch64 support is experimental is a boundary, not a footnote. If your team develops on Apple Silicon and deploys on x86, you are testing on a configuration the project does not treat as fully supported.

## Spark NLP versus spaCy, and when the Spark dependency pays off

The comparison people search for is Spark NLP against spaCy, and the difference is architectural rather than a matter of accuracy. spaCy is a Python library that runs in your process. You load a model, call it on a document, get a Doc object back. Scaling means more processes, more containers, or a framework like Ray or Dask layered on top. Spark NLP runs inside a Spark job. You define a Pipeline, call fit or transform on a DataFrame, and Spark decides how to partition the work. Scaling means more executors.

That has consequences. With spaCy, the unit of work is a document and the failure mode is a slow loop. With Spark NLP, the unit of work is a partition and the failure modes are serialization errors, executor OOM and skew, where one partition holds a disproportionate share of the text. You get fault tolerance and a scheduler for free, and you pay for it with JVM tuning and a heavier dependency graph.

There is also a language dimension. The README makes a point of exposing the same transformers to Python and R as well as to the JVM ecosystem (Java, Scala, Kotlin). If your organisation has Scala services that need NLP, that is a genuine advantage over a Python-only library, because the annotators are usable from the same runtime as the rest of the service.

## Licence, maintenance and upgrade cost

The library itself is Apache-2.0, and the repository carries a LICENSE file at the top level. That covers the code. It does not automatically cover every pretrained model the library can download, and the README does not enumerate model-level terms in the sections available here. Treat the licence of the code and the licence of the weights as two separate checks before you ship.

Maintenance signals are mixed in a way worth stating plainly. The last push to the default branch was on 2026-09-09, and the most recent release listed is 6.4.2 from 2026-06-24. So the repository is being touched, but the release cadence and the commit cadence are not the same thing, and nothing here tells you what those commits contain.

Upgrade cost is dominated by the version matrix. The cheatsheet ties Spark NLP packages to Apache Spark major versions, and the pip example pins both `spark-nlp` and `pyspark`. Upgrading Spark means re-checking the package row, and upgrading Spark NLP means re-checking that your pretrained pipelines still load. The changelog file exists at the repository root, which is where you would look before moving a production pipeline between minor versions.

## Conclusion

Adopt Spark NLP if your text already lives in Spark, Parquet or a lakehouse and you want NER, classification or translation as a stage in that job rather than a second service. Skip it if your workload is a few thousand documents a day on one machine: the JVM, the Spark session and the model downloads cost more than a single-process library would. Before committing, verify three things on your own cluster: that your Spark version has a matching package row in the cheatsheet, that the specific pretrained pipeline you need exists for your language, and that your licence entitlement covers the models you plan to load, since the Apache-2.0 licence on the library is not the same thing as the terms attached to every model it can pull down.

## FAQ

### What is Spark NLP?

It is a natural language processing library built on top of Apache Spark, providing annotators such as tokenization, part-of-speech tagging, named entity recognition, classification and machine translation as Spark ML pipeline stages. The README describes it as offering pretrained pipelines and models in more than 200 languages.

### Is Spark NLP free?

The repository is licensed under Apache-2.0 and carries a LICENSE file. The README does not state the licence terms attached to the individual pretrained models, so the code licence and the model terms should be checked separately.

### What are the alternatives to Spark NLP?

The architectural alternative is a single-process Python NLP library such as spaCy, where you load a model and call it on a document directly instead of defining a Spark ML pipeline over a DataFrame. The trade is scaling model: processes and containers versus Spark executors.

### Is Spark still written in Scala?

The repository's primary language is Scala, and the README states that Spark NLP exposes its transformers to the JVM ecosystem (Java, Scala and Kotlin) as well as to Python and R by extending Apache Spark natively.

### Is Spark the same as Python?

No. Spark is a distributed processing engine, and Spark NLP is built on top of it. The Python interface is PySpark: the README's quick start installs both `spark-nlp` and `pyspark`, and starts a session with `sparknlp.start()`.

## Sources

- [JohnSnowLabs/spark-nlp on GitHub](https://github.com/JohnSnowLabs/spark-nlp)
- [License: Apache-2.0](https://github.com/JohnSnowLabs/spark-nlp/blob/master/LICENSE)
- [Project website](https://sparknlp.org/)
- [README](https://github.com/JohnSnowLabs/spark-nlp/blob/master/README.md)
- [Releases](https://github.com/JohnSnowLabs/spark-nlp/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/johnsnowlabs-spark-nlp
