SpanMarker returns spans with character offsets and a score
SpanMarker for Named Entity Recognition
At a glance
- What is it?
- A named entity recognition framework built on Transformers, from the PL-Marker paper, that accepts any common label scheme without declaration and hands back character offsets rather than token indices. Its dependency comments record exactly which library breakages forced each floor, and its setup.py is one spaCy factory registration.
- Who is it for?
- Adopt SpanMarker if you need entity extraction with positions you can use against the original text, since character offsets are what make the output usable for highlighting and linking rather than only for counting. Do not adopt it if you need a bare token classifier with no training framework, because everything here is arranged around Transformers and its training loop.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The prediction is a list of dictionaries
The inference example is short enough to read as an API description.
A model loads from the Hugging Face Hub by name, then a single call takes a string:
from span_marker import SpanMarkerModel
# Download from the 🤗 Hub
model = SpanMarkerModel.from_pretrained("tomaarsen/span-marker-bert-base-fewnerd-fine-super")
# Run inference
entities = model.predict("Amelia Earhart flew her single engine Lockheed Vega 5B across the Atlantic to Paris.")What comes back is a list of dictionaries, one per entity. The visible entries show the fields: a `span` holding the matched text, a `label`, a `score` as a float, and `char_start_index` and `char_end_index`.
The character offsets are the design decision worth noticing. Token-based entity output forces the caller to map token positions back onto the original string, which is a source of off-by-one bugs whenever tokenisation and the source disagree. Returning character positions makes the result directly usable for highlighting, linking and redaction.
The labels in the example are also finer than a coarse entity type. `person-other` and `product-airplane` are the kind of label a fine-grained dataset such as FewNERD produces, and the score is a plain float, so a caller can threshold on it rather than taking whatever the model felt like asserting.
The label scheme is detected, not declared
Most entity extraction libraries require you to tell them which annotation scheme your data uses, and getting that wrong produces training runs that look fine and evaluate badly.
SpanMarker removes the declaration. It works automatically with datasets using the IOB scheme, IOB2, BIOES, BILOU, or no scheme at all.
That list covers most of what exists in practice. IOB2 is the common convention in the CoNLL style corpora, BIOES and BILOU give finer boundary information, and some datasets mark only the first token of an entity with no continuation marker at all. Handling the last case matters because it is the one where a model trained on IOB2 data and evaluated on bare-offset data will lose points at every entity boundary.
The encoders are equally unremarkable by design. It works out of the box with common encoders including `bert-base-cased`, `roberta-large` and `bert-base-multilingual-cased`, and the approach comes from the PL-Marker paper rather than from a new architecture.
The inheritance is stated as the reason the library is small. Being tightly built on Transformers means it inherits model loading and saving, hyperparameter optimisation, logging into various tools, checkpointing, callbacks, mixed precision training and 8-bit inference, rather than reimplementing each of those.
Dependency floors exist because of named upstream bugs
Two lines in the dependency list carry a comment explaining the reason, and they are more useful than the version numbers themselves.
Transformers is required at 4.23.0 or newer and below 5, with the note that it is required for `EvalPrediction.inputs`. That attribute is how the evaluation result carries its inputs back, so a lower floor means the evaluation path cannot work.
Datasets is required at 2.20.0 or newer, with a comment recording an `AttributeError` about a pyarrow attribute that no longer exists. That is the fix for a specific breakage rather than a feature requirement.
The remaining runtime dependencies are torch, accelerate, evaluate, seqeval for sequence labelling evaluation, scikit-learn, jinja2 for the model card templates, packaging and huggingface_hub at 0.24.0 or newer.
The documentation extra shows the other style of pinning. Sphinx is pinned to an exact version, as is lxml, alongside markdown and notebook tooling with upper bounds on nbconvert and pandoc. A docs extra that pins the exact documentation toolchain is normal for a project whose published site has to keep building.
Two more extras exist for integrations rather than function: one for Weights and Biases and one for CodeCarbon, which measures the emissions of a training run. The fact that carbon accounting is a first-class optional extra tells you what kind of user this library has in mind.
setup.py exists to register one spaCy factory
The entire contents of the setup file is a single call, and it is worth seeing because it explains how the library meets the wider NLP ecosystem.
It registers an entry point in the group spaCy uses for its components, pointing at a factory function inside the package's initialisation module. In other words, installing SpanMarker makes a named entity component available inside spaCy, and it is discovered like any other registered factory rather than being imported by hand.
Everything else about the package is declared in the modern project file: the name `span_marker`, the description, the readme, the Python floor of 3.10 or newer, the Apache-2.0 licence, keywords, a single author and maintainer, and the dependency sets above. The version is read dynamically from the package's own `__version__` attribute rather than written in the metadata, which is how a single-source version works with an installed package.
The presence of both files is a transitional pattern rather than a mistake. The modern file does the packaging; the legacy file exists to add an entry point that some tools still read from it.
The test configuration is similarly plain: tests live in one directory, coverage is measured for the package with the ten slowest durations reported, and a tensorboard deprecation warning is filtered out.
Every published model ships the script that made it
The pretrained models section makes a promise that is unusual for a model library, and it is the part to check before you trust a checkpoint.
Every model in the list contains a training script file showing the training that produced it, and all the training scripts used are stored in a directory in the repository.
That is a stronger claim than a model card. A model card describes what a model does; a training script lets you reproduce it, including the arguments, the dataset identifiers and the label mapping. If a published checkpoint cannot be regenerated, you cannot tell whether it was trained on the data it claims.
The Hub integration is the other half of the distribution story. SpanMarker models are integrated with the Hugging Face Hub and the Inference API, and any model on the Hub can be tested through a widget on its model page. Each public model also gets a free API for fast prototyping, which can be moved to production through hosted inference endpoints.
That is a sensible ladder: prototype against a free hosted endpoint, and when the cost or the latency matters, deploy the same model. The documentation for the integration lives on the Hugging Face site rather than only in this repository.
Four notebook platforms, one getting started file
The training story is told through notebooks rather than scripts, and the project goes out of its way to make the same notebook runnable in four places.
A getting started notebook is linked from four hosts at once: Google Colab, Kaggle, Paperspace Gradient and Amazon SageMaker Studio Lab. Each link points at the same file in the repository, so the notebook is versioned with the code and each platform opens it directly from GitHub.
That is a deliberate choice about the audience. The framework's stated appeal is accessibility, and the first person through it is running a notebook without a local GPU. Supporting Colab and its competitors from one file avoids maintaining four copies that drift.
The notebook explains the training snippet in more detail, and the snippet itself is short. A dataset is loaded by identifier, the original tag column is dropped, a fine-grained column is renamed to the standard name, and the label list is read out of the dataset features. The rest of the imports name three things from the package: the model class, a trainer, and a model card data class.
The repository root also carries a thesis PDF, a changelog and a citation file, so the academic basis, the version history and the citation are all in the same place.
Editorial conclusion
Adopt SpanMarker if you need entity extraction with positions you can use against the original text, since character offsets are what make the output usable for highlighting and linking rather than only for counting. Do not adopt it if you need a bare token classifier with no training framework, because everything here is arranged around Transformers and its training loop. Verify two things first. Run the inference example and check the offsets against your own text, since the character index fields are the part that has to be right for downstream highlighting. Then read the dependency comments before upgrading Transformers or datasets, because two version floors exist because of specific upstream breakages.
Frequently asked questions
What is SpanMarker used for?
It is a framework for training named entity recognition models on top of Transformers, using encoders such as BERT, RoBERTa and ELECTRA. The approach comes from the PL-Marker paper, and the package is named span_marker on PyPI.
What does SpanMarker return from predict?
A list of dictionaries, one per detected entity, each holding the matched span text, a label, a float score, and the char_start_index and char_end_index offsets into the original string.
Which entity label schemes does SpanMarker support?
It handles IOB, IOB2, BIOES, BILOU and datasets with no continuation label at all, automatically, without being told which scheme the data uses. Encoders such as bert-base-cased, roberta-large and bert-base-multilingual-cased work out of the box.
How do I install SpanMarker?
Run pip install span_marker. Python 3.10 or newer is required. The install also registers a spaCy factory entry point, so the component becomes available inside spaCy without any further configuration.
Can I reproduce a published SpanMarker model?
The project states that every model in its pretrained list contains a training script showing the training used to generate it, and all training scripts are kept in a training_scripts directory in the repository.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tomaarsen-spanmarkerner)