GLiNER: zero-shot named entity recognition that runs on a CPU
Generalist and Lightweight Model for Named Entity Recognition (Extract any entity types from texts)
At a glance
- What is it?
- GLiNER is an Apache-2.0 Python framework for training and deploying small NER models that accept entity labels at inference time. It installs with pip, runs on consumer hardware, and adds a Ray Serve layer for production throughput.
- Who is it for?
- Adopt GLiNER when your entity types change often, when you want NER on a CPU or a single consumer GPU, or when you need relation extraction and PII detection from the same model family. Do not adopt it if you need a hosted API with an uptime commitment, or if your labels are stable enough that a small supervised classifier would do the job with less machinery.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 10 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem GLiNER solves: labels that arrive after the model is trained
Classic NER is a closed-set problem. You pick a tag set, annotate thousands of sentences, train a classifier, and redeploy whenever the tag set changes. That works when the tag set is stable: people, organizations, locations, dates. It breaks when the entity types are defined by whoever is reading the text that week.
GLiNER inverts the order. The label list is an argument to the prediction call, not a property of the checkpoint. The README example passes ["person", "date", "organization", "location"] to model.predict_entities and prints matches back as text and label pairs. Nothing is retrained. The entity types are supplied per request.
The intended audience follows from that. Teams doing information extraction over documents whose schema is still moving, PII detection where the categories are set by a compliance team rather than an ML team, and knowledge graph work that needs relations as well as spans. It is a Python library first; the repository also ships a serving layer and a training script, so it can be used as an application dependency or as a fine-tuning base.
How the bi-encoder split makes 100+ entity types practical
The README describes a bi-encoder design: label embeddings are pre-computed, which is what lets the model scale to 100+ entity types, in the README's wording, without degradation. The practical consequence is that adding labels does not mean re-encoding the document for every label. If you have run a cross-encoder NER model, you know the alternative: each label costs a full pass over the text.
The repository layout reflects a broader scope than plain span tagging. The README lists zero-shot NER, streaming NER, joint entity and relation extraction through what it calls the RelEx architecture, and multi-task token classification. There is a streaming_demo.py under examples/, which matches the incremental streaming claim rather than contradicting it.
Two optional dependency groups hint at where the hard parts are. The tokenizers extra pulls in langdetect, python-mecab-ko, janome, jieba3, camel_tools, indic-nlp-library, spacy and stanza, which is the dependency footprint of segmenting Korean, Japanese, Chinese, Arabic and Indic scripts correctly. The onnx and openvino extras point at export paths for non-PyTorch runtimes, and examples/convert_to_onnx.ipynb is the worked example. None of this is free: the base install is small, and the multilingual path is not.
Installing GLiNER and extracting your first entities
The README gives two install routes. The plain pip route is the one most readers want:
pip install glinerThere is also a uv variant, uv pip install gliner, and a serving extra, uv pip install gliner[serve], which the README equates with pip install gliner ray[serve]. The package requires Python 3.10 or newer according to pyproject.toml, and it pins torch>=2.0.0 and transformers>=4.51.3,<5.17.0. Note that requirements.txt carries a different transformers floor, 4.57.3, so the two files are not interchangeable as a source of truth.
The smallest working example is four lines of setup and a call. This is the README's own snippet, with a checkpoint the README names:
from gliner import GLiNER
model = GLiNER.from_pretrained("gliner-community/gliner_small-v2.5")
text = "Cristiano Ronaldo dos Santos Aveiro (born 5 February 1985) is a Portuguese professional footballer."
labels = ["person", "date", "organization", "location"]
entities = model.predict_entities(text, labels, threshold=0.5)
for entity in entities:
print(entity["text"], "=>", entity["label"])The README shows the expected output as lines such as Cristiano Ronaldo dos Santos Aveiro => person and 5 February 1985 => date. The threshold argument is the dial you will actually turn: 0.5 is the README's value, and lower values return more spans with more false positives. Expect the first call to spend time downloading weights from Hugging Face before any text is processed.
If you want the HTTP path instead of the in-process one, the README gives this command, which starts a server on port 8000:
python -m gliner.serve --model gliner-community/gliner_small-v2.5 --dtype fp16The client then connects to http://localhost:8000/gliner, and client.predict accepts a list of texts plus a labels list. The compose.yml in the repository runs the same service in Docker, publishing 8000:8000 and reading GLINER_MODEL, GLINER_DEVICE, GLINER_DTYPE and GLINER_MAX_BATCH_SIZE from the environment. Its healthcheck hits http://localhost:8000/-/healthz. Note that compose.yml sets GLINER_MODEL to urchade/gliner_small-v2.1, not the v2.5 checkpoint used elsewhere in the README.
Quantization and compilation: what the README claims and what it costs
The README makes specific numbers here, and they are the project's numbers, not measured independently. torch.compile is described as yielding up to about 1.5x speedup with no quality loss. FP16 quantization via quantize=True is said to halve model memory, and combined with compilation to give up to about 1.9x faster GPU inference with virtually no quality loss. INT8 is described as cutting memory by another 2x on top of FP16, with the caveat that models need Quantization-Aware Training to preserve accuracy at that precision.
The caveat is the important part. INT8 is not a free switch you flip on a downloaded checkpoint; the README states the model has to have been trained with QAT for accuracy to hold. If you enable INT8 on a checkpoint that was not, you are accepting an accuracy change the README does not quantify.
The configuration is passed to the loader, which is also where device placement happens:
model = GLiNER.from_pretrained(
"gliner-community/gliner_small-v2.5",
map_location="cuda",
quantize=True,
compile_torch_model=True,
)Compilation has a first-call cost, since the graph is built before it is reused. For a long-running service that is amortized; for a one-shot script it is a loss. The README's own framing is that these optimizations matter on edge devices, at high throughput, or when keeping GPU costs low.
Where GLiNER is the wrong tool
Zero-shot is a trade, not a free upgrade. A model that accepts arbitrary labels at inference time has not seen your label definitions, and the label string is the entire specification. "person" and "date" behave well because they are conventional. A label like "termination clause" or "adverse event" carries domain meaning that a short English phrase cannot convey, and there is no mechanism in the README for attaching a definition or an example to a label. If your categories need definitions, you are in fine-tuning territory, and the repository's train.py with a config under configs/ is the documented path.
The version pin is the second constraint. transformers is capped below 5.17.0 in pyproject.toml. If another part of your stack requires a transformers release outside that range, the resolver will fail and you will be choosing which dependency to move. This is a real friction point in shared environments, not a theoretical one.
The third is operational. There is no hosted endpoint here. You run the model, you own the GPU or the CPU time, and you own the failure modes. The serving layer is Ray Serve based, which means adopting Ray's deployment model along with GLiNER's. For a team that just wants a text-in, JSON-out endpoint, that is more infrastructure than the task requires.
Finally, threshold tuning is unavoidable and undocumented in any quantitative way. The README shows 0.5 and nothing about precision and recall trade-offs at other values. You will be measuring that yourself on your own data.
GLiNER against a fine-tuned transformer and against an LLM prompt
The obvious alternative is a fine-tuned encoder such as a BERT or DeBERTa token classifier trained on your tag set. The difference is where the flexibility lives. A fine-tuned classifier encodes your labels in its weights, so inference is a single forward pass and there is no label prompt at all. It will typically beat a zero-shot model on the exact tag set it was trained on. What it cannot do is accept a new tag without retraining, and that is precisely the case GLiNER is built for. If your tag set has not changed in two years and you have annotations, the fine-tuned classifier is the more boring and probably better choice.
The other alternative is prompting a large language model for JSON output. The README positions GLiNER as running on CPUs and consumer hardware with performance competitive with LLMs several times its size, naming ChatGPT and UniNER. Treat that as the project's claim rather than a settled comparison; the README does not present the evaluation setup. What is structurally different is the cost model. A prompted LLM bills per token and per call, and label changes are free but latency is not. GLiNER is a local forward pass with a fixed memory footprint, and label changes are free there too. For high-volume extraction, the local model removes the per-call cost entirely, which is the argument the project makes in its own optimization section. The counter-argument is that a prompted LLM can be given a paragraph explaining what "termination clause" means, and GLiNER cannot.
Maintenance, upgrades and the Apache-2.0 licence
The repository is not archived, and the last push was on 2026-09-23. Releases have been regular: v0.2.27 on 2026-05-11, v0.2.28 on 2026-07-24, v0.2.29 on 2026-09-08. That cadence is visible in the release list, and it is the only maintenance signal available here; the README does not publish a support policy or a deprecation timeline.
Upgrade cost is dominated by two things. First, the transformers upper bound in pyproject.toml means a major transformers release can require a coordinated upgrade of GLiNER itself. Second, the version is dynamic, read from gliner.__version__, so the installed version is whatever the package exposes rather than a static string in the build file. Pin the package version in your own lockfile and read RELEASE.md before moving.
The licence is Apache-2.0, declared both in pyproject.toml and in the LICENSE file at the repository root. For most commercial use that is a permissive licence with an explicit patent grant. Two things to check yourself rather than assume: model checkpoints are downloaded from Hugging Face and carry their own licence terms, which may differ from the code licence, and the optional dependency groups pull in packages under their own terms. This is not legal advice; read the LICENSE file and each checkpoint page.
Editorial conclusion
Adopt GLiNER when your entity types change often, when you want NER on a CPU or a single consumer GPU, or when you need relation extraction and PII detection from the same model family. Do not adopt it if you need a hosted API with an uptime commitment, or if your labels are stable enough that a small supervised classifier would do the job with less machinery. Before committing, verify three things: that a checkpoint on Hugging Face matches your language coverage, that your threshold gives acceptable precision on your own text, and that your torch and transformers versions satisfy the pins in pyproject.toml. The project's own training script is the escape hatch if zero-shot quality falls short, and that path is documented in examples/finetune.ipynb.
Frequently asked questions
What is GLiNER?
GLiNER is a framework for training and deploying small named entity recognition models with zero-shot capabilities, according to the README. It also supports incremental streaming NER, joint entity and relation extraction, and multi-task token classification.
How do I install GLiNER?
The README gives pip install gliner, or uv pip install gliner for the faster uv path. Serving support comes from the serve extra, which the README equates with pip install gliner ray[serve]. The package requires Python 3.10 or newer.
How do I use GLiNER to extract entities?
Load a checkpoint with GLiNER.from_pretrained, then call model.predict_entities with the text, a list of labels and a threshold. Each returned entity is a dict with text and label keys, as shown in the README's quick start.
Is GLiNER an LLM?
No. The README describes GLiNER as a framework for small NER models, fine-tunable and optimized to run on CPUs and consumer hardware, with performance the project describes as competitive with LLMs several times its size.
Is GLiNER open source?
Yes. The repository declares Apache-2.0 in pyproject.toml and ships a LICENSE file at the root. Model checkpoints are downloaded separately from Hugging Face and carry their own terms.
Is GLiNER multilingual?
The README describes state-of-the-art multilingual PII models covering major entity types across 100+ languages, and the tokenizers extra installs language-specific segmenters for Korean, Japanese, Chinese, Arabic and Indic scripts. Coverage depends on the checkpoint you choose.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/urchade-gliner)