Open-source project
fastino-ai/GLiNER2 avatar
fastino-ai/GLiNER2

GLiNER2: one schema, entity extraction and classification on CPU

Unified Schema-Based Information Extraction

2,065 stars184 forksPythonApache-2.0

At a glance

What is it?
A schema-conditioned encoder family where the label list is an input rather than a training artefact, split across a span architecture and a new boundary architecture, both behind one loader.
Who is it for?
GLiNER2 is worth evaluating when your entity labels change faster than you want to retrain a model, since passing a label list at inference is the whole premise, and when your data cannot leave the machine. It is the wrong tool when you need a guarantee that no two schemas in one process can collide, or when you want GPU-scale throughput on a large batch, since the boundary architecture still has to fit the encoded window.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 18 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The label list is an argument, not a checkpoint

Most named entity recognition models have the labels baked in. You train on `PERSON` and `ORG`, the checkpoint belongs to those labels, and adding a third means training again. GLiNER2 inverts that: it is a schema-conditioned encoder, and the schema is passed in at inference. The README's summary is that entities, classification, structured records, relations and span attributes all come out of a single forward pass against one schema.

That covers five task families that normally mean five different systems: Named Entity Recognition, Text Classification, Structured Data Extraction, Relation Extraction and span attributes. The claim is not that one call does everything, it is that the same conditioning mechanism covers all of them, and that a `Schema` object describes what you want the way a JSON schema would.

The project is Apache 2.0 licensed, Python, requires 3.10 or newer, and the last push was on 2026-09-18. `pyproject.toml` names Urchade Zaratiana as maintainer and pulls the version from `gliner2.__version__`, so the package is the source of truth rather than a hardcoded string. The base dependency set is small and telling: pydantic, requests, tqdm and urllib3. No torch, no transformers, no numpy.

The repository tree explains where the project's attention goes: `gliner2/` for the package, `docs/` for the architecture guides, `bench/` and `benchmarks/` plus a `benchmark_statistical.py`, a `tutorial/` directory, `tools/`, `scripts/`, `tests/`, an `image/` folder and a `RELEASE.md` at the root.

Four pip profiles, and what the base install deliberately omits

Installation is split so that a data engineer validating schemas never pulls a deep learning stack onto their laptop:

bash
pip install gliner2
pip install gliner2[local]
pip install gliner2[train]
pip install gliner2[test]

The base install still gives you `Schema`, `SchemaInput`, `RegexValidator`, `GLiNER2API` (aliased as `API`), `InputExample`, `TrainingDataset` and the JSONL validation tooling. That is enough to build schemas, validate a dataset and talk to the hosted API. The `[local]` extra adds numpy, peft, safetensors, torch and transformers.

This split is recent engineering, not the original design. Release v1.3.1 on 2026-05-06 made torch optional and migrated the LoRA internals to the PEFT backend, and v1.3.2 a month later fixed a missing `local` extra in `pyproject.toml`, added an LRU cache for the tokenizer, and removed the old `gliner` dependency. A dependency being optional and an extra being missing in the same release cycle is normal churn; the direction of travel is clearly away from a heavy import.

The base package is also a client for a remote API that reads `PIONEER_API_KEY`, and it does something mildly clever with long documents:

python
client = API()  # reads PIONEER_API_KEY
results = client.batch_extract_entities(
    documents,
    ["company", "person"],
    batch_size=8,
)
long_result = client.extract_entities_long(
    annual_report,
    ["company", "person"],
    chunk_size=384,
    chunk_overlap=64,
    include_spans=True,
)

Batching and chunking with overlap happen on the client side, so scanning a document that exceeds the server's window does not require a different protocol or a different account tier.

Span grid versus sparse boundary pairing

Two extraction architectures sit behind one public API, and picking the wrong loader is the first mistake to avoid.

The `span` architecture, exposed as `GLiNER2` and `SpanExtractor`, uses a fixed-width span grid. It is what legacy checkpoints and the specialty fine-tunes (GLiGuard for safety, PII for personal data) use. The `boundary` architecture, exposed as `BoundaryExtractor` and shipped as GLiNER2.5, uses sparse start and end pairing, so a span can be any length that fits inside the encoded window. That is a structural difference, not a tuning difference: with a fixed grid, spans longer than the grid width are not representable at all.

`AutoExtractor.from_pretrained(...)` dispatches on the `architecture` field saved in the checkpoint, so it loads span, boundary, GLiGuard and PII checkpoints correctly. `GLiNER2.from_pretrained(...)` remains span-only and the README states in bold that it will not load GLiNER2.5 boundary checkpoints. `GLiNER2` is also just an alias for `SpanExtractor`.

The basic call looks like this:

python
from gliner2 import AutoExtractor  # requires gliner2[local]

# Default English GLiNER2.5 boundary checkpoint
model = AutoExtractor.from_pretrained("fastino/gliner2.5-base-v1")

text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday."
result = model.extract_entities(text, ["company", "person", "product", "location"])

One consequence of schema conditioning is worth stating before you get attached to it. Because labels arrive at runtime, the model has no way to guarantee your `company` label and a different pipeline's `organization` label mean the same thing. Consistency is your schema's job. The README's answer to more complex output rules is constrained decoding rather than prompting: `Classifier` for cross-task label rules and `JointIE` for typed entity and relation graphs, both introduced around version 2.0.0.

Word splitters decide quality on Chinese and other unspaced text

The detail that most affects real accuracy is documented carefully, which is to say that GLiNER2 first splits text into word tokens, then encodes those tokens with the model's subword tokenizer. The default `"whitespace"` splitter is what the public checkpoints were trained with.

For a language without whitespace-delimited words, the README directs you to the character-level splitter:

python
model = AutoExtractor.from_pretrained(
    "fastino/gliner2.5-base-v1",
    word_splitter="char",
)

# Or after loading
model.set_word_splitter("char")

The built-in table has two entries. `"whitespace"` maps to `WhitespaceTokenSplitter`, for space-delimited languages, and matches the public checkpoints. `"char"` maps to `CharLevelSplitter`, for languages such as Chinese, and the README describes its behaviour precisely: it keeps Latin words and emails intact while splitting other non-space characters. You can also pass any custom callable yielding `(token, start, end)` with exclusive-end offsets into the original text, which is what you would do for Thai, Khmer or Japanese with a morphological segmenter.

The warning attached to it is the part to take seriously: changing a pretrained model's word boundaries can affect quality unless the model was trained with the same splitter. And the setting is runtime only, because saved checkpoints reload with `"whitespace"` unless you pass `word_splitter` again. So a splitter change that improves your results disappears silently on reload unless you handle it in your own loading code.

Inference speed options are set the same way, with no extra dependencies:

python
# fp16
model = AutoExtractor.from_pretrained("fastino/gliner2.5-base-v1", map_location="cuda", quantize=True)

# torch.compile (fused GPU kernels, first call triggers tracing)
model = AutoExtractor.from_pretrained("fastino/gliner2.5-base-v1", map_location="cuda", compile=True)

Both require a CUDA map location, so the CPU-first claim in the README applies to the default path rather than to the accelerated one.

What version 2.0.0 actually changed

Version 2.0.0 was published on 2026-08-24 and its release notes are a list of pull request titles, which makes the scope unusually legible. The boundary architecture arrived as two consecutive pulls. Joint entity relation decoding arrived as `add joint ie decoding`, constrained classification decoding as `Cls constr`, and entity attributes as two more pulls. Chunking arrived three times over, once as code, once as a tutorial, and multi-GPU training support was contributed separately by @Yuvrajxms09.

Several entries in that same release say `some change` or `update engine`, which tells you the internal engine was moving while the public features landed. Read that as a caution: 2.0.0 is a large refactor with new capabilities, and examples written against 1.3.x may not match the new decoding paths.

The two headline capabilities are worth separating, because they solve different problems. Span attributes let you attach a classification to an extracted span, which is how you get from `this person` to `this person is a customer of tier two`. JointIE produces a typed graph of entities and relations in one structure, which is a different output contract from a dict of lists and needs a different consumer downstream.

Fine-tuning is pointed at a hosted product rather than a local script: the README says to fine-tune via Fastino at fastino.ai. The local `[train]` extra and the `tutorial/` directory exist, and LoRA is supported through the PEFT backend after v1.3.1, so local training is possible; the docs for boundary models, including record and relation decoding, save and load, LoRA aliases, export mode, loss and imbalance controls and the gold-capacity policy, live in `docs/boundary_architecture.md` and `docs/gliner2_5_boundary_architecture.md` rather than in the README.

Where the CPU-first claim holds and where it does not

The README's argument is that inference runs locally on standard hardware with no GPU required, and that processing is entirely local with zero external dependencies. The first half is credible from the package design: the model is an encoder loaded through transformers, the label set is small, and there is no generation step. The second half needs reading carefully, because the same README documents a hosted API client that reads `PIONEER_API_KEY`. Local inference is local; the API client is a separate option, and its existence is why the base install has no torch.

Against the alternatives, the comparison is clearest against a fine-tuned encoder per label set. A model trained on your twenty entity types will usually beat schema conditioning on those twenty types, because it has specialised. GLiNER2 wins when the label set is unstable, when several pipelines share one deployment, or when the cost of a retraining cycle exceeds the accuracy difference.

The other comparison is a hosted extraction API, and there the trade is privacy and latency against model quality on hard documents. GLiNER2 keeps the text on your machine, which matters for anything under a data processing agreement. The README does not publish accuracy tables or latency numbers against competitors, and the `bench/`, `benchmarks/` and `benchmark_statistical.py` in the tree suggest the work exists without being summarised in the README, so treat any performance expectation as something you should establish on your own text.

The real constraint on the boundary architecture is structural: any span must fit inside the encoded window. Long documents are handled by chunking, which version 2.0.0 added, and the boundary spans that cross a chunk boundary are the cases where extraction quality degrades. `docs/boundary_architecture.md` is where the project documents its own capacity policy, and it is the file to read before you run an annual report through this.

Editorial conclusion

GLiNER2 is worth evaluating when your entity labels change faster than you want to retrain a model, since passing a label list at inference is the whole premise, and when your data cannot leave the machine. It is the wrong tool when you need a guarantee that no two schemas in one process can collide, or when you want GPU-scale throughput on a large batch, since the boundary architecture still has to fit the encoded window. Version 2.0.0 landed on 2026-08-24 with the boundary architecture, joint entity relation decoding and multi-GPU training, so read the release notes before following a 1.x tutorial. Install the base package first, load `fastino/gliner2.5-base-v1` with `AutoExtractor`, and switch to `word_splitter="char"` before you judge quality on any language without whitespace words.

Frequently asked questions

What is the difference between GLiNER and GLiNER 2?

GLiNER2 is the schema-conditioned encoder family in this repository, where the label list is supplied at inference time and covers entities, classification, records, relations and span attributes. Version 1.3.2 of the package also removed the older gliner dependency, so GLiNER2 no longer builds on it.

How do I install GLiNER2 for local model inference?

Use the local extra: `pip install gliner2[local]`, which adds numpy, peft, safetensors, torch and transformers. The plain `pip install gliner2` gives you schema validation, the API client and the JSONL tooling with no torch at all. Python 3.10 or newer is required.

Which GLiNER2 checkpoint should I load?

`AutoExtractor.from_pretrained("fastino/gliner2.5-base-v1")` loads the default English boundary checkpoint and dispatches on the saved architecture field. Legacy span checkpoints and GLiGuard or PII fine-tunes also load through `AutoExtractor`, while `GLiNER2.from_pretrained` stays span-only and will not load a 2.5 checkpoint.

How do I run GLiNER2 on Chinese text?

Pass the character-level word splitter, either at load time with `word_splitter="char"` or afterwards with `model.set_word_splitter("char")`. The README warns that changing word boundaries can affect quality unless the model was trained with that splitter, and the setting is runtime only, so saved checkpoints reload as whitespace.

Official sources

  1. fastino-ai/GLiNER2 on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/fastino-ai-gliner2.svg)](https://hysenlabs.com/projects/fastino-ai-gliner2)