# presidio-research: evaluating PII detection models before you trust them

> The presidio-evaluator package provides synthetic PII data generation, a standardized sample format, and precision/recall scoring for Presidio recognizers. It is a research workbench, not a production detector.

**data-privacy-stack/presidio-research** — This package features data-science related tasks for developing new recognizers for Presidio. It is used for the evaluation of the entire system, as well as for evaluating specific PII recognizers or PII detection models.

- Repository: https://github.com/data-privacy-stack/presidio-research
- Website: https://data-privacy-stack.github.io/presidio-research/
- Stars: 313 · Forks: 83
- Language: Jupyter Notebook
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/data-privacy-stack-presidio-research

## Who needs a PII evaluation harness

Presidio itself detects personally identifiable information. presidio-research answers the question that comes after detection works at all: how well does it work, and on which entities does it fail? The README names two audiences. The first is anyone developing or evaluating PII detection models, an existing Presidio instance, or a specific Presidio recognizer. The second is anyone generating new data from existing datasets or sentence templates, for example to widen the coverage of entity values for a named entity recognition model. Both groups share a problem. A recognizer that catches names in one phrasing may miss them in another, and without a labeled set you cannot see the difference. The package supplies the labeled set, the splitting logic, and the scoring. It is aimed at data scientists and NLP engineers, not at application developers who simply want to redact text. The repository's primary language is Jupyter Notebook, which tells you how the maintainers expect the work to be done: interactively, notebook by notebook, with each notebook covering one stage of the pipeline.

## How the data generation and evaluation pipeline fits together

The generation step starts from a template file. A template such as `My name is {{name}}` is sampled, PII values are drawn from the bundled fake data, and the result is a synthetic sentence. The generator then tokenizes that sentence and produces tags in IO, BIO or BILUO scheme plus spans for the entity mentions. The output is not a loose list of strings. It is a standardized structure, `List[InputSample]`, defined in `presidio_evaluator/data_objects.py`, and that single representation is what makes the rest of the pipeline possible. From it you can emit CoNLL, spaCy v3 binary format, or JSON, and you can read CoNLL2003 files back in through `presidio_evaluator.dataset_formatters.CONLL2003Formatter`. The evaluation side takes the same `InputSample` list and compares a model's predictions against the gold spans, reporting precision, recall and an error analysis. One detail in the README deserves attention: when several models each return a different set of entity labels, the results are not directly comparable, so the package includes an entity mapping process, demonstrated in Notebook 6, that normalizes labels before scoring. Without that step you are comparing two different label spaces and the numbers mean little.

## Installing presidio-evaluator and running a first evaluation

The package is published on PyPI under the name presidio-evaluator, which differs from the distribution name in pyproject.toml (presidio_evaluator). Install it and the two spaCy pipelines the documentation calls for. The small model is used for tokenization, the large one for NER and for the default Presidio configuration.

```bash
pip install presidio-evaluator
python -m spacy download en_core_web_sm
python -m spacy download en_core_web_lg
```

If you prefer to work from a clone, the README uses uv. The sync command installs the package with its dev extras, which include pytest and Jupyter, and the final command runs the test suite as an installation check.

```bash
pip install uv
uv sync --extra dev
uv run python -m spacy download en_core_web_sm
uv run python -m spacy download en_core_web_lg
uv run pytest
```

With the package in place, the fastest path is to open the notebooks in order. Notebook 1 covers generation, Notebook 3 splits the generated data into train, test and validation folds while keeping each template in exactly one fold, and Notebook 4 scores the vanilla Presidio analyzer. Notebook 5 shows a custom configuration that the README says raises the f score by roughly 30 percent. If you would rather load a dataset in a script than in a notebook, the conversion API is small. This snippet reads a generated dataset and writes it out in CoNLL format, tab separated.

```python
from presidio_evaluator import InputSample
dataset = InputSample.read_dataset_json("data/synth_dataset_v2.json")
conll = InputSample.create_conll_dataset(dataset)
conll.to_csv("dataset.csv", sep="\t")
```

The same object converts to spaCy v3 with `InputSample.create_spacy_dataset(dataset, output_path="dataset.spacy")` or to JSON with `InputSample.to_json(dataset, output_file="dataset_json")`.

## The template leakage problem and why split-by-pattern exists

Synthetic data has a failure mode that real data does not. If the same template appears in both the training fold and the test fold, the model can memorize the sentence frame and score well without having learned anything about entity boundaries. The package addresses this directly. Notebook 3 is titled around splitting by pattern number, and the README states the purpose plainly: splitting while avoiding leakage due to the same pattern appearing in multiple folds, and it notes that this applies only to synthetically generated data. That restriction is worth taking seriously. The technique is a property of how the data was built, not a general splitting strategy, so you cannot carry it over to a human-annotated corpus. The practical consequence is that your reported scores are only meaningful when the split respects template boundaries. A random split of generated data will inflate recall, and nothing in the evaluation code will warn you. If you generate data with few distinct templates, the leakage risk is larger, and the fix is more templates rather than a different splitter.

## Dependency weight, dropped integrations and the maintenance picture

The dependency list in pyproject.toml is long and opinionated. It pins spaCy at 3.8 or newer, numpy at 2.4 or newer, pandas at 3.0 or newer, transformers at 5.3 or newer, and pulls in presidio-analyzer and presidio-anonymizer at 2.2.364 or newer, along with faker, scikit-learn, plotly, requests and xmltodict. Those lower bounds are aggressive: pandas 3.x and transformers 5.x are recent major lines, so installing into an environment that already holds older versions of either will force an upgrade and may break unrelated code. Python support is bounded at >=3.11,<3.15, which excludes 3.10 and earlier. The README also carries a warning that some dependencies, such as Flair and Stanza, are no longer supported, and redirects users to Presidio Analyzer directly for adding custom NER models. Anyone following an older tutorial that wires a Flair or Stanza model into the evaluator will find that path abandoned. On maintenance, the last push to the repository was on 2026-09-10 and the most recent release listed is 0.3.2 on 2026-08-04, with 0.3.1 and 0.3 in the weeks before that. The repository is not archived. The project is MIT licensed, which permits commercial use and modification; the bundled fake identities come from Fake Name Generator under a Creative Commons Attribution-Share Alike 3.0 United States licence, and the README states that attribution requirement, so redistributing the raw data carries terms that the MIT licence on the code does not.

## Where presidio-research is the wrong tool

This package does not detect PII. It measures detectors. If your goal is to redact names and account numbers from a stream of documents, you want Presidio Analyzer and Anonymizer, and presidio-research adds nothing to that runtime path. A second boundary is language and domain. The bundled fake data and the spaCy pipelines named in the installation instructions are English, and the README does not describe a non-English data source, so a team working on German or Japanese text would be generating its own PII values and templates before any of the evaluation machinery becomes useful. Third, the evaluation is only as good as the gold labels. The generator produces spans automatically from templates, which is clean but synthetic; a recognizer tuned to score well on templated sentences may behave differently on messy real text, and the README does not present a benchmark against a human-annotated corpus. Finally, the README itself cautions that the vanilla Presidio results in Notebook 4 are not very accurate, which is an honest framing: the baseline is a starting point for comparison, not a quality bar to ship against.

## How presidio-research differs from a general NER evaluation library

A general-purpose NER evaluation library, such as seqeval or the evaluation utilities in a framework like spaCy's own scorer, expects token-level or span-level predictions against a gold corpus and computes metrics. That works fine when both sides already agree on a label set. presidio-research sits a layer above that, in a pipeline built for PII specifically. It ships a fake PII value bank and a template-driven generator, so you can manufacture labeled data without annotators. It defines `InputSample` as the interchange format across generation, conversion and scoring. And it includes the entity mapping step, which a generic scorer has no reason to provide, because the problem of reconciling different entity sets only appears when you compare several recognizers that each label the world differently. The trade-off is scope. seqeval will evaluate any tagging model in any language with no data-generation opinions; presidio-research will evaluate PII recognizers and help you build the data, but it assumes English, assumes Presidio, and expects you to work through its notebooks to get oriented.

## Conclusion

Adopt presidio-research if you are building or tuning a Presidio recognizer and need a reproducible way to generate synthetic PII sentences, split them without template leakage, and score precision and recall against a labeled set. Skip it if you need a detector that runs in production: this package evaluates detectors, it does not ship one, and the README says the vanilla Presidio configuration it scores in Notebook 4 is not very accurate. Before committing, verify that your Python version falls inside the >=3.11,<3.15 range in pyproject.toml, that you can install en_core_web_lg because the default Presidio configuration requires it, and whether the entity mapping in Notebook 6 covers the label sets your analyzer actually returns.

## FAQ

### What is presidio-research used for?

It provides evaluation and data-science capabilities for Presidio and PII detection models in general, including a fake data generator that builds synthetic sentences from templates and fake PII. It is meant for developing or evaluating PII detection models and for generating new data for NER models.

### How do I install presidio-research?

Install the PyPI package with pip install presidio-evaluator, then download the spaCy pipelines en_core_web_sm for tokenization and en_core_web_lg for NER. From a clone, the README uses uv sync --extra dev and verifies the install with uv run pytest.

### Does presidio-research detect PII itself?

No. The package evaluates Presidio as a system, a NER model, or a specific PII recognizer for precision, recall and error analysis. Detection is done by Presidio Analyzer, which is a dependency rather than the subject of this package.

### Why does the splitting notebook avoid putting the same pattern in two folds?

Because synthetically generated sentences share templates, and a template appearing in both training and test folds lets a model score well without learning entity boundaries. The README notes this splitting approach applies only to synthetically generated data.

### Can I still use Flair or Stanza models with presidio-research?

The README states that some dependencies, such as Flair and Stanza, are no longer supported, and points users to Presidio Analyzer directly for adding custom NER models.

## Sources

- [data-privacy-stack/presidio-research on GitHub](https://github.com/data-privacy-stack/presidio-research)
- [License: MIT](https://github.com/data-privacy-stack/presidio-research/blob/main/LICENSE)
- [Project website](https://data-privacy-stack.github.io/presidio-research/)
- [README](https://github.com/data-privacy-stack/presidio-research/blob/main/README.md)
- [Releases](https://github.com/data-privacy-stack/presidio-research/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/data-privacy-stack-presidio-research
