Library / SDK
data-privacy-stack/presidio avatar
data-privacy-stack/presidio

Presidio: PII detection you run yourself, with no guarantee it finds everything

An open-source framework for detecting, redacting, masking, and anonymizing sensitive data (PII) across text, images, and structured data. Supports NLP, pattern matching, and customizable pipelines.

11,112 stars1,319 forksPythonMIT

At a glance

What is it?
data-privacy-stack/presidio is an MIT-licensed Python SDK that detects and anonymizes sensitive data across text, images and structured data. It splits detection from anonymization into separate services, and its own README warns that automated detection cannot be relied on to find all sensitive information.
Who is it for?
Adopt Presidio if you need PII detection and redaction running inside your own boundary, across text, images including DICOM, and structured data, and if you are prepared to tune thresholds and add recognizers for entity types it does not ship with.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Detection and anonymization are two services

Presidio is a data protection and de-identification SDK. The README describes it as providing fast identification and anonymization modules for private entities in text, naming credit card numbers, names, locations, social security numbers, bitcoin wallets, US phone numbers and financial data.

The split matters more than the feature list. Detection lives in presidio-analyzer and the rewriting lives in presidio-anonymizer, and they deploy as separate services. That separation is what lets you log what was found before deciding what to do about it, run detection in one environment and redaction in another, or swap the anonymization operator without touching the detection logic.

The repository shows five components, each a directory at the root: presidio-analyzer, presidio-anonymizer, presidio-image-redactor, presidio-structured and presidio-cli. The README's component table tracks downloads and coverage separately for analyzer, anonymizer, image-redactor and structured, which tells you they release independently.

Scope has grown past text. The README lists structured data and images, including DICOM medical images, which is the format hospital imaging actually uses and the reason this shows up in healthcare compliance work.

How recognizers work

The README says recognizers are predefined or custom, and lists four mechanisms they use: named entity recognition, regular expressions, rule based logic and checksum, with relevant context, in multiple languages.

That combination is the sensible design. A credit card number is found by regex plus a checksum. A person's name needs an NER model plus context, because a surname alone is not identifying and a surname next to the word "patient" is a different decision. Rule based logic covers the cases where neither works, such as a local identifier format your organisation invented.

The repository topics name the NLP backends: spacy and transformers. So detection quality is bounded by whichever model you run underneath, and swapping in a different one is a configuration decision rather than a rewrite.

The README also mentions the option of connecting to external PII detection models, and the docker-compose file in the repository points the analyzer at an Ollama host, so a local model server is one of the supported backends rather than something you bolt on.

Custom recognizers are the part most teams end up needing. No shipped set knows your internal customer IDs, and the README's stated goal of extensibility for a specific business need is what that refers to.

Running it in Docker

The README lists four installation routes: pip, Docker, from source, and a V1 to V2 migration guide, all documented on the project site rather than inline. The repository ships a docker-compose.yml that shows how the pieces fit together.

The analyzer service in that file exposes port 5002 on the host mapped to 5001 in the container, and takes the model backend from an environment variable:

yaml
  presidio-analyzer:
    environment:
      - PORT=5001
      - OLLAMA_HOST=http://ollama:11434
    ports:
      - "5002:5001"
    depends_on:
      ollama:
        condition: service_healthy

The anonymizer service runs on host port 5001 mapped to container port 5001, and both services define a healthcheck that curls /health, so the compose file is written for orchestration rather than for a laptop demo. A presidio-image-redactor service is defined too, and an ollama service pinned to image 0.32.0 provides the model backend that the analyzer depends on.

Each service also carries a restart: unless-stopped policy. The practical read is that this is a service you deploy, not a library you import, at least at this scale.

Images and structured data

The image path is a separate package, presidio-image-redactor, and the README says it handles standard image types and DICOM medical images. Redacting an image means finding text in it, deciding which regions are sensitive and then covering them, which is a different pipeline from text and needs OCR somewhere in the chain.

The structured path is presidio-structured, for data that already has columns and types. When a field is named email, you do not need a model to tell you it holds an email address, so structured detection is largely a matter of applying the right operators per column, and the interesting question becomes what you replace values with.

Both share the anonymizer's operator concept, so the same masking behaviour applies across text, image and table. Whether the operators preserve referential integrity, so that the same person maps to the same pseudonym in every row, is not something the README states, and it is the first thing to check if you are de-identifying a dataset you intend to join later.

The guarantee it does not make

This is the most important paragraph in the README, and it is a warning rather than a feature. Presidio can help identify sensitive data in unstructured and structured text, but because it uses automated detection mechanisms, there is no guarantee it will find all sensitive information, and the README says additional systems and protections should be employed.

That is the honest framing for the whole category. A detector has a recall ceiling, and a missed social security number in a document you published is a breach, not a bug report. Recall and precision also pull against each other through the confidence threshold: turn it down to catch more and you redact half the document, turn it up and you miss things.

The practical consequence is that Presidio belongs in a defence-in-depth design, not as the control. Treat its output as one signal, keep a human review path for anything leaving the organisation, and measure it against your own data rather than against a demo.

The release history shows this being worked at the edges: release 2.2.364 includes a fix to the UK_NINO regex matching a numeric suffix character, which is exactly the kind of false positive that erodes trust in a detector.

Cloud DLP services as the alternative

The alternative most organisations weigh is a managed classification service, such as Google Cloud's DLP or AWS Macie, and the difference is where your data goes.

A cloud DLP service is an API. You send it content, it returns findings, and you pay per unit scanned. You get a maintained detector covering many entity types with no model to run and no pipeline to operate, and you accept that the content leaves your network and that the bill scales with volume.

Presidio runs where you run it. Nothing leaves, the cost is your own compute, and you can add recognizers for entity types no cloud service knows about, such as internal identifiers. In exchange you own the model, the threshold tuning, the upgrade cycle and the recall measurement.

The decision usually turns on policy rather than technology. If your data cannot leave a jurisdiction or a network boundary, a local detector is the only option. If it can, and you have no appetite for tuning, a managed service will find more out of the box.

Governance change and upkeep

Something is happening to the project's home that anyone adopting it should read about first. The README opens with a notice that Presidio is moving to a new home and links to a project transition document. Release 2.2.364 contains a change replacing [email protected] with [email protected] in PyPI metadata, and another updating the licence copyright to Presidio Contributors.

So the project has moved out of Microsoft's org into an independent structure, the repository confirms it, and the contact address has changed. The documentation is now hosted at data-privacy-stack.github.io. If you are adopting this on a compliance timeline, read that transition document rather than assuming continuity.

On the positive side the signals are strong: MIT licence with a NOTICE file, an OpenSSF best practices badge, SECURITY.md, SUPPORT.md, CODE_OF_CONDUCT.md and CONTRIBUTING.md, plus e2e-tests/ and per-component coverage tracking.

Release activity is steady. Version 2.2.364 was published on 2026-07-22 and 2.2.363 on 2026-06-28, and the last push was on 2026-09-15. The 2.2.363 notes include German PII recognizers under a DE_ prefix, which is what incremental language coverage looks like in practice.

Editorial conclusion

Adopt Presidio if you need PII detection and redaction running inside your own boundary, across text, images including DICOM, and structured data, and if you are prepared to tune thresholds and add recognizers for entity types it does not ship with. Do not adopt it as your only control, because the README states plainly that automated detection carries no guarantee of finding all sensitive information, and do not adopt it if you want a detector someone else tunes, in which case a managed cloud DLP service will find more with less work. If you are evaluating on a compliance timeline, read the project transition document first: the contact address and the licence copyright have already changed as the project moved out of Microsoft into the data-privacy-stack organisation.

Frequently asked questions

What does Presidio mean?

The README gives the origin as Latin praesidium, meaning protection or garrison, which fits its role as a data protection and de-identification SDK.

What is a Presidio tool?

It is a context aware, pluggable PII de-identification service for text and images, split into an analyzer that finds sensitive entities and an anonymizer that rewrites them, plus image redaction and structured data components.

Does Presidio find all sensitive data?

No. The README warns that because detection is automated there is no guarantee Presidio will find all sensitive information, and says additional systems and protections should be used alongside it.

Official sources

  1. data-privacy-stack/presidio on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/data-privacy-stack-presidio.svg)](https://hysenlabs.com/projects/data-privacy-stack-presidio)