# rizzo-pii: local, reversible PII anonymization for Italian legal text

> rizzo-pii is a 0.3B-parameter, CPU-only token-classification model plus a Flask workflow that replaces 22 categories of personal data with type-aware placeholders, so documents can be sent to a frontier LLM and re-identified locally. The privacy guarantee is structural, but the model is Italian-first and the re-identification step puts the dictionary on your disk.

**Rizzo-AI-Academy/rizzo-pii** — Local-first privacy guard: anonymize your documents before sharing with LLMs.

- Repository: https://github.com/Rizzo-AI-Academy/rizzo-pii
- Website: https://rizzo-ai-academy.github.io/rizzo-pii/
- Stars: 1,126 · Forks: 82
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/rizzo-ai-academy-rizzo-pii

## The problem rizzo-pii solves, and for whom

The README opens with a blunt observation: people summarize contracts and draft replies by pasting the document into a chat window, and that quietly moves names, addresses, tax codes, IBANs and health details to servers the user does not control. For a law firm or a hospital, the README argues, this is a direct GDPR compliance failure rather than a hypothetical risk.

The obvious fix is to stop sending data out and run an open model locally. The README rejects that too, on cost grounds: it states that a frontier-grade open model needs roughly EUR 9,000 to EUR 10,000 of hardware, while the small models that fit a laptop are not competitive on legal reasoning and dense contracts. So the project keeps the frontier model and removes the data instead.

The intended audience is narrow and explicitly named: law firms, accountants, notaries, and anyone bound by the GDPR who wants to keep using ChatGPT, Claude or Gemini on sensitive documents. The model is described as Italian-first, and the README claims coverage of Italian-legal identifiers (codice fiscale, partita IVA, dati catastali) that, in its words, no other open model covers. That claim is the project's own and I have not verified it against other models.

## How the placeholder round trip actually works

The mechanism is a pseudonymization loop with three local steps and one remote step. Locally, the model tags every span of personal data and replaces each span with a stable, type-aware placeholder such as [FULLNAME_1], [IBAN_1] or [CF_1]. The mapping from placeholder to real value goes into a dictionary that stays on disk. Identical values receive the same placeholder, so the frontier model still sees a coherent document and can reason about repeated parties across a contract.

The anonymized text is what crosses the network. When the answer comes back, a local pass swaps the placeholders for the true values. The README's framing is that the cloud provider never receives a single real name, code or number.

The design choice worth noting is reversibility. Classic redaction destroys information, which makes the output useless for real work because you cannot tell who the parties were. rizzo-pii pseudonymizes instead, so the model's answer can be reconstructed. The trade-off is that a mapping file now exists on your machine, and the README does not document how that dictionary is encrypted, rotated or deleted. The privacy boundary moved; it did not disappear.

## Installing rizzo-pii with Docker and a first anonymization run

The repository ships a docker-compose.yml for the Flask webapp. Its comments are in Italian and it maps the service to port 5005 on localhost only, with the note that the network boundary is enforced by Docker and that you should remove the 127.0.0.1: prefix only if you really need LAN access. The compose file builds from the Dockerfile and passes MODEL_REPO and MODEL_REVISION as build arguments, defaulting to rizzoaiacademy/rizzo-pii-0.3B at revision v1.5.0. Its header comment gives the intended start command:

```bash
docker compose up -d --build     # build + avvio in background
docker compose logs -f           # segui i log
docker compose down              # stop (il volume delle preferenze resta)
```

The build downloads the model into the image. At runtime the container sets HF_HUB_OFFLINE: 1 and TRANSFORMERS_OFFLINE: 1, so the README's promise that documents do not leave the machine is enforced by the container environment rather than by convention. The Dockerfile comments state that the image is CPU-only and that libgomp1 is installed because torch's CPU kernels need the OpenMP runtime.

The compose file also sets OMP_NUM_THREADS: 4 with a comment that this prevents torch from saturating every core, and you should raise it on a dedicated machine. A named volume, rizzo-pii-home, is mounted at /home/app so UI preferences survive container recreation.

The healthcheck in the compose file polls /health every 30 seconds and treats a non-200 response as failure. The comment states that /health returns 503 until the model has loaded, and start_period is set to 120s to cover a cold start. If the request returns a 503 immediately after launch, the model is still loading rather than broken.

The README also points to a desktop route: Windows installer, macOS (Apple Silicon) and Linux AppImage builds linked from the releases page, with v2.0.0 published on 2026-08-08 and v1.0.0 titled "rizzo-pii 1.0.0, Windows-Linux-macOS" on 2026-06-28. The README does not document a pip install path for the application itself; requirements.txt is the dependency list for the Python code in the repository, not a published package.

## Where rizzo-pii is the wrong tool

The model is Italian-first, and that is a constraint rather than a footnote. If your documents are in English, German or French, the Italian-legal identifiers that justify the project's existence are irrelevant to you, and you are paying the Italian-language bias for nothing. The README offers no multilingual claim.

The second limitation is structural. The guarantee depends on detection and re-identification running locally on a CPU, so any deployment that moves the Flask server off the machine holding the data breaks the property the project is selling. The compose file's localhost-only port mapping is the visible expression of that rule.

Third, the README says nothing about rollback. If the anonymization pass misses a span, the real value goes to the cloud provider and there is no documented way to retract it. The README also does not describe a confidence threshold, a manual review queue, or a dry-run mode that shows what would be replaced before it is replaced. For a workflow whose entire value is that nothing identifiable crosses the boundary, the absence of a documented pre-flight check is the gap I would want closed before relying on it for client material.

Finally, the dictionary on disk is the new sensitive artifact. The README does not document encryption at rest, a retention policy, or a purge command for it.

## How it differs from running a local LLM instead

The realistic alternative is to keep everything on the machine and run an open-weight model locally, which is what the README considers and rejects on hardware cost: it puts a frontier-grade local setup at roughly EUR 9,000 to EUR 10,000. The difference in approach is what each side gives up. A local LLM keeps the data and downgrades the model, so you never send anything out but you accept weaker legal reasoning on dense contracts. rizzo-pii keeps the frontier model and downgrades the data instead, sending pseudonymized text that the provider can still process at full quality.

That reframes the failure modes. With a local LLM, the risk is a worse answer. With rizzo-pii, the risk is a missed entity, and the consequence of a miss is that a real identifier leaves the machine. The two approaches also differ in what they leave behind: a local LLM leaves nothing on a third party's servers, while rizzo-pii leaves a placeholder-to-value dictionary on your own disk.

There is a third path the README does not discuss, which is simply not using an LLM on the document at all. That is worth naming because rizzo-pii is only the right answer if the frontier model's output is genuinely worth the extra moving parts.

## Licence, dependencies and the upgrade bill

The repository is MIT licensed, and there is a separate THIRD_PARTY_LICENSES.md file at the top level, which suggests the maintainers have at least inventoried the transitive licences. The model weights are hosted on Hugging Face under rizzoaiacademy/rizzo-pii-0.3B, and the compose file pins MODEL_REVISION to v1.5.0. Model weights and the MIT-licensed code are separate artifacts with potentially separate terms; the README does not state the licence of the weights, so check that before commercial use. This is not legal advice.

On maintenance: the last push was on 2026-08-09, and v2.0.0 was released on 2026-08-08. The repository is not archived.

The upgrade cost is visible in the Dockerfile comments, which state that torch and transformers are pinned to the versions the image was verified against, because without pins the same build six months later pulls a different major version. The Dockerfile also notes that it uses transformers 5.x while Dockerfile.linux stays on 4.57.3 for PyInstaller, so the container and the Linux bundle do not track the same transformer version. Anyone maintaining a fork should expect to rebuild and re-verify after bumping those pins rather than assuming a drop-in upgrade. The requirements.txt adds its own warning: PyTorch must match your GPU, and NVIDIA Blackwell (RTX 50xx, sm_120) training requires the cu128 build because the cpu and cu121 wheels do not support sm_120.

## Conclusion

Adopt rizzo-pii if you handle Italian legal, tax or notarial documents and want to keep using ChatGPT, Claude or Gemini without sending real identifiers out. Skip it if your documents are mostly English, or if you cannot run a local process on the machine that holds the data, because the entire guarantee is that detection and re-identification happen there. Verify first that the 22 categories cover the identifiers in your own corpus, and confirm that the placeholder dictionary file is stored where your retention policy expects it.

## FAQ

### Does rizzo-pii send my documents to a cloud provider?

No. The README states that detection and re-identification run locally on a CPU with no API key, no telemetry and no upload, and the Docker image sets HF_HUB_OFFLINE and TRANSFORMERS_OFFLINE at runtime. Only the placeholder text is sent to the frontier model, and the real values are restored locally from the dictionary.

### Which languages does the rizzo-pii model support?

The README describes rizzo-pii:0.3B as Italian-first and highlights Italian-legal identifiers such as codice fiscale, partita IVA and dati catastali. It makes no multilingual claim, so treat it as a tool for Italian documents.

### How much RAM does rizzo-pii need?

The README states a footprint of roughly 0.5 GB of RAM on CPU for the approximately 0.3B-parameter model, and notes that a normal laptop is enough. The Dockerfile installs libgomp1 because torch's CPU kernels need the OpenMP runtime.

### What happens to the placeholder dictionary rizzo-pii creates?

The README says the mapping from placeholder to real value is recorded in a dictionary that stays on disk. It does not document encryption, rotation or a purge command for that file, so plan for it in your own retention policy.

## Sources

- [License: MIT](https://github.com/Rizzo-AI-Academy/rizzo-pii/blob/main/LICENSE)
- [Project website](https://rizzo-ai-academy.github.io/rizzo-pii/)
- [README](https://github.com/Rizzo-AI-Academy/rizzo-pii/blob/main/README.md)
- [Releases](https://github.com/Rizzo-AI-Academy/rizzo-pii/releases)
- [Rizzo-AI-Academy/rizzo-pii on GitHub](https://github.com/Rizzo-AI-Academy/rizzo-pii)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/rizzo-ai-academy-rizzo-pii
