Open-source project
ttttccxxui/DataInfra-RedactionEverything avatar
ttttccxxui/DataInfra-RedactionEverything

DataInfra-RedactionEverything: a local redaction workbench for scanned and bilingual documents

DataInfra Series. Redact EVERYTHING with local llms and vlms.

1,143 stars169 forksPythonNOASSERTION

At a glance

What is it?
RedactionEverything runs semantic NER, OCR and visual grounding on your own machine to find names, seals, faces and signatures in PDFs, Word files and images. The catch is its licence: personal use is free, everything else needs a commercial agreement.
Who is it for?
Adopt RedactionEverything if you handle Chinese or bilingual contracts, scans and stamped documents and you need the raw files to stay on your own hardware, and if you are an individual doing unpaid work or a company ready to buy a commercial licence. Do not adopt it as a general-purpose PII filter for clean English text, and do not drop it into production on the strength of the Apache-licensed parts alone.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 29 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What RedactionEverything solves that a token-level PII filter does not

Most open privacy filters assume the input is clean text. Real business documents are not. They arrive as scanned PDFs, Word contracts with tracked changes, screenshots, phone photos of paper, and pages where a company seal sits directly on top of the party name you need to redact. RedactionEverything is built for that second category. The README describes it as a "local-first redaction workbench" and is explicit that the goal is not a narrow fixed-rule PII scanner: detection is driven by configurable schemas rather than one hardcoded label set.

The intended user is someone working with Chinese or bilingual material in legal, finance, healthcare or general office settings, where the sensitive items are names, organizations, ID numbers, accounts, addresses, amounts, dates, seals, faces and signatures. The repository positions this against OpenAI Privacy Filter, which it calls a valuable high-throughput baseline for token-level PII detection, and draws the line at scope: messy documents, visual privacy regions, human review, batch delivery and local deployment. That is a fair framing. A regex pass over extracted text will not see a face in a scanned ID card, and it will not see a name hidden under red seal ink.

Two detection paths: semantic text NER and visual feature grounding

The architecture splits every file into a text path and a visual path. On the text side, HaS Text performs semantic NER directly against the configured tags, and the README notes that regex is kept only as a user-defined fallback rather than the primary mechanism. For images and scanned documents, PP-StructureV3 running on PP-OCRv6 engines converts the page into text blocks, with PaddleOCR-VL available as an optional supplement. HaS Text then recognizes entities in those blocks and the results are mapped back to glyph-exact coordinates, which is why a label like 户名: can stay outside the mask while the value after it is covered.

The visual path is a single LocateAnything-3B service that grounds a fixed preset list (faces, fingerprints, IDs, bank cards, seals, screens, QR and barcodes, signatures) plus any user-defined visual label. A local OpenCV detector runs alongside it to recover red binding and edge seals, with deduplication against boxes LocateAnything already produced. There is also a stamp-crushed text recovery pass: red-ink suppression whitens the seal and re-detects, which the README gives as the case of a party name under a company seal.

What is honest about this design is that it is not one model doing everything. It is four services with different failure modes, which is also where the operational cost lives.

Installing RedactionEverything with Docker and running a first file

The repository ships a docker-compose.yml with two modes. The default brings up the API and frontend on CPU; the gpu profile adds the model stack. The compose file documents both directly in its header comment.

bash
docker compose up -d
docker compose --profile gpu up -d

The first command starts the backend and frontend only. The second adds the model services. The backend listens on port 8000 by default, controlled by BACKEND_PORT, and the compose file points OCR_BASE_URL at http://ocr:8082, HAS_TEXT_VLLM_BASE_URL at http://ner:8080/v1 and VISUAL_FEATURES_BASE_URL at http://visual-features:8090.

Before starting, copy the environment template and set a secret, because authentication is on by default.

bash
cp .env.example .env
# then edit .env: set JWT_SECRET_KEY and keep AUTH_ENABLED=true

The .env.example file states that AUTH_ENABLED defaults to true and that JWT_SECRET_KEY is a separate value. It also warns that DATA_DIR, UPLOAD_DIR and OUTPUT_DIR should stay commented out for local Windows development, because /app/data would otherwise resolve to a drive root.

For local development without Docker, the root package.json exposes a single script that starts the frontend, backend, PaddleOCR-VL, PP-StructureV3, HaS Text and LocateAnything together. The engines field requires Node 24.x.

bash
npm run dev

The same .env.example file says this path requires VENV_DIR and related variables to point at WSL Linux paths, and that a missing value fails immediately with a config error rather than degrading. Once the services are up, the workflow is upload, recognize, review, redact, export, and the README lists TXT, DOCX, PDF, scanned PDF, PNG and JPG as accepted inputs.

GPU memory is the real deployment constraint

The README devotes a section to limitations and GPU memory, and that is where a prospective adopter should start rather than the feature table. Running HaS Text through vLLM, PP-StructureV3 on PP-OCRv6, and LocateAnything-3B at the same time is a multi-model deployment. The compose defaults are conservative for a reason: JOB_CONCURRENCY defaults to 3, but BATCH_RECOGNITION_PAGE_CONCURRENCY, HAS_NER_MAX_PARALLEL_REQUESTS and VISUAL_FEATURES_CONCURRENCY each default to 1, and VISION_DUAL_PIPELINE_PARALLEL defaults to false. Raising those numbers is how you get throughput, and it is also how you exhaust VRAM.

There is a second failure mode that has nothing to do with hardware. The task center requires a running task to be cancelled before it can be deleted, which means a stuck recognition job blocks cleanup until it is stopped. OCR_REQUIRE_GPU defaults to false, so you can run the OCR path on CPU, but the visual feature service is a separate concern and the README does not document a CPU fallback for LocateAnything-3B.

The wrong tool judgement is simpler. If your documents are already clean digital text in English and your labels are a small fixed set, this is a lot of infrastructure for a problem a token-level classifier handles. The project's own positioning says as much.

How RedactionEverything differs from OpenAI Privacy Filter

The README names OpenAI Privacy Filter as the comparison point and describes it as a high-throughput baseline for token-level PII detection in text. The difference is not accuracy on a shared benchmark; the two tools are answering different questions. Privacy Filter takes text and returns labeled spans. RedactionEverything takes a file, decides whether it needs an OCR pass at all, extracts text blocks with coordinates, runs semantic NER against a schema, runs a separate visual model for regions that have no text representation, merges and deduplicates the boxes, then hands the result to a review interface with batch task state and export packaging.

The practical consequence is that RedactionEverything carries a model-serving footprint and a review workflow, while a token-level filter can be embedded in a pipeline with far less. If your input really is text and your privacy boundary allows a hosted model, the lighter tool is the better fit. If your input is a photographed contract with a stamp over the signature line, the lighter tool cannot see the problem at all.

Licence terms and the third-party components you inherit

The licence is the part most likely to surprise an adopter, so read it before anything else. The badge and the README both point to a custom Personal Use License, not an OSI-approved one. Individuals may use it free for personal, non-commercial purposes. The README states that paid work, consulting delivery, companies, institutions, government agencies, teams, hosted services, production deployments, OEM redistribution and commercial integrations all require a separate commercial licence, with contact at [email protected]. The repository metadata reports the licence as NOASSERTION, which is consistent with a custom text.

There is a second layer. The README says commercial deployments must clear third-party component licences themselves: the LocateAnything-3B weights are released under an NVIDIA non-commercial licence, and PyMuPDF is AGPL-3.0, dual-licensed commercially by Artifex. That means even a paid RedactionEverything licence does not by itself settle your obligations for the visual model or the PDF library. The README points to a component table in its License section for the full list. This is not legal advice; the point is that two independent licence questions exist and only one of them is answered by buying a commercial licence from the author.

Maintenance status and what upgrading costs you

The repository is not archived and the last push was on 2026-09-02. There are no retrieved releases, so there is no versioned upgrade path to reason about: you track the main branch or you pin a commit. The README does not document rollback, migration between schema versions, or a compatibility matrix for the model services, and the README does not state a support policy for the free personal tier beyond the commercial contact address.

Upgrade cost is therefore mostly environmental. The stack pins expectations in several places at once: Node 24.x in package.json, a vLLM runtime for HaS Text, WSL paths for local development, and a PaddleOCR stack for the OCR service. Bumping any one of those means re-testing the coordinate mapping between OCR text blocks and the masks drawn from them, because that mapping is what keeps labels outside the redaction box. The .env.example file is the closest thing to a compatibility document, and it is a template rather than a changelog.

Editorial conclusion

Adopt RedactionEverything if you handle Chinese or bilingual contracts, scans and stamped documents and you need the raw files to stay on your own hardware, and if you are an individual doing unpaid work or a company ready to buy a commercial licence. Do not adopt it as a general-purpose PII filter for clean English text, and do not drop it into production on the strength of the Apache-licensed parts alone. Before you commit, read the LICENSE file and the component table, confirm whether your GPU memory clears the thresholds in the limitations section, and check that LocateAnything-3B and PyMuPDF are cleared for your use case, because the repository states that commercial deployments must resolve those third-party terms themselves.

Frequently asked questions

What is DataInfra-RedactionEverything and what is it for?

It is a local-first redaction workbench that finds and anonymizes sensitive content in documents, scanned PDFs, images, Word files and plain text. It combines semantic NER, OCR, visual feature grounding, configurable industry schemas, human review, batch processing and export workflows.

How do you redact data with DataInfra-RedactionEverything?

You upload files such as TXT, DOCX, PDF, scanned PDF, PNG or JPG, select a schema, run recognition, review and correct the detected items, then export the redacted output. Batch processing follows the same sequence across a mixed queue, with results tracked in the task center.

What is the difference between data redaction and data obfuscation in RedactionEverything?

The repository describes redaction as finding and anonymizing sensitive content in files and exporting redacted artifacts; it does not document an obfuscation mode, so the distinction between the two terms is not addressed in the README.

Official sources

  1. Issues
  2. README
  3. ttttccxxui/DataInfra-RedactionEverything on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ttttccxxui-datainfra-redactioneverything.svg)](https://hysenlabs.com/projects/ttttccxxui-datainfra-redactioneverything)