# Sensitive-lexicon: a plain-text Chinese word list for content filtering

> Sensitive-lexicon is an MIT-licensed repository of Chinese sensitive-word lists in plain text, plus an optional Go detection service. It is a data source, not a filtering engine, and the README is explicit that the final judgement on what counts as sensitive is yours.

**konsheng/Sensitive-lexicon** — 一个持续更新的中文敏感词库，帮助开发者和内容审核者快速识别并过滤不当文本，即将迎来重大更新

- Repository: https://github.com/konsheng/Sensitive-lexicon
- Website: https://github.com/konsheng/Sensitive-lexicon
- Stars: 4,170 · Forks: 439
- Language: Unknown
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/konsheng-sensitive-lexicon

## What Sensitive-lexicon actually is: a word list, not a filter

The repository describes itself as a continuously updated Chinese sensitive-word list that helps developers and content reviewers identify and filter inappropriate text. That sentence is the whole product. What you get is text files, organised into directories, that you read from your own code. The README lists the coverage as political, pornographic and violent vocabulary, and says the list runs to tens of thousands of entries.

The intended audience is narrow and specific. It is for engineers building a moderation pipeline who need a Chinese lexicon to feed into it, and for content reviewers who want a reference list they can inspect by hand. It is not for someone who wants a hosted moderation endpoint, and it is not for anyone working in a language other than Chinese. The README's own notes say sensitive-word definitions depend on culture, region and context, and that you should evaluate and adjust them against your own business needs. That is an unusually honest disclaimer for a list like this, and it tells you the maintenance burden does not disappear when you clone the repository.

## Repository layout: Vocabulary, Organized and ThirdPartyCompatibleFormats

The top level of the repository contains four directories and the README. Vocabulary/ holds the word lists themselves. Organized/ holds lists that have already been tidied up. ThirdPartyCompatibleFormats/ holds versions shaped for other tools' import formats. The README's directory tree shows exactly these three data directories plus LICENSE and README.md.

The split between Vocabulary/ and Organized/ is the interesting design decision. Raw vocabulary accumulation and curated output are kept apart, which means a contributor can add an entry without having to match an existing schema, and a consumer can pick the tidier directory if they do not want to clean the data themselves. The cost is that you have to decide which directory is authoritative for your use case, and the README does not spell out the difference in detail. If you are integrating this, read both directories before you pick one.

Contributions go through the Vocabulary/ directory: the README asks for a pull request that adds or modifies entries there, and asks that the PR include a source or a use case so the maintainer can review it. That requirement is the main quality control on the list, and it depends entirely on reviewers actually applying it.

## Installing Sensitive-lexicon and running a first filter

There is no package to install. The README's quick start is a git clone, and after that you read the .txt files from wherever you put them.

```bash
git clone https://github.com/Konsheng/Sensitive-lexicon.git
```

After cloning, the repository root contains Vocabulary/, Organized/ and ThirdPartyCompatibleFormats/. The README tells you to read the .txt files in the lexicon (or the branch file you need) and then choose a matching algorithm for your business case, naming DFA, Trie and regular expressions as examples. It does not ship a reference implementation for those, so the matching step is yours to write.

The repository does also describe a Go detection service, with source under ./cmd/server and REST endpoints at /detect, /contains, /reload and /health. The README gives a Docker invocation:

```bash
docker run -p 8080:8080 ghcr.io/<您的用户名>/sensitive-lexicon-server:latest
```

The image name in the README still contains a placeholder for your own username, so the published image path is not fixed there; check the repository before relying on that exact reference. The service reads environment variables named PORT, LEXICON_DIR, FUZZY_MIN_NGRAM, FUZZY_MAX_NGRAM and FUZZY_MAX_DISTANCE. The n-gram and distance variables are the knobs for the fuzzy matching the README mentions, and nothing in the README states their default values or recommended ranges, so you will be guessing until you read the server source.

## The fuzzy-matching server and its undocumented knobs

The Go service is the most ambitious part of the project and the least documented. The README says it supports fuzzy matching and hot reloading of the lexicon, and it names the endpoints: /detect, /contains, /reload and /health. Hot reload matters in practice, because a word list that changes on a community schedule is useless if applying an update requires a restart.

What the README does not give you is any request or response schema for those endpoints, any statement of what the fuzzy parameters mean, or any guidance on what values produce acceptable precision. FUZZY_MIN_NGRAM, FUZZY_MAX_NGRAM and FUZZY_MAX_DISTANCE are exposed, which tells you the matching is n-gram based with an edit-distance threshold, but the README stops there. Fuzzy matching on a sensitive-word list is a precision problem as much as a recall problem: widen the distance and you start flagging ordinary text. Without documented defaults, the only way to calibrate is to read ./cmd/server and test against your own corpus.

The README also points to a dev branch that it describes as having more frequent service and engineering updates. That is a normal arrangement, but it means the main branch and the dev branch can diverge on service behaviour, and you should know which one you are deploying.

## Where Sensitive-lexicon is the wrong tool

The clearest limitation is the one the README states itself: sensitive-word definitions vary by culture, region and context, and you are expected to assess and adjust them for your own needs. A word list cannot make that judgement. A term that is a slur in one community is ordinary vocabulary in another, and a static list will produce both false positives and false negatives no matter how large it grows.

The second limitation is coverage. The README claims tens of thousands of entries across political, pornographic and violent domains, but it gives no per-domain counts and no accuracy figures, and there is no published evaluation. You cannot tell from the documentation whether a specific category you care about is well covered or thinly covered. That is a gap you have to close by reading the files.

The third is that this is not a complete moderation system. There is no normalisation step documented for the classic evasions (spacing, homophones, character substitution), no scoring, no context window, no appeal path. If your requirement is a service that returns a confident moderation decision, this repository gives you an ingredient, not the dish. Teams without the capacity to build and tune the matching layer should look elsewhere first.

## How it compares with a hosted moderation API

The obvious alternative is a hosted content-moderation service, such as those offered by large cloud providers or by platforms that publish Chinese word lists of their own. The difference is where the work sits. A hosted API takes your text, applies a model and returns a label with a confidence score, and the vendor owns the list, the model and the updates. Sensitive-lexicon takes the opposite position: it hands you the list and nothing else, so you own the matching, the thresholds and the tuning, and you can read every entry before it reaches production.

That trade is real in both directions. A hosted API gets you running in an afternoon and improves without your involvement, but you cannot audit why a specific string was flagged, you cannot remove an entry that is wrong for your community, and your text leaves your infrastructure. Sensitive-lexicon gives you auditability and local control, and charges you the engineering time to build the matcher. If your corpus is small, or your tolerance for false positives is low, the hosted route is usually cheaper. If you handle text you cannot send to a third party, or you need to explain each decision, the local list wins.

Within the same family of open lists, the distinguishing feature here is the plain-text format. There is no database, no binary index, no build step. That makes the data easy to diff in review and easy to merge into whatever structure your matcher wants.

## Maintenance, versioning and the MIT licence

The repository is not archived, and the last push was on 2026-08-17. Releases are infrequent and unevenly spaced: 1.0 on 2024-02-26, 1.1 on 2025-07-28, and 1.2 on 2025-08-12. The README's own description says a major update is coming, which suggests the current release line is not the end state. For a word list, release cadence matters less than commit cadence, because the value is in the entries; the last push date is the better signal of whether the list is still moving.

Upgrade cost is low by design. Because the data is plain text, pulling a new version is a git pull, and the diff shows you exactly which entries changed. The exception is the Go service: if you deploy it, you inherit its interface, its environment variables and the dev-branch divergence. Budget review time for the diff, not migration time.

The licence is MIT. Under MIT you may use, modify and distribute the project provided you keep the copyright and licence notice. That is permissive and compatible with commercial use, but it says nothing about the legal status of the words themselves or about your obligations as a platform operator. The README's notes direct you to comply with local laws and platform policies, and that obligation sits with you, not with the licence. This is not legal advice; if your use case is regulated, take your own.

## Conclusion

Adopt Sensitive-lexicon if you need a starting Chinese word list you can read, diff and edit, and you already have or plan to write the matching layer. Skip it if you expect a drop-in moderation API, or if your content is not Chinese. Before committing, open the Vocabulary/ directory, count what is actually there, check whether the words you care about are present, and read the notes on fuzzy matching to see whether the server's approach fits your traffic.

## FAQ

### What is the Sensitive-lexicon word list used for?

It is a Chinese sensitive-word list intended to be embedded in a text-review pipeline so that developers and content reviewers can identify and filter inappropriate text. The README describes coverage of political, pornographic and violent vocabulary.

### What are the different directories in Sensitive-lexicon?

The repository has Vocabulary/ for the word lists, Organized/ for lists that have already been tidied up, and ThirdPartyCompatibleFormats/ for versions shaped for other tools' import formats. Contributions are made through Vocabulary/.

### Does Sensitive-lexicon include a detection service?

Yes. The README describes a Go detection service under ./cmd/server with REST endpoints at /detect, /contains, /reload and /health, supporting fuzzy matching and hot reloading of the lexicon.

### How do I run the Sensitive-lexicon server with Docker?

The README gives a docker run command publishing port 8080 and using the image ghcr.io/<your-username>/sensitive-lexicon-server:latest. It configures the service through the environment variables PORT, LEXICON_DIR, FUZZY_MIN_NGRAM, FUZZY_MAX_NGRAM and FUZZY_MAX_DISTANCE.

### What licence does Sensitive-lexicon use?

It uses the MIT License, so you may use, modify and distribute the project as long as you keep the copyright and licence notice. The README separately asks users to comply with local laws and platform policies.

## Sources

- [konsheng/Sensitive-lexicon on GitHub](https://github.com/konsheng/Sensitive-lexicon)
- [License: MIT](https://github.com/konsheng/Sensitive-lexicon/blob/main/LICENSE)
- [Project website](https://github.com/konsheng/Sensitive-lexicon)
- [README](https://github.com/konsheng/Sensitive-lexicon/blob/main/README.md)
- [Releases](https://github.com/konsheng/Sensitive-lexicon/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/konsheng-sensitive-lexicon
