# Eight language bindings, two install commands, and a version that skipped three

> A Persian profanity wordlist shipped as two data files plus a helper class per language, with a DOI and a Kaggle page for provenance. The engineering is thin and the packaging is uneven, and the README says outright that production users should customise it.

**amirshnll/Persian-Swear-Words** — Persian Swear Dataset - you can use in your production to filter unwanted content.  دیتاست کلمات نامناسب و بد فارسی برای فیلتر کردن متن ها

- Repository: https://github.com/amirshnll/Persian-Swear-Words
- Website: https://www.kaggle.com/amirshnll/persian-swear-words
- Stars: 315 · Forks: 34
- Language: C#
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/amirshnll-persian-swear-words

## Two data files and eight bindings of different shapes

The data is two files at the repository root, one JSON and one plain text. Everything else in the repository is a wrapper.

Six of the wrappers are single files sitting beside the data: one each for Java, Go, PHP, Python, JavaScript, and Swift. Two are directories with their own build, one for C# and one for TypeScript. So the eight languages are not on equal footing as projects; six are a file you copy and two are something you build.

The API surface is broadly the same idea in every language: check whether a single string is in the list, check whether a longer piece of text contains something in the list, and return the text with matches replaced. Add and remove methods exist in most of them, which means the list is mutable at runtime rather than frozen at build time. The PHP example makes the point explicitly: a multi-word phrase returns false, the same phrase is added, and it then returns true.

The naming is not consistent, which is normal for ported helpers and worth expecting. Java and Go use lower camel case, Python and PHP use snake case, and the C# helper uses Pascal case. The C# version also splits the single-word check from the whole-sentence check into two named methods rather than one.

The Swift binding is the odd one structurally: it exposes a shared singleton, so the list is process-global rather than an instance you construct.

## Two install commands for eight languages

The installation section is short enough to read in one breath and covers exactly two package managers:

```bash
composer require amirshnll/persian-swear-words
```

```bash
npm i persian-swear-words
```

There is no pip instruction, no Maven coordinate, no Go module path, and no Swift Package Manager entry. Six of the eight languages are therefore install-by-copy, and the readme does it implicitly by linking each helper file directly from its usage section rather than by explaining the install.

That is a workable arrangement for a small wrapper set, since the file you need is one click away in every case. It does mean there is no dependency manager pinning the data file version alongside the code version, so a project using three of the eight languages ends up with three copies of the list at three versions.

The npm package narrows this further. Its file list contains the JavaScript helper and the two data files and nothing else, so installing from npm gives you one binding regardless of how many languages the readme lists. The C# and TypeScript directories have their own builds and their own dependency on the data, and neither is published under this package name as far as the manifest shows.

For the other six, the honest description is that you vendor the file.

## The manifest is one release behind and the test script always fails

Two fields in the npm manifest are worth reading before you rely on them.

The version field reads 3.0.0. The current release is 3.1.0. So the published manifest and the release tag disagree, and installing from the registry gives you the older one unless something overrides it.

The test script is the second field:

```json
"scripts": {
  "test": "echo \"Error: no test specified\" && exit 1"
}
```

That is the placeholder the package generator inserts when you do not write a test script. It echoes a message and exits non-zero, which means `npm test` fails by design. There is a tests directory in the repository and there are test files described in the TypeScript section, but they are not wired into the manifest.

So for the JavaScript binding, nothing verifies it in the packaging. If you take that binding, the first thing to do is run the repository's own tests against the file you copied and then fix your own copy of the manifest.

The dependency list, by contrast, is empty, which is correct. The helper is one file with no runtime dependencies, so there is nothing to resolve.

Everything is licensed under Apache 2.0, and the manifest records the author's name, email, and site.

## The readme title contains a zero width joiner and the fence tags are camel case

The first line of the readme is a heading that ends with a formatting artefact. Between the word and the backticked file extension there is a zero width joiner, an invisible character that renders as nothing and occupies no width.

It matters for two practical reasons. A heading that is copied into an issue or a search loses or keeps the character depending on the tool, so the same title matches or does not match itself. And any tooling that strips or normalises whitespace will not remove it, because it is not whitespace.

The code fences have a related habit. The JavaScript example is tagged with a capitalised language name and the TypeScript example with another, rather than the conventional lower-case identifiers. Most highlighters are case-insensitive and cope. Some are not, and a renderer that does not recognise the tag drops to plain text, which is how the TypeScript example ends up unstyled in some viewers.

Both are one-character problems in a document whose other quality problems are spelling. The functions heading is misspelled in the source, and the Swift section heading is misspelled in a way that breaks its anchor link, which means the table of contents entry for it does not resolve.

None of this affects the data. All of it affects the first impression.

## Five of the eight usage examples are cut off mid statement

The usage section is one block per language, and most of the blocks are truncated in the readme as published. The Java block ends after the first filtering call. The Go block ends partway through a method name. The PHP block ends inside a conditional. The Python block ends inside a function call. The Swift block ends after a variable assignment.

Only the JavaScript example is complete, and the TypeScript one is a different thing entirely: it is a unit test file rather than a usage example, with the test runner's imports at the top and a set of assertions.

What this means practically is that the readme gives you the shape of the API in each language and nothing more. You get the method names, the argument order, and one call's result; you do not get a runnable program.

For a wrapper this thin that is survivable, since the implementations are short. But it is also the difference between copying a working example and reconstructing one.

The C# section is the sparsest of all. It has no code block. It tells you to construct a filter object, tells you the constructor accepts an optional path to a JSON file, and then lists four method signatures as one-line fragments with no example of any of them.

So of eight languages, one has a complete example, one has a test file, one has fragments, and five have truncated blocks.

## The notes block says customise it, and the search terms say what is missing

The most useful paragraph in the readme is the one written in Persian, inside a right-to-left block near the top.

It says the dataset contains words that may need filtering in some cases, that users with specific uses should adapt it to their own needs, and that contributions are encouraged. It then asks for substantial contributions rather than small pull requests, and notes that adding a class or a function in other languages using this dataset is possible.

That first sentence is the honest framing and it is the one to hold on to. A profanity list is not a moderation decision; it is a starting vocabulary. The readme says so, in the language of its authors, near the top, where it is easy to miss.

The search terms attached to the project are a useful inventory of what people expect from it and do not get. They include pronunciation, audio, translation into English, and a query about usage in Afghan Persian, which is Dari rather than Iranian Persian and is a different vocabulary in several respects. None of those are in the repository, which contains two flat word lists.

So the gap is not hidden. It is just that the gap lives in what people search for rather than in what the readme promises.

## The repository is classified as C# and carries an IDE folder

The repository's primary language is recorded as C#, which is a reasonable classification given that the C# binding is a directory with its own project, while the other seven are single files or a small package.

It is also the only binding with a project directory, so the classification is following the shape of the code rather than the amount of it.

The root directory tells the same story about how this is worked on. There is a Visual Studio folder committed at the top level, which is an IDE's per-user settings directory and normally belongs in the ignore file. There is a Jupyter notebook at the root next to the source files, so at some point the list was explored interactively and the notebook was kept.

And there is a standalone HTML file at the root, which is almost certainly the demo page, checked in next to the package manifests rather than hosted.

Alongside those sit the two governance files the readme links, a contributing guide and a code of conduct, and a directory of tests that the manifest does not call.

The provenance side is better handled. The dataset has a DOI and a page on a public dataset host, so there is a citable artefact separate from the code, which is the right arrangement for something meant to be reused in production.

## Conclusion

Use this dataset as a starting point, not as a finished moderation list. The readme itself says so, in Persian, in the notes block: it contains words that may need filtering in some cases, and users with specific needs should adapt it. That caveat is the correct reading of a single-maintainer wordlist with no coverage claim for Dari usage, no pronunciation or audio data, and no false-positive testing that the documentation mentions. Two engineering things to fix before you ship it. The npm package's test script exits with an error unconditionally, so nothing verifies the JavaScript helper in that packaging. And the manifest version is one release behind the current tag, so pin by tag rather than by the manifest if you install from a registry. Also note that only two of the eight language bindings have documented installation at all.

## FAQ

### What is the Persian-Swear-Words dataset?

It is a list of Persian profanity terms and phrases published as a JSON file and a plain text file, with a DOI and a page on a public dataset host for provenance. The readme describes it as a list that is still to be completed, intended as a starting point for filtering unwanted content.

### Which programming languages does Persian-Swear-Words support?

Eight: Java, Go, PHP, Python, JavaScript, TypeScript, C#, and Swift. Six are single files to copy, while the C# and TypeScript bindings are directories with their own builds. Naming follows each language's conventions, so method names are camel case, snake case, or Pascal case depending on the port.

### How do I install Persian-Swear-Words?

Only two installers are documented: a composer require for the PHP package and npm i for the JavaScript package. There is no pip, Maven, Go module, or Swift Package Manager instruction, so the other six bindings are installed by copying the single source file into your project.

### Can I add my own words to the Persian-Swear-Words list at runtime?

Yes. Add and remove methods exist in most of the bindings, and the PHP example shows a phrase returning false, being added, and then returning true. The Swift binding exposes the list as a process-wide shared instance rather than an object you construct.

### Is the Persian-Swear-Words list ready for production moderation?

The readme's own Persian notes say it contains words that may need filtering in some cases and that users with specific needs should adapt it. There is also no coverage claim for Dari usage, no pronunciation or audio data, and the npm package's test script exits with an error unconditionally.

## Sources

- [amirshnll/Persian-Swear-Words on GitHub](https://github.com/amirshnll/Persian-Swear-Words)
- [License: Apache-2.0](https://github.com/amirshnll/Persian-Swear-Words/blob/master/LICENSE)
- [Project website](https://www.kaggle.com/amirshnll/persian-swear-words)
- [README](https://github.com/amirshnll/Persian-Swear-Words/blob/master/README.md)
- [Releases](https://github.com/amirshnll/Persian-Swear-Words/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/amirshnll-persian-swear-words
