Model or dataset
amirshnll/Persian-Swear-Words avatar
amirshnll/Persian-Swear-Words

Persian-Swear-Words: A JSON Persian profanity list with helpers in eight languages

Persian Swear Dataset - you can use in your production to filter unwanted content. دیتاست کلمات نامناسب و بد فارسی برای فیلتر کردن متن ها

315 stars34 forksC#Apache-2.0

At a glance

What is it?
amirshnll/Persian-Swear-Words ships a Persian (Farsi) wordlist as data.json plus small is_bad / has_swear / filter_words helpers in Java, Go, PHP, Python, JavaScript, TypeScript, C# and Swift. It is a wordlist with wrappers, not a moderation service, and the README says the list is 'to-be-complete'.
Who is it for?
Adopt it if you need a Persian profanity list you can read, edit and ship inside your own codebase, and if you are prepared to treat the list as a starting point rather than a finished filter. Do not adopt it as a hosted moderation API or as a classifier: there is no service, no model and no scoring, only matching.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly C#, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Persian-Swear-Words actually is, and who it is for

This repository is a dataset first and a library second. The payload is data.json, a JSON wordlist of Persian (Farsi) swear words and phrases, with data.txt as a plain-text companion. Around that payload sit small helper implementations in eight languages: Java, Go, PHP, Python, JavaScript, TypeScript, C# and Swift. Each helper exposes roughly the same surface: is_bad for a single word, has_swear for a sentence, filter_words to replace matches, and add_word / add_words / remove_word / remove_words to change the in-memory set.

The intended user is a developer building Persian-language user-generated content: comment sections, chat, forum posts, usernames. The README frames the goal as filtering unwanted content in production. That framing matters, because the project does not do the filtering for you. It gives you the matching primitives and the list, and you decide where to call them. If your product has no Persian text, there is nothing here for you.

The matching mechanism: exact tokens, optional dot-skipping, and a replace step

The core is string matching against the loaded wordlist, not a model and not a score. Two behaviours are visible in the README examples. The first is plain membership: is_bad('خر') is true, is_bad('امروز') is false. The second is the ignoreOT flag, which the Python and JavaScript examples use. With ignoreOT=True, is_bad('خ.ر') returns true while is_bad('ام.روز') returns false. The name reads as 'ignore other things' and the observable effect is that a separator character inside a word does not defeat the match. That is aimed at the common evasion where a writer drops a dot or similar character into a banned word.

The text-level functions are a scan plus a replace. has_swear('تو هیز هستی') is true, has_swear('تو دوست من هستی') is false, and filter_words('تو هیز هستی') returns 'تو * هستی'. filter_words takes an optional replacement argument: the Java and Go examples pass "&" and get 'تو & هستی', the PHP example passes "&" and gets 'تو & هستی'.

Two behaviours are worth flagging before you build on this. The README examples show that has_swear and filter_words do not always agree: in the Python block, filter_words('تو هی.ز هس.تی', ignoreOT=True) returns 'تو * هس.تی', so the second token is left alone even though the first is replaced. The README does not explain the boundary rules. The other is that the list is deliberately blunt. خر and گاو are both listed, and the PHP example walks through removing گاو because it is a normal word for 'cow'. A wordlist built for recall will produce false positives in ordinary conversation.

Installing Persian-Swear-Words and running a first filter

The README gives two package installs: composer for PHP and npm for JavaScript. The npm package publishes PersianSwear.js, data.json and data.txt, per package.json. The composer command is:

bash
composer require amirshnll/persian-swear-words

The npm command is:

bash
npm i persian-swear-words

For Python, Java, Go, C#, Swift and TypeScript there is no package install in the README. You copy the corresponding file (PersianSwear.py, PersianSwear.java, PersianSwear.go, PersianSwear.swift) or the PersianSwear-CSharp and PersianSwear-TypeScript directories into your project. The TypeScript directory has a dist folder and the README's TypeScript example is written as Jest tests importing from ../src, so the TypeScript side is the most build-tooling-dependent of the eight.

A first run in Python, following the README's own example, loads the class and checks a sentence:

python
persianswear = PersianSwear()

print(persianswear.has_swear('تو دوست من هستی' , ignoreOT=False )) # False

print(persianswear.has_swear('تو هیز هستی' , ignoreOT=False )) # True

print(persianswear.filter_words('تو هیز هستی' , ignoreOT=False )) # تو * هستی

After that, the obvious next step is to add your own domain words, because the README treats the shipped list as a base. The Python and PHP examples both show add_word followed by a fresh is_bad check, and the README's Persian note asks users to customise the dataset for their own use case. If you would rather not use any of the helpers, data.json is the whole product and you can load it however you like.

Where the wordlist approach breaks down

The README calls this a 'to-be-complete list', and that phrasing is honest about the main limitation. Coverage is whatever contributors have added. A wordlist cannot catch a misspelling it has never seen, a transliteration into Latin script, or a phrase assembled from individually innocent words. The ignoreOT flag widens matching for one kind of evasion, a separator inside a word, but the README documents nothing about spacing variants, character substitution, or zero-width characters. If your abuse pattern is creative misspelling rather than the plain words in data.json, this project is the wrong tool and you want a trained classifier instead.

False positives are the second cost, and they are structural rather than a bug. خر and گاو appear in the list and the README's own examples treat them as swear words; گاو means 'cow'. Any product where livestock, food or idioms appear in normal text will need a custom removal pass, which the API supports through remove_word. There is also no severity or context signal: is_bad returns a boolean, so you cannot route a mild word to a warning and a severe one to a block without maintaining your own second list.

The third limitation is operational. There is no service, no API endpoint and no rate limiting, because there is nothing to host. The matching runs in your process against a list you ship. That is a feature for latency and data residency, and a cost for anyone who wanted a drop-in moderation endpoint.

How it compares with a general moderation API

The natural alternative is a hosted content-moderation API from a cloud provider or a dedicated vendor. The difference in approach is not the language coverage, it is who owns the decision. A hosted API takes text and returns labels and confidence scores, and the vendor updates the model behind it without you shipping anything. Persian-Swear-Words takes text and returns a boolean from a list you can open in an editor. You get auditability and no network call; you give up recall on novel phrasing and you take on the maintenance of the list yourself.

There is a middle option worth naming: a general-purpose profanity library such as better-profanity, which also uses wordlists but ships English-centric dictionaries and its own matching rules. The practical difference here is the Persian payload and the eight language ports. If your stack is Python-only and your text is English, better-profanity is the shorter path. If your text is Persian and your services are split across Go and PHP, having the same list and the same four functions in both is the reason to pick this repository.

One thing this project is not is a translation resource. The related searches around this topic lean toward meaning, pronunciation and translation of Persian swear words, and data.json is a flat wordlist. It does not carry glosses, transliterations or explanations, so it will not answer those questions.

Maintenance, releases and the Apache-2.0 terms

The last push to the default branch was on 2026-03-26, and the most recent release is 3.1.0 from 2026-02-21, following 3.0.0 in September 2024 and 2.2.0 in August 2024. The repository is not archived. The release history shows the list does change, and 3.1.0 landed about a month before the last commit, so upgrades are real events rather than a frozen file. Note that package.json in the repository still declares version 3.0.0 while the releases page shows 3.1.0; if you pin by package version, check which one you actually resolved.

The upgrade cost is mostly review time. A new release can add words that fire on legitimate text in your product, and the README gives no changelog of which words were added. The safe pattern is to keep your own removal list in code, apply it after loading data.json, and re-run your false-positive samples whenever you bump the version. That is a few lines with remove_word or remove_words, and it survives list updates.

The licence is Apache-2.0, which permits commercial use and modification and requires that you keep the licence and attribution notices. The README also lists a DOI, 10.34740/kaggle/dsv/2094967, pointing at a Kaggle dataset page, and a homepage on Kaggle. If you redistribute the list inside a product, read LICENSE and the Kaggle terms yourself rather than relying on a summary; this is not legal advice. Contributing is handled through CONTRIBUTING.md, and the README's Persian note asks for substantive contributions rather than many small pull requests.

Editorial conclusion

Adopt it if you need a Persian profanity list you can read, edit and ship inside your own codebase, and if you are prepared to treat the list as a starting point rather than a finished filter. Do not adopt it as a hosted moderation API or as a classifier: there is no service, no model and no scoring, only matching. Before you ship, read data.json end to end, decide whether ignoreOT matches the evasion you actually see, and check how the false positives in the README examples (خر, گاو) sit against your own content. The list is versioned by release, so pin the version you audited and re-audit on upgrade.

Frequently asked questions

How do I install Persian-Swear-Words?

The README gives two package installs: composer require amirshnll/persian-swear-words for PHP and npm i persian-swear-words for JavaScript. For the other languages you copy the matching source file, such as PersianSwear.py or PersianSwear.go, into your project.

What is the #1 cuss word?

The repository does not rank or annotate its entries, so it cannot say which word is the most severe. data.json is a flat list of Persian swear words and phrases with no severity field, and is_bad returns only true or false.

How does Persian-Swear-Words handle a swear word written with a dot inside it?

The Python and JavaScript examples use an ignoreOT flag. With ignoreOT=True, is_bad('خ.ر') returns true while is_bad('ام.روز') returns false, so the separator does not defeat the match. The README does not document the exact boundary rules beyond these examples.

Can Persian-Swear-Words filter swear words in text, not just single words?

Yes. has_swear returns whether a sentence contains a listed word, and filter_words replaces matches, so filter_words('تو هیز هستی') returns 'تو * هستی'. Both accept an optional replacement character, such as "&" in the Java, Go and PHP examples.

Official sources

  1. amirshnll/Persian-Swear-Words on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/amirshnll-persian-swear-words.svg)](https://hysenlabs.com/projects/amirshnll-persian-swear-words)