The Big List of Naughty Strings ships four copies of the same data
GitHub describes it as The Big List of Naughty Strings is a list of strings which have a high probability of causing issues when used as user-input data.. The repository metadata lists Python as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.
At a glance
- What is it?
- A corpus of user-input strings for QA, published as a commented text file plus JSON, base64 variants and a package manifest, kept in sync by hand with no script that checks it. The list is deliberately missing whole classes of hostile input, which is a design choice you inherit when you test with it.
- Who is it for?
- Use blns.json when your harness can load a JSON array, and treat the file as a fixed snapshot dated to the 2024-04-18 push rather than a feed that follows your software. Before wiring it into a suite, confirm your four copies still agree, decide explicitly whether you need the excluded inputs such as a null byte, and keep every string inside software you own, because the README's own disclaimer puts third-party targets outside its intended use.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Probably not. The repository last received commits 29 months ago, on April 18, 2024.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 23, 2026, and from our analysis. They are not legal advice.
Editorial analysis
blns.txt carries the section comments, blns.json is a bare array
Two shapes of the same data ship in the repository, and picking the wrong one produces a broken harness. blns.txt is newline-delimited, with comments introduced by # that divide the strings into sections for manual reading and for copy and paste into input forms. blns.json is the machine-facing form: an array with all of the comments stripped out, generated by a Python script that lives in the scripts folder. If your test loads blns.txt line by line, every one of those section comments becomes a test case, and your QA report fills up with lines that were never strings. If you read blns.json instead, you lose the section labels, which are the only thing telling a human reading a failure report which family of input produced it. The JSON is the right default for automation, the text file is the right default for a person poking at a form.
One added string means four files edited by hand, with nothing checking
The contribution rule is blunt: when adding or removing a string, update all files in the same pull request. The top-level tree holds blns.txt, blns.json, blns.base64.txt and blns.base64.json, so a single insertion is four edits in two encodings, one of which is a data blob nobody reads by eye. Nothing in the repository enforces the rule. package.json has no scripts block, no dependencies and no devDependencies, so there is no test command, no generator exposed as an npm script and no continuous integration configuration in the tree to diff blns.json against blns.txt. The only automation mentioned anywhere is the Python script in scripts/ that produces blns.json. The consequence for a consumer is silent drift: you pin a commit, pull blns.json, and never learn that the file disagrees with the text version or that a base64 consumer is reading a stale copy.
The 255 character cap and the null byte ban remove inputs you may need
The list states what it will not accept, and the exclusions are the interesting part. Pull requests with strings of 255 or more characters are refused because they make the file hard to view, which sets a deliberate size ceiling on the corpus. The EICAR test string is named as unacceptable because it can cause the file to be flagged by antivirus scanners, and anything that alters the encoding of blns.txt is out for the same reason. A null character, U+0000, is excluded specifically because it changes the file format on GitHub to binary and renders the file unreadable in pull requests. So a QA suite built from this list will never feed your application a null byte, will never see the long-payload truncation path exercised by a multi-hundred-character string, and will not trip the antivirus scanner. Manual usability won, and the missing coverage is now your problem.
The base64 copies sit in the tree with no explanation of their purpose
blns.base64.txt and blns.base64.json exist beside the plain files, and no document explains them. The usage section describes blns.txt and blns.json only, and the contribution section refers to updating all files without saying what the base64 pair is for or which script produces it. That gap matters for a consumer deciding what to read. If you were planning to use a base64 copy, you cannot tell whether it is current, whether it is generated or hand-maintained, or what downstream tool expects it. If you are adding a string, you have no way to confirm you regenerated the right thing, and a base64 consumer reading a stale copy will quietly test a smaller corpus than you think. Treat the pair as undocumented and verify it against blns.txt yourself before relying on it.
The disclaimer is an ownership test, and the list is not a security test
The project states its own boundary in plain terms: the list is intended for software you own and manage. Some of the strings can indicate security vulnerabilities, and feeding such strings to third-party software may be a crime, with the maintainer declining responsibility for negative actions that follow. The second half of the disclaimer is the one teams tend to skip. It states that the list is not a fully comprehensive substitute for formal security or penetration testing. So it is a QA input corpus, not an attack corpus: it helps you find the input your validation forgot, and it does not stand in for someone trying to break your service. The correct use is a pass over your own form, API or import path, not a probe aimed at a system you do not run.
package.json points main at blns.json but the npm packages live elsewhere
The repository carries a package.json named big-list-of-naughty-strings at version 1.0.0, with main set to blns.json, license MIT and a bugs URL pointing at the issue tracker. Reading that file alone suggests a consumable Node package. The library section undercuts it, listing npm packages blns and big-list-of-naughty-strings and stating that those implementations are maintained by outside parties. Nothing here states that the npm artifact is published from this repository, and there are no GitHub releases to compare a version against, so 1.0.0 is a static field in a manifest rather than a signal that the corpus changed. A Node consumer pinning 1.0.0 gets no promise about which strings are inside it, and no deprecation signal when the corpus grows. If content freshness matters, pin a commit in this repository rather than a package version.
The .NET, PHP and C++ ports are wrappers somebody else maintains
Five entry points are listed for consuming the data outside a plain text or JSON read, and only two of them are npm packages. The others are a .NET implementation at SimonCropp/NaughtyStrings, a PHP one at mattsparks/blns-php and a C++ one at eliabieri/blnscpp. Each is maintained by an outside party, and the project asks for a pull request to list others rather than vetting them. That arrangement keeps the languages covered without the maintainer writing ports, and it moves the sync burden onto whoever owns each wrapper. No cadence is stated for any of them, and the wrappers are not required to mirror this repository exactly. Before you depend on a port, check when that port last updated against the 2024-04-18 push here, because a wrapper that stopped syncing gives you a corpus with a familiar name and quietly older contents.
The last push was 2024-04-18, so the corpus is a dated snapshot
The repository is not archived, and the last push was on 2024-04-18 to the master branch, with no GitHub releases. A string corpus ages differently from code, because what counts as a nasty string depends on what software exists: new formats, new normalization edge cases and new client stacks appear outside the repository and never enter blns.txt. The social threads the project links run from June 2015 to November 2018, so the visible discussion record ends years before the last commit, and there is no changelog to say what a commit added. If you need a wider set of payloads than one file of user-input strings, SecLists is the common alternative, covering many payload categories at once for security testing rather than one QA-oriented list. The two are not substitutes: this project itself says the list is not a comprehensive replacement for formal security testing.
Editorial conclusion
Use blns.json when your harness can load a JSON array, and treat the file as a fixed snapshot dated to the 2024-04-18 push rather than a feed that follows your software. Before wiring it into a suite, confirm your four copies still agree, decide explicitly whether you need the excluded inputs such as a null byte, and keep every string inside software you own, because the README's own disclaimer puts third-party targets outside its intended use.
Frequently asked questions
What is the Big List of Naughty Strings used for?
It is a list of strings with a high probability of causing issues when used as user-input data, intended for both automated and manual QA testing. The stated example is an input your product rejects with an internal server error when a multi-billion dollar company's automated testing still missed it.
How do I load the Big List of Naughty Strings in a test?
Use blns.json, an array with all the comments stripped out, and submit one element at a time as input. blns.txt is the manual form, newline-delimited with comments introduced by # that separate the strings into sections for reading and copy and paste.
Is it safe to run the Big List of Naughty Strings against other people's software?
The project states the list is intended for software you own and manage, notes that some of the strings can indicate security vulnerabilities, and says using such strings with third-party software may be a crime, with the maintainer not responsible for resulting actions. It also states the list is not a fully comprehensive substitute for formal security or penetration testing.
How do I add my own string to the Big List of Naughty Strings?
Open a pull request adding the string or a new section, and update all files in that same pull request. The project refuses strings of 255 or more characters, the EICAR test string, anything that alters the encoding of blns.txt, and a null character (U+0000) because it turns the file binary on GitHub.
How does the Big List of Naughty Strings differ from SecLists?
SecLists collects many categories of payloads for security testing, while this project is one file of user-input strings aimed at QA on your own product. The project states it is not a comprehensive substitute for formal security or penetration testing, so the two serve different steps in the same process.