dariusk/corpora: a CC0 JSON corpus collection for bot prototypes
A collection of small corpuses of interesting data for the creation of bots and similar stuff.
At a glance
- What is it?
- dariusk/corpora ships small JSON word lists for prototyping bots and other odd internet projects. The data is CC0, the files stay under roughly 1000 items, and the project is explicit that it is not a dictionary.
- Who is it for?
- Adopt dariusk/corpora when you need a few hundred or a thousand interesting strings to get a prototype running today, and when you are prepared to swap in a larger source later. Skip it if you need exhaustive coverage, metadata, part-of-speech tags, or a stable API with a versioning policy, because the README points to Wordnik and the MediaWiki API for those needs and the repository is a flat collection of JSON files.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 14 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What dariusk/corpora solves, and who it is for
The README opens with a confession that will sound familiar to anyone who has built a bot: the author kept copy/pasting an adjs.json file from project to project. Corpora exists so that one repository holds those small lists instead of scattering them across half-finished experiments. The stated goal is rapid prototyping. You might start with nouns.json to see whether an idea is any good, then rip it out and replace it with a more complex or exhaustive data source once the prototype proves itself.
The second audience is teaching. The README notes that a three-hour workshop on making Twitter bots leaves no time for students to find, scrape, clean and parse interesting data, so students can be pointed at this repository and pick files to meld together.
That framing matters when you evaluate the project. This is not a data platform. It is a shelf of small JSON files with a narrow, honest purpose: get a prototype off the ground before you invest in real data plumbing.
How the repository is laid out and what the data looks like
The top level holds .editorconfig, .github/, .gitignore, Gruntfile.js, README.md, data/ and package.json. All the content lives under data/, and the README describes the repository as a collection of JSON files meant to be language-neutral. Any language that can parse JSON can read them, which is why the README says the author is happy for someone else to publish an NPM package but that this repository will remain a collection of data files.
There is no schema, no manifest and no index documented in the README beyond the directory itself. You navigate by path. The README gives one concrete size rule: a list of resources should contain somewhere in the vicinity of 1000 items. That rule has a visible consequence. Corpora will not contain complete dictionary-style files; instead it hosts a sampling of 1000 common nouns, adjectives and verbs. Some categories are small enough to be complete anyway, and the README's example is a list of heavily populated U.S. cities that might hold only 75 entries and still be considered complete.
So the data model is deliberately flat: files of strings you can load whole into memory. There is no pagination, no query layer, and no attempt to model relationships between lists.
Cloning the repo and a first real use
There is no install command in the README, because there is nothing to install. You clone the repository and read the JSON. The package.json is for the repository's own tooling (grunt, grunt-jsonlint, grunt-cli as devDependencies) rather than for consumers, so treating it as an installable library is a mistake.
Start by cloning the repository, as its package.json repository URL indicates:
git clone https://github.com/dariusk/corpora.gitAfter that you have a data/ directory of category folders full of .json files. The README does not document a manifest or key convention, so open a file and check its shape before writing code against it. The README does say the files are language-neutral and that any language able to parse JSON can read them, so the loading step in your own project is a standard JSON parse.
If you would rather not vendor the data yourself, the README lists three community tools rather than an install path of its own: corpora-project, described as a Node.js NPM package for accessing corpora data offline; pycorpora, described as a simple Python interface for corpora; and corpora-api, described as a Node.js server that offers up the corpora as a JSON API, which the README says is live at https://corpora-api.glitch.me. Those are separate projects, not part of this repository, and the README does not state their maintenance status.
To check your own data before submitting it, the repository's test script is grunt --verbose, which runs grunt-jsonlint. That is the same check CI applies to pull requests.
The CC0 licence is the real feature here
The README explains the choice directly: since Corpora is more data than code, the author chose CC0 rather than MIT or a similar licence. The text waives copyright and related rights to the extent possible under law, and the contribution guidelines make the same demand of anyone submitting data: by opening a pull request you agree to CC0 being applied, meaning anyone can use the data for any reason without attribution in perpetuity.
For a bot builder this removes the licensing homework that usually accompanies a scraped word list. You can ship the data inside a commercial product without an attribution page. The README notes that if you want credit, your name can be added to the README, while making clear that nobody using the data is obligated to credit you.
The package.json also declares "license": "CC0", so the machine-readable metadata agrees with the prose. That said, this is a description of the project's own terms, not legal advice, and the repository has no LICENSE file at the top level in the listing above, only the README section and the package.json field.
Where Corpora is the wrong tool
The README devotes a section to this, which is unusually candid. Corpora is not meant to replace exhaustive APIs. If you want nouns and you want every noun in English with metadata, the README points to Wordnik. If you want the title of every Wikipedia article, it points to the MediaWiki API.
The size cap is the mechanism behind that boundary. Around 1000 items per file is a design decision, not an accident, and it means any application that needs long-tail coverage will hit a wall. A generator that must avoid repeating itself across thousands of outputs will exhaust a 1000-item list quickly. A search or classification feature that needs labelled examples will find no labels here at all.
The second limitation is structural. Because the repository is a flat set of JSON files with no manifest and no documented key convention, there is no versioning story for individual lists. If a list changes, your prototype sees the change the next time you pull. The README does not document rollback, deprecation or a changelog for the data, and there are no releases, so pinning a commit hash is the only way to freeze what you depend on.
Finally, do not confuse this project with the linguistic meaning of the word. Search traffic for "corpora" mostly concerns corpora cavernosa, corpora amylacea and academic corpus linguistics. This repository is a folder of word lists for bots, and nothing more.
Alternatives and the difference in approach
The README names two alternatives itself, and the contrast is instructive. Wordnik is an API: you query it, it returns nouns with metadata, and the coverage is intended to be exhaustive. Corpora is a static file you download once and read from disk. That difference shows up in latency, in offline behaviour, in rate limits and in whether your prototype works on a plane.
The MediaWiki API follows the same pattern at a larger scale. If your project needs every Wikipedia article title, no local file will do, and the README says so plainly.
Among the community tools, corpora-project, pycorpora and corpora-api are wrappers rather than alternatives: they read the same data through a different interface. If you are working in Node, corpora-project saves you writing the file-loading code. If you want an HTTP endpoint instead of vendored files, corpora-api is the shape you want, with the caveat that it is a separate project the README does not describe in operational detail.
The honest summary is that Corpora competes with copying a word list out of a gist, not with a dictionary service. Its advantage over the gist is that the lists are collected in one place and validated by CI.
Contributing, CI and the maintenance picture
Submissions go through pull requests. The README asks for JSON in a .json file, run through JSONLint before submitting, and notes that CI testing will jsonlint your pull request automatically. If you see a test failure notification after submitting, the JSON is malformed. That CI runs through grunt: package.json defines the test script as grunt --verbose, with grunt-jsonlint among the devDependencies, so the check is a lint pass over the data rather than a semantic validation of content.
The only other contribution rule is the size cap of about 1000 things per file, with fewer being fine.
On activity, the repository is not archived and the last push was on 2026-09-16, which is recent. The absence of releases is worth noting for anyone planning to depend on it: there are no tagged versions, so your upgrade path is a git pull or a pinned commit, and the README does not describe a deprecation process for lists that fall out of use. Because the data is CC0, forking a frozen copy is a legitimate and simple way to insulate yourself from upstream edits.
Editorial conclusion
Adopt dariusk/corpora when you need a few hundred or a thousand interesting strings to get a prototype running today, and when you are prepared to swap in a larger source later. Skip it if you need exhaustive coverage, metadata, part-of-speech tags, or a stable API with a versioning policy, because the README points to Wordnik and the MediaWiki API for those needs and the repository is a flat collection of JSON files. Before you build on it, verify three things: that the specific file you want exists under data/ and is valid JSON, that the CC0 terms suit your redistribution plans, and that the file's size is enough for your use case.
Frequently asked questions
What is dariusk/corpora?
It is a repository of small JSON files holding word lists and similar data, described in the README as a collection of static corpora useful for creating weird internet stuff such as bots. The data is licensed CC0 and is meant to be readable by any language that can parse JSON.
How do I install dariusk/corpora?
There is nothing to install. You clone the repository and read the JSON files under data/, or use one of the community tools the README lists, such as the corpora-project NPM package or pycorpora.
Does dariusk/corpora contain every English noun or adjective?
No. The README states that Corpora will not contain complete dictionary-style files and instead hosts a sampling of about 1000 common nouns, adjectives and verbs, and it points to Wordnik for exhaustive coverage with metadata.
What licence applies to the data in dariusk/corpora?
The README says the author chose CC0 rather than MIT because the project is more data than code, waiving copyright to the extent possible under law, and the contribution guidelines require submitters to agree to the same terms.
How do I submit data to dariusk/corpora?
Open a pull request with your data as JSON in a .json file, run it through JSONLint first, and keep the file to about 1000 things maximum. The README notes that CI will jsonlint the pull request automatically.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dariusk-corpora)