Self-hosted service
benkeen/generatedata avatar
benkeen/generatedata

generatedata: an engine with three extension points

A powerful, feature-rich, random test data generator.

2,285 stars610 forksTypeScriptGPL-3.0

At a glance

What is it?
The downloadable version of generatedata.com describes itself as an engine rather than an application: about thirty data types, twelve export formats and around thirty-two country data sets, and an explicit invitation to extend it any way you want. The three axes are orthogonal, which is why those numbers are the ones worth remembering.
Who is it for?
Adopt generatedata if you want plausible reference data in whatever format your fixtures need, and treat the country data sets as the reason to use it rather than a placeholder generator. Read V5_CHANGES.md before you move off 4.x, because the 5.0 line has been in beta since March 2026 and the rearchitecture is described as major.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 50 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Three axes, and the numbers that matter

The README describes the project as an engine rather than an application, and that word does the work. It is not a fixture library with a fixed API; it is a generator with three separate things you can extend.

The first axis is data types, the things it can generate. The count given is about thirty, which is the ceiling of that axis. The second is export types, the formats it can write, given as twelve and illustrated with CSV, SQL and JSON. The third is data sets, described as around thirty-two for specific countries, providing city names and regions and the like.

Those axes are orthogonal, which is why the numbers mean something. A new data type multiplies against every export format and every country's data set, which is why adding one is worth more than its own count suggests. And the third axis is the one that separates this from a placeholder generator.

That distinction is worth being concrete about, because it changes what your fixtures are good for. A generator that produces random strings gives you a row that has the right shape and no meaning, which is enough to exercise a parser and useless for anything that displays the data, sorts it, groups by country, or renders a name in an address. Thirty-two country data sets mean the generated rows have plausible places in them, so a test that filters by region or renders an address is testing something.

The invitation to extend it in any way you want is the third thing, and it is the reason this is described as an engine. Thirty data types is not a ceiling, it is what shipped.

The practical consequence for an evaluator is that the feature list is not the interesting number. The interesting question is how many axes you will need to add to, and for most projects the answer is one: a new data type, or a new country.

Six months into a beta, with a changes document

Version 5 is described as a major rearchitecture of the script, and the release history shows what that has cost in time.

Two of the early betas shipped within about an hour of each other, at the end of March 2026, and the fourth came in July. The last push was on 2026-08-15. So the 5.0 line has been in beta for roughly six months, with four published betas and no stable release, on a project whose 4.x line people are presumably running in production.

That is a normal pattern for a rewrite and it is still a decision for the adopter. The tool is a test utility, so the blast radius of a major version is your fixture generation, and the migration is a configuration change rather than a code change if you were only ever calling the engine. But if you extended it, which the README invites, then your extensions were written against the 4.x engine and the 5.0 engine is a different one.

The repository carries a document dedicated to this, named for the version and its changes, alongside a changelog and separate documents for development and production setup. A dedicated migration document for one version is a good sign, and it is the first file to read if you are deciding whether to upgrade now or in six months.

What it is not is a statement that the migration is trivial. Six months of betas on a rearchitecture of the core suggests the engine's shape is still settling, and beta users are the mechanism by which it settles. If you need a stable tool today, the version line is a legitimate reason to stay where you are.

Three checks before the build starts

The production script is the most instructive thing in the package manifest, and it is instructive because of what happens before the build.

json
"prod": "run-s precheck extract-config check-ports buildCmd prodUp",

Five steps in order. A precheck script, an extract-configuration step, a port check, the actual build, and then bringing the stack up. The build is the fourth thing that happens.

That ordering is a decision about where failures are cheap to diagnose. A port conflict is the classic Docker Compose failure: the compose command exits with something unhelpful, and the cause is a process on your machine that you did not know was running. Checking it in a named step before building turns that into a failure with a name, and it does so before you have spent two minutes compiling. The same argument applies to the configuration extraction and the environment precheck, which both fail early and legibly.

The final step carries two details worth copying:

json
"prodUp": "COMPOSE_HTTP_TIMEOUT=150 docker compose --env-file .env.prod -f docker-compose.prod.yml up --build",

The compose HTTP timeout is raised, which is the fix for a build that dies halfway through pulling images on a slow or throttled connection. And the environment comes from a file passed explicitly rather than from whatever is in the shell, which is what stops a developer's local variables leaking into a production-shaped run.

The development script mirrors the structure, with the precheck in front and a filtered parallel task run behind it, so the fast path has the same guard as the slow one. That consistency is the point: the check that catches a misconfiguration should run in both modes or it will only run in one.

The database bootstrap is a sentence in a scripts block

One entry in the scripts block is not a command, and it is worth reading carefully because it describes a step that has not been written.

json
"initApp": "-- build config, create database, output some sort of message summarizing what just happened + handle re-runs gracefully --",

That is the whole value: a sentence in the form of a comment, wrapped in dashes so it will not execute. It describes four responsibilities, building configuration, creating the database, reporting what happened, and handling the case where it is run a second time.

So somewhere between the 4.x line and the 5.0 rearchitecture, application initialisation stopped being implemented and became a note about what it should do. That is an ordinary kind of debt and it is usually invisible until someone tries to deploy from the manifest rather than from the production document, at which point the missing step surfaces as a database that does not exist.

The last of the four responsibilities is the one that shows the author was thinking about the failure mode rather than just the happy path. Handling re-runs gracefully is the difference between a bootstrap you can run on every deploy and one you run once and then maintain. It is a reasonable requirement to have written down.

For an evaluator, the practical note is that the documented installation path is the production document, not the scripts block. The README points at one for installation and the other exists for development. If you are deploying, follow the former and treat the manifest as a development convenience.

A destructive command sitting among the build scripts

There is one script in that block that deserves a warning label, and it is not the one you would guess.

json
"dockerCleanup": "docker system prune -a && docker volume rm $(docker volume ls -qf dangling=true)"

Two operations. The first removes all unused Docker objects, not just stopped containers, so every image on the machine that is not currently referenced by a container. The second removes every dangling volume, which is the set of volumes no container is attached to.

On a developer's own machine this is a genuine convenience, and it is especially welcome in a project whose production script rebuilds the stack, because repeated rebuilds accumulate images quickly. The neighbourhood supports that reading: there is a down script and a logs script next to it, which is where lifecycle commands belong.

On a shared machine, or a machine with anything else running in Docker, the same command is destructive in a way that is not obvious from the name. Images for a database or a queue that you are not currently running will go, and a dangling volume may be the only copy of something. The name reads like routine housekeeping, and that is exactly the problem.

None of this is a criticism of the project; it is a note about npm scripts in general. They are convenient to write and easy to run without reading, which is why a maintenance command and a build command sharing a block deserves a second look before you run it on anything you did not build yourself.

Node 24, declared three times

The requirements are short: Docker, Node 24 with a version manager suggested, and pnpm. What is interesting is how firmly the Node version is pinned in the project itself.

It appears in an engines block requiring Node 24 or newer and pnpm at a specific version, in a package manager field naming the same pnpm version, and in a version file at the repository root for the version manager to read. There is also an enforcement flag that turns those requirements into a hard failure at install time rather than a warning.

Triple declaration is what makes a JavaScript project reproducible across four developers' machines and a build agent, and the enforcement flag is the part that makes it stick. A version requirement that only warns is a version requirement that will be violated on somebody's machine in March.

The workspace layout follows the current convention. There is a workspace file, a lock file, and separate directories for applications and for packages, with a task runner at the root dispatching to both. The development script filters to one package and appends the marker that expands the filter to include its dependencies, which is the standard way to run a client and everything it needs without starting the server.

The lint and format configuration is split across ignore files and a shared configuration, which is what a monorepo looks like once more than one package needs the same rules. And the cleanup of the metadata is handled with a dedicated script rather than left to whichever release happened to be current, which for a project shipping as a downloadable application rather than a library is a deliberate and slightly unusual choice.

Two licences, one of them in the wrong field

There is a small factual problem in the package manifest, and it is the kind that matters.

The package description says the tool is free and MIT-licensed. The licence field in the same file says GPL-3.0-plus. The README says the script is freely available under the GPL 3 licence and adds that all contributors agree that all code is released under that licence.

So the declared licence is the copyleft one, in three places, and the package's own one-line description is the outlier. Those two fields are both displayed by a package index, which means the metadata a reader sees contradicts itself, and for anyone deciding whether they can vendor this into a closed-source test harness, the description is the field most likely to be read first.

The contributor point deserves its own note, because it is unusual to see stated plainly. The README says all contributors agree that all code is released under the licence, which is a licensing position taken on behalf of every contributor rather than left ambiguous. For a project that has existed long enough to have accumulated a long contributor list, that removes a real question about who can grant what.

Neither point changes whether the tool is worth using. The licence for a self-hosted test data generator is a normal choice either way, and the fix is a one-line change in a manifest. It is on the list because metadata contradictions are cheap to fix and annoying to discover later, and because a reviewer skimming a package listing will read the description before the licence.

Editorial conclusion

Adopt generatedata if you want plausible reference data in whatever format your fixtures need, and treat the country data sets as the reason to use it rather than a placeholder generator. Read V5_CHANGES.md before you move off 4.x, because the 5.0 line has been in beta since March 2026 and the rearchitecture is described as major. Check the manifest's licence field before you embed it, since the package description claims MIT while the declared licence is GPL-3.0-plus. And do not run the docker cleanup script on a shared machine without reading it first.

Frequently asked questions

What is generatedata and what does it generate?

It is the downloadable, self-hosted version of generatedata.com, described as an engine that generates random data in any format. It ships with around 30 data types, 12 export formats such as CSV, SQL and JSON, and around 32 data sets for specific countries providing city names and regions.

What are the requirements for running generatedata?

Docker, Node 24 using a version manager, and pnpm. The project pins those versions itself, declaring Node 24 or newer and a specific pnpm version in an engines block, again as the package manager, and again in a version file, with enforcement turned on so the requirement fails the install rather than warning.

Is generatedata 5.0 stable?

Not yet. Version 5 is described as a major rearchitecture and has been in beta since March 2026, with the fourth beta in July and the last push in August. There is a document dedicated to the version's changes in the repository, which is the place to start if you are upgrading from 4.x.

What licence is generatedata released under?

GPL 3. The README states the script is freely available under that licence and that all contributors agree their code is released under it. Note that the package manifest's own description says MIT while its licence field says GPL-3.0-plus, so the declared licence field is the authoritative one.

How do I deploy generatedata?

Through the production script, which runs an environment precheck, extracts configuration, checks that the required ports are free, builds, and then brings up the compose stack with a production environment file and a raised compose HTTP timeout. The repository's production document holds the full installation instructions.

Official sources

  1. benkeen/generatedata on GitHub
  2. License: GPL-3.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/benkeen-generatedata.svg)](https://hysenlabs.com/projects/benkeen-generatedata)