json-schema-faker: generating fake data from JSON Schema in JavaScript
JSON-Schema + fake data generators
At a glance
- What is it?
- json-schema-faker turns a JSON Schema document into a valid instance, with a seeded PRNG, a synchronous API for in-memory schemas, and a CLI. It suits JavaScript teams that already treat JSON Schema as the source of truth for their fixtures.
- Who is it for?
- Adopt json-schema-faker if your schemas already live in the repository and you want fixtures that stay valid as those schemas change; install it with npm install json-schema-faker and start from generate(). Skip it if you need a general-purpose faker library with rich locale data, or if your schemas resolve $ref over the network and you were counting on generateSync().
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 61 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap json-schema-faker fills between a schema and a usable instance
A JSON Schema describes what a valid document looks like. It does not produce one. Teams that keep schemas in the repository for validation still hand-write fixture objects next to them, and those fixtures go stale the moment a required property is added. json-schema-faker reads the schema and generates an instance that satisfies it, so the fixture and the contract come from the same file.
The audience is narrow but real. This is a JavaScript and TypeScript package: package.json sets "type": "module" and exposes ./dist/index.js with types at ./dist/index.d.ts. If your test suite is Node-based and your schemas are already JSON Schema, the fit is direct. If you are looking for a general-purpose fake data library with names, addresses and locale-specific data, this is not that tool. The README points at faker and chance as extensions rather than bundling them, and the build marks both as external. The package description in package.json is the whole promise: "Generate valid JSON data from JSON Schema definitions."
How generation works: seeded PRNG, schema traversal, and extension hooks
The generator walks the schema and emits a value for each keyword it understands. Randomness comes from a seeded PRNG, which the README names as Mulberry32. That matters more than it sounds: with a fixed seed the same schema produces the same output, so a failing test can be reproduced. createGenerator() and createGeneratorSync() take a seed and auto-increment it on each call, which gives you repeatable runs without every call returning identical data.
Coverage is broad. The README lists all standard types and the composition keywords allOf, anyOf, oneOf, not, and if/then/else under JSON Schema 2020-12, plus dependentSchemas and dependentRequired for conditional properties. $ref resolution includes cycle detection and a configurable recursion depth, and createRemoteResolver() handles HTTP and file refs. Built-in format generators cover date-time, email, uri, hostname, ipv4, ipv6, uuid and json-pointer. Regex pattern support is described as AST-based rather than a naive character-class shuffle.
Two extension points carry most of the customisation. registerFormat(name, generator) adds a format globally, and define(name, callback) adds a keyword. The define callback receives the keyword value, the schema node and a GenerateContext, and the README documents ctx.path as the schema traversal path and ctx.outputPath as the final generated-output JSON pointer. ctx.proceed(subschema, overrides?) lets an extension recurse back through the normal generator. reset(name?) clears registered keywords, either one or all.
One detail worth noting from the release history: v0.6.3 changed composition merging so nested allOf branches are preserved instead of the earlier branch being discarded. The README shows a schema with two nested allOf branches, each contributing a const property and a required entry, and states that both are now retained. If you pinned an older version and relied on the previous merge behaviour, that is a behavioural change, not a bug fix you can ignore.
Installing json-schema-faker and generating your first object
The README gives two install commands, npm and bun. Both pull the same published package.
npm install json-schema-faker
# or
bun add json-schema-fakerAfter that, import generate and pass a schema. The README's opening example is a small object with a name, an age, an email and a capped tag array, with name and email marked required. The call returns a Promise, so it must be awaited.
import { generate } from "json-schema-faker";
const data = await generate({
type: "object",
properties: {
name: { type: "string", minLength: 2 },
age: { type: "integer", minimum: 0, maximum: 120 },
email: { type: "string", format: "email" },
tags: { type: "array", items: { type: "string" }, maxItems: 3 }
},
required: ["name", "email"]
});The result is an object whose email field is a string matching the email format and whose tags array holds at most three strings. For a deterministic run, createGenerator takes a seed and advances it on each call, so the first call uses seed 42 and the second uses 43.
const gen = createGenerator({ seed: 42 });
const a = await gen.generate(schema); // seed 42
const b = await gen.generate(schema); // seed 43If your schemas are in memory and you do not want to await anything, generateSync() returns the value directly. The README is explicit about the boundary: a remote $ref that is not already registered and has no refResolver throws, and an async refResolver, extension or outputTransform causes generateSync() to throw immediately with messages such as Cannot use async refResolver in generateSync(). A refResolver that returns plain values, for example one reading from an in-memory map, works in sync mode. There is also a jsf binary declared in package.json, so the CLI is available once the package is installed.
Where json-schema-faker is the wrong tool
The synchronous API is the sharpest constraint. generateSync() is documented as supporting only schemas that are fully resolvable without I/O. If your schema tree leans on remote $refs that are not pre-registered, sync mode is closed to you and you are back to the Promise-based generate(). That is a design boundary, not a defect, but it decides the shape of your test helpers before you write them.
Extension authors pay a second cost. Because define() callbacks receive a GenerateContext and can call ctx.proceed(), a custom keyword can recurse through the generator. That is flexible, and it also means an extension can be written in a way that only works under generate() and not under generateSync(). The README warns about this in the sync section rather than in the define() section, which is the wrong place for the warning if you are reading the API top to bottom.
The third limitation is the one the README never states. The package description is "Generate valid JSON data from JSON Schema definitions", and the feature list covers types, composition, formats, patterns and refs. It does not promise semantic plausibility. A generated email is a string in email format; it is not a realistic address. If your tests assert on business rules over the data rather than on shape and type, this generator will produce values that satisfy the schema and nothing else. Teams that need realistic-looking values usually combine it with faker or chance through the documented extension support, which means adding a dependency the core package deliberately does not carry.
json-schema-faker against faker and chance
The obvious alternative is a fake data library such as faker or chance, and the difference in approach is the direction of the dependency. With faker you write the shape yourself and call faker.person.fullName() or faker.internet.email() at each field. The generator is the human, and the schema, if one exists, is a separate artefact that may or may not match what you produced. With json-schema-faker the schema drives generation: you describe the object once and the library fills it in.
That trade favours schema-driven projects. If you already validate request and response bodies against JSON Schema, generating fixtures from the same documents removes a class of drift, because a new required field breaks generation until the schema is satisfied. A faker-based fixture would simply omit the field and pass until something downstream rejects it. The cost is control: faker gives you locale-aware, human-readable values and a large catalogue of methods, while json-schema-faker gives you whatever the schema permits. The README treats them as complementary rather than competing, listing extension support for faker and chance as a feature. The practical setup is json-schema-faker for structure and an extension for the fields where realistic values matter.
Maintenance, licence and what an upgrade actually costs
The repository is not archived, and the last push was on 2026-08-01. The release cadence visible in the history is roughly every two months: v0.6.1 on 2026-04-07, v0.6.2 on 2026-05-25, v0.6.3 on 2026-08-01. The version in package.json is 0.6.3, matching the most recent release. This is a pre-1.0 package, so the maintainers have not committed to a stable public API by semver convention, and a minor bump can carry a behavioural change. The allOf merge fix in v0.6.3 is exactly that kind of change.
There is a MIGRATION.md file at the repository root, which is where upgrade instructions live. The README itself does not document rollback, and it does not list breaking changes per release, so the migration file is the document to read before moving between 0.6.x versions. The Makefile shows the maintainer workflow: make test runs bun test, make test_all sets NO_SKIP=1 to include skipped tests, make typecheck runs tsc --noEmit, and make ci chains typecheck, test_all and build. If you vendor a fork, those targets are the ones to run.
On licensing, the package is MIT, declared in both package.json and the LICENSE file. MIT permits commercial use and modification with the copyright notice retained. That is a summary of the identifier, not legal advice; if your organisation has a policy on dependency licences, the LICENSE file is the authoritative text. Note also that the README asks for sponsorship through GitHub Sponsors. Sponsorship is optional and does not change the MIT terms.
Editorial conclusion
Adopt json-schema-faker if your schemas already live in the repository and you want fixtures that stay valid as those schemas change; install it with npm install json-schema-faker and start from generate(). Skip it if you need a general-purpose faker library with rich locale data, or if your schemas resolve $ref over the network and you were counting on generateSync(). Before committing, check the Options section of the README on the master branch, since the version published to npm and the documentation can drift between releases.
Frequently asked questions
Is there a schema generator for JSON?
json-schema-faker goes the other direction: it takes an existing JSON Schema and generates data that satisfies it. The README lists support for JSON Schema 2019-09 and 2020-12, including allOf, anyOf, oneOf, not and if/then/else.
How do I use json-schema-faker in JavaScript?
Import generate from the package, pass a schema and await the result. The README's example builds an object with name, age, email and tags properties and marks name and email as required.
What is the difference between JSON and JSON Schema?
A JSON document is data; a JSON Schema is the definition that describes what valid data looks like, including types, required properties and constraints such as minimum and maxItems. json-schema-faker reads the second and produces the first.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/json-schema-faker-json-schema-faker)