datamodel-code-generator: Pydantic v2 Models from OpenAPI, JSON Schema and More
Project brief: Generate Pydantic v2 models, dataclasses, TypedDict, and msgspec.Struct from OpenAPI, JSON Schema, GraphQL, Avro, Protobuf, and raw JSON/YAML/CSV.
At a glance
- What is it?
- It turns schema definitions into Python models with one CLI call, and the hard part is not generation but choosing the preset, formatter and output style you can live with in version control.
- Who is it for?
- Adopt it if you keep a schema as the source of truth and want generated Pydantic v2, dataclass, TypedDict or msgspec modules checked into a repository, and if you are willing to pin the generator version and regenerate on every schema change. Do not adopt it if you need a stable, frozen output format across upgrades, because the project labels itself Development Status 4 - Beta and the README shows preset names carrying a date.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The hand-written model problem this CLI removes
If a service publishes an OpenAPI document, someone still has to write the Python classes that mirror it. That work is mechanical, and it drifts: a field is renamed in the spec, the model keeps the old name, and the mismatch surfaces at runtime instead of at review time. datamodel-code-generator exists to delete that step. You point datamodel-codegen at a schema file and it writes a Python module. The README lists the accepted inputs: OpenAPI 3, AsyncAPI, JSON Schema, Apache Avro, XML Schema, Protocol Buffers and gRPC, GraphQL, MCP tool schemas, and raw JSON, YAML or CSV. The output side is equally explicit: Pydantic v2, Pydantic v2 dataclass, plain dataclasses, TypedDict, or msgspec. The audience is Python teams that already treat a schema as the contract, particularly FastAPI projects where the Pydantic models are the runtime validation layer and not just type hints. A second, less obvious audience is teams migrating a hand-written Pydantic v1 or dataclass module: the README documents --input-model path/to/file.py:ClassName, which reads an existing Pydantic, dataclass or TypedDict class from another Python file and retargets it to a different output type. That is a conversion tool, not just a generator.
How generation actually flows: input type, output type, preset
The mechanism is a single pass with three decisions. First, --input-file-type tells the generator how to parse the source, for example jsonschema or openapi. Second, --output-model-type selects the emitted class style, for example pydantic_v2.BaseModel. Third, --preset bundles a set of options for a target Python version, and preset names encode that version: the README states that py312 means Python 3.12. Schema features that normally break naive generators are handled in this pass: $ref, allOf, oneOf, anyOf, enums and nested types. The emitted module carries a header comment recording the source filename, which is what makes regeneration reviewable in a diff. Formatting is a separate stage. The README says generated Python is currently formatted with black and isort by default, and that --formatters builtin produces standard generated model modules without those external formatter dependencies. That distinction matters more than it looks: the default path pulls black and isort into the environment, while the builtin path keeps the pipeline self-contained but may not match your repository's existing formatting rules. The README also notes that in a future version the default is expected to change, which is a signal to pass the formatter flag explicitly rather than rely on the default.
Install and generate your first model
The README recommends uv for standalone CLI use. This installs the tool outside your project environment, which is the right shape if you invoke it from a Makefile or a CI job rather than importing it.
uv tool install datamodel-code-generatorIf the generator version must be pinned alongside your application dependencies, the README suggests adding it as a development dependency instead. That keeps the version in the lockfile.
uv add --dev datamodel-code-generatorpip and conda-forge are both documented as alternatives, and the README notes community packages for Debian, Ubuntu, nixpkgs and openSUSE Tumbleweed, warning that availability and versions vary by distribution. Once installed, the quick-start command takes a JSON Schema file and writes a Pydantic v2 module. Note that the preset name includes a date, so copy it from the README or the presets documentation rather than guessing it.
datamodel-codegen \
--input schema.json \
--input-file-type jsonschema \
--output-model-type pydantic_v2.BaseModel \
--preset standard-py312-20260826 \
--output model.pyThe README's own example schema declares a Pet object with a required name, an enum of species with a default, an integer age with a minimum, and a boolean. The documented output is a StrEnum subclass named Species plus a Pet class using Annotated field types, a ConfigDict with populate_by_name, and a default of Species.dog. If your generated file instead shows a different enum base class or drops the description strings, you are on a different preset or output type than the quick start. For output that preserves schema-authored names, reuses models and embeds generated documentation, the README points at practical-py312-20260826 rather than the standard preset.
Docker, remote refs and the permission trap
The published image is koxudaxi/datamodel-code-generator, and the Dockerfile shows the entrypoint is datamodel-codegen, so arguments are passed directly to the CLI. Two details from the Dockerfile are worth knowing before you wire it into a pipeline. The image installs the http extra, and it runs as a non-root appuser. The README spells out the consequence: when writing generated files into a bind-mounted directory, that directory must be writable by the container user, or you pass an explicit user such as --user "$(id -u):$(id -g)". This is the most common first failure with the container, and it looks like a generator error when it is a filesystem permission problem. Remote $ref resolution is a separate installation decision. The README states that stable HTTP support comes from the http extra, that the extra is supported and not deprecated, and that an experimental HTTPX2 backend exists behind the httpx2 extra plus --http-backend httpx2. It also states that the experimental extra is not included in datamodel-code-generator[all], so installing all does not silently give you it. If your schema references remote documents, install the http extra explicitly rather than assuming the base package resolves URLs.
Where this tool is the wrong choice
The project classifies itself as Development Status 4 - Beta in pyproject.toml. Treat that as a real constraint, not boilerplate. Preset identifiers embed a date, which means the project's own recommended baseline is versioned by release, and a preset you rely on today is tied to a specific generator version. If your workflow assumes generated files stay byte-identical across tool upgrades, you will spend time on regeneration diffs rather than on features. Pin the generator and regenerate deliberately. A second boundary is the formatter default. Because generated code is currently formatted with black and isort by default, and the README says that default is expected to change in a future version, a repository that never passes --formatters explicitly can see its output shift without any schema change. Third, the tool generates models; it does not generate the HTTP layer, routing, or client code. Teams looking for a full API scaffold will find that the output stops at the data classes. Finally, the README documents the Playground's privacy model in detail: generation runs locally in the browser via Pyodide, schemas and options are not sent to a backend, and shared repro URLs encode state in the URL fragment, which browsers do not send to the server. The README adds the caveat that the full URL can still be stored in browser history or wherever it is shared. That is an honest disclosure, and it also means the Playground is a scratchpad for exploring options, not a place to paste production credentials or internal schemas you would not put in a URL.
Alternatives and the difference in approach
The closest comparison is FastAPI's own tooling, which people search for as Fastapi-code-generator. The difference is direction and scope. FastAPI code generation typically starts from an OpenAPI document and produces a runnable application skeleton: route functions, dependency wiring, and the models alongside them. datamodel-code-generator produces only the models, but accepts far more input formats, including Avro, Protobuf and GraphQL, and can emit TypedDict or msgspec instead of Pydantic. If you want a service scaffold, the FastAPI-oriented generators match that goal. If you want typed models for a data pipeline, a Kafka schema registry, or a gRPC service, this tool covers ground the app scaffolders do not. Compared with generic OpenAPI generators, the distinguishing feature here is the output-type matrix and the --input-model flag that retargets an existing Pydantic, dataclass or TypedDict class. A team sitting on a large hand-written Pydantic v1 codebase gets a path to v2 output without rewriting the class definitions by hand, which is a narrower but more concrete benefit than a general-purpose generator offers.
Licence, maintenance and what upgrades cost you
The licence is MIT, declared in pyproject.toml and in the LICENSE file at the repository root, and the Docker image labels carry org.opencontainers.image.licenses="MIT". MIT is permissive, so generated files carry no copyleft obligation from the generator itself; the usual caveat applies that the licence of the input schema is a separate question, and that is not something the project can settle for you. On maintenance, the last push to main was on 2026-08-29, and releases 0.76.0, 0.75.1 and 0.75.0 all landed between 2026-08-24 and 2026-08-29, so the release cadence around that date was tight. The repository is not archived. The upgrade cost is concentrated in two places: preset names, which carry dates and therefore change with releases, and the formatter default, which the README says is expected to change. Both are handled the same way, by pinning the generator version in your lockfile and passing --preset and --formatters explicitly in your invocation rather than relying on defaults. The README also notes the http extra is supported and not deprecated while the httpx2 backend is experimental and excluded from the all extra, so an upgrade that touches HTTP ref resolution deserves a separate check against the HTTP backend selection documentation.
Editorial conclusion
Adopt it if you keep a schema as the source of truth and want generated Pydantic v2, dataclass, TypedDict or msgspec modules checked into a repository, and if you are willing to pin the generator version and regenerate on every schema change. Do not adopt it if you need a stable, frozen output format across upgrades, because the project labels itself Development Status 4 - Beta and the README shows preset names carrying a date. Before rolling it out, verify the exact command your schema needs: run the quick-start command against one real schema, inspect the emitted header comment, and decide whether --formatters builtin or the default black and isort pipeline fits your CI.
Frequently asked questions
How can I convert a JSON schema to a Pydantic model with datamodel-code-generator?
Run datamodel-codegen with --input pointing at the schema, --input-file-type jsonschema, --output-model-type pydantic_v2.BaseModel, and --output naming the target file. The README's quick start adds a --preset such as standard-py312-20260826 for a Python 3.12 baseline. The generated module starts with a comment recording the source filename.
What is a Pydantic datamodel in the context of datamodel-code-generator?
It is a Python class generated from a schema rather than written by hand. The README shows the output as a pydantic.BaseModel subclass with Annotated field types, a ConfigDict, and enum classes such as StrEnum for schema enums. The project can also emit Pydantic v2 dataclass, plain dataclasses, TypedDict or msgspec instead.
What is a simple code generator and how does datamodel-code-generator work?
It parses the input according to --input-file-type, then emits classes according to --output-model-type, with --preset bundling options for a target Python version. It handles $ref, allOf, oneOf, anyOf, enums and nested types in that pass. Generated Python is currently formatted with black and isort by default, or with --formatters builtin to avoid those dependencies.
How do code generators work in datamodel-code-generator?
The README describes a single pass from a schema file to a Python module: the input file type decides parsing, the output model type decides the class style, and the preset bundles options for a target Python version such as py312 for Python 3.12. The emitted module carries a header comment recording the source filename.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/koxudaxi-datamodel-code-generator)