What is Structured output?
Structured output (often called JSON mode) is a way of making a language model return data in a fixed, machine-readable shape instead of free prose. The model is constrained by a schema or format so the response can be parsed by code without guessing.
How structured output works
The mechanism has two layers. At the model layer, the provider is told what shape to produce. In JSON mode, the constraint is usually just that the output must be valid JSON. In schema mode, the constraint is a JSON Schema document that describes required fields, types, enums and nesting. Some providers enforce the schema during decoding, so a token that would break the grammar is rejected. Others only prompt the model and then validate afterwards, which is why malformed output still happens.
At the application layer, the JSON text is parsed and validated. A validator such as Ajv or python-jsonschema checks the parsed object against the same schema. Ajv compiles JSON Schema and JSON Type Definition documents into JavaScript validation functions, so the check runs as ordinary code. python-jsonschema validates JSON documents against JSON Schema in Python, with lazy error iteration and programmatic error inspection. If validation fails, the caller either retries, repairs or rejects the response.
A third pattern skips raw JSON entirely. Instructor turns a Pydantic class into a validated response model for OpenAI, Anthropic, Google, Ollama and Groq clients. The class is the contract; the library handles parsing and retries. Promptify wraps named NLP tasks in Python classes that return Pydantic models, with LiteLLM as the provider layer. In both cases the schema lives in application code, not in a separate document.
JSON Schema itself is specified by json-schema-org/json-schema-spec. That repository is not a validator. It is the Markdown source of the JSON Schema IETF Internet-Drafts, plus the Remark build that turns it into HTML and the meta-schema tests that keep it honest.
When you need structured output, and when you do not
You need it when a downstream program must read the model's answer. Extraction is the clearest case: pulling names, dates or line items out of text and writing them to a database column. Classification is another: the model must return one label from a fixed set, not a paragraph explaining its choice. Tool calling and function calling also depend on the same idea, because the arguments passed to a function have to match its parameter types.
You do not need it for open-ended writing, summarisation or chat. Forcing a schema onto a prose task adds constraint without adding value, and it can make the output worse by removing the model's freedom to explain. If a human reads the answer and no parser touches it, plain text is simpler.
There is also a cost question. Schema-constrained decoding can slow generation and can push the model toward awkward phrasing to satisfy the grammar. A schema with many optional fields and deep nesting is harder for a model to fill correctly than a flat one with few required fields. The practical rule is to keep the schema as small as the task allows.
Finally, structured output is not a substitute for validation. A model can return JSON that matches the schema and still be wrong. The schema guarantees shape, not truth.
Common pitfalls and limits
Malformed JSON is the classic failure. Even with JSON mode, a model may emit trailing commas, unescaped quotes or a partial object if it hits a token limit. Promptify's README does not document how it handles malformed JSON, and its evaluation module's exact measurements are not described in its public documentation, so those remain open questions for anyone adopting it.
Schema mismatch is subtler. A provider may accept a schema but silently ignore parts of it, or support only a subset of JSON Schema keywords. Ajv supports JSON Schema draft-04/06/07/2019-09/2020-12 and JSON Type Definition (RFC8927), but its strict mode and schema-language choices need a decision before you commit. A schema that validates in one tool may not validate in another.
Retries and cost are another limit. Instructor removes manual JSON parsing and retries, but retries still consume tokens and time. A retry loop without a cap can turn a cheap call into an expensive one.
Versioning is easy to overlook. Schemas change, and a stored response validated against an old schema may fail against a new one. json-schema-org/json-schema-spec is versioned as IETF Internet-Drafts, so the spec itself moves. Pinning a draft version is a decision, not a default.
Finally, structured output does not solve hallucination. A model can fill every required field with a plausible but invented value. The schema only tells you the shape is correct.
How it shows up in open-source projects
The projects below are not all model-side tools. Some generate schemas, some validate them, and some turn them into forms or code. That split is the point: structured output is a chain, and each project covers a different link.
Ajv (ajv-validator/ajv) is the validator link. It compiles JSON Schema and JSON Type Definition documents into JavaScript validation functions. It suits API boundary checks and config validation, but its strict mode and schema-language choices need a decision before you commit.
python-jsonschema (python-jsonschema/jsonschema) is the Python validator link. It validates JSON documents against JSON Schema in Python, with lazy error iteration and programmatic error inspection. It is a validator, not a data model, and that distinction decides when you should reach for it.
Instructor (567-labs/instructor) is the model-side link. It turns a Pydantic class into a validated response model for OpenAI, Anthropic, Google, Ollama and Groq clients. It removes manual JSON parsing and retries, but it is an extraction and classification library, not an agent runtime.
Promptify (promptslab/Promptify) is a task wrapper. It wraps named NLP tasks (NER, classification, QA, summarization, SQL generation) in Python classes that return Pydantic models, with LiteLLM as the provider layer. The API is small; the interesting questions are what happens when a model returns malformed JSON and what the evaluation module actually measures.
On the schema-generation side, datamodel-code-generator (koxudaxi/datamodel-code-generator) turns schema definitions into Python models with one CLI call, and the hard part is not generation but choosing the preset, formatter and output style you can live with in version control. jsonschema2pojo (joelittlejohn/jsonschema2pojo) does the same for Java, turning JSON Schema or sample JSON into annotated Java classes for Jackson or Gson. Its build-plugin route is the one that scales; the online generator is the one people try first.
For human-facing forms, rjsf-team/react-jsonschema-form turns a JSON Schema document into a React form, with a validator package and a theme package deciding how the fields look. It fits teams that already own a schema; it is a poor fit for people who want to draw a form by hand. json-editor/json-editor turns a JSON Schema into a working HTML form, with no runtime dependencies and optional integrations for Bootstrap, Tailwind and Spectre. Its README now states the library is in maintenance mode, with active development moved to Jedison.
SchemaStore (SchemaStore/schemastore) hosts JSON schemas that editors and AI clients fetch by URL. The repository itself is a catalog, not a validator, and its MCP server at mcp.schemastore.org is the part worth understanding before you point a tool at it.
Finally, json-schema-org/json-schema-spec is the specification source. It is not a validator. It is the Markdown source of the JSON Schema IETF Internet-Drafts, plus the Remark build that turns it into HTML and the meta-schema tests that keep it honest.
In practice
Structured output is a contract between a model and a parser, and the contract is only as strong as the validation behind it. Pick the link you actually need: a validator such as Ajv or python-jsonschema, a model-side library such as Instructor or Promptify, or a schema generator such as datamodel-code-generator or jsonschema2pojo. Read the JSON Schema specification at json-schema-org/json-schema-spec before writing a schema you intend to keep.