Model or dataset
mangiucugna/json_repair avatar
mangiucugna/json_repair

json_repair: repairing malformed JSON from LLMs and logs in Python

Repair malformed JSON from LLMs, APIs, logs, and user input in Python.

5,118 stars218 forksPythonMIT

At a glance

What is it?
json_repair is a Python parser that turns broken JSON from LLMs, APIs and logs back into usable objects. It is a small, focused library, and its value depends on whether your input is actually malformed.
Who is it for?
Adopt json_repair if your pipeline regularly receives JSON from an LLM or a third-party API that you do not control, and you want a single call instead of hand-written regex cleanup. Do not adopt it if your input is machine-generated by a serializer you control; strict json.loads is the correct tool there, and the repair parser will silently accept data you would rather reject.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem json_repair targets: JSON that is almost right

A language model asked for structured output usually returns something close to JSON. It may drop a closing bracket, leave a trailing comma, wrap the object in a sentence of prose, or emit Python-style True instead of true. Strict parsers reject all of these, and the failure is not a data problem: the content is there, the syntax is not.

The README states the motivation directly: the author looked for a lightweight Python package that could reliably fix this and could not find one, so he wrote one. The library is aimed at that gap. It is for Python developers who consume JSON they did not produce, whether the source is an LLM response, a third-party API, a log line, or user input.

json_repair is not a validator and not a schema enforcement tool by default. It assumes the payload is worth recovering and does its best to recover it. That assumption is the whole design, and it is also where the trade-offs live.

How the repair parser decides what to keep

The mechanism is visible in the README's handling of edge cases rather than in an architecture diagram. The parser walks the input and repairs as it goes: it adds missing quotes and commas, closes unterminated brackets, strips comments and stray non-JSON characters, and fills missing values with defaults such as null or an empty string.

Some rules are more opinionated than others. Python-style tuples, meaning comma-separated parenthesized sequences, become JSON arrays, but a single parenthesized value stays a scalar. Inside arrays, objects and tuples, true, false, null and None are recognized case-insensitively as JSON booleans or null. That is a deliberate choice to accept input that no strict parser would take.

Performance-wise, the default path tries the standard-library JSON loader first and only falls back to the repair parser when strict parsing fails. The README describes skip_json_loads=True as an explicit tradeoff for inputs you already know are invalid. If you already ran a strict parse, wrapping json_repair in your own try/except around json.loads is redundant work, and the README calls that pattern out as an antipattern.

Installing json_repair and a first repair

The package is on PyPI under the name json-repair, and the README gives a single install command. Python 3.10 or newer is required, according to pyproject.toml.

bash
pip install json-repair

The README's quick example repairs an object with a trailing comma and a missing closing brace. The repaired result drops the trailing comma and closes the object.

python
import json_repair

bad_json = '{"users":[{"name":"Ada","role":"admin",}],"ok":true'
decoded_object = json_repair.loads(bad_json)

# {'users': [{'name': 'Ada', 'role': 'admin'}], 'ok': True}

If you want a string back rather than a decoded object, repair_json returns one. The README notes that a string that is too broken will come back as an empty string, so check for that case rather than assuming a dict.

python
from json_repair import repair_json

good_json_string = repair_json(bad_json_string)

For files, json_repair.load takes a file descriptor and json_repair.from_file takes a path. The README is explicit that IO errors are not caught by the library and remain your responsibility. There is also a live demo at mangiucugna.github.io/json_repair if you want to check a specific payload before writing code.

Where json_repair is the wrong tool

The library repairs by guessing, and guessing has a cost. A missing value becomes null or an empty string. If your downstream code treats a null field as meaningful, the repair has converted a parse error into a silent data error, which is harder to notice than an exception.

The README's own warning about skip_json_loads=True is the clearest boundary: it is only for inputs you already know are invalid. Using it on valid input skips the fast path and sends everything through the repair parser for no benefit.

There is also a text-encoding trap. Non-Latin characters are escaped by default. The README's example shows repair_json("{'test_chinese_ascii':'统一码'}") returning the escaped form \u7edf\u4e00\u7801, and only with ensure_ascii=False does the output keep 统一码. If you are processing Chinese, Japanese or Korean text and do not pass that flag, the output is valid JSON but not what you expected.

Finally, this is not a validator. If your requirement is to reject malformed input rather than accept it, json_repair is pointed the wrong way.

json_repair against strict parsing and schema validation

The natural alternative is the standard library. json.loads is strict, fast, and part of Python, with no dependency to add. The difference in approach is philosophical: json.loads tells you the input is wrong, json_repair tries to make it right. For data you generate yourself, strict parsing is strictly better because it catches bugs at the boundary.

A second comparison is schema validation. The project ships an optional extras group that pulls in jsonschema and pydantic, and the examples directory contains pydantic_schema.py and fastapi_app.py. That suggests a two-stage pattern: repair the syntax, then validate the shape. The repair step makes the text parseable; the schema step decides whether the parsed result is acceptable. Neither replaces the other, and the README does not claim json_repair validates against a schema on its own.

Compared with writing your own regex cleanup, the difference is maintenance. Hand-rolled fixes for trailing commas and missing braces accumulate edge cases; this library centralizes them, at the cost of accepting input you might prefer to reject.

Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-03, which is recent. Releases have been frequent: v0.63.4 on 2026-08-25, v0.63.3 on 2026-08-19 and v0.63.2 on 2026-08-14. The version number still sits below 1.0, and the README describes the library as maintained as a side project with a sponsorship link. Frequent point releases at that cadence mean you should pin a version in your dependency file rather than tracking the latest.

The licence is MIT, declared both in pyproject.toml and as a LICENSE file at the repository root. MIT is permissive, so commercial use is generally unproblematic, but this is not legal advice and your own counsel should confirm anything that matters to you.

The dependency surface is small. The core install has no required runtime dependencies listed in pyproject.toml; jsonschema and pydantic live behind the optional schema extra. Upgrading is therefore mostly about behaviour changes in the repair rules, not about a dependency tree. Because the library accepts malformed input by design, a change in repair rules can change your parsed output without raising an error, so a pinned version plus a test on a real sample of your broken payloads is the cheap insurance.

Editorial conclusion

Adopt json_repair if your pipeline regularly receives JSON from an LLM or a third-party API that you do not control, and you want a single call instead of hand-written regex cleanup. Do not adopt it if your input is machine-generated by a serializer you control; strict json.loads is the correct tool there, and the repair parser will silently accept data you would rather reject. Before wiring it in, verify two things: how it handles your specific failure mode, using the live demo at mangiucugna.github.io/json_repair, and whether you need ensure_ascii=False for non-Latin text, because the default escapes it.

Frequently asked questions

What is json_repair and what does it do?

json_repair is a Python library that repairs malformed JSON from LLMs, APIs, logs and user input. It fixes missing quotes, commas and brackets, strips comments and stray prose, and can complete missing values with defaults.

How do I install json_repair?

The README gives one command: pip install json-repair. Python 3.10 or newer is required according to pyproject.toml, and the package name on PyPI is json-repair.

How do I use json_repair?

Import it and call json_repair.loads on the string, which returns a decoded object, or repair_json if you want a repaired string back. The README notes that a string that is too broken returns an empty string.

What causes a JSON error?

The README lists the failures json_repair handles: missing quotes, misplaced commas, unescaped characters, incomplete key-value pairs, unterminated arrays and objects, and extra non-JSON characters such as comments or prose. These are the syntax problems the library was written to fix.

Official sources

  1. License: MIT
  2. mangiucugna/json_repair on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mangiucugna-json-repair.svg)](https://hysenlabs.com/projects/mangiucugna-json-repair)