Graphtage: a semantic diff for JSON, YAML, XML and CSV
A semantic diff utility and library for tree-like files such as JSON, JSON5, XML, HTML, YAML, and CSV.
At a glance
- What is it?
- Graphtage compares tree-structured files by rewriting one tree into the other and colouring the edits in place. It is useful when a line diff drowns in reformatting noise, and it is overkill for a one-line config change.
- Who is it for?
- Adopt Graphtage when you compare structured config, data dumps or API payloads and need to see what actually changed rather than which lines moved. Skip it for prose, source code, or any file whose meaning depends on line order, and skip it for very large unordered lists unless you can accept the cost the README describes.
- Can I use it commercially?
- Yes, with conditions. LGPL-3.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Graphtage solves: line diffs lie about structured files
Run a conventional line differ over two JSON documents that differ by one key and you get a wall of changed lines, because the tool is comparing text, not meaning. Graphtage takes the opposite approach. It parses both inputs into trees, computes the edits that turn one tree into the other, and prints the result in the syntax of the original format with the changes marked inline. The README's opening example shows a JSON object where an element is added to a list, a key is renamed and a new key appears; the output keeps the shape of the file and marks the removed and added characters.
The name is a portmanteau of graph and graftage, the horticultural practice of joining two trees so they grow as one, which is a fair description of what the edit computation does. The audience is narrow but real: engineers who review generated configuration, compare API responses across environments, or track changes in data files that are stored as YAML, TOML or CSV rather than as source code. If you have ever looked at a diff of a Kubernetes manifest or a Terraform state file and could not tell whether anything meaningful changed, that is the gap this tool addresses.
How the matching works: an intermediate tree, three dict strategies and list order
Graphtage parses each input into an intermediate representation that is independent of the file format. That is why the README can state that you may diff a JSON file against a YAML file, and why the output format can differ from either input. The comparison then operates on that representation, not on text.
Matching is where the interesting choices live. By default Graphtage pairs key/value entries that share a key and then tries to match the remaining elements against each other. The --dict-strategy flag exposes three behaviours. match tries every pair of keys and values to reach a minimum edit distance and is the most computationally expensive. none only pairs entries with exactly the same key, and is equivalent to the --no-key-edits or -k option. auto, the default, pairs identical keys first and then falls back to match for what remains.
Lists get their own controls. --ignore-list-order treats a list as an unordered collection, so moving an element is not an edit; the README notes that [1, 1, 2] matches [2, 1, 1] but not [1, 2, 2], because duplicates still count. Two node types ignore that option and also ignore --no-list-edits: the rows of a CSV file and the children of an XML or HTML element stay ordered. That is a sensible boundary, since row and element order usually carries meaning in those formats, but it is a boundary you have to remember when you point the tool at a CSV export and wonder why reordering rows shows up as edits.
Graphtage also sorts the keys of every dictionary it reads, so output key order is alphabetical regardless of input order. That makes output deterministic and comparable, at the cost of not reproducing the file you fed it.
Installing Graphtage and running a first semantic diff
Graphtage requires Python 3.10 or later and the README states it is tested against Python 3.10 through 3.14. Installation is a single pip command, which puts two executables on your PATH: graphtage and graphtage-git-diff.
pip3 install graphtageAfter that, point the tool at two files. Type is inferred from the file extension; the README's own example diffs two JSON files named original.json and modified.json.
graphtage original.json modified.jsonThe output is the first file's structure with the edits marked in place. For the README's example, where an element was added to a list, a key was renamed and a new key appeared, you should see the original text with the removed characters struck through and the added characters underlined, rather than a pair of separate hunks.
If your file does not carry a recognised extension, state the format explicitly for each side. Every supported format has both flags: csv, html, ini, json, json5, pickle, plist, toml, xml and yaml.
graphtage --from-json config.txt config.jsonWhen the output is too verbose, the formatting flags compact it. --join-lists removes line breaks after list items, --join-dict-items removes them after key/value pairs, and --condensed applies both.
graphtage --condensed original.json modified.jsonIf you want the changes as a list rather than applied to the input, use --only-edits, or --edit-digest for a shorter human-readable context per edit.
graphtage --only-edits original.json modified.jsonWhere Graphtage gets expensive or simply wrong for the job
The unordered list mode is the clearest limitation, and the README states it plainly rather than hiding it. Two lists whose elements all match each other cost nothing to compare however long they are, but comparing two lists that differ is more expensive than the ordered comparison and grows faster. The README reports one measurement: matching 30 dictionaries where every element differs took about 30 seconds with --ignore-list-order, against 0.7 seconds without it. Graphtage logs a warning when a comparison is large enough for this to matter, which is the right behaviour but does not make the cost go away.
There are also combinations that are rejected outright. --ignore-list-order cannot be combined with --no-list-edits or --no-list-edits-when-same-length, because those two options only make sense for ordered lists. If you were hoping to treat a list as unordered and simultaneously suppress interstitial edits, the tool will not let you.
Beyond the documented flags, the deeper limitation is conceptual. Graphtage compares trees. If your file's meaning depends on line order, whitespace, comments or textual layout, a tree comparison is the wrong instrument, and a line differ will tell you more. The README does not document rollback behaviour, conflict resolution, or how to merge two files automatically; the name evokes grafting, but the documented surface is comparison and edit reporting, not a three-way merge. Anyone expecting a merge tool should check the documentation before assuming one exists.
Graphtage against Difftastic and delta: different layers of the same problem
The searches that bring people here often mention Difftastic and delta, so it is worth being precise about what separates them. Difftastic is a structural differ for source code: it parses a file according to its programming language and compares syntax trees, which is why it can show that a function body moved without flagging every line as changed. Delta is a pager and highlighter for git output; it improves how an existing diff is displayed rather than computing a different one.
Graphtage sits in a third position. Its inputs are data formats (JSON, JSON5, XML, HTML, YAML, TOML, INI, CSV, plist, pickle), not programming languages, and its output is the data structure with edits marked, not a syntax-highlighted patch. The distinction matters at the level of the matching algorithm: Graphtage exposes dictionary matching strategies and a list-order switch because data formats have ambiguity that source code does not, and it lets you diff across formats, which a syntax-tree differ for source code has no reason to do. If your problem is reviewing a Python refactor, none of Graphtage's dictionary strategies help you. If your problem is a 4,000-line YAML manifest that was regenerated with reordered keys, a source-code differ will not help you either.
Maintenance, licence and what upgrading costs you
The repository is not archived, and the last push was on 2026-09-24, four days before this writing. Two releases landed in September 2026: v0.5.0 on 2026-09-16 and v0.4.0 on 2026-09-09. The previous release, v0.3.1, was on 2024-01-08, so the project spent roughly twenty months between releases before this recent burst. That pattern is worth knowing if you are planning to depend on it: the code is current, but the cadence has not been steady.
The package is licensed LGPL-3.0-or-later according to pyproject.toml, and the PyPI classifier says GNU Lesser General Public License v3 or later. The LGPL is a copyleft licence with a linking exception; using Graphtage as a command-line tool is a different situation from importing it as a library into your own application, and the two are worth separating before you ship. This is not legal advice, and if you plan to embed the library the licence text is the thing to read, not this paragraph.
The dependency list is not trivial. Graphtage pulls in numpy, scipy, intervaltree, PyYAML, json5, toml, colorama, tqdm and fickling. The scipy and numpy requirements are the ones most likely to constrain your environment, since they pin lower bounds rather than exact versions. There is also a conditional dependency on typing_extensions for Python below 3.12, and pyproject.toml carries a comment explaining why: fickling imports typing_extensions.Buffer on Python under 3.12 without declaring the dependency, so importing graphtage.pickle fails on 3.10 and 3.11 unless something else installs it. That is a concrete reason to install through pip rather than vendoring the source, since the workaround lives in the package metadata.
Development Status in the classifiers is 4 - Beta. There is a dev extra (pip3 install 'graphtage[dev]') that adds pytest, Ruff and Sphinx if you intend to work on the tool itself.
Editorial conclusion
Adopt Graphtage when you compare structured config, data dumps or API payloads and need to see what actually changed rather than which lines moved. Skip it for prose, source code, or any file whose meaning depends on line order, and skip it for very large unordered lists unless you can accept the cost the README describes. Before relying on it, verify two things yourself: that your Python is 3.10 or later, and that the default --dict-strategy auto produces the matching you expect on a representative pair of files.
Frequently asked questions
What Python version does Graphtage need?
Graphtage requires Python 3.10 or later, and the README states it is tested against Python 3.10 through 3.14. The pyproject.toml classifiers list each of those versions individually.
Can Graphtage diff a JSON file against a YAML file?
Yes. Graphtage performs its analysis on an intermediate representation divorced from the input filetypes, so the two inputs may be in different formats and the output format may differ from both. The README gives diffing a JSON file against a YAML file as the example.
Why does Graphtage reorder the keys in my output?
Graphtage sorts the keys of every dictionary it reads, so the output orders keys alphabetically regardless of how the input files ordered them. The README notes this while showing examples of a file diffed against itself.
Is Graphtage a merge tool?
The documented surface is comparison and edit reporting: the README describes flags for formatting, matching and match constraints, and the package installs graphtage and graphtage-git-diff. The README does not document rollback or automatic merging of two files.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/trailofbits-graphtage)