# Diaphora: an IDA plugin for diffing two binaries function by function

> Diaphora is an AGPL-licensed program diffing tool that runs as an IDA plugin and compares two exported databases of functions. It is aimed at reverse engineers doing patch analysis, and its heuristics go past assembler comparison into pseudo-code, microcode and type porting.

**joxeankoret/diaphora** — Diaphora, the most advanced Free and Open Source program diffing tool.

- Repository: https://github.com/joxeankoret/diaphora
- Stars: 4,412 · Forks: 418
- Language: Python
- License: AGPL-3.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/joxeankoret-diaphora

## The patch analysis problem Diaphora was built for

Comparing two builds of the same program is one of the few ways to find out what a vendor actually changed. A security patch that fixes a memory corruption bug usually touches a handful of functions out of thousands, and the release notes rarely say which. Doing that comparison by eye in a disassembler is slow and error-prone, and the interesting result is often a single added bounds check inside a function whose address moved because the linker shifted everything around.

Diaphora targets exactly that task. It works as an IDA plugin, and the README describes it as "the most advanced program diffing tool (working as an IDA plugin)". The audience is reverse engineers and vulnerability researchers who already have IDA and need to answer the question of which function corresponds to which across two versions. The README's screenshots are all patch-diffing cases: CVE-2020-1350, CVE-2023-28231, CVE-2023-21768, the PEGASUS iOS kernel fix in iOS 9.3.5, and the older MS15-034 and MS15-050 bulletins. That selection tells you what the author considers the primary use case.

## How the diffing engine matches functions

Diaphora does not diff two binaries directly in one pass. The workflow is export then diff. You load one binary in IDA and export a database describing its functions; you do the same for the second binary; then you run the diffing step over the two exported databases. The repository layout reflects this split: diaphora_ida.py handles the IDA-side export, diaphora_import.py and diaphora_load.py handle loading, and diaphora_heuristics.py holds the matching logic.

The README lists the heuristics as "dozens of heuristics based on graph theory, assembler, bytes, functions' features, etc..." plus call graph matching and similarity ratio calculation. The point of having many independent signals is that no single one survives compilation differences. Addresses change, register allocation changes, and instruction selection changes, but the shape of a control flow graph and the constants a function compares against tend to survive. Diaphora also exposes pseudo-code based heuristics and microcode support, which means matching can happen above the assembler level where the compiler's choices matter less.

One capability worth singling out is type porting. The README claims the ability to "port structs, enums, unions and typedefs" as a feature not available in other public tools. If that works on your target, it means a struct you reconstructed in the old binary carries over to the new one instead of being rebuilt by hand. The README also mentions parallel diffing and scripting support for both the exporting and diffing processes, which matters once you are diffing more than a pair of small files.

## Installing Diaphora and running a first diff

The README gives two installation routes. The first is the Hex-Rays plugin manager:

```bash
hcli plugin install https://github.com/joxeankoret/diaphora/archive/refs/tags/3.4.2.zip
```

That command installs the tagged 3.4.2 archive. The second route is manual, and the README is explicit that Diaphora "requires no installation": you download the code and run diaphora.py from within IDA or from the command line, the latter only for diffing databases that were already exported. To get the plugin menu inside IDA, the README says to copy plugins/diaphora_plugin.py and plugins/diaphora_plugin.cfg into IDA's plugins directory, then edit diaphora_plugin.cfg and set the path value to the Diaphora directory.

Before any of that, the Python dependencies have to be present in the interpreter IDA uses. The repository's requirements.txt lists them:

```text
cdifflib
scikit-learn
joblib
pandas
numpy
```

Note what that list implies: scikit-learn, pandas and numpy are heavy packages, and they must be importable from inside IDA's embedded Python, not from a separate virtual environment. That is the step most likely to go wrong on a fresh machine.

Once the plugin is loaded, the sequence is: open the first binary in IDA, run the export from the Diaphora menu to produce a database, open the second binary and export it the same way, then run the diff over the two databases. The README points to the project wiki for automating the export and diffing steps and for speeding up operations, which is where you should look if you intend to run this over many pairs rather than interactively.

## Where Diaphora stops being the right tool

The dependency on IDA is the first boundary. Diaphora is an IDA plugin and the README states it supports IDA versions 7.4 and above because the code only runs on Python 3. If your team works in Ghidra, radare2 or Binary Ninja, Diaphora is not an option regardless of how good the heuristics are, and the README does not describe a standalone mode that would change that. The command-line path exists, but only for diffing databases that were already exported, and the export itself is an IDA operation.

The second boundary is the licence. Since version 2.0 Diaphora is under AGPL-3.0, and the README explains the change was made so that companies modifying and adapting it "cannot offer web services based on these modified versions without contributing back the changes". The README also states that for most users the change has no effect, and that commercial licences are available for companies whose policy disallows AGPL. Whether your organisation falls into that group is a question for your legal team, not something this article can settle.

The third boundary is scale and setup cost. Every binary you want to compare has to be opened and exported in IDA first, and the exported database is a real artifact you have to store and track. For a one-off look at two small files, that overhead may exceed the value. The README does not document rollback or undo for the symbol and comment porting step, so if you let Diaphora write names and comments into a database you care about, keep a copy of the original.

## Diaphora against BinDiff, and the case for each

BinDiff is the comparison people reach for first, and the difference in approach is architectural rather than a matter of which heuristics are better. BinDiff is a standalone toolchain: you export from your disassembler, run the comparison, and inspect results in a separate viewer. Diaphora keeps the whole loop inside IDA. The practical consequence is that when Diaphora finds a match, you are already in the disassembler with the pseudo-code, the cross-references and the type information in front of you, and porting a name or a comment is a menu action rather than a manual copy between applications.

The second difference is depth of representation. Diaphora's README lists microcode support, pseudo-code based heuristics and pseudo-code patches generation. Comparing compiler intermediate representations rather than instruction bytes is a different bet about what survives recompilation: it tolerates more instruction-level noise but depends on IDA's decompiler being able to produce a usable pseudo-code for the function in the first place. For heavy obfuscation or hand-written assembler, that assumption breaks down and the lower-level heuristics carry the work.

Neither tool removes the need to understand the code you are looking at. A high similarity ratio tells you two functions correspond; it does not tell you which line is the fix. The README's CVE-2023-28231 screenshot is instructive here, because the vulnerability was fixed by a comparison against 32 (0x20), and finding that means reading the pseudo-code diff, not trusting a score.

## Licence terms and the cost of staying current

Diaphora is AGPL-3.0. The README is unusually direct about why: the licence was changed so that companies adapting the tool cannot run modified versions as a network service without contributing the changes back. For an individual researcher or a team using Diaphora internally to analyse binaries, the README says the change does not affect them. For a company that wants to embed Diaphora in a product, or expose it through a hosted service, the AGPL is a real constraint and the README points to commercial licences and consultancy via admin@joxeankoret.com. This is a description of what the project states, not legal advice; if the distinction matters to your employer, that conversation belongs with counsel.

On maintenance, the repository is not archived and the last push was on 2026-09-04, with releases 3.4.0, 3.4.1 and 3.4.2 landing between May and August 2026. The README makes a stronger claim than most projects can: it says Diaphora has been ported to every minor version of IDA since 6.8 up to 9.4, over 11 years. That track record is the upgrade cost argument. The plugin depends on IDA's Python API and on the decompiler, both of which Hex-Rays changes between versions, so a diffing tool that does not track IDA releases becomes unusable quickly. Diaphora's release cadence suggests that tracking is happening, but the README also notes Python 3.11 was the last version tested, which is the number to check against your IDA installation before you plan an upgrade.

## Conclusion

Adopt Diaphora if you already work in IDA and your job involves comparing two builds of the same target, especially for patch analysis where the fixed function has to be located quickly. Do not adopt it if you need a diffing workflow outside IDA, or if your organisation cannot accept AGPL-3.0 terms and will not buy a commercial licence. Before committing, verify that your IDA version is at least 7.4 on Python 3, that the dependencies in requirements.txt install cleanly in the interpreter IDA uses, and that the exported database for your largest target completes in a time you can live with.

## FAQ

### How do I install Diaphora?

The README gives two routes. You can install it with the Hex-Rays plugin manager using hcli plugin install against the 3.4.2 archive, or install it manually by copying plugins/diaphora_plugin.py and plugins/diaphora_plugin.cfg into IDA's plugins directory and setting the path value in the cfg file to the Diaphora directory.

### How do I use Diaphora with IDA?

Diaphora works as an IDA plugin. You open the first binary in IDA and export a database, do the same for the second binary, then run the diff over the two exported databases. The README points to the project wiki for automating the export and diffing steps.

### How does Diaphora compare with BinDiff?

BinDiff is a standalone toolchain with a separate viewer, while Diaphora runs the whole loop inside IDA, so matches open directly in the disassembler with pseudo-code and cross-references available. Diaphora also lists microcode support, pseudo-code based heuristics and type porting among its features.

## Sources

- [Issues](https://github.com/joxeankoret/diaphora/issues)
- [joxeankoret/diaphora on GitHub](https://github.com/joxeankoret/diaphora)
- [License: AGPL-3.0](https://github.com/joxeankoret/diaphora/blob/master/LICENSE)
- [README](https://github.com/joxeankoret/diaphora/blob/master/README.md)
- [Releases](https://github.com/joxeankoret/diaphora/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/joxeankoret-diaphora
