Reverify: an MCP server and CLI that makes deterministic tools the judge of every AI claim
Stop your AI from making things up — it proposes, deterministic tools decide, every claim checked against ground truth with evidence. Grounded facts and context survive resets. Reverse engineering is the proving ground. MCP server + CLI.
At a glance
- What is it?
- Reverify is a Python anti-hallucination toolkit that pairs a language model with a deterministic reverse-engineering core. The model proposes a claim, the tools check it against the actual bytes, and only what survives is reported as fact.
- Who is it for?
- Adopt Reverify if you already run an agent against binaries or source and want a deterministic gate between the model's prose and your notes: the CLI exits non-zero on any refuted claim, which is exactly the hook a CI job or an agent loop needs. Do not adopt it if you want a Ghidra replacement or a general-purpose disassembler with a GUI; the semantic layer needs the large angr extra, and the pure-Python core is a fallback, not a full engine.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 23 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Reverify is for, and who should care
Language models read code well and reconstruct binaries badly. Ask one to describe a struct field, an offset or a function prologue from a compiled artifact and it will produce an answer with the same confidence whether or not the bytes support it. Reverify's premise is that this is not a prompting problem but an authority problem: the model should not be allowed to assert a structural fact at all. It proposes, a deterministic tool checks, and the result comes back as VERIFIED, REFUTED or INCONCLUSIVE with the bytes that were actually observed.
The audience is narrow on purpose. The README frames the project as a tool for authorized reverse engineering: malware analysis, CTF work, interoperability research, and software you own or are permitted to analyze, with SECURITY.md covering that scope. In practice the people who get value are agent users who already have a model reading disassembly and want a mechanical check before a claim enters their notes, plus anyone running an automated pipeline where a wrong offset silently poisons everything downstream. If you never let a model near a binary, the verification loop has nothing to judge.
The claim-and-judge loop, and the deterministic core underneath it
The mechanism is a claim object. A claim is a hypothesis about the artifact, expressed as a JSON document with a kind and the parameters that kind needs. The verifier dispatches on the kind, runs the corresponding deterministic check, and returns a verdict plus evidence. Claim kinds listed in the README include bytes_at, the typed reads u16_at, u32_at and u64_at, pattern_present, string_present, instructions, emulate_result, behavior_equiv, prove_equiv, protobuf_field, import_present, export_present and section_present, plus a semantic group: function_at, calls, references and reachable_from_entry. Offsets default to file offsets; a claim can set space to rva or va, and the verifier translates through the section table and echoes all three addresses.
The engine layer is what makes the verdicts mean anything. Out of the box the toolkit is pure Python and parses PE, ELF and Mach-O, disassembles x86, x64, ARM and ARM64, scans AOB patterns, emulates a CPU, dissects Protobuf and TLV, and generates Frida hooks. The optional extras upgrade that in place: capstone for full-fidelity disassembly, unicorn for real CPU emulation, lief for object-format parsing, and Z3 for proof-grade expression equivalence. angr is a separate extra, deliberately excluded from the full bundle, and it is what supplies function boundaries, the call graph and cross-references for the semantic claim kinds. The README is explicit that a missing engine causes a fallback to the pure-Python core rather than an error, and reverify backends reports which engines are active. That fallback is a design trade-off worth naming: a verdict produced by the pure-Python path and the same verdict produced by unicorn are not the same evidence, and the tool does not hide that, but it also does not stop you from treating them as equivalent.
Installing Reverify and verifying your first claim
The package is on PyPI and requires Python 3.8 or later. The plain install pulls no dependencies at all; the full extra adds capstone, unicorn, lief and Z3.
pip install reverify
reverify backendsThe first command installs the CLI and MCP server. The second prints which engines are active, so you know whether you are on the pure-Python core or the upgraded path. If you want the mature engines, install the full extra instead and re-run reverify backends to confirm the change.
pip install "reverify[full]"There is also a no-install path straight from a checkout, which the README describes as pure standard library:
python reverify/cli.py auto sample.bin --json
python reverify/cli.py parse-pe sample.exe --jsonThe auto subcommand runs the toolkit against a sample and emits JSON. The real work, though, is the verification loop. Here is the README's own example of checking a claimed function prologue at file offset 4096:
reverify verify sample.bin --claim '{
"kind": "instructions", "offset": 4096,
"mnemonics": ["push", "mov", "sub"], "note": "function prologue"
}'What you should see is a verdict, not prose. If the bytes at that offset do not disassemble to those mnemonics, the claim comes back REFUTED with the observed bytes attached. A second example checks that a small routine actually computes what the model said it computes, by emulating it and asserting a register value:
reverify verify - --claim '{
"kind": "emulate_result", "code": "b805000000b90300000001c8c3",
"arch": "x86", "expect_registers": {"eax": 8}
}'Claims can also be batched from a file with --claims-file claims.json. The behaviour that matters for automation is that the CLI exits non-zero if anything is refuted, which is what lets an agent or a CI job gate on a grounded reconstruction rather than a plausible one.
Context rollover, and the claim that it is lossless
The second feature is unrelated to binaries. Long agent sessions degrade: the model starts drifting, and the usual fix is an auto-summary that quietly drops detail or a manual /clear that drops everything. reverify rollover instead writes the session state to a file and starts a fresh session from it. The v0.11.0 release notes describe this as lossless context rollover across Claude Code, Codex, Gemini CLI and OpenCode.
Lossless is a strong word and the README does not document what happens when the handoff file itself exceeds the new session's budget, nor is there a documented rollback if the resumed session turns out worse than the original. Treat the label as the project's claim about its own mechanism, not as a guarantee you can rely on without testing it against your workload. The design choice is still defensible: a file on disk is inspectable in a way that a summary embedded in a prompt is not, and if the new session goes wrong you can read what was handed over.
Where Reverify is the wrong tool
The project is classified as Development Status 3, Alpha, in its own pyproject.toml. That is the honest signal, and it should shape how you deploy it.
The larger constraint is that verification only covers claims the toolkit can express. The claim kinds are a fixed vocabulary. If your question about an artifact does not map onto bytes_at, instructions, emulate_result, protobuf_field or one of the other listed kinds, Reverify has no opinion, and INCONCLUSIVE is not the same as safe. A model can be wrong in ways the schema cannot describe, and the tool will not catch that. The semantic kinds narrow this further, because function_at, calls, references and reachable_from_entry depend on angr, which the README describes as large and deliberately not part of the full extra. On a machine without angr, those claims fall back rather than fail, which is convenient and also means the strength of your evidence depends on an install decision you may not have made consciously.
Finally, this is not a disassembler you would use interactively. There is no GUI, no decompiler output to browse, and no replacement for Ghidra as an analysis environment. It is a judge that sits between a model and your notes. If you want to explore a binary by hand, Reverify is not the tool you reach for.
How Reverify differs from Ghidra and from a plain MCP tool server
The obvious comparison is Ghidra. Ghidra is an interactive reverse-engineering platform with a decompiler, a scripting API and a GUI; it is where a human analyst works. Reverify does not try to be that. It is a headless verifier whose output is a verdict on a specific claim, designed to be called by an agent rather than driven by a person. The overlap is the parsing and disassembly layer, and Reverify's own README notes it installs clean with no Ghidra dependency, which is the point: the deterministic core is pure Python and the mature engines are pip extras.
The second comparison is a generic MCP tool server that exposes a disassembler to an agent. Those give the model a function to call and trust the model to interpret the return value. Reverify inverts the direction. The model states what it believes, and the tool decides whether that belief survives contact with the bytes. The difference is not the disassembly, it is who holds authority over the conclusion. That inversion is also why the CLI exit code matters: a refuted claim is a hard failure an automated pipeline can act on, not a sentence in a transcript.
Licence, maintenance and what an upgrade costs
Reverify is MIT licensed, and the pyproject.toml declares license = { text = "MIT" } with the OSI classifier. MIT is permissive: you can use it commercially, modify it and redistribute it, provided the copyright notice and permission notice are preserved. The optional engines are separate packages with their own licences, and installing reverify[full] or reverify[angr] pulls them in, so the licence terms of capstone, unicorn, lief, z3-solver and angr apply to those components independently. That is a fact about the dependency graph, not legal advice; if your organisation has a policy on bundled licences, check the extras you actually install.
The repository is not archived and the last push was on 2026-09-07, which is recent. The release history shows v0.9.1, v0.10.0 and v0.11.0 all dated 2026-09-04, with v0.11.0 adding the rollover feature; the pace suggests the interfaces are still moving. The upgrade cost follows from that: the claim schema is the contract your automation depends on, and an alpha project can extend or rename claim kinds between minor versions. Pinning reverify to a specific version and reading CHANGELOG.md before bumping is the cheap insurance here. The extra installs are also not free in size, particularly angr, which the README flags as large.
Editorial conclusion
Adopt Reverify if you already run an agent against binaries or source and want a deterministic gate between the model's prose and your notes: the CLI exits non-zero on any refuted claim, which is exactly the hook a CI job or an agent loop needs. Do not adopt it if you want a Ghidra replacement or a general-purpose disassembler with a GUI; the semantic layer needs the large angr extra, and the pure-Python core is a fallback, not a full engine. Before committing, run reverify backends on your machine to see which engines are actually active, and read BENCHMARK.md and EXAMPLE.md to confirm the 71-file corpus matches the kind of artifact you analyze.
Frequently asked questions
Is it re-verify or reverify?
Both spellings appear in the project's own documentation. The package name, the CLI command and the repository name are all reverify, while the README also writes re-verify in prose when describing the act of checking a claim.
Is reverify a word?
The project's documentation does not discuss the English word. It documents reverify as the name of a Python package installable from PyPI, with a CLI command of the same name and an MCP server mode.
What is Reverify and what does it verify?
Reverify pairs a language model with a deterministic reverse-engineering toolkit and makes the toolkit the judge. The model proposes a claim, the tools check it against the actual artifact, and the result is VERIFIED, REFUTED or INCONCLUSIVE with the observed bytes as evidence.
How do I install Reverify?
Install it from PyPI with pip install reverify for the pure-Python core, or pip install "reverify[full]" to add capstone, unicorn, lief and Z3. The README also documents running it straight from a checkout with python reverify/cli.py.
Does Reverify work without Ghidra or angr?
Yes for the core. The README states the toolkit is pure Python out of the box and installs clean with no Ghidra, falling back to that core when an optional engine is absent. The semantic claim kinds such as function_at, calls and references need the separate angr extra, which is not included in the full bundle.
Can Reverify check ordinary source code, not just binaries?
The README documents reverify equiv, which runs a candidate implementation and a reference over shared inputs and checks that they agree, with --lang python or C. A refutation comes back with the input and both outputs.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/2akouwu-reverify)