Model or dataset
Dryxio/reagent avatar
Dryxio/reagent

ReAgent: reconstruct C/C++ functions from binaries with an LLM reverser and checker

Reconstruct and validate C/C++ code from compiled programs with AI.

1,996 stars203 forksPythonMIT

At a glance

What is it?
ReAgent (PyPI package auto-re-agent) drives Ghidra evidence through independent LLM reverser and checker roles, then gates every candidate C/C++ function behind build, test and parity checks. It is a conservative assistant, not a semantic equivalence proof.
Who is it for?
Adopt ReAgent if you already have Ghidra, a Ghidra Bridge export, and an authenticated Claude or Codex CLI, and you want candidate C/C++ functions with explicit evidence gaps rather than a finished patch. Do not adopt it if you expect it to edit your source tree automatically, or if you cannot accept that a PASS from the LLM checker is not a proof of semantic equivalence.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 22 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem ReAgent targets: recovering a C/C++ function when only the binary survives

Reverse engineering a stripped binary into readable C/C++ is slow, and the parts that are slow are the parts that repeat: pull the decompilation, collect cross-references and structs, guess at types, write a candidate, compile it, and check whether the result behaves like the original. ReAgent packages that loop as a command. The README describes it as an "open-source AI reverse-engineering agent" that uses Ghidra and LLMs, including Claude, Codex, and OpenAI-compatible models, to reconstruct and validate C/C++ functions from compiled binaries. The audience is narrow and specific: engineers working on decompilation projects where a build and a test suite already exist, so a generated candidate can be compiled and exercised rather than merely read. The repository lists topics such as binary-analysis, decompilation and source-code-recovery, and the pyproject classifiers place it at Development Status 3 - Alpha. Treat that label as accurate. The tool generates candidate implementations; the README states plainly that it does not patch the original source tree automatically, so the last step stays manual.

How the reverser and checker loop works, and why candidate parity is a separate gate

The pipeline in the README is a sequence, not a single model call. Configuration comes from YAML plus supported environment overrides and CLI flags. Function selection offers dependency-order, easiest-first and high-impact. Context gathering pulls decompilation, xrefs, structs, enums, vtables, globals and strings, plus normalized high P-code, a CFG, assembly and nearby project source. Then a reverser and a checker run in a bounded fix loop, with a limit on rounds and investigations. The separation matters: the checker can be a different provider and model from the reverser, which is why the sample configuration shows a Claude CLI reverser and a Codex checker. After the loop, a conservative structural verifier looks for strong mismatches, and the candidate is placed in an overlay where configured build, test and runtime gates run. A parity gate then returns GREEN, YELLOW or RED. The README is explicit that a successful reversal can require four independent conditions: the LLM checker returns PASS, the objective verifier finds no strong structural mismatch, candidate validation satisfies the configured acceptance policy, and parity is not blocked by the configured RED/YELLOW policy. That is a conservative verification design, and the README says so: it is not a proof of semantic equivalence. The four-condition framing is the most useful thing in the document, because it tells you where a run can stall even when the generated code looks correct.

Installing ReAgent and running a first reversal on one function

Requirements are Python 3.10+, Git for the source install, Ghidra with a configured Ghidra Bridge, and at least one LLM setup: an ANTHROPIC_API_KEY, an OPENAI_API_KEY, an authenticated local claude command, or an authenticated local codex command. Install the agent and its Ghidra query bridge from PyPI. The extra name and version floor are as the README gives them.

bash
python3 -m pip install --upgrade "auto-re-agent[ghidra-bridge]>=0.4.0"

For headless Ghidra exports, the README points to the headless extra instead, which pulls the bridge with its PyGhidra support.

bash
python3 -m pip install --upgrade "auto-re-agent[headless]>=0.4.0"

Run the Ghidra evidence commands from the project you want to reverse. The first creates ghidra-bridge.yaml, which you then edit to point at your Ghidra project and program paths. The export step requires the bridge headless extra and a local Ghidra installation.

bash
ghidra-bridge init
ghidra-bridge export all
ghidra-bridge info

Then create a configuration in the target project. The README recommends the portable default profile for new projects, and notes that running re-agent init without --profile preserves the original GTA-reversed defaults.

bash
re-agent init --profile generic-cpp

Edit re-agent.yaml so it selects an LLM, points the backend at the installed bridge executable, sets source paths and configures validation. The README's minimum shape looks like this.

yaml
llm:
  provider: claude-cli
  model: sonnet

backend:
  type: ghidra-bridge
  cli_path: ghidra-bridge

project_profile:
  name: generic-cpp
  language_standard: C++20
  source_root: src

The README's own suggested first run is a single small function, with re-agent doctor run first to check setup, and a small model-call limit set before starting. Version 0.4.0 adds re-agent plan, which builds bounded function manifests without model calls, and re-agent reverse --manifest, which reconstructs selected functions across classes in an isolated project copy. After a run, re-agent status --manifest reports coverage, stale results and individual validation checks, which is where you find out which of the four conditions held.

Windows argument arrays, shell strings, and the failure the doctor command catches

The 0.4.0 release notes describe a change that will break existing configurations on native Windows. Build, test and runtime validation now support argument arrays that execute directly on Windows and POSIX. On native Windows, shell strings must be converted to arrays; legacy strings still require /bin/sh, and re-agent doctor reports a missing shell. If you are upgrading a project that was configured with shell-string validation commands on Windows, the acceptance policy can fail for reasons that have nothing to do with the generated C/C++. The README points to docs/configuration.md#portable-validation-commands for migration. The same release also notes that Clang indexing handles CRLF offsets and that Codex CLI requests use UTF-8 stdin for large prompts, both of which read as fixes for platform-specific breakage rather than new capability. The lesson is to run re-agent doctor after any upgrade before blaming the model.

Where ReAgent is the wrong tool

The README does not document rollback, and it does not claim the tool edits your source. If your goal is to take a binary and get a drop-in replacement file, ReAgent stops one step short by design: it produces a candidate overlay and a report, and you apply the result yourself. The parity gate's RED and YELLOW outcomes are policy decisions, not verdicts, so a team that wants a binary yes or no will be configuring acceptance policy instead of receiving an answer. There is also a real cost surface the README names directly: the setup prompt tells the user to ask which AI provider will be used and any API costs, and to set a small model-call limit. The bounded rounds and investigations exist because model calls are the expensive part of the loop. If you cannot supply build and test gates, the candidate validation condition cannot be satisfied, and the run loses one of its four legs. And if you have no Ghidra project or cannot run ghidra-bridge export all, there is no evidence for the agent to reason over. Finally, the Alpha classifier is not decoration: the interface changed between 0.3.0 and 0.4.0, and the Windows validation change shows that configuration compatibility is not guaranteed across minor versions.

How ReAgent differs from running Ghidra's decompiler on its own

Ghidra's decompiler is the baseline alternative, and it is already inside ReAgent's loop. The difference is what happens after the decompilation appears. A decompiler gives you pseudo-C and leaves type recovery, naming and verification to you. ReAgent adds a second model in a checker role, an objective structural verifier that looks for strong mismatches, and a build and test gate that compiles the candidate and runs it. That is the whole trade: you accept a Python dependency, a Ghidra Bridge install, an LLM provider and API cost, in exchange for a candidate that has been compiled and exercised rather than only printed. The README's own framing of four independent conditions is the honest description of what you get. If your project has no build or no test suite, a plain decompiler session is cheaper and gives you the same pseudo-C without the configuration surface. ReAgent only pays off where a candidate can actually be compiled and run.

Licence, maintenance and upgrade cost

ReAgent is MIT licensed, and pyproject.toml declares license = "MIT" with the OSI Approved :: MIT License classifier. The practical implication is that you can read, modify and redistribute the code, including in commercial settings, provided you keep the licence notice; this is a description of the licence text, not legal advice, and your own counsel should confirm how it interacts with any Ghidra or LLM provider terms you are bound by. Maintenance signals are current: the last push to main was on 2026-09-09, and v0.4.0 was released the same day, with v0.3.0 on 2026-09-04 and v0.2.1 on 2026-07-23. That is a fast release cadence for an Alpha project, which cuts both ways. You get fixes such as the CRLF and UTF-8 stdin handling, and you inherit configuration churn. The upgrade path is pip against the version floor, so pinning is straightforward, but the 0.4.0 validation change means a pinned upgrade can invalidate existing validation commands on Windows. Budget for re-running re-agent doctor and re-checking acceptance policy at each minor version rather than assuming the YAML carries over.

Editorial conclusion

Adopt ReAgent if you already have Ghidra, a Ghidra Bridge export, and an authenticated Claude or Codex CLI, and you want candidate C/C++ functions with explicit evidence gaps rather than a finished patch. Do not adopt it if you expect it to edit your source tree automatically, or if you cannot accept that a PASS from the LLM checker is not a proof of semantic equivalence. Before committing to a workflow, run re-agent doctor, run re-agent init --profile generic-cpp, and reverse one small function to see which of the four acceptance conditions actually pass on your binary.

Frequently asked questions

What is ReAgent?

ReAgent is an open-source AI reverse-engineering agent that uses Ghidra and LLMs to reconstruct and validate C/C++ functions from compiled binaries. It combines independent reverser and checker models, agentic evidence gathering, candidate build and test gates, structural verification and parity analysis in one autonomous workflow. It generates candidate implementations and does not patch the original source tree automatically.

How do I install ReAgent?

Install the agent and its Ghidra query bridge from PyPI with python3 -m pip install --upgrade "auto-re-agent[ghidra-bridge]>=0.4.0". For headless Ghidra exports, the README gives the headless extra instead. You also need Python 3.10+, Git, Ghidra with a configured Ghidra Bridge, and at least one LLM setup such as an ANTHROPIC_API_KEY, an OPENAI_API_KEY, or an authenticated claude or codex command.

How do I use ReAgent on a binary?

Run ghidra-bridge init, edit ghidra-bridge.yaml with your Ghidra project and program paths, then run ghidra-bridge export all and ghidra-bridge info from the project you want to reverse. Create a configuration with re-agent init --profile generic-cpp, edit re-agent.yaml to select an LLM, point the backend at the bridge executable and set source paths, then start with one small function. The README suggests running re-agent doctor first and setting a small model-call limit.

How do I enable ReAgent?

There is no enable step in the README; the tool is used by running its commands. After installing auto-re-agent and exporting Ghidra evidence with ghidra-bridge export all, you create a configuration with re-agent init --profile generic-cpp and edit re-agent.yaml to select an LLM and point the backend at the ghidra-bridge executable.

What are the requirements for running ReAgent?

Python 3.10+, Git for the current source installation, Ghidra plus a configured Ghidra Bridge, and at least one LLM setup: an ANTHROPIC_API_KEY, an OPENAI_API_KEY, an authenticated local claude command, or an authenticated local codex command. Headless Ghidra exports need the bridge installed with its PyGhidra extra.

Official sources

  1. Dryxio/reagent on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/dryxio-reagent.svg)](https://hysenlabs.com/projects/dryxio-reagent)