# GhidrAssist: an LLM extension that puts a knowledge graph inside Ghidra

> GhidrAssist connects any OpenAI v1-compatible endpoint to Ghidra for function explanation, chat and agentic investigation, and layers a SQLite plus JGraphT semantic graph over the binary. Here is what it does, how to install it, and where it stops being the right tool.

**symgraph/GhidrAssist** — An LLM extension for Ghidra to enable AI assistance in RE.

- Repository: https://github.com/symgraph/GhidrAssist
- Stars: 728 · Forks: 69
- Language: Java
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/symgraph-ghidrassist

## The gap GhidrAssist fills in a Ghidra session

Ghidra gives you a decompiler, a disassembler and a pile of analysis passes. What it does not give you is a way to ask a question about a function in plain language and get an answer that already knows which binary, which address and which cross-references you are looking at. GhidrAssist is built for that moment. It is a Ghidra extension, written in Java, that adds tabs to the CodeBrowser for explaining code, running multi-turn chat, proposing bulk analysis actions, and browsing a semantic graph of the program.

The audience is narrow and specific: reverse engineers who already work inside Ghidra and who have access to an LLM endpoint. The README states that it supports any OpenAI v1-compatible API, which covers local servers such as Ollama, LM-Studio and Open-WebUI as well as cloud providers including OpenAI, Anthropic and Azure. That matters for anyone analyzing binaries they cannot upload. A local model keeps the decompiled pseudo-C on your own machine, and the plugin does not care which one you pick as long as the endpoint speaks the same protocol.

## How the Graph-RAG backend is actually structured

The interesting part of GhidrAssist is not the chat box. It is the knowledge layer underneath. The README describes a five-level semantic hierarchy: Statement, Block, Function, Module, Binary. Summaries are pre-computed by an LLM through a component called SemanticExtractor, which does batch processing over functions. Those summaries are then stored so that later queries can be answered without calling a model at all. The README calls this a "LLM-free query engine", which is the design bet: pay the model cost once at indexing time, then query cheaply.

Storage is hybrid. BinaryKnowledgeGraph combines SQLite with JGraphT graph algorithms, and full-text search runs over summaries and security annotations through SQLite FTS5. Module discovery is not manual. A CommunityDetector runs the Leiden algorithm to group related functions into logical modules with their own hierarchical summaries. Alongside that, SecurityFeatureExtractor performs static analysis for network APIs (POSIX sockets, WinSock, DNS, SSL/TLS, WinHTTP, WinINet), file I/O APIs, crypto APIs from OpenSSL and Windows, and string patterns such as IP addresses, URLs, domains, file paths and registry keys. The output is a risk level of LOW, MEDIUM or HIGH plus an activity profile.

That combination is what separates GhidrAssist from a thin chat wrapper. A wrapper sends a function to a model and prints the reply. GhidrAssist builds a persistent artifact you can search, re-index and explore through a Semantic Graph tab with N-hop depth traversal. The trade-off is that the artifact only exists after indexing, and indexing costs model calls proportional to the number of functions summarized.

## Installing the extension and getting one real answer

The README gives a short quickstart. Copy the binary release ZIP archive into Ghidra_Install/Extensions/Ghidra if it is not already there, then install and enable it from Ghidra itself. The steps below follow that order.

First, install the extension from the Ghidra launcher. The README says to use File then Install Extension, and to enable GhidrAssist. No shell command is given for this step; it is a GUI action.

Next, load a binary and open the CodeBrowser, then enable the plugin. The README lists two separate toggles, which is easy to miss:

```
CodeBrowser -> File -> Configure -> Miscellaneous -> Enable GhidrAssist
CodeBrowser -> Window -> GhidraAssistPlugin
```

The first toggle makes the plugin available in the tool; the second opens it. After that, the README says to check that the RLHF and RAG database paths are appropriate for your environment, and to point the API host at your provider and set the API key. Those are configuration fields in the plugin, not environment variables; the README does not name them beyond describing the API host and API key.

If you want the extended thinking controls, the README places them in the Analysis Options tab, where you set the Reasoning Effort level to None, Low, Medium or High for models that support it. The README notes that this setting persists per program, so a slow binary can stay at High while a quick triage target stays at Low.

Finally, open the plugin from the Windows menu and start exploring. The first useful action is the Explain tab on a function you already understand, so you can judge the model's output before trusting it on one you do not.

## Agentic mode, MCP and what the plugin does without asking

Beyond chat, GhidrAssist offers a ReAct agentic mode. The README describes a Think-Act-Observe loop: the model proposes investigation steps from your query, tools execute systematically with progress tracked through todo lists, iteration history is preserved, and the run ends with a synthesis and key findings plus metrics for iterations, tool calls and duration. Function calling lets the model rename functions and variables, navigate to addresses and cross-references, and execute Ghidra commands.

Read that list carefully. An autonomous loop that can rename symbols and run Ghidra commands is editing your analysis state, not just reading it. The README does not document an undo path for agent-applied renames, and it does not describe a confirmation gate before tool execution. If you run agentic mode on a database you care about, the safe assumption is that the changes land. Work on a copy, or keep the loop pointed at functions you are still triaging rather than ones you have already cleaned up.

The MCP integration follows the same pattern. GhidrAssist acts as a Model Context Protocol client and is documented to work with GhidrAssistMCP, a separate repository by the same author, for Ghidra-specific tools. The README mentions conversational tool calling with automatic function execution and support for SSE transport. Automatic execution is the same trade-off in a different place.

## Where GhidrAssist is the wrong tool

The plugin lives inside the Ghidra GUI. There is no documented headless mode, no CLI, and no batch entry point in the README. If your workflow is a CI job that decompiles a sample and files a report, GhidrAssist does not fit; you would be scripting against the Ghidra API directly, not driving this extension.

The second constraint is the model endpoint. GhidrAssist is a client. It ships no model and no inference. Every explanation, every summary during indexing and every agentic step depends on an endpoint you supply. Point it at nothing and the tabs stay empty. Point it at a small local model and the quality of the five-level summaries is whatever that model produces, because the graph is built from LLM output, not from Ghidra's own analysis. The SecurityFeatureExtractor is static analysis and does not depend on a model, but the semantic summaries do.

Third, indexing has a cost that scales with binary size. The README describes batch processing of function summarization and pre-computed summaries as the mechanism for fast queries. A large stripped binary means a large number of functions to summarize before the Semantic Graph tab is useful. The README does not publish a per-function token or time estimate, so the only way to know your cost is to index a sample and watch.

## GhidrAssist against GhidraGPT and other Ghidra LLM plugins

The related searches around this project point at GhidraGPT and GhidraMCP-adjacent tooling, so the comparison is worth making concrete. GhidraGPT, as the name suggests, is in the same family: an LLM assistant inside Ghidra for explaining and renaming. The difference in approach is what happens to the model's output. A chat-style plugin produces text in a panel and, at most, applies a rename you accept. GhidrAssist persists the output. Summaries go into SQLite, get linked into a JGraphT graph across five levels, get clustered into modules by the Leiden algorithm, and become searchable through FTS5 without another model call.

That is a heavier design with a real cost. You pay for indexing before you get the fast queries, and the graph is only as good as the summaries that built it. A plain chat plugin has no indexing step and no persistence, which is exactly what you want for a one-off look at a single function. GhidrAssist is aimed at the case where you return to the same binary repeatedly and want your earlier analysis to still be there.

The MCP angle is separate. GhidrAssist's MCP client support is documented to work with GhidrAssistMCP, which is a different repository and a different install. If your interest is tool-based analysis through MCP rather than the graph, that companion project is the thing to read, not this one.

## Licence, maintenance and upgrade cost

GhidrAssist is MIT licensed. For most users that means the usual permissions to use, modify and redistribute with attribution, and no copyleft obligation on your own code. If you plan to bundle the extension into a commercial product, read the LICENSE file in the repository rather than relying on the identifier alone; this is not legal advice.

The repository is not archived. The most recent push recorded for the default branch is 2026-05-29, and the newest release listed is 2.2.0 on the same date, following 2.1.0 on 2026-05-25 and 2.0.0 on 2026-04-19. The 2.0.0 release notes describe a new SymGraph service tab, so the project has been adding surface area rather than only fixing bugs. The README lists a future roadmap covering model fine-tuning from the collected RLHF dataset, more MCP tool integrations, multi-agent collaboration and embedding-based similarity search. None of those are shipped; treat them as intent.

Upgrade cost is where the data layer matters. The README names a SchemaMigrationRunner described as versioned database migrations for transparent upgrades. That is the mechanism that keeps your chat history, RLHF feedback and knowledge graphs usable across versions, and it is worth knowing it exists before you build a long-lived analysis database on top of the plugin. The README does not document how to roll back a migration if one goes wrong, so back up the SQLite files before upgrading.

## Conclusion

Adopt GhidrAssist if you already live in Ghidra and want LLM explanations, a persistent semantic graph and an agentic investigation loop without leaving the CodeBrowser. Skip it if you need a headless pipeline, or if you cannot point it at a model endpoint you control, since the plugin is a client and ships no model. Before committing, verify the RLHF and RAG database paths, confirm your provider speaks the OpenAI v1 API, and check whether the Reasoning Effort control is meaningful for the model you selected.

## FAQ

### What is GhidrAssist?

It is a Ghidra extension written in Java that adds LLM assistance to the CodeBrowser: code explanation, interactive chat, custom queries, an agentic ReAct mode, MCP tool calling and a Graph-RAG knowledge system. The README describes it as an advanced LLM-powered plugin for interactive reverse engineering assistance.

### Which LLM providers does GhidrAssist support?

The README states that it works with any OpenAI v1-compatible API, naming local options such as Ollama, LM-Studio and Open-WebUI alongside cloud providers including OpenAI, Anthropic and Azure. Setup details are provider-specific and the README links out to each provider's own documentation.

### What is the semantic knowledge graph in GhidrAssist?

It is a five-level hierarchy of Statement, Block, Function, Module and Binary, built from pre-computed LLM summaries and stored in SQLite with JGraphT graph algorithms. The README says this enables fast queries without calling an LLM, and that full-text search runs over the summaries through SQLite FTS5.

### Can GhidrAssist modify my Ghidra analysis?

Yes. The README lists function calling that lets the LLM rename functions and variables, navigate to addresses and cross-references, and execute Ghidra commands, and it describes MCP tool calling with automatic function execution. The README does not document an undo path for those changes.

### Does GhidrAssist need an internet connection?

Not necessarily. Because it supports any OpenAI v1-compatible API, the README lists local providers such as Ollama, LM-Studio and Open-WebUI, which keep the binary and the decompiled code on your own machine. It does require a reachable model endpoint of some kind, since the plugin ships no model.

## Sources

- [Issues](https://github.com/symgraph/GhidrAssist/issues)
- [License: MIT](https://github.com/symgraph/GhidrAssist/blob/master/LICENSE)
- [README](https://github.com/symgraph/GhidrAssist/blob/master/README.md)
- [Releases](https://github.com/symgraph/GhidrAssist/releases)
- [symgraph/GhidrAssist on GitHub](https://github.com/symgraph/GhidrAssist)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/symgraph-ghidrassist
