Model or dataset
416rehman/DeepZero avatar
416rehman/DeepZero

DeepZero's quickstart analyses two text files, not a kernel driver

Find zero-days while you sleep. DeepZero is an automated vulnerability research framework that parses, decompiles, and analyzes thousands of Windows kernel drivers for exploitable IOCTLs natively using AI agents.

735 stars92 forksPythonMIT

At a glance

What is it?
A YAML pipeline engine for vulnerability research, with driver analysis living in an external docs site and a Ghidra headless processor. The interesting parts are what the manifest installs, what it deliberately does not, and two dependency pins written to stop a tool upgrade from breaking CI.
Who is it for?
DeepZero is a pipeline engine with a research project attached, and the two halves are unevenly documented. If you only want the orchestration, resumability and reporting, the demo runs from a clone with Python 3.11 and nothing else.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 26 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The quickstart filters two text files and calls it a demo

The five commands that make up the quickstart need Python 3.11 or later and nothing else:

sh
git clone https://github.com/416rehman/DeepZero.git
cd DeepZero
python -m pip install -e .
deepzero run pipelines/demo/samples -p pipelines/demo/pipeline.yaml
deepzero report -p pipelines/demo/pipeline.yaml --open

No API keys, no Ghidra, no driver corpus, and the README says so twice: the quickstart works in PowerShell and POSIX shells, and it does not run vulnerability analysis. What the demo actually does is discover two harmless text files, keep one and filter out the smaller one, then produce a browsable HTML report that `--open` launches. Run the same `run` command again and it resumes from saved state instead of starting over.

So the advertised product, which the repository description frames as thousands of Windows kernel drivers parsed for exploitable IOCTLs, is not reachable from what ships in the box. That path is behind the setup guide and pipeline prerequisites pages on the external documentation site.

The stage kinds in the bullet list do not match the stage directory

The feature bullets name four kinds of stage: ingest, filter, transform and LLM-assess. The source layout names three: `src/deepzero/stages/` holds the built-in processors and is annotated as map, reduce and ingest. `src/deepzero/engine/` covers orchestration, state persistence and pipeline execution, while `src/deepzero/api/` is the Starlette REST layer.

The external processors sit outside the package, in a `processors/` directory described as shipped examples, and each carries the base class it implements: `ghidra_decompile` is a MapProcessor wrapping the Ghidra headless decompiler, `loldrivers_filter` is a MapProcessor doing loldrivers.io hash exclusion, `pe_ingest` is an IngestProcessor parsing PE headers and driver metadata, and `semgrep_scanner` is a BulkMapProcessor wrapping the semgrep batch scanner. That gives four base classes across the two directories, so filter, transform and LLM-assess are configured through something the README does not show. The pipeline schema page is on the documentation site, not in the repository.

The repository map itself stops there too. The block ends mid-word, at `pipelin`, immediately after the processors list.

The serve extra installs a server for an API called incomplete

The extras in pyproject.toml are small and explicit:

toml
llm = ["litellm>=1.0"]
serve = ["starlette>=0.27", "uvicorn>=0.22"]
pe = ["lief>=0.14.0"]
full = ["deepzero[llm,serve,pe]"]

Against that, the REST API bullet in the feature list is labelled WIP and described as experimental and incomplete, and its stated purpose is narrow: query run state and sample data over HTTP. The consequence is a package you can install a web server into for an endpoint the project itself does not present as finished. `full` and `dev` both pull `serve` in, so anyone installing the development extra gets Starlette and Uvicorn whether or not they intend to serve anything.

The base install is five dependencies: click for the command line, rich for output, pyyaml, jinja2 for prompt templates and python-dotenv. The console entry point is a single script, `deepzero = "deepzero.cli:main"`, and packages are discovered from `src`.

The no-API-key route is a signed-in claude binary on PATH

The environment template offers two backends, chosen per pipeline through `model:` in the YAML or overridden per run with `deepzero run -m ...`. The first is called Claude Code, requires no API key, and uses a locally installed, already signed in Claude Code CLI on the machine running the pipeline. The template says nothing needs configuring for it beyond setting the model, which can be `claude-code` as the default, `claude-code/sonnet`, `claude-code/opus`, or a full model name. It also states the requirement plainly: `claude` on PATH and signed in.

The second is LiteLLM, metered and described as the right choice for CI and servers, with keys for whatever endpoints your pipelines target. The template ships them commented out: `GEMINI_API_KEY`, `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `OPENROUTER_API_KEY`, and a Vertex pair of `VERTEXAI_PROJECT` and `VERTEXAI_LOCATION`.

Note what the first route moves. Choosing the no-key option replaces a Python dependency with an external binary, an active login and a PATH entry, none of which the manifest declares or CI provisions.

Ghidra, semgrep and JDK 17 are all outside the manifest

Three of the four external processors depend on software the package does not install. The Ghidra section of the environment template is the only part with required keys, and both are placeholders:

bash
# GHIDRA_INSTALL_DIR=/path/to/ghidra_11.X_PUBLIC
# JAVA_HOME=/path/to/java/jdk-17

Those two paths are what `ghidra_decompile` needs, so a headless decompile run depends on a Ghidra install and a JDK that pip cannot provide. The semgrep batch scanner has no entry at all: the manifest lists `lief>=0.14.0` under the `pe` extra for PE header work and nothing for semgrep, so that processor expects the tool on your machine. The same is true of the `claude` CLI from the previous section.

So the full set for the driver path is: the package, one of the extras, a Ghidra installation, a JDK, a semgrep installation, a corpus of drivers, and a model route. The README defers the details to the setup guide and pipeline prerequisites pages on the documentation site rather than collecting them here.

ruff and bandit are pinned to a minor because their rules move

Two dependency pins carry comments explaining themselves, which is the clearest sign of what has already gone wrong:

toml
lint = ["ruff==0.16.*", "bandit==1.9.*"]

The comment says these tools change their own rules between releases, so an unpinned upgrade can fail the build with nothing in the repository having changed, and it asks for deliberate bumps plus local checks that match CI. The same reasoning shows up a second time under `[tool.ruff]`, where `extend-exclude = ["docs"]` keeps documentation out of the formatter because newer ruff versions format python snippets embedded in markdown, which made `ruff format --check` fail in CI purely from a ruff upgrade.

The contributor command runs all three checks in sequence:

bash
ruff check . && ruff format --check . && bandit -ll -ii -c pyproject.toml -r .

Ruff is configured with a line length of 100, target version py311, select E, F, W and I, and E501 ignored. Formatting is double quotes with space indentation, and the file that carries these settings ends mid-word in the middle of that line.

Four interpreters in CI, and Ghidra is opted into by marker

CI runs Python 3.11, 3.12, 3.13 and 3.14, and the reason Ghidra does not run on every one of them is written into the pytest configuration as a comment. The `ghidra` marker drives a local Ghidra install and requires `GHIDRA_INSTALL_DIR`, and it is selected on its own rather than making every interpreter under test pay to download one. That single marker is the difference between a four-interpreter matrix and a four-interpreter matrix that installs a decompiler four times.

Pytest runs against `testpaths = ["tests"]` with `pythonpath` set to both `src` and the repository root, so tests import both the installed package and the top-level scripts. The development extra is `pytest>=7.0`, `pytest-asyncio>=0.21.0` and `deepzero[full,lint]`, which chains the extras together, and asyncio support is there because parts of the engine are exercised asynchronously.

Versioning runs through release-please, with its manifest and config at the top level. The package version in pyproject.toml is 0.4.0, matching the `deepzero: v0.4.0` release dated 2026-08-02, and the previous one is v0.3.0 from 2026-07-26.

Per-sample atomic state is what makes a Ctrl+C cheap

The engine's three stated properties are worth separating, because only one is a design decision you have to pay for. Parallel execution is a ThreadPoolExecutor with concurrency configured per stage. Extension is Python classes for custom processors, referenced by path in the YAML, which is why the processors directory can hold four different base classes without touching the package.

Resumable runs are the interesting one: state is atomic per sample and written to disk, so interrupting with Ctrl+C and running the same command again picks up where you left off. That is what makes the demo safe to rerun, and it is also what lets a long driver corpus survive a crash mid-way. The trade is state on disk that has to be cleaned up between unrelated runs, and the README does not describe where that state lives or how to discard it.

Nothing in the README claims throughput numbers or a corpus size. The count of drivers in the description is the only scale claim available, and it comes with no benchmark attached.

Editorial conclusion

DeepZero is a pipeline engine with a research project attached, and the two halves are unevenly documented. If you only want the orchestration, resumability and reporting, the demo runs from a clone with Python 3.11 and nothing else. If you want the driver analysis the repository description promises, you are reading a site outside this repository for the schema and prerequisites, and you supply Ghidra with a JDK, a PE parser via the pe extra, semgrep from your own PATH, and an LLM route that is either a signed-in claude binary or a metered key. Check the license of any driver corpus you point it at and stay inside the scope you are authorized to test; the README states no testing scope of its own. Two version pins and the `docs` formatter exclusion tell you what has already broken in CI, and the last push was 2026-09-10, between the v0.4.0 release on 2026-08-02 and today.

Frequently asked questions

What does DeepZero need to run its demo pipeline?

Python 3.11 or later. Clone the repository, run `python -m pip install -e .`, then `deepzero run pipelines/demo/samples -p pipelines/demo/pipeline.yaml` and `deepzero report -p pipelines/demo/pipeline.yaml --open`. No API keys, Ghidra or driver corpus are required, and the demo does not run vulnerability analysis.

Can DeepZero call a model without an API key?

Yes. Setting `model: claude-code` in the pipeline YAML uses a locally installed, already signed in Claude Code CLI, which must be on PATH. The alternative is the metered LiteLLM route with keys for Gemini, OpenAI, Anthropic, OpenRouter or Vertex, and `deepzero run -m ...` overrides the model for one run.

What external tools do the DeepZero processors need?

The `ghidra_decompile` processor needs `GHIDRA_INSTALL_DIR` and `JAVA_HOME` pointing at a Ghidra install and a JDK. The `semgrep_scanner` processor expects semgrep on the machine, and the manifest lists no dependency for it. The `pe_ingest` processor uses lief, which comes from the `pe` extra.

Is DeepZero's REST API ready to use?

Not by its own account. The feature list labels it WIP and experimental and incomplete, with the scope of querying run state and sample data. The `serve` extra still installs starlette and uvicorn, and `full` and `dev` both include it.

Official sources

  1. 416rehman/DeepZero on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/416rehman-deepzero.svg)](https://hysenlabs.com/projects/416rehman-deepzero)