# Joern: building code property graphs for cross-language static analysis

> Joern turns C/C++, Java, JavaScript, Python, Kotlin, bytecode and binaries into code property graphs stored in a custom graph database, then lets you query them with a Scala-based DSL. It is built for vulnerability discovery and static analysis research, not for drop-in CI linting.

**joernio/joern** — Open-source code analysis platform for C/C++/Java/Binary/Javascript/Python/Kotlin based on code property graphs. Discord https://discord.gg/vv4MH284Hc

- Repository: https://github.com/joernio/joern
- Website: https://joern.io/
- Stars: 3,540 · Forks: 466
- Language: Scala
- License: Apache-2.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/joernio-joern

## What Joern solves, and for whom

Most static analysis tools are built around one language and one report format. Joern takes a different position: it converts source code, bytecode and binary executables into a single graph representation called a code property graph (CPG), stores those graphs in a custom graph database, and exposes them through a Scala-based domain-specific query language. The README states the project is developed with the goal of providing a useful tool for vulnerability discovery and research in static program analysis.

The audience follows from that goal. Security researchers who want to trace taint from an input to a sink, engineers auditing a C library they did not write, and teams that need to compare a Java artifact against its compiled bytecode all fit. Someone who wants a preconfigured rule set that fails a pull request is not the target user, because Joern hands you a graph and a query language rather than a verdict.

## Code property graphs and the query layer

The pipeline has three visible stages. A language frontend parses input into a CPG; the CPG is written to a custom graph database; queries run against that database in a Scala-based DSL. The repository names the frontends by convention, for example javasrc2cpg for Java source, and the topics list covers C, C++, Java, JavaScript, Python, Kotlin, bytecode and binary, with Ghidra and LLVM appearing among the binary-related topics.

Cross-language analysis is the point of the graph model. Because control flow, data flow and the syntax tree are edges and nodes in one structure, a query does not care whether the node originated in source or in a decompiled binary. The specification lives at cpg.joern.io, separate from the user documentation at docs.joern.io. That separation is worth noting: the schema is versioned and detailed enough to be its own document, which tells you the graph model is treated as a public interface rather than an implementation detail.

The storage layer changed in v4.0.0, which the changelog describes as a migration from overflowdb to flatgraph. Earlier, v1.2.0 removed the overflowdb.traversal.Traversal class, and the changelog explicitly says that change is not completely backwards compatible. Queries written against pre-1.2 APIs will not run unchanged.

## Installing Joern and running a first query

The README gives a quick installation path for Linux and macOS. JDK 21 is listed as a requirement, with the note that other versions might work but have not been properly tested. The install script downloads the platform-specific zip for your machine.

```bash
wget https://github.com/joernio/joern/releases/latest/download/joern-install.sh
chmod +x ./joern-install.sh
sudo ./joern-install.sh
joern
```

After the script finishes, the joern command opens an interactive shell that prints the ASCII banner and a version line, then a joern> prompt. The README shows help as the command to type to begin. If the script fails, the documented fallback is to rerun it interactively.

```bash
./joern-install --interactive
```

Windows users are told to skip the script and download the appropriate zip directly from the releases page, then extract it manually. Releases are created automatically once per day, and the published zips are named by platform, including joern-cli-linux-x86_64.zip, joern-cli-macos-arm64.zip and joern-cli-windows-x86_64.zip.

Docker is the other documented route, and it is the one that avoids installing a JDK locally. The image mounts the current directory at /app and runs the interactive shell there.

```bash
docker run --rm -it -v /tmp:/tmp -v $(pwd):/app:rw -w /app -t ghcr.io/joernio/joern:master joern
```

Adding --server to the same command starts Joern in server mode instead of the interactive shell. One platform caveat is documented: Almalinux 9 requires a CPU supporting SSE4.2, and for kvm64 virtual machines the README points to the joern-alma8 image tag instead.

## Where Joern gets in the way

The binary and bytecode frontends rely on fuzzy parsing and decompilation, and that is a real accuracy boundary. A CPG built from a stripped binary will not carry the same fidelity as one built from source with full type information, and the topics list itself names fuzzy-parsing as a characteristic of the project. If your question depends on precise type resolution in a binary you cannot rebuild, the graph may not support the answer.

Build and integration cost is the second constraint. The development setup expects sbt, and the Dockerfile pins SBT_VERSION 1.12.1 and openjdk-21-jdk, so the toolchain is not incidental. The unit test path is sbt test, and the integration path stages a distribution and runs a Python driver against it. That is a heavier loop than editing a YAML rule file.

Query language churn is the third. The v1.2.0 removal of overflowdb.traversal.Traversal was documented as not completely backwards compatible, and v4.0.0 changed the underlying graph store. Anyone maintaining a private query library should expect periodic rewrites, and the changelog directory exists precisely because those migrations need explanation.

## Joern compared with Semgrep-style pattern matching

The closest mental alternative for many teams is a pattern-matching scanner such as Semgrep, which matches syntactic patterns and lightweight dataflow within a single language and is designed to run in CI with low setup cost. The difference in approach is structural. Semgrep evaluates rules against parsed source and returns findings; Joern builds a persistent graph and hands you a query language over it.

That means Joern can answer questions a pattern matcher cannot easily express, such as following a value through a decompiled binary or joining facts from a Java source frontend with facts from the compiled bytecode. It also means Joern will not give you a rule library and a pass/fail gate out of the box. Choosing between them is a question of whether you want a detector or a workbench. The README's own framing, calling Joern a workbench, is accurate about which one this is.

## Licence, releases and the cost of staying current

Joern is licensed under Apache-2.0, a permissive licence that permits commercial use and modification, with the usual obligations around notices and the absence of warranty. This is not legal advice; if you redistribute a modified Joern or embed it in a product, have counsel review the notice requirements.

The release cadence is the practical cost. The README states a new release is created automatically once per day, and contributors can trigger the workflow manually if they need it sooner. Three releases appeared within roughly two days in the recent history. That cadence is good for fixes and awkward for pinning: an organisation that needs reproducible analysis output should record the exact version it used rather than tracking latest, because the graph store changed at v4.0.0 and the traversal API changed at v1.2.0. The repository also carries a .scala-steward.conf, so dependency updates arrive continuously as well.

## Conclusion

Adopt Joern if you need to reason about data flow and control flow across languages, or if you want to query binaries alongside source in one graph model. Do not adopt it as a lightweight linter or as a substitute for a build-integrated SAST tool: it needs JDK 21, produces its own graph database, and expects you to write queries. Before committing, verify that the CPG frontend for your target language parses your codebase without excessive fuzzy parsing, and check the current query language against the v4.0.0 flatgraph migration notes, since the storage layer changed and older traversal APIs were already removed in v1.2.0.

## FAQ

### What is Joern?

Joern is an open-source platform for analyzing source code, bytecode and binary executables. It generates code property graphs, stores them in a custom graph database, and lets you query them with a Scala-based domain-specific query language.

### How do I install Joern?

The README gives a quick install path: download joern-install.sh from the latest release, make it executable, run it with sudo, then start the joern command. It requires JDK 21, and Windows users are told to download the platform zip from the releases page and extract it manually.

### Is Joern open source?

Yes. The repository is licensed under Apache-2.0 and the source is published on GitHub under joernio/joern.

### How do you pronounce the name Joern?

The README and the documentation it links to do not cover pronunciation, so there is nothing in the project's own pages to confirm this.

### What are the alternatives to Joern?

The project's own pages do not list alternatives. The structural contrast is with pattern-matching scanners that evaluate rules against parsed source and return findings, whereas Joern builds a persistent code property graph and exposes a query language over it.

## Sources

- [joernio/joern on GitHub](https://github.com/joernio/joern)
- [License: Apache-2.0](https://github.com/joernio/joern/blob/master/LICENSE)
- [Project website](https://joern.io/)
- [README](https://github.com/joernio/joern/blob/master/README.md)
- [Releases](https://github.com/joernio/joern/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/joernio-joern
