# Kunlun-M: static code analysis for PHP, Java, Go and C/C++ with an AI agent skill

> Kunlun-M (昆仑镜) is an MIT-licensed Python static analysis system that scans PHP, Nodejs/JavaScript, Python, Java, Go and C/C++ with AST-based rules and ships a skill directory so an AI agent can drive it. It is a maintained research tool, not a drop-in replacement for a commercial SAST platform.

**LoRexxar/Kunlun-M** — KunLun-M — Open-source static code analysis for PHP, Nodejs/JavaScript, Python, Golang, Java and C/C++, with AST-based semantic scanning and one-click AI Agent integration (OpenClaw, Codex, Claude Code, Hermes, and more).

- Repository: https://github.com/LoRexxar/Kunlun-M
- Stars: 2,416 · Forks: 317
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/lorexxar-kunlun-m

## The audit problem Kunlun-M is built around

Most static analysis tools are sold as a service. You upload code, or you install an agent, and the rules live somewhere you cannot read. Kunlun-M takes the opposite position. It is a local scanner with rules and tampers stored in a database you initialize yourself, and the repository keeps rules/ and core/ as ordinary directories. The README frames the project's history plainly: Cobra-W was a fork of Cobra 2.0 that shifted focus "from discovering as many threats as possible to improving the accuracy and precision of vulnerability detection", and Kunlun-Mirror evolved from Cobra-W 2.0 to serve security researchers rather than to maximize finding counts.

That framing matters when you decide whether to adopt it. The intended user is a white-box auditor or a security engineer who wants to read a rule, understand why a line was flagged, and write a new rule for a framework the tool does not know. The README states the tool primarily supports semantic analysis for PHP, Nodejs/JavaScript, Python, Java, Go and C/C++, with basic scanning for Chrome extensions and Solidity. If your codebase is none of those, the rule engine has nothing to match against.

## How the AST and tree-sitter analysis stack fits together

The dependency list in requirements.txt is the clearest description of the architecture. Two parsing generations coexist. The older layer uses lphply for PHP, lesprima, pyjsparser and jsbeautifier for JavaScript, and ljavalang for Java. The newer layer is tree-sitter: the requirements pin tree-sitter>=0.23,<1 together with grammars for Python, Java, JavaScript, TypeScript, Go, C, C++, Rust, Ruby, C#, Kotlin and Lua. python-igraph>=1.0,<2 is listed as the graph engine, which is what a semantic scan needs when it follows a value from a source to a sink rather than matching a single line.

So a scan is not a regex pass. The tool parses a file into a syntax tree, builds relationships between functions and calls, and applies rules against that structure. The practical consequence is that a rule can express "this function's argument reaches this sink" in a way a grep-based scanner cannot. The cost is that a grammar gap becomes a blind spot: if a language construct does not parse, the rule never fires, and the scan output will not necessarily tell you that it skipped something. The README does not document a coverage report for unparsed files, so treat clean output as "no rule matched", not as "no vulnerability exists".

## Installing Kunlun-M and running a first scan

The README gives a three-step install. First install the Python dependencies, then copy the settings template, then initialize the database. SQLite is the default backend, and mysqlclient is in requirements.txt for the MySQL path used with Docker. Python 3.10 or newer is recommended, with 3.13 preferred, and the README notes Python 2.7 has reached end-of-life.

```bash
pip install -r requirements.txt
cp Kunlun_M/settings.py.bak Kunlun_M/settings.py
python kunlun.py init
```

The init command creates the database that holds rules and tampers. Until it runs, scan and show have nothing to read from.

A first scan points at a directory. The README's own examples use the bundled test tree, which is the safest way to confirm the install before aiming it at real code.

```bash
python3 kunlun.py scan -t ./tests/vulnerabilities/
```

To write a machine-readable report instead of console output, the README shows the -f and -o flags with json, md and html as formats.

```bash
python3 kunlun.py scan -t ./tests/vulnerabilities/ -f json -o /tmp/report.json
```

Before scanning your own project, list what rules are loaded. The show command takes a rule or tamper argument, and -k filters by language.

```bash
python3 kunlun.py show rule -k php
```

If that list is thin for your target language, the scan will be thin too. This is the step most people skip.

## The AI agent skill and the CI scan driver

Two integrations in the repository are worth separating from the core scanner. The first is skills/kunlun-m-general/. The README instructs the user to send an AI agent a single line: download the repository and load its skill. The README states the agent will typically recognize the skills/kunlun-m-general/ directory and follow the documentation to complete initialization and scanning, and points to docs/skill_kunlunm_general.md for a scripted workflow with test and report commands. The README names OpenClaw, Codex, Claude Code and Hermes as examples.

This is a convenience layer, not a new analysis engine. The agent still runs python kunlun.py init and a scan; the skill supplies the instructions. Whether that is useful depends on whether you trust an agent to run a scanner over a checkout, which is a policy question about your own environment rather than a property of the tool.

The second integration is tools/ci_scan.py, which the README describes as producing stable JSON reports with clear exit codes and a gating flag.

```bash
python tools/ci_scan.py --target . --output artifacts/kunlun-ci.json --fail-on high
```

The README points to docs/ci.md for exit codes, report structure and GitHub Actions, GitLab CI and Jenkins examples. The repository layout backs this up: .gitlab-ci.yml, Jenkinsfile, .travis.yml and a ci/ directory are all present at the top level. If you want Kunlun-M to fail a build, this is the entry point, and the --fail-on threshold is the knob that decides how noisy that gate becomes.

## Where Kunlun-M is the wrong tool

The honest limitation is rule coverage, and the README does not claim otherwise. Rules ship in the repository and live in a database you populate at init. There is no managed rule feed described in the README, so a framework released after your last rules update is scanned with whatever the engine can still infer structurally. For an auditor writing custom rules, that is an acceptable trade. For a team that wants a vendor to keep pace with new CVEs in a third-party library, it is not.

The second limitation is scale and reporting. SQLite is the default backend, and the web mode is a local dashboard. The README documents an API token and a task-oriented API surface (/api/task/list, /api/task/<int:task_id>/result, and so on), which suggests the intended deployment is one instance you query, not a shared service with per-team isolation. Nothing in the README describes scheduling, retention, or multi-user access control.

The third is maintenance shape. The project is not archived, and the last push was on 2026-09-10, so the repository is current. The README is unusually direct about why: the maintainer says AI tooling lets them handle basic maintenance at low cost, describes the project's concepts as possibly not cutting-edge by today's standards, and states an intent to iterate rapidly and experiment. Read that as a signal about stability expectations. A tool that iterates boldly is a poor fit if you need rule behaviour to stay byte-identical across upgrades, and the README does not document a rollback path for rule or schema changes.

## Semgrep and CodeQL: the difference in approach

The closest alternatives are Semgrep and CodeQL, and the difference is where the rules live and who writes them.

Semgrep's model is pattern-based rules written in a YAML-like syntax that resembles the target language, distributed through a registry, with a CLI that runs locally. The barrier to writing a rule is low, and the ecosystem of community rules is the product. Kunlun-M's rules live in a database populated by init and inspected with python kunlun.py show rule, and the engine leans on tree-sitter grammars plus an igraph-based call graph for semantic matching. If your need is "write a quick pattern for our internal framework", Semgrep's authoring loop is shorter. If your need is "follow a value through function calls in PHP code we did not write", Kunlun-M's semantic layer is the more relevant design.

CodeQL takes the other extreme: you compile the target into a queryable database and write queries in a dedicated language, which gives deep dataflow analysis at the cost of a build step and a steeper learning curve. Kunlun-M does not require a build. You point scan at a directory. That is a real advantage for PHP and JavaScript codebases that have no build system at all, and a real disadvantage when you want the precision a compiled database provides.

None of these three is strictly better. The choice is whether you want rules you can read and edit locally (Kunlun-M), a large community rule registry (Semgrep), or maximum dataflow precision with a build step (CodeQL).

## Licence, upgrade cost and what to check before adopting

Kunlun-M is MIT licensed, which is permissive: you can use, modify and redistribute it, including in commercial settings, provided the copyright notice and permission notice are retained. That is a statement about the licence text, not legal advice, and if you plan to redistribute a modified version inside a product, have your own counsel read the LICENSE file at the repository root.

The upgrade cost is dominated by two things. First, the dependency pins. requirements.txt constrains tree-sitter to >=0.23,<1 and python-igraph to >=1.0,<2, and pins a long list of individual tree-sitter grammars. Grammar packages track upstream parser changes, so a routine pip install -r requirements.txt on a fresh machine can pull grammar versions that parse differently from the ones your rules were tuned against. Pin them in your own lockfile rather than relying on the ranges. Second, the settings file. Kunlun_M/settings.py is created by copying settings.py.bak, which means it is untracked local state. If a release adds a new setting, your copy will not have it, and the README does not describe a migration command for that file.

The practical checks before adoption are concrete. Run python kunlun.py show rule -k <language> and count what applies to your stack. Run a scan against a directory containing a vulnerability you already know about, and confirm the rule fires. Then wire tools/ci_scan.py into a non-blocking job first, look at what --fail-on high produces on your real code, and only then make it a gate.

## Conclusion

Adopt Kunlun-M if you audit source you already have on disk, want rules you can read and edit, and are comfortable with a Python 3.10+ toolchain plus a SQLite database created by python kunlun.py init. Do not adopt it expecting a hosted dashboard, a managed rule feed, or a polished multi-tenant service; the web mode is a local dashboard on port 9999, and the README does not document rollback for rule or schema changes. Before rolling it into a pipeline, verify three things on your own codebase: that the PHP, Java or Go parser handles your framework's syntax, that the rules you care about exist under python kunlun.py show rule -k <language>, and that the exit codes in tools/ci_scan.py match what your CI job treats as failure.

## FAQ

### What does Kunlun-M mean, and what is the project's relationship to Cobra?

The README states that since Cobra-W 2.0 the project was officially renamed to Kunlun-M (昆仑镜), and that Kunlun-Mirror evolved from Cobra-W 2.0, which was itself a fork of Cobra 2.0. The rename reflects a shift in focus toward serving security researchers.

### Which programming languages does Kunlun-M support?

The README states the tool primarily supports semantic analysis for PHP, Nodejs/JavaScript, Python, Java, Go and C/C++, with basic scanning for Chrome extensions and Solidity. The tree-sitter grammars in requirements.txt also cover TypeScript, Rust, Ruby, C#, Kotlin and Lua.

### How do I install and initialize Kunlun-M?

The README gives three steps: pip install -r requirements.txt, then cp Kunlun_M/settings.py.bak Kunlun_M/settings.py, then python kunlun.py init. SQLite is used by default, and Python 3.10 or newer is recommended with 3.13 preferred.

### Can Kunlun-M be run from an AI agent instead of the command line?

Yes. The README says that if you use an AI agent such as OpenClaw, Codex, Claude Code or Hermes, you can send it the instruction to download the repository and load its skill (kunlun-m-general), and the agent will typically recognize the skills/kunlun-m-general/ directory and follow the documentation to initialize and scan.

## Sources

- [Issues](https://github.com/LoRexxar/Kunlun-M/issues)
- [License: MIT](https://github.com/LoRexxar/Kunlun-M/blob/master/LICENSE)
- [LoRexxar/Kunlun-M on GitHub](https://github.com/LoRexxar/Kunlun-M)
- [README](https://github.com/LoRexxar/Kunlun-M/blob/master/README.md)
- [Releases](https://github.com/LoRexxar/Kunlun-M/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lorexxar-kunlun-m
