# Semgrep CLI: pattern-based static analysis for 30+ languages

> Semgrep Community Edition is a local, open source scanner whose rules look like the code they match. It is fast to set up and useful for one-off searches, but the README itself warns that single-file analysis misses many security true positives.

**semgrep/semgrep** — Lightweight static analysis for many languages. Find bug variants with patterns that look like source code.

- Repository: https://github.com/semgrep/semgrep
- Website: https://semgrep.dev
- Stars: 16,815 · Forks: 1,080
- Language: C
- License: LGPL-2.1
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/semgrep-semgrep

## What Semgrep Community Edition actually searches for

The README describes Semgrep as "semantic grep for code" and gives the clearest possible illustration: `grep "2"` matches the literal string 2, while Semgrep matches `x = 1; y = x + 1` when you search for the value 2. The tool parses source into a syntax tree and matches your pattern against that tree, so whitespace, variable names and formatting stop mattering. That is the whole pitch, and it is a real difference from regex-based scanning.

The audience is anyone who already greps a codebase and wishes the query understood syntax: developers enforcing a house style, reviewers hunting a known bug shape, maintainers auditing a dependency before adopting it. Rules are written in the target language itself rather than a separate query DSL, which lowers the entry cost. The README claims support for 30+ languages, listing Apex, Bash, C, C++, C#, Clojure, Dart, Dockerfile, Elixir, HTML, Go, Java, JavaScript, JSX, JSON, Julia, Jsonnet, Kotlin, Lisp, Lua, OCaml, PHP, Python, R, Ruby, Rust, Scala, Scheme, Solidity, Swift, Terraform, TypeScript, TSX, YAML, XML and a generic mode for templates such as ERB and Jinja. The list matters: a language absent from it cannot be scanned semantically.

## How a pattern becomes a finding: parsing, matching, and the single-file boundary

The repository is split between a compiled core and a Python wrapper. The Dockerfile explains the split directly: semgrep-core is built as a fully static binary on Alpine, where Musl allows static linking, and is then copied into a second Alpine image that carries the Python CLI, called pysemgrep in the file. The build stage needs ocamlc, gcc and make; the runtime image does not. That is a standard multi-stage build, and it explains why the distributed CLI is small even though the build toolchain is not.

The important architectural consequence is in the README's own warning. Semgrep Community Edition "can only analyze code within the boundaries of a single function or file." Cross-file and cross-function data-flow reachability are listed as capabilities of the AppSec Platform, not of the open source engine. So a rule that matches a call to a dangerous function will fire, while a rule that needs to know where the argument came from will not. For style enforcement and local pattern checks this is fine. For taint-style security analysis it is a hard ceiling, and the README says so before any competitor does.

The README also states that analysis runs locally and that by default code is never uploaded. That claim is about the default, not about every mode; the platform options in the README involve a hosted account, and the document does not describe what leaves the machine in those flows.

## Installing the Semgrep CLI and running a first scan

The README shows two installation routes from the CLI section. macOS users get a Homebrew formula, and Ubuntu, WSL, Linux and macOS users get a pip path. The README's second command is truncated mid-line, so use the canonical pip form rather than guessing flags:

```bash
# For macOS
brew install semgrep

# For Ubuntu/WSL/Linux/macOS
python3 -m pip install semgrep
```

A Docker image is also published, and the Dockerfile in the repository builds `semgrep/semgrep:canary`, so container-based runs are a supported distribution channel rather than an afterthought.

For a first real scan, change into a repository and let Semgrep pick a ruleset. The README does not print a full scan command in the excerpt, and the CLI's documented entry point is `semgrep scan`, with the registry-backed ruleset selected through `--config`. Expect findings printed to the terminal with file, line and the rule that matched. Read them one by one. A pattern-based scanner reports what the pattern says, not what the code intends, so the first run is usually a mix of real issues and rules that do not fit your conventions. To scan without network access to the registry, point `--config` at a local rule file instead; the repository itself keeps rules in `semgrep.yml` at the top level, which is a working example of the format.

If you want Semgrep to run before every commit, the repository ships `.pre-commit-hooks.yaml` and a root `setup.py` whose comment says it exists "for pre-commit since it expects a setup.py in repo root." The package it defines is named `semgrep_pre_commit_package` and depends on `semgrep==1.176.0`. That version pin is worth noticing: the hook package tracks a specific release, so a pre-commit setup can lag the CLI you installed by hand.

## Where the community edition stops being the right tool

The README is unusually direct about this. In security contexts, it says, Semgrep Community Edition "will miss many true positives" because analysis stops at function and file boundaries. If your goal is SAST, SCA or secrets scanning, the README recommends the AppSec Platform instead, citing cross-file and cross-function data-flow reachability, AI post-processing to reduce noise, and 20,000+ proprietary rules maintained by a security research team.

Those figures come from the vendor's own README and should be read as vendor claims, not independent measurements. The design point behind them is not in dispute, though: reachability analysis needs a whole-program view, and a scanner that parses one file at a time cannot have one. A taint rule that tracks a user-controlled value from an HTTP handler into a database call spans files by definition.

Two smaller limits are visible in the repository layout. One is language coverage: the generic mode handles templates such as ERB and Jinja, but anything outside the listed languages has no parser. The other is rule quality. The open source engine is a matching machine; the rules you feed it determine whether the output is signal or noise, and the README does not claim the community ruleset is curated to the same standard as the proprietary one.

## Semgrep vs SonarQube: pattern matching against a quality platform

The comparison people search for is Semgrep versus SonarQube, and the difference is architectural rather than a matter of rule counts. Semgrep is a scanner you point at a tree with a config; the config is the product. Rules are written as code-shaped patterns, so a team can add a rule for its own internal API in a few lines and run it immediately. There is no server to stand up and no project model to configure.

SonarQube is a platform with a persistent server, a database of historical results, and a fixed catalogue of quality and security rules per language. You get dashboards and trend data out of the box, and you give up the ability to express a new check as a snippet of the language you are already writing. If your need is "flag every call to this internal helper outside the auth module," Semgrep is the shorter path. If your need is a maintained dashboard of code quality over time with no rule authoring, a server-based platform fits better. They are not substitutes in every case, and running both is common because the overlap is partial.

The same distinction applies to the hosted Semgrep AppSec Platform: it is the same engine plus the cross-file analysis and rule catalogue, so choosing between community and platform is choosing between a matching library and a managed service.

## Licence, releases and the cost of staying current

Semgrep is distributed under LGPL-2.1, as stated in the repository metadata and in the header of `setup.py`, which carries the standard LGPL notice and points to the LICENSE file for details. LGPL is a copyleft licence with a linking exception relative to the GPL, but the obligations depend on how you distribute the software and any modifications, so read the licence text rather than treating this paragraph as guidance. Running the CLI inside a company to scan your own code is a different situation from redistributing a modified binary.

Release cadence is fast. The recent tags are v1.177.0 on 2026-09-10, v1.176.0 on 2026-09-01 and v1.175.0 on 2026-08-26, roughly weekly. The last push to the default branch, develop, was on 2026-09-21, so the repository is under active development. The practical cost is not the upgrade itself but rule drift: a ruleset pinned to an older CLI can behave differently after an upgrade, and the pre-commit package pins an exact version (`semgrep==1.176.0` in the root `setup.py`) precisely because floating versions break reproducibility. If you run Semgrep in CI, pin the version and bump it deliberately rather than tracking latest.

## Conclusion

Adopt Semgrep Community Edition if you want fast, local, pattern-based checks and are willing to write or review rules yourself; the README states it is not sufficient for security work because it only analyzes within a single function or file. If SAST, SCA or secrets scanning is the goal, the same README points to the Semgrep AppSec Platform for cross-file and data-flow analysis. Before committing, run a scan on one repository and read the findings rather than the count, then check whether your language appears in the supported list.

## FAQ

### Can I use Semgrep for free?

Yes. The repository is the open source Semgrep Community Edition, licensed under LGPL-2.1, and the README describes scanning code locally with the CLI. The README also points to a separate hosted AppSec Platform for cross-file analysis and a proprietary rule set.

### Is Semgrep better than SonarQube?

They take different approaches. Semgrep matches code-shaped patterns against parsed source and runs as a local CLI with a config, while a server-based platform like SonarQube keeps persistent project results and a fixed rule catalogue. Pick based on whether you need to author your own checks quickly or want a maintained dashboard.

### What is Semgrep used for?

The README describes it as static analysis that searches code, finds bugs, and enforces secure guardrails and coding standards. It can run in an IDE, as a pre-commit check, and in CI/CD workflows, and the README notes that scans run locally with code not uploaded by default.

### How do I install Semgrep on Ubuntu or Linux?

The README gives the pip route for Ubuntu, WSL, Linux and macOS: python3 -m pip install semgrep. A Homebrew formula covers macOS, and a Docker image is published for container-based runs.

### What does the Semgrep name mean?

The repository does not state the origin of the name. The README only frames the tool as semantic grep for code, contrasting it with a literal grep match.

## Sources

- [License: LGPL-2.1](https://github.com/semgrep/semgrep/blob/develop/LICENSE)
- [Project website](https://semgrep.dev)
- [README](https://github.com/semgrep/semgrep/blob/develop/README.md)
- [Releases](https://github.com/semgrep/semgrep/releases)
- [semgrep/semgrep on GitHub](https://github.com/semgrep/semgrep)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/semgrep-semgrep
