# sem: entity-level diffs, blame and impact analysis on top of Git

> sem parses a repository with tree-sitter and reports changes as functions, classes and methods rather than line ranges. It is aimed at coding agents and reviewers who need to know what a change touches, and its cross-file impact graph is the part worth evaluating first.

**Ataraxy-Labs/sem** — Semantic version control => entity-level diffs, blame, and impact analysis on top of git. 28 languages via tree-sitter. Built for coding agents.

- Repository: https://github.com/Ataraxy-Labs/sem
- Website: https://ataraxy-labs.github.io/sem/
- Stars: 3,368 · Forks: 102
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ataraxy-labs-sem

## The problem sem solves: line diffs hide what actually changed

A line diff answers a question nobody asks. When a pull request moves a function, renames a parameter and adds a branch, Git reports a block of added and removed lines, and the reviewer reconstructs the intent by hand. sem takes the position that the unit of change should be the entity: a function, a method, a class. The README puts it plainly, saying sem tells you what entities changed instead of lines changed, and gives the example of seeing that a function was modified rather than a line range.

The audience is narrower than "all developers". The repository topics list ai-agents, coding-agents and llm-tools, and the README places sem inside an Ataraxy Labs stack alongside weave (an entity-level git merge driver), inspect (semantic code review) and opensessions (a tmux sidebar for coding agents). The tool is built for the case where a program, not a person, consumes the diff. A line-oriented diff forces an agent to re-read surrounding code to learn which function it is looking at. An entity diff hands it the name.

That framing also explains the output formats. sem diff --format json exists so a pipeline or an agent can parse the result, and the README lists JSON output as being for AI agents and CI pipelines. Markdown output is listed for pull requests and reports. The plain format is described as git status style.

## How tree-sitter parsing turns a Git repo into an entity graph

The mechanism is a parse step in front of Git. sem reads file contents, runs them through tree-sitter grammars, and extracts every function, class and method as a named entity. The README states support for 32 languages in its badge, while the repository description says 28 languages via tree-sitter; the two numbers disagree, and the README is the more recent source. Either way, the grammar set is the hard boundary of what sem can see. A file in a language without a grammar falls back to something closer to text.

Entities are compared using structural hashing, which the README lists as a feature of sem diff alongside rename detection and word-level inline highlights. That combination is what makes a rename visible as a rename rather than as a delete plus an add. Verbose mode adds word-level inline diffs inside each entity.

The results are cached in SQLite. The README is explicit about where: the cache lives outside the repository, under the OS cache directory by default, and SEM_CACHE_DIR can override the cache root. Repo-local overrides are ignored so that cache files do not dirty the working tree. That is a deliberate trade: sem never writes into your checkout, so a .gitignore entry is unnecessary, but the cache is also not shared when you clone the same repo onto another machine, and the first run on a large tree pays the parse cost again.

sem impact reads a cross-file dependency graph rather than a single file. The README describes it as showing what breaks if an entity changes, with flags to narrow the result: --deps for direct dependencies, --dependents for direct dependents, --tests for affected tests only. sem log works the other way in time, tracking one entity through history, and with no entity argument it reports hotspots (most-changed functions and classes, with author counts) and co-change pairs, entities that repeatedly change in the same commits.

## Installing sem and running a first entity diff

The README gives several install paths. The shell installer is the shortest, and it pipes a script from the repository's main branch into sh.

```bash
curl -fsSL https://raw.githubusercontent.com/Ataraxy-Labs/sem/main/install.sh | sh
```

Homebrew users get a formula, and Windows users get winget or Scoop.

```bash
brew install sem-cli
```

```powershell
winget install AtaraxyLabs.sem
```

There is also an npm wrapper, published as @ataraxy-labs/sem, which downloads the matching release binary and exposes sem in node_modules/.bin. It requires Node 20 or newer. With Bun, the README says to trust the package first so its postinstall script can download the binary.

```bash
bun add -d @ataraxy-labs/sem
bun pm trust @ataraxy-labs/sem
```

If you installed through npm or Bun, the README notes that the binary lives in node_modules/.bin/sem and is invoked through npx sem or bunx sem. That detail matters for the name collision described below.

Once installed, the first useful command is a diff of your working changes. Run it inside any Git repo; the README states no setup is required.

```bash
sem diff
sem diff --staged
sem diff --commit abc1234
sem diff --from HEAD~5 --to HEAD
```

The first command shows unstaged working changes at the entity level. The second restricts the diff to staged changes, which is what you would run before committing. The third and fourth take a single commit or a range. For machine consumption, switch the format.

```bash
sem diff --format json
sem diff --format markdown
```

You should see a list of entities with their change status rather than a hunk of lines. If the output looks like a line diff, you are probably running the GNU Parallel binary of the same name; check with sem --version. The README documents that GNU Parallel ships /usr/bin/sem as a symlink to parallel, and that the two collide. The suggested fixes are an alias pointing at $HOME/.cargo/bin/sem, putting that directory first in PATH, or, for Homebrew installs, prepending $(brew --prefix)/bin to PATH.

sem also works outside Git. The README shows comparing two files directly, and reading a JSON array of file changes from stdin, which is how a host application would feed sem without a checkout on disk.

```bash
echo '[{"filePath":"src/main.rs","status":"modified","beforeContent":"...","afterContent":"..."}]' \
  | sem diff --stdin --format json
```

Docker is another route. The repository ships a Dockerfile that builds the sem-cli crate in a rust:1-slim-bookworm stage and copies the binary into a Debian slim image with the entrypoint set to sem.

```bash
docker build -t sem .
docker run --rm -it -u "$(id -u):$(id -g)" -v "$(pwd):/repo" sem diff
```

The -u flag matters: without it the container writes cache files as root. Note that the published npm wrapper in package.json reports version 0.16.2 while the latest release is v0.24.0, so the wrapper and the CLI binary are versioned separately.

## Where sem stops: merge conflicts, grammars and impact scope

sem analyzes. It does not merge. The Ataraxy Labs stack lists weave as a separate entity-level git merge driver, which means the conflict-resolution half of the problem lives in another tool. If your pain is repeated merge conflicts in generated or heavily edited files, sem will describe the conflict in entity terms but will not resolve it.

Language coverage is the second boundary. Structural hashing and rename detection depend on a tree-sitter grammar producing named entities. The README's badge says 32 languages and the repository description says 28. Whichever figure is current, the count is finite, and a file in an unsupported language cannot yield entity-level results. The README does not document what sem does for those files, so the fallback behaviour is something to check on your own repository before trusting a diff in a mixed-language tree.

The impact graph has a documented escape hatch that doubles as a warning. sem impact --no-default-excludes includes paths that are excluded by default: generated, fixture, vendor, benchmark and build trees. That default exclusion list is sensible for most repos, but it means an entity that is only referenced from a fixture or a vendored copy will not appear as a dependent. If your codebase leans on generated clients, the default impact answer will be incomplete, and you need the flag to see the full picture.

Finally, cloud-backed queries are opt-in per repo. The README states that logging in does not upload a repo or send a query, and points to docs/cloud-consent.html for the public and private repo states, the preview screen, the local audit log and the forget controls. That page is the place to read before anyone on the team authenticates, not after.

## sem against git diff and Lazydiff: different layers, not rivals

The honest comparison is with git diff itself. Git's diff is line-oriented, universal and free of any parser dependency. It works on a file Git has never seen a grammar for, which is why it remains the fallback for configuration, prose and any language sem does not support. sem is not a replacement for git diff; it is a second view that costs a parse and a cache. Teams that keep both get the entity summary for review and the raw hunk when something looks wrong.

Lazydiff appears in the related searches for this project, and the difference in approach is worth stating. A TUI diff viewer changes how a line diff is presented: navigation, side-by-side layout, keyboard-driven review. sem changes what the diff contains. One is a rendering layer over Git's output; the other re-parses the source and redefines the unit of change. If your complaint is that review is tedious, a viewer helps. If your complaint is that the diff does not name the function that changed, a viewer cannot help, because the information was never in the line diff.

The same distinction applies to sem blame. Git blame attributes lines to commits. sem blame attributes an entity to the commit that last modified it, which is a coarser and usually more useful answer for a function that has been edited across twenty commits. Neither is wrong; they answer different questions.

## Maintenance, licence and the cost of upgrading

The repository is not archived and the last push was on 2026-09-09, eight days before this writing. Releases are frequent: v0.24.0 landed on 2026-08-31, following v0.23.1 and v0.23.0 earlier in August. That cadence is a real cost. A CLI that moves through minor versions monthly will change output shapes, and anything parsing sem diff --format json is coupled to those shapes. The repository keeps a CHANGELOG.md at the top level, which is the file to read before bumping a pinned version in CI.

The project ships an update command, sem update, which fetches the latest release. Convenient for a laptop, less so for a build image where you want a fixed version. Pin the version in CI and run sem update only deliberately.

Licensing is dual. The repository carries LICENSE-APACHE and LICENSE-MIT, and package.json declares "MIT OR Apache-2.0". The repository metadata supplied here lists Apache-2.0 as the licence, and the README's badge points at LICENSE-MIT. Both files exist, so the dual grant is the accurate reading; the discrepancy is in how the metadata is reported, not in what is shipped. For a commercial user, the MIT option is the permissive one, but this is a description of the files present, not legal advice, and anyone embedding sem in a distributed product should read both licence texts.

The upgrade surface is small. There is no server, no schema migration to run and no config file in the repository root; the state that matters is the SQLite cache under the OS cache directory, which can be deleted and rebuilt. The npm wrapper adds its own version line, currently 0.16.2 in package.json against a CLI at v0.24.0, so a team installing through npm should verify which binary version the postinstall script actually fetched.

## Conclusion

Adopt sem if you want entity-level diffs and a dependency graph in CI or inside a coding agent, and if you can accept a SQLite cache that lives outside the repo. Skip it if your team needs merge conflict resolution rather than analysis, or if a GNU Parallel install already owns the sem name on your machines. Before rolling it out, run sem diff --format json on one real branch and check whether the entity names match how your team talks about the code, then read docs/cloud-consent.html before anyone logs in.

## FAQ

### What is Ataraxy-Labs/sem and how is it different from git diff?

sem is a semantic version control tool that runs on top of Git. It parses code with tree-sitter and diffs at the entity level, so it reports that a function, method or class changed rather than which lines changed.

### How do I install sem and run a first diff?

The README offers a shell installer, Homebrew as sem-cli, winget as AtaraxyLabs.sem, Scoop, an npm wrapper at @ataraxy-labs/sem, cargo install sem-cli, prebuilt release binaries and Docker. After installing, run sem diff inside any Git repo; no setup is required.

### Why does sem collide with GNU Parallel?

GNU Parallel ships a sem binary at /usr/bin/sem as a symlink to parallel, so both tools claim the same command name. Run sem --version to see which one you are invoking, then alias sem to $HOME/.cargo/bin/sem or put that directory first in PATH.

### Does sem send my repository or queries to the cloud?

The README states that cloud-backed queries are opt-in per repo and that logging in does not upload a repo or send a query. docs/cloud-consent.html documents the public and private repo states, the preview screen, the local audit log and the forget controls.

### Which languages does sem support?

The README's badge says 32 languages via tree-sitter, while the repository description says 28. The README does not document fallback behaviour for files in unsupported languages, so verify on a mixed-language tree before relying on entity output there.

### Where does sem store its cache?

sem stores a SQLite entity cache outside the repository, under the OS cache directory by default. Set SEM_CACHE_DIR to override the cache root; the README notes that repo-local overrides are ignored so cache files do not dirty the working tree.

## Sources

- [Ataraxy-Labs/sem on GitHub](https://github.com/Ataraxy-Labs/sem)
- [License: Apache-2.0](https://github.com/Ataraxy-Labs/sem/blob/main/LICENSE)
- [Project website](https://ataraxy-labs.github.io/sem/)
- [README](https://github.com/Ataraxy-Labs/sem/blob/main/README.md)
- [Releases](https://github.com/Ataraxy-Labs/sem/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ataraxy-labs-sem
