Three of Semia's four stages are deterministic, and the fourth one writes the fix
Semia, security audit for AI agent skills.
At a glance
- What is it?
- berabuddies/Semia audits agent skill files by reading them as data, mapping what they could do, and evaluating rules over that map. The audit never runs the skill; the repair path hands a language model the job of rewriting it.
- Who is it for?
- Use it to decide whether to trust a skill, not as a gate that certifies one, since the deterministic half produces findings and the other half decides what the prose around them means. Read the SARIF output rather than only the markdown if you plan to act on the results, and treat a patched file as untrusted input again.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 34 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The audit is deterministic except for the step that reads the skill
The scan is four stages, prepare, synthesize, detect and report, and only one of them needs a model. The other three are described as deterministic and needing no key. What lands on disk makes that split explicit. The normalized skill text is written with stable line anchors, a units file records what the evidence text is aligned against, and the model's own contribution is confined to one file holding behaviour-mapping facts in Datalog form, described as re-queryable. A second Datalog file holds findings derived by rule evaluation, so the conclusions are separated from the model's output and can be re-derived without calling a model again. A metadata file records provider, model, retries, score and stop reason, and a run manifest ties the whole thing together. The user-facing output is a markdown report whose findings are ranked by severity and each tied to a specific source line, which is also what lets the SARIF form annotate a skill pull request directly in code scanning rather than pointing at a file.
The repair step is the one that writes a new skill with a model
This is the tension in the design. `semia repair` takes a run directory and, with a flag pointing back at the scan, reads the findings and the synthesized facts, traces each violation back through the Datalog rules to work out a root cause, and then calls a language model to generate a patched skill file. The patch either rewrites the problematic content directly or adds explicit security constraints. The result is written to a `patched` subdirectory of the run rather than over your original. Two things are not said anywhere in the documentation. Nothing states that the patched output is scanned again, and nothing states that the added constraints are checked against the same rules that produced the findings. So the tool built on the principle that you should not trust a file because its prose looks fine ends by generating new prose with a model and handing it back to you.
Three hosts, three key models, and one path from another registry
Three agent hosts are supported and they are not equivalent. Claude Code gets the shortest documented path, two shell commands that add the marketplace and install the plugin, or an interactive panel where you press the right arrow twice to reach the marketplace entry:
claude plugin marketplace add berabuddies/Semia
claude plugin install semia@semiaCodex is subtly different: the shell command only adds the marketplace, and enabling the plugin means editing the Codex configuration file by hand to add an enabled flag under the plugin's name. OpenClaw is different again, a single install command from a registry namespace that is not this repository's name:
openclaw plugins install clawhub:semiaAfter that you ask the host agent in chat to run an audit on a skill path, the host's own model does the synthesis step, and no API key of yours is involved. That last path also runs a bundled zipapp for the deterministic stages, and it is the one path this repository's own build does not produce: the make targets for bundling plugins cover the Codex and Claude Code hosts only, and the shell-outs to a local agent CLI are treated the same way elsewhere, as a delegation of the synthesis step rather than as a bundled component.
The two documented run paths disagree on the run directory name
The CLI section says a scan writes to a runs directory keyed by a skill slug, and that an output flag overrides the location. The section describing the artifacts says a run writes everything under a runs directory keyed by a run identifier. Both cannot be the placeholder for the same directory, and the commands that come later take a path argument whose shape follows the first version. Nothing in the documentation says which one a script should expect. A smaller detail in the same area: the project excludes the dot-prefixed run directory from its own linting, which makes sense since it is generated output, and the same exclusion list covers build artefacts and an output directory.
Two wire formats and two shell-outs, and only the model survives on the shell-outs
Provider selection has four values: an HTTP path against an OpenAI-style Responses API, which is the default and whose base URL can be pointed at other compatible endpoints, an HTTP path against an Anthropic-style Messages API, and two shell-outs that invoke locally installed command-line agents and reuse the login you already have there. The asymmetry is stated in the sample environment file rather than the README: for the two shell-outs only the model argument is honoured. So a timeout, a retry count and a temperature setting that work on the HTTP paths have no effect when you delegate synthesis to a local CLI. The defaults are a model name in the current generation for the default provider and a different one for the Anthropic-family options, with a request timeout of six hundred seconds and two retries per call. Model names are free-form strings and the tool says outright that it does not validate them.
Temperature is chosen by matching model names, and no dependency list exists
Two details explain how little the tool assumes about your endpoint. The temperature setting is not a number you supply by default; an unset value triggers a rule that inspects the model name and sends zero for chat-style model families and omits the parameter entirely for reasoning-style families, and you can override that by setting it empty to force omission. Since model strings are free-form and unvalidated, an endpoint with an unusual name simply gets the default treatment. A separate model variable acts as a fallback when you pass none on the command line, which is what makes the provider defaults reachable without editing the invocation each time. The second detail is the package metadata. The runtime dependency list is empty, the build system declares no requirements at all, and the build backend is an in-tree module referenced by path rather than something installed from an index. The keyword list still names the Datalog engine the detector is built around, so a zero-dependency install is carrying a rule evaluator with it somehow.
The build actively checks that development state does not leak
The make file is longer than the release tooling usually needs, and the interesting targets are the guards. One target asserts that the source distribution does not contain the test directory or coverage output. Another smoke-tests the installed package, and another smoke-tests the zipapp bundles after they are built, which is the artifact the host plugins actually run. A third validates the plugin manifests, and the default check target compiles the sources, runs the tests and validates those manifests in sequence. Type checking covers the package tree and the build backend together. Secret scanning runs against the full commit history in continuous integration, and the sample environment file says so explicitly and warns against committing real keys. The test fixtures are described in the lint configuration as external skills used as untrusted input, which is a nice admission for a scanner's own test corpus. Two documentation files sit next to the readme with the worked example and the trust model in one and the repository layout in the other, so the reasoning behind the four-stage split is written down rather than left to the code.
Three tags in five weeks, then three months of commits with no tag
The release history is short. The first two tags are a day apart in mid-May, the third arrives in early June, and that is where publishing stops. The last push to the default branch is dated 2026-09-01, roughly three months later, so the working tree has moved well past the newest release while the version in the package metadata still reads the same three-part number as that tag. The two surfaces do at least agree on the number, which is more than several projects in this batch manage. The release titles are less tidy: the two older entries carry a prefix naming the release while the newest does not, which is the kind of inconsistency that makes a changelog filter miss entries. Around the code sits a full governance file set for a package at version zero point one: a licence file, a notices file that the packaging metadata lists alongside it as a licence file, a security policy, a privacy policy, a trademark policy, a contributing guide and a changelog. The declared author is a different organisation from the account that owns the repository, and the licence is the permissive Apache 2.0 rather than anything restrictive, which fits a tool whose argument is that you should be able to read its rules.
Editorial conclusion
Use it to decide whether to trust a skill, not as a gate that certifies one, since the deterministic half produces findings and the other half decides what the prose around them means. Read the SARIF output rather than only the markdown if you plan to act on the results, and treat a patched file as untrusted input again. Before wiring a host plugin, check which one you are installing: two are built from this repository and the third is fetched from a separate registry namespace.
Frequently asked questions
Does Semia execute the agent skills it audits?
No. It reads a skill as data and never executes it. The scan runs four stages where only the synthesize step needs a language model, and the bundled zipapp handles prepare, detect and report deterministically.
Which model providers can Semia use for its synthesis step?
Four: an OpenAI-style Responses API as the default, whose base URL can point at other compatible endpoints, an Anthropic-style Messages API, and two shell-outs to locally installed command-line agents that reuse your existing login there. Model names are free-form strings and the tool states that it does not validate them.
What files does a semia scan produce?
A report in markdown is always written, and a SARIF report and a JSON report can be generated on request for code scanning and for programmatic consumers. The run directory also holds a re-queryable Datalog behaviour map, findings from rule evaluation, normalized skill text with stable line anchors, and a metadata file recording provider, model, retries, score and stop reason.
Can Semia repair a skill for me?
Yes. The repair command reads findings and synthesized facts from an existing scan, traces each violation back through its rules to a root cause, then calls a language model to generate a patched skill file, either editing the problematic content or adding security constraints. The patch is written under a patched directory in the run.
Which agent hosts can Semia be installed into?
Codex, Claude Code and OpenClaw. The repository's bundling targets cover the Codex and Claude Code plugins, and the OpenClaw route is a single install command from a separate registry namespace. Codex also needs a manual configuration edit to enable the plugin after adding its marketplace.
What Python versions and dependencies does Semia need?
Python 3.11 and above, with classifiers for 3.11 and 3.12 and a linting target of 3.11. The package is classified as Alpha and declares no runtime dependencies, and its build system requires nothing from an index because the build backend lives in the repository.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/berabuddies-semia)