Open-source project
slowmist/openclaw-security-practice-guide avatar
slowmist/openclaw-security-practice-guide

slowmist/openclaw-security-practice-guide: a defense matrix you hand to the agent, not a checklist you run yourself

This guide is designed for OpenClaw itself (Agent-facing), not as a traditional human-only hardening checklist.

2,855 stars194 forksShellMIT

At a glance

What is it?
The repository packages an OpenClaw-facing hardening guide in two versions, v2.7 and v2.8 Beta, plus a red teaming validation document. Its premise is that the agent reads and deploys its own security controls, which is also where its limits sit.
Who is it for?
Adopt this if you already run OpenClaw with terminal or root-level access and you want the agent itself to carry the hardening work, starting from the v2.8 Beta document if your engine is version 2026.4 or later, or v2.7 if you are still on 2026.3 and earlier. Do not adopt it as a substitute for host-level controls, and do not expect it to make OpenClaw fully secure; the guide says so directly and puts final judgment on the human operator.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 177 days ago.
What is it written in?
Mainly Shell, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem: an agent with root access and a habit of installing things

Most hardening material assumes a human runs the commands. This repository assumes the opposite. Its stated target scenario is OpenClaw running with high privileges in a terminal or root-capable environment, continuously installing and using Skills, MCPs, scripts and tools, with the goal of maximizing capability while keeping risk controllable and auditable. The README frames the shift as moving from host-based static defense to what it calls Agentic Zero-Trust Architecture, aimed at destructive operations, prompt injection, supply chain poisoning and high-risk business logic execution.

The audience is narrow on purpose. If you run an agent that cannot touch a shell, or one whose tool set is fixed and reviewed by a human before every change, the threat model here does not match your setup. The guide is written to be handed to the agent in chat rather than read cover to cover by a person, and it says as much: you can send it directly to OpenClaw, let it evaluate reliability, and deploy the defense matrix with minimal manual setup. That is a deliberate inversion of the usual trust direction, and it is the single design decision everything else depends on.

How the 3-Tier Defense Matrix is structured

The README describes three layers. Pre-action covers behavior blacklists and a strict Skill installation audit protocol, which the guide positions against supply chain poisoning. In-action covers permission narrowing and cross-skill pre-flight checks, aimed at business risk control. Post-action covers a nightly automated audit reporting 13 core metrics, plus what the README calls Brain Git disaster recovery.

The four core principles underneath are more informative than the tiers. Zero-friction operations means the guide tries to keep manual security setup out of the user's day except when a defined red line is crossed. High-risk requires confirmation means irreversible or sensitive actions must pause for human approval. Explicit nightly auditing means every core metric gets reported, including the healthy ones, so there is no silent pass. Zero-Trust by default means the guide assumes prompt injection, supply chain poisoning and business-logic abuse are always possible.

The fourth principle is the one worth arguing with. A nightly audit that reports healthy metrics as well as failures produces a lot of output, and the guide's answer is that silence is worse than noise. That is a defensible position for an autonomous agent, since a missing report is indistinguishable from a crashed cron job. It also means the audit trail grows at a steady rate regardless of whether anything happened. The README does not describe a retention or rotation policy for those reports.

Installing it: you send a markdown file, the agent does the work

There is no package to install. The README's Zero-Friction Flow is a five-step chat procedure, and it explicitly states that the scripts/ directory exists for open-source transparency and human reference only, and that you do not need to manually copy or run it. Step one is choosing a version: the classic document at docs/OpenClaw-Security-Practice-Guide.md (v2.7) or the enhanced docs/OpenClaw-Security-Practice-Guide-v2.8.md (v2.8 Beta).

Step two is dropping that markdown file into your chat with the agent. Step three is the evaluation prompt, which the README gives verbatim:

text
Please read this security guide. Identify any risks or conflicts with our current setup before deploying.

What you should see is the agent reporting conflicts between the guide and your existing configuration before it changes anything. Treat that response as the real gate. If the agent proposes deployment without naming any conflict, it has not done the evaluation step.

Step four is the deployment command, and it differs by version. For v2.8 the README gives:

text
Follow the Agent-Assisted Deployment Workflow in this guide.

For v2.7 the README gives a longer instruction that names the red and yellow line rules, permission tightening and the nightly audit cron job. Step five is optional: the Red Teaming Guide at docs/Validation-Guide-en.md lets you simulate an attack and check that the agent correctly interrupts the operation. The README does not document a rollback procedure for any of this, which is a gap worth noting before you run the deployment step on a machine you care about.

A note on model choice, since the guide raises it: the README recommends a strong, latest-generation reasoning model from families such as Gemini, Opus, Kimi or MiniMax, on the grounds that better models handle long-context security constraints, detect hidden instruction patterns, and execute deployment steps with fewer mistakes. That is the project's own recommendation, not an independent measurement.

The version split is the sharpest limitation

The README is unusually direct about this. v2.7 underwent extensive production validation on OpenClaw version 2026.3 and earlier, and is now archived as a classic legacy version for users on older engines. v2.8 Beta targets OpenClaw version 2026.4 and later. The risk warning that follows is the part to read twice: because OpenClaw iterates quickly and its engine is unstable, some practices in the guide may become incompatible or ineffective after official updates, and testing only covers versions up to the date of the latest repository update. The README recommends thorough testing against your current OpenClaw version before production deployment.

That warning has teeth given the repository history. The last push was on 2026-04-06. If your OpenClaw build is newer than that, you are deploying a defense matrix whose compatibility with your engine is untested by the people who wrote it. The guide does not ship a version-detection step that would catch a mismatch for you.

There is a second limitation in the framing itself. The guide states plainly that it does not make OpenClaw fully secure, that absolute security does not exist, and that final responsibility and last-resort judgment remain with the human operator. An agent-facing hardening document that the agent can also misinterpret is a real failure mode: prompt injection is named as an assumed-possible threat, yet the deployment path runs through the same model that is the injection target. The guide's answer is human confirmation on red lines and a nightly audit, not a claim that the loop is closed.

Finally, this is the wrong tool if your agent has no shell access, or if you need a host-level control baseline. The README itself notes that traditional measures such as chattr +i and firewalls are either incompatible with agentic workflows or insufficient against LLM-specific attacks. That is an argument for adding this guide, not for replacing what you already run.

What it does not replace: host hardening and container isolation

The natural alternative for someone in this situation is a conventional host or container hardening baseline: an immutable filesystem attribute on sensitive paths, a firewall policy, a container runtime with a read-only root, or a seccomp profile. Those tools operate on the process and the kernel, and they do not care what the agent was asked to do or whether an instruction came from a poisoned skill description.

The difference in approach is where the control lives. A container policy decides what the process can reach no matter what the model believes. This guide decides what the agent should refuse to do, and then asks the agent to enforce that refusal on itself. The first is enforced by the kernel and cannot be talked out of it. The second is enforced by a model reading a document, which is why the guide leans on nightly audits and human confirmation for irreversible actions rather than claiming prevention.

Neither approach covers the other's blind spot. A seccomp profile will not notice that a newly installed skill quietly rewrote the agent's own instructions. The guide's Skill installation audit protocol is aimed exactly at that, and it is the part of the design with no equivalent in a standard container baseline. If you are choosing one, you are choosing which class of failure you would rather absorb.

Maintenance cost and what the MIT licence means here

The repository carries the MIT licence, so you can copy the documents, modify them for your environment, and redistribute them, subject to the licence terms. Since the deliverable is markdown rather than code, the practical implication is that forking and editing the guide to match your own red lines is permitted and probably expected. This is not legal advice; read the LICENSE file in the repository root if the distinction matters to your organisation.

Upgrade cost is the harder question. The project ships two parallel documents plus a red teaming guide, with English and Simplified Chinese versions of each, so a version bump means reconciling up to six files. The v2.8 workflow adds six steps (Assimilate, Harden, Pre-check Operator Scope, Deploy Cron, Configure Backup, Report), and the backup step is marked optional in the README. Because the agent performs the deployment, an upgrade is not a diff you review; it is a conversation you have with a model that then edits your configuration. The README does not describe an upgrade path from v2.7 to v2.8, and it does not describe how to revert a deployment. Plan for that gap before you treat either version as something you can move between casually.

Editorial conclusion

Adopt this if you already run OpenClaw with terminal or root-level access and you want the agent itself to carry the hardening work, starting from the v2.8 Beta document if your engine is version 2026.4 or later, or v2.7 if you are still on 2026.3 and earlier. Do not adopt it as a substitute for host-level controls, and do not expect it to make OpenClaw fully secure; the guide says so directly and puts final judgment on the human operator. Before deploying, verify two things in your own environment: which OpenClaw version you are actually running, since the guide warns that practices may become incompatible after official updates, and whether your operator scope matches the pre-check step in the v2.8 workflow. The last push to the repository was on 2026-04-06, so anything you deploy rests on testing that stopped there.

Frequently asked questions

Do I need to run the scripts in the scripts/ directory of slowmist/openclaw-security-practice-guide?

No. The README states that the scripts/ directory exists strictly for open-source transparency and human reference, and that you do not need to manually copy or run it. The agent is expected to extract the logic from the guide and handle deployment itself.

Which version of the OpenClaw security practice guide should I use, v2.7 or v2.8 Beta?

The README says v2.7 was validated on OpenClaw version 2026.3 and earlier and is now archived as a classic legacy version, while v2.8 Beta targets version 2026.4 and later. Pick based on the engine version you actually run, and test before deploying to production.

Does the OpenClaw security practice guide make OpenClaw fully secure?

The guide states directly that it does not make OpenClaw fully secure, that absolute security does not exist, and that it is built for a specific threat model and operating assumption. Final responsibility and last-resort judgment remain with the human operator.

How do I check whether the guide's controls still work after an OpenClaw update?

The README warns that practices may become incompatible or ineffective after official updates, and that testing only covers OpenClaw versions up to the date of the latest repository update. It recommends thoroughly testing and validating the procedures against your current OpenClaw version before deploying them in production.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. slowmist/openclaw-security-practice-guide on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/slowmist-openclaw-security-practice-guide.svg)](https://hysenlabs.com/projects/slowmist-openclaw-security-practice-guide)