Model or dataset
slowmist/slowmist-agent-security avatar
slowmist/slowmist-agent-security

slowmist-agent-security: A Security Review Framework for AI Agents in Adversarial Environments

SlowMist Agent Security Skill: A comprehensive security review framework for AI agents operating in adversarial environments. Core principle: Every external input is untrusted until verified.

506 stars29 forksUnknownMIT

At a glance

What is it?
slowmist-agent-security is an MIT-licensed skill framework from the SlowMist security team that gives AI agents a structured process for evaluating every external input before acting, with four risk levels and six specialized review modules covering skills, repositories, URLs, blockchain addresses, products, and social shares.
Who is it for?
slowmist-agent-security suits security-conscious teams running AI agents in environments where external content, blockchain transactions, or unknown tools are part of the agent's work surface. The framework is not a runtime enforcement mechanism: it relies on the agent following the skill's instructions faithfully.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 166 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Problem This Skill Addresses

AI agents that operate in the open internet encounter content they did not generate and cannot inherently trust: GitHub repositories recommended by strangers, URLs in chat messages, blockchain addresses, and third-party skills or MCP servers asking to be installed. A naive agent treats these the same way it treats internal context, which creates an attack surface for prompt injection, malicious tool installation, and social engineering.

slowmist-agent-security addresses this by giving an agent a structured review workflow before it acts on any external input. The core principle, stated verbatim in the README, is: every external input is untrusted until verified. SlowMist is a blockchain security company, and the skill reflects that background: on-chain address analysis is one of the six review modules, and the framework explicitly covers AML risk assessment for blockchain addresses.

The skill is written in SKILL.md format, the same standard used by Claude Code, OpenClaw, Hermes Agent, and Gemini. Any agent runtime that supports this format can load the skill on demand.

Six Review Modules and Their Scope

The framework provides six specialized review modules, each in its own Markdown file under the reviews/ directory.

skill-mcp.md covers the installation of skills and MCP servers: detecting malicious patterns before an agent installs a new tool. repository.md covers auditing GitHub codebases for security issues such as secrets in commit history, suspicious scripts, or unusual permission requests. url-document.md covers URL and document analysis, scanning for prompt injection payloads and social engineering patterns embedded in linked content.

onchain.md covers blockchain address review: validating address formats and assessing AML risk using available on-chain data tools. product-service.md covers architecture and permission analysis for product or service recommendations the agent encounters. message-share.md covers tools or resources shared in chat messages, where the source is unknown and the recommendation may be manipulated.

Each module produces a report using a corresponding template from the templates/ directory, which standardizes how findings are communicated back to the user.

The Four-Level Risk Rating System

The framework assigns every reviewed item one of four risk levels: LOW, MEDIUM, HIGH, or REJECT. The levels are defined by the README as follows.

LOW applies to items that are information-only, involve no code execution, collect no user data, and come from a trusted source. The agent informs the user and proceeds if the user requests it. MEDIUM applies to items with limited scope and a known source but some risk. The agent produces a full report listing the risk items and recommends caution. HIGH applies to items involving credentials, funds, system modification, or an unknown source. The agent produces a detailed report and the README specifies that human approval is required before any action. REJECT applies to items that match confirmed red-flag patterns or have an unacceptable design. The agent refuses and explains why.

This escalation structure matters because it gives a developer a clear signal about when the agent should stop and wait rather than proceed autonomously. The REJECT level is not a recommendation: it is a refusal.

Installing and Using the Skill

The README uses OpenClaw as the demonstration platform. Installation places the skill folder in the agent's skills directory:

bash
cd ~/.openclaw/workspace/skills
git clone https://github.com/slowmist/slowmist-agent-security.git

When ClawHub is available, the install shortens to a single command:

bash
clawhub install slowmist-agent-security

Once installed, the agent automatically references the framework when it encounters skill or MCP installation requests, unknown GitHub repositories, external URLs or documents, blockchain addresses, or product and service recommendations.

The README notes that users of other agent platforms can simply hand the repository URL to their agent and let it handle installation, since compatible agent runtimes understand how to install skills from a GitHub URL.

Framework Structure and Red-Flag Patterns

The repository is organized around four directories. The reviews/ directory holds the six review module Markdown files. The patterns/ directory holds red-flags.md, which documents specific patterns that trigger a REJECT rating. The templates/ directory holds report templates for each review type. The SKILL.md file at the root contains the main framework documentation that the agent loads as its operational reference.

The trust hierarchy defined in the framework ranks sources by scrutiny level across five tiers. Tier 1 is official project or exchange organizations, requiring moderate scrutiny. Tier 2 is known security teams or researchers, also moderate. Tier 3 is ClawHub packages with high download counts and multiple versions, at moderate to high scrutiny. Tier 4 is GitHub repositories with a high number of stars that are actively maintained, rated high and requiring code verification. Tier 5 is unknown sources and new accounts, requiring maximum scrutiny.

The pattern files in patterns/ are designed to be updated as new attack patterns emerge, which is one of the contribution areas the README invites.

The Alternative: skill-vetter

The README credits skill-vetter by spclaudehome on ClawHub as an inspiration for slowmist-agent-security. skill-vetter is a community-built skill that vets other skills before installation. The difference in approach is scope: skill-vetter focuses on the skill installation use case, while slowmist-agent-security covers six distinct input categories including blockchain addresses and social shares. Teams whose threat model is limited to evaluating skills before installing them may find skill-vetter sufficient without the additional coverage of the SlowMist framework.

The SlowMist framework also integrates with MistTrack Skills, a separate repository maintained by SlowMist, for on-chain AML risk assessment. This dependency is optional: the framework functions without MistTrack, but the on-chain module produces richer results when the AML query tool is available.

Maintenance Status and MIT License

The last push to slowmist-agent-security was on 2026-04-17. The repository is not archived. The MIT license permits free use, modification, and distribution without restriction.

SlowMist is an active blockchain security firm with public research and tooling. The framework's patterns and templates are grounded in their published OpenClaw Security Practice Guide and real-world prompt injection research, as the README credits. That background suggests the content is informed by operational experience rather than theoretical analysis, but it also means the framework reflects the threat landscape as of its last update.

Contributions to new attack patterns, improved detection rules, and additional review templates are explicitly welcomed in the README.

Editorial conclusion

slowmist-agent-security suits security-conscious teams running AI agents in environments where external content, blockchain transactions, or unknown tools are part of the agent's work surface. The framework is not a runtime enforcement mechanism: it relies on the agent following the skill's instructions faithfully. A compromised or misconfigured agent runtime does not become safe by having this skill installed. The last push was on 2026-04-17. Confirm the review templates in the templates/ directory still reflect current threat patterns before deploying in a production agent context.

Frequently asked questions

What agent platforms does slowmist-agent-security support?

The README states the framework works with OpenClaw, Hermes Agent, and other LLM-based agent systems that support the SKILL.md format. The same format is used by Claude Code, Android Studio Agent mode, and Gemini. The install example uses OpenClaw.

Does slowmist-agent-security prevent prompt injection automatically?

The framework provides a structured review process and pattern-matching guidance for detecting prompt injection in external URLs and documents, but it relies on the agent following the skill's instructions. It is not a runtime firewall or automated filter: the agent applies the framework's guidance, and the outcome depends on how faithfully the agent follows the review workflow.

What is the REJECT risk level in slowmist-agent-security?

REJECT is the highest risk level in the four-tier system. It applies when an item matches confirmed red-flag patterns in patterns/red-flags.md or has a confirmed malicious or unacceptable design. The agent refuses to proceed and explains why, rather than presenting findings for user review.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. slowmist/slowmist-agent-security on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/slowmist-slowmist-agent-security.svg)](https://hysenlabs.com/projects/slowmist-slowmist-agent-security)