Model or dataset
superagent-ai/superagent avatar
superagent-ai/superagent

Superagent SDK: runtime guardrails for AI agents, and what its TypeScript and Python clients actually do

Superagent protects your AI applications against prompt injections, data leaks, and harmful outputs. Embed safety directly into your app and prove compliance to your customers.

6,751 stars963 forksTypeScriptMIT

At a glance

What is it?
Superagent is an MIT-licensed SDK that classifies prompts, redacts PII and secrets, and scans repositories for agent-targeted attacks. The guard and redact paths are documented; the red team module is marked coming soon, and the runtime calls depend on a hosted API key.
Who is it for?
Adopt Superagent SDK if you already run an LLM-backed agent in TypeScript or Python and want a classification step in front of user input, or a redaction pass before text reaches a model or a log. Skip it if you need the red team module today, since the README marks Test as coming soon, or if you cannot send prompt text to a hosted endpoint.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 36 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Superagent SDK is for, and who ends up using it

The README frames the project as an SDK for AI agent safety, and the four documented capabilities map to four different jobs. Guard classifies an incoming message and returns a classification you can branch on, which is the piece you would put between a user and an agent loop. Redact strips PII, PHI and secrets from text, so it belongs on both the input side and wherever you persist transcripts. Scan analyzes a repository for repo poisoning and malicious instructions, which is a supply-chain check rather than a runtime one. Test runs red team scenarios against a live endpoint, and the README marks it as coming soon.

That split matters for adoption. A team building a customer-facing chat agent will use Guard and Redact. A team that pulls in third-party agent skills or tool repositories will care about Scan. The README states the SDK works with OpenAI, Anthropic, Google, Groq, Bedrock and other providers, so it is positioned as a layer that sits beside your model choice rather than replacing it. The repository is TypeScript-first, with a Python SDK, a CLI, and an MCP server for Claude Code and Claude Desktop listed as separate integration options.

How the guard and redact calls work in practice

The mechanism visible in the README is a client object plus a method call that returns a structured result. You create a client, pass input text to guard, and read result.classification. The documented values include "block", and a blocked result carries violation_types. That is the whole runtime contract: a classification string and a list of violation categories. Everything else, including how you respond to a block, is your code.

Redact follows the same shape but takes an extra parameter. The README example passes input and a model identifier, "openai/gpt-4o-mini", and returns a redacted string with placeholders such as <EMAIL_REDACTED> and <SSN_REDACTED> substituted in place. The model parameter is worth noticing. Redaction here is model-backed, not a fixed regex table, which is why the call takes a model name at all. That also means redaction quality varies with the model you point it at, and the README does not document which models are supported for this path beyond the single example.

Scan takes a repo URL and returns result.result as a security report plus result.usage.cost. The presence of a cost field on the response tells you scans are metered per call. Guard and Redact responses in the README do not show a usage field, so the cost model for those two is not documented in the README.

Installing the TypeScript SDK and running a first guard check

The README points you to superagent.sh to sign up and get an API key, and the key is read from the SUPERAGENT_API_KEY environment variable. Install the TypeScript package with npm, then export the key in the same shell before running any code.

bash
npm install safety-agent
export SUPERAGENT_API_KEY=your-key

With the package installed and the key exported, the smallest useful program creates a client and guards one message. The README gives this example, and the branch on classification is the part you would actually wire into an agent loop.

typescript
import { createClient } from "safety-agent";

const client = createClient();

const result = await client.guard({
  input: userMessage
});

if (result.classification === "block") {
  console.log("Blocked:", result.violation_types);
}

A blocked message prints its violation categories. A message that is not blocked falls through the if, and the README does not show what classification value it carries in that case, so log the raw result once before you rely on the branch. The Python path is the same shape: install with uv, then call client.guard(input=user_message) and compare result.classification against "block".

bash
uv add safety-agent
export SUPERAGENT_API_KEY=your-key

Where the SDK stops being the right tool

Two limits stand out from the README alone. First, Test is marked coming soon. If your reason for evaluating Superagent is red teaming an existing agent endpoint, the documented scenario list ("prompt_injection", "data_exfiltration") describes an interface that is not yet shipping, and the README does not give a date. Second, the runtime calls go through a client that needs an API key obtained from superagent.sh. The README does not document an offline mode for guard or redact, so the default path sends prompt text to a hosted service. For teams with data-residency constraints, that is a blocker before any accuracy question comes up.

The README does mention open-weight models and running Guard on your own infrastructure with 50-100ms latency, and links to HuggingFace. That claim is stated in the README, not demonstrated there. The README does not document the self-hosting procedure, the model names, or the hardware requirements, so treat the self-hosted path as something to confirm against the HuggingFace model cards and the docs site before you plan around it.

A third limit is subtler. Guard returns a classification, not a policy. There is no documented rate limit, timeout, or fail-open versus fail-closed behaviour when the guard call itself errors. If the guard request fails and your code treats a missing result as "allow", the safety layer disappears exactly when the system is under stress. The README is silent on this, and it is the first thing to test in your own integration.

Superagent SDK versus a self-hosted classifier like Llama Guard

The closest comparison is Meta's Llama Guard, which is a model you run yourself and prompt with a taxonomy. The difference in approach is where the policy lives. Llama Guard puts the category definitions in your prompt, so you control the taxonomy and you own the weights and the inference cost. Superagent SDK puts the taxonomy behind a method call and returns violation_types as a fixed set, with an option the README describes for running open-weight Guard models on your own infrastructure.

That trade is real in both directions. With a self-hosted classifier you get no vendor dependency and no per-call billing, but you own the prompt engineering, the serving stack, and the evaluation. With Superagent you get a defined interface and a redaction path that a raw classifier does not give you, but you inherit whatever categories the SDK exposes and, on the default path, a hosted dependency. If your requirement is "classify this text into my own five categories", a prompted model is less machinery. If your requirement is "block injections, redact PII, and scan repos without building three separate pipelines", the SDK is the shorter path.

Licence, maintenance and what upgrading costs you

The repository is MIT licensed, which permits commercial use and modification. The README does not state a separate licence for the hosted service or the API key, and an MIT licence on the SDK code says nothing about the terms of the API you call with it. If you plan to ship the SDK inside a product, read the terms attached to superagent.sh separately from the LICENSE file.

On maintenance, the last push to the default branch was on 2026-08-25. The most recent releases listed are rust-v0.0.9 and node-v0.0.9, both dated 2025-09-14. The version numbers are pre-1.0, and the README itself carries a coming-soon marker on one of four advertised features. Pre-1.0 packages can change method signatures between minor versions, so pin the version in package.json or pyproject.toml rather than tracking the latest tag. The repository layout shows a CHANGELOG.md at the top level, which is where to check before bumping. The README does not document a deprecation policy or a version support window.

Editorial conclusion

Adopt Superagent SDK if you already run an LLM-backed agent in TypeScript or Python and want a classification step in front of user input, or a redaction pass before text reaches a model or a log. Skip it if you need the red team module today, since the README marks Test as coming soon, or if you cannot send prompt text to a hosted endpoint. Verify two things first: whether the guard path can run fully on your own infrastructure with the open-weight models listed on HuggingFace, and what the scan call costs, because the README returns a usage.cost field per scan rather than a flat price.

Frequently asked questions

What is Superagent AI?

Superagent is an open-source SDK for AI agent safety, MIT licensed and written primarily in TypeScript. It provides guard, redact and scan capabilities, and the README describes a test module for red team scenarios as coming soon.

How much does Superagent cost?

The README does not state a price for the SDK or the API. It does show that a scan response includes a usage.cost field, which indicates scans are metered per call, and it directs you to superagent.sh to sign up for an API key.

How to use Superagent?

Install the package, export SUPERAGENT_API_KEY with the key from superagent.sh, then create a client and call guard, redact or scan. The README shows TypeScript and Python examples for each of those three calls.

What is Superagent in Node.js?

In Node.js it is the safety-agent npm package, which exposes createClient and the guard, redact and scan methods. The README installs it with npm install safety-agent.

What is a superagent in AI?

In this project the term refers to the Superagent SDK, which embeds guard, redact and scan calls into an AI application. Guard returns a classification such as "block" along with violation_types.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. superagent-ai/superagent on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/superagent-ai-superagent.svg)](https://hysenlabs.com/projects/superagent-ai-superagent)