# OPFOR: adversary emulation for AI agents and MCP servers

> Agent-opfor (OPFOR) is an Apache 2.0 TypeScript tool that fires OWASP-mapped attacks at LLM apps, agents and MCP servers, then judges each reply with an LLM. The CLI is the solid part; the browser extension and MCP server are what set it apart from probe libraries.

**KeyValueSoftwareSystems/agent-opfor** — Open-source adversary emulation for AI agents and MCP servers. 

- Repository: https://github.com/KeyValueSoftwareSystems/agent-opfor
- Stars: 582 · Forks: 33
- Language: TypeScript
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/keyvaluesoftwaresystems-agent-opfor

## The gap OPFOR was built to fill: agent surfaces nobody was attacking

The README is unusually direct about the origin. The maintainers write that they have shipped 130 products for 90 startups over ten years, and that in the last 18 months almost every one had an AI agent in it, with every team hitting the same testing wall. OPFOR is the internal answer, released under Apache 2.0.

The name is the thesis. OPFOR stands for Opposition Force, a military term the README defines as the unit that plays the enemy in training. The stated goal is that to defend AI agents better, you have to attack them first.

The intended audience is broader than security engineers. The README lists product managers, designers, QA and security analysts as users of the browser extension, alongside engineers who want the CLI and developers who want slash commands in an IDE. That split is the interesting design decision. Most red-team tools assume a Python notebook and a target URL. OPFOR assumes several people on one team, only some of whom will write code.

## How a scan actually runs: plan, emulate, judge

The README describes a four-step flow. First, OPFOR fetches target info: it connects to your agent and detects available tools, MCP endpoints and capabilities. Second, it plans attacks per category, generating targeted prompts for each evaluator in the selected suite. Third, it emulates the attack by running multi-turn adversarial conversations with real requests and real responses. Fourth, an LLM judge evaluates the replies.

That ordering matters. Target introspection comes before prompt generation, so attacks are shaped by what the agent can actually reach rather than drawn from a static list. The README claims OPFOR is "trace-aware", integrating with Langfuse and Netra so the judge sees what the agent did internally, not only what it said. That is the part worth checking in docs/ if you run agents that call tools, because a response that looks clean can hide a dangerous tool invocation.

Coverage is the other half of the pitch: OWASP LLM Top 10, OWASP Agentic AI Top 10, OWASP MCP Top 10, OWASP API Security, plus EU AI Act bias suites. The repository layout backs this up with top-level evaluators/, suites/ and data/ directories, and the README states that all five run modes share the same evaluators, attack templates and judge logic. The README also says every attack prompt, request, response and judge verdict is logged.

One caveat: the README excerpt does not name the judge model, nor say whether judging can be swapped independently of the attacker. If you need a judge you can audit, that is a question for the docs.

## Installing the CLI and running a first scan against your own agent

The README gives a global npm install and one provider key. Node 20 or later is required, per the engines field in package.json.

```bash
npm install -g @keyvaluesystems/agent-opfor-cli
export OPENAI_API_KEY=your-key    # or GEMINI_API_KEY, ANTHROPIC_API_KEY, etc.
```

The env var name matters. The .env.example file lists GROQ_API_KEY, OPENAI_API_KEY, ANTHROPIC_API_KEY and GOOGLE_GENERATIVE_AI_API_KEY, and states that any one of them is enough to run a scan. It also documents OPFOR_API_KEY for OpenAI-compatible endpoints such as LiteLLM, OpenRouter, Azure or Ollama, with the base URL set in the opfor config rather than the environment.

There are two ways to start. The one-shot form runs a setup wizard and immediately begins the scan:

```bash
opfor run
```

The two-step form saves a config you can reuse or commit to CI:

```bash
opfor setup                               # wizard saves a config to .opfor/configs/
opfor run --config .opfor/configs/<file>  # run any time against the saved config
```

Use the two-step form for anything you intend to repeat. The wizard writes into .opfor/configs/, and the README explicitly frames that saved file as something you can commit, which is what makes a scan reproducible across machines and reviewers.

For a first real use, point a saved config at a staging agent rather than production. The flow opens multi-turn adversarial conversations and records every request and response, so a production target will accumulate real attack traffic and real logs.

## Five entry points, and the one with a hard dependency

OPFOR ships five ways to run: CLI, browser extension, MCP server, skills, and an SDK. The README's table maps each to a user. The CLI is for engineers and CI/CD. The browser extension is for anyone who cannot or will not write code. The MCP server registers OPFOR in Cursor or Claude Desktop so a coding agent can red-team your other agents through chat. Skills are slash commands (/opfor-setup, /opfor-run, /opfor-mcp-setup, /opfor-mcp-run) for in-IDE testing. The SDK installs as @keyvaluesystems/agent-opfor-sdk and exposes run and hunt for programmatic use.

The dependency worth flagging is opfor hunt, the autonomous mode. The .env.example states plainly that hunt does not use the general provider keys, because its commander, operator and scout agents run on the Claude Agent SDK and are Claude-only. Your target can still be any model or agent. Credentials resolve in a documented order: ANTHROPIC_API_KEY first, then ANTHROPIC_BASE_URL paired with ANTHROPIC_AUTH_TOKEN, then CLAUDE_CODE_OAUTH_TOKEN from claude setup-token, then ~/.claude/.credentials.json from claude login.

The file also warns about a real failure mode. Setting ANTHROPIC_AUTH_TOKEN alone is ignored, because it is indistinguishable from a token inherited from a parent Claude Code session, and the run silently falls through to option 3 or 4. The consequence the file names is billing your personal subscription instead of the gateway. That is the kind of misconfiguration that shows up on a statement rather than in a log, so set both variables or neither.

## Where OPFOR is the wrong tool

If you need a deterministic, offline regression suite, this is not it. The judging step is an LLM, the README says so without hedging, and that means verdicts can vary between runs. For a CI gate that must produce the same answer twice, you would need to pin the judge and the attacker and still accept some variance. Nothing in the README excerpt describes a deterministic mode or a golden-output comparison.

If your attacker budget is a non-Anthropic key and you want autonomous hunting, the Claude-only constraint on opfor hunt is a hard stop, not a preference. You can still run standard scans on any supported provider, but the autonomous path is closed.

If your target is a plain model endpoint with no tools, memory or MCP surface, you are paying for capability you will not exercise. The README's own framing is that OPFOR was built for agents rather than just models, designed around tool calls, MCP, memory and multi-turn state from day one. A static prompt-injection scanner will be cheaper and simpler for that case.

Finally, the repository metadata reports the licence as NOASSERTION while the README badge and package.json both say Apache-2.0. That discrepancy is not a reason to avoid the tool, but it is a reason to read the LICENSE file before you depend on it.

## OPFOR against a probe library or a programmatic framework

The README makes its own comparison, and it is worth quoting the shape of it rather than the wording: most red-team tooling in this space is excellent at one thing, whether that is a probe library, a developer evaluator, or a programmatic framework, and OPFOR's claim is coverage in one tool.

Take a probe library such as garak as the contrast. A probe library gives you a catalogue of attack strings and a harness that fires them at a model endpoint and scores the raw text. It is fast, deterministic, easy to run offline, and it does not care what your application does. What it cannot see is the agent layer: which tools were called, what an MCP server returned, what got written to memory between turns. OPFOR's answer is target introspection before planning and trace integrations with Langfuse and Netra, so the judge can consider internal behaviour.

A programmatic framework such as promptfoo inverts the trade-off differently. It is a test runner you configure and extend, with assertions you write yourself, which makes it a good fit for teams that want their own scoring logic and their own CI wiring. OPFOR hands you the evaluators and the judge instead. You get OWASP-mapped suites out of the box and less control over how a verdict is reached.

Neither approach is strictly better. If your team already writes its own assertions, OPFOR's built-in judge is a layer you did not ask for. If your team has no security testers, that same judge is the reason the tool gets used at all.

## Maintenance, upgrades and what the licence metadata does not settle

The last push to the default branch was on 2026-09-10, and the most recent release in the list is v0.10.2 from 2026-08-10, following v0.10.1 on 2026-07-20 and v0.10.0 on 2026-07-06. The repository is not archived. On that evidence the project is being pushed to, and the release cadence over that window is roughly one minor release per month at the 0.x line.

The 0.x versioning is the practical upgrade risk. The workspace root package.json marks the root private and lists core, runners/cli, runners/mcp, runners/extension, runners/sdk and tests/e2e/sdk as workspaces, which means the published artifacts are the workspace subpackages rather than the root. A minor bump can therefore change a subpackage's behaviour, and the SDK surface (run and hunt) is the part most likely to move. The repository carries a .release-please-manifest.json and release-please-config.json, so releases are automated from commit messages; the CHANGELOG.md at the root is where the actual breaking changes will be recorded. Pin the CLI and SDK versions in CI rather than tracking latest.

On licensing: the README badge says Apache 2.0, and package.json declares "license": "Apache-2.0". The repository metadata reported alongside the project says NOASSERTION, which typically means an automated classifier could not match the LICENSE file to a known template. The README does not discuss licence implications for commercial use, and this is not legal advice. Read the LICENSE file itself, and if you redistribute the CLI or embed the SDK in a product, have someone who can read it confirm the terms.

## Conclusion

Adopt OPFOR if you already ship an agent or MCP server and need repeatable, logged attack runs you can commit to CI; the CLI and the saved config under .opfor/configs/ are the parts worth building on. Do not adopt it if your attacker budget is a non-Anthropic key and you intend to use opfor hunt, because the .env.example states the commander, operator and scout agents run on the Claude Agent SDK regardless of which provider key you set. Before committing, verify the four OWASP suite names in suites/ against the coverage you actually need, and check whether the repository's LICENSE file matches the Apache-2.0 badge the README shows, since the repository metadata reports NOASSERTION.

## FAQ

### Why is OPFOR called Opfor?

OPFOR is short for Opposition Force, the military term the README defines as the unit that plays the enemy in training. The README says the tool was named after that idea: to defend AI agents better, you have to attack them first.

### What is OPFOR in simple terms?

It is an open-source adversary emulation tool for AI agents, LLM apps and MCP servers. It generates targeted attacks for OWASP suites, fires them at your target, and judges each response with an LLM.

### How do I install the OPFOR CLI and run a first scan?

Install it globally with npm install -g @keyvaluesystems/agent-opfor-cli, set one provider key such as OPENAI_API_KEY, then run opfor run for a one-shot scan. For a reusable setup, run opfor setup to save a config under .opfor/configs/ and then opfor run --config with that file.

### Does OPFOR work with LLM providers other than Anthropic?

Standard scans do. The .env.example lists Groq, OpenAI, Anthropic and Google Gemini keys, and states that any one of them is enough to run a scan, plus OPFOR_API_KEY for OpenAI-compatible endpoints. The autonomous opfor hunt mode is the exception: its agents run on the Claude Agent SDK and are Claude-only.

### Which OWASP suites does OPFOR cover?

The README lists OWASP LLM Top 10, OWASP Agentic AI Top 10, OWASP MCP Top 10, OWASP API Security, and EU AI Act bias suites. The repository has top-level evaluators/ and suites/ directories that correspond to this coverage.

### Can non-developers use OPFOR?

Yes. The README ships a browser extension, available from the Chrome Web Store, aimed at product managers, designers, QA and security analysts. The README says anyone on the team can red-team a deployed chatbot through it with no code, no env vars and no YAML.

## Sources

- [Issues](https://github.com/KeyValueSoftwareSystems/agent-opfor/issues)
- [KeyValueSoftwareSystems/agent-opfor on GitHub](https://github.com/KeyValueSoftwareSystems/agent-opfor)
- [README](https://github.com/KeyValueSoftwareSystems/agent-opfor/blob/master/README.md)
- [Releases](https://github.com/KeyValueSoftwareSystems/agent-opfor/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/keyvaluesoftwaresystems-agent-opfor
