Model or dataset
mrwadams/attackgen avatar
mrwadams/attackgen

AttackGen's narrative is generated, its detection reference is not

AttackGen is a cybersecurity incident response testing tool that leverages the power of large language models and the comprehensive MITRE ATT&CK framework. The tool generates tailored incident response scenarios based on user-selected threat actor groups and your organisation's details.

1,248 stars169 forksPythonGPL-3.0

At a glance

What is it?
AttackGen builds incident response exercise material out of MITRE ATT&CK, ICS and ATLAS data with a language model in the loop, and its release notes are unusually precise about which half of the output is model-written and which half is copied straight out of the data. The packaging tells the same story: dependency floors with reasons written beside them, an unprivileged container, and two dependency files that disagree.
Who is it for?
AttackGen fits a defender who needs a tabletop exercise built from real technique data rather than invented lore, and who will read the generated narrative critically rather than run it as a script. Before adopting it, plan for two model calls per scenario rather than one, keep the deterministic detection and mitigation reference under your own change control since the model never touches it, and pin your Streamlit version to at least the floor the requirements file names.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The generated narrative and the detection reference have different authors

The mechanism that makes this usable for a real exercise is a split between what the model writes and what it is allowed to write. A scenario can be generated from a chosen threat actor group, from an ATLAS case study, from a hand-picked selection of ATT&CK or ATLAS techniques, or from templates covering common incident types including AI and machine learning specific patterns. Alongside them sits a set of AI insider threat scenarios built from a published threat model for frontier AI agents, shaped by the agent's deployment archetype in other words its autonomy level, a threat category, STRIDE threats and an optional free-text seed. The latest release then draws the line explicitly: applying a refinement from the Assistant chat rewrites only the narrative, while the detection and mitigation reference drawn from the ATT&CK data is carried over byte for byte, and the Navigator layer is not regenerated at all but captioned as reflecting the original scenario's techniques.

A second model call is optional, priced in the interface, and recoverable

Cost is visible before you commit, which is unusual and useful. The AI-enhanced adversary and purple-team options sit above the Generate button, and the purple-team option states in the interface that it makes a second model call. If that optional narrative fails or is stopped part way, the base scenario, its downloads, the Navigator layer and the deterministic reference all stay usable, with a message saying what failed, and a retry targets only the failed piece. The Assistant has its own failure mode handled too: an empty reply from the model is refused with no changes made. Every refined result is captioned with the time of the last apply, and a Revert to original action restores the scenario and the narrative together, so a bad refinement is one click from undoing rather than a regeneration.

Every provider goes through one wrapper, so a new model is a config line

Model access is deliberately provider-agnostic. The supported list covers the OpenAI, Anthropic, Google AI, Mistral and Groq APIs, and then any OpenAI compatible endpoint, with Ollama, LM Studio, Azure OpenAI and OpenRouter named as examples. All of them are routed through LiteLLM behind a single internal wrapper, and the stated consequence is that adding a new model is a one-line change. In practice that means you can run the whole console against a local model with no code change, which changes what the tool is useful for in an environment where prompts cannot leave the building. Two supporting pieces sit around it: the retry count is fixed at three on every call in the internal layer, and LangSmith is an optional integration for debugging, testing and monitoring model performance rather than a requirement.

Dependency floors are pinned, each with its reason written next to it

The packaging comments are the most useful documentation in the repository, because each floor exists for a stated reason. Python 3.11 is the minimum rather than 3.10 because the LiteLLM version in use imports two typing constructs that only exist from 3.11 on the plain import path. The MCP server pins its package to 2.x or newer because the server module imports a name that does not exist on 1.x, which would fail at import rather than at run time. Tenacity is declared explicitly because LiteLLM imports it lazily inside its retry helper and, when it is missing, replaces the real provider error with a false instruction to install tenacity, and Streamlit supplied tenacity only up to 1.61. Setuptools carries a floor as a guard so a future resolve cannot land below two named vulnerability fixes, with an honest note that the floor is not what cleared them.

Two dependency files disagree, and the comment cites an older version than the pin

The same dependency list is written twice, in the packaging manifest and in the requirements file, and the two are not identical. The requirements file pins Streamlit to 1.64 or newer, while the manifest lists it without a floor at all. The comment explaining that floor also does not agree with it: it says the deprecated theme fields that caused the browser to log an invalid colour error on every rerun were removed in 1.55, and a named smoke test fails below the floor. So the reason points at 1.55, the pin says 1.64, and a manifest-driven install resolves something else entirely. This is the kind of drift that matters because the failure it prevents is a browser console error on every rerun rather than a hard crash, which is exactly the kind of bug that survives a release.

The container drops every capability, mounts the data read only, and runs unprivileged

For a tool that handles API keys and adversary data, the compose file is unusually explicit:

yaml
    security_opt:
      - no-new-privileges:true
    cap_drop:
      - ALL
    cap_add:
      - NET_BIND_SERVICE

Beyond that, temporary and cache directories are writable tmpfs mounts with execution and set-user-ID bits disabled and size caps, the MITRE data directory is mounted read only, the process runs as a fixed non-root user matching the image, cross-site request forgery protection is enabled in both the environment and the entrypoint, and CPU and memory limits are set alongside smaller reservations. The health check calls the application's own health endpoint over Python rather than shelling out to curl, and the image pins its base distribution by digest and upgrades the installer first to pick up a symlink extraction fix. For a security tool that is a defensible baseline, and none of it is described in the readme.

The entry point is a welcome script with an emoji in its filename

The container's entrypoint starts the application from a file whose name carries an emoji, which is a small thing to rediscover when writing your own health check, container command or documentation, and it is the only such filename at the root. The rest of the root explains the project's shape: a Model Context Protocol server with its own console script, a skills directory, agent instruction files for two assistants, a secret-scanning configuration, an example environment file and a dev container definition. The test suite includes a browser marker that launches a real Streamlit server and a real Chromium through Playwright and skips itself when either is unavailable, which is a reasonable trade against a suite that silently passes without testing the interface. Security scanning is configured in the manifest rather than in a config file that was silently doing nothing.

Editorial conclusion

AttackGen fits a defender who needs a tabletop exercise built from real technique data rather than invented lore, and who will read the generated narrative critically rather than run it as a script. Before adopting it, plan for two model calls per scenario rather than one, keep the deterministic detection and mitigation reference under your own change control since the model never touches it, and pin your Streamlit version to at least the floor the requirements file names. Anyone needing a shared multi-user server, an audit trail of generated output, or an offline install without a language model should look elsewhere.

Frequently asked questions

What does AttackGen generate?

Incident response exercise scenarios based on a chosen threat actor group or an ATLAS case study, on a selection of ATT&CK or ATLAS techniques, or on templates for common incident types. You can supply your organisation's size and industry so the generated scenario is tailored to it, and results download as Markdown.

Which frameworks does AttackGen use?

MITRE ATT&CK Enterprise, ATT&CK for ICS, and ATLAS, the adversarial threat landscape for AI systems. There is also a set of AI insider threat scenarios built from a published threat model for frontier AI agents, shaped by the agent's autonomy level, a threat category, STRIDE threats and an optional free-text seed.

Which model providers can AttackGen use?

The OpenAI, Anthropic, Google AI, Mistral and Groq APIs, plus any OpenAI compatible endpoint, with Ollama, LM Studio, Azure OpenAI and OpenRouter named as examples. All of them are routed through LiteLLM behind a single internal wrapper, so adding a model is described as a one-line change.

How many model calls does generating an AttackGen scenario take?

One by default, and a second when the optional purple-team narrative is enabled, which the interface states before you generate. If that narrative fails or is stopped, the base scenario and its downloads remain usable and only the failed part can be retried.

How is AttackGen deployed?

As a Docker container built from the repository and published on port 8501, running as a non-root user with all capabilities dropped except one, temporary directories mounted with execution disabled, the MITRE data directory mounted read only, and credentials supplied through a .env file. A docker compose file is included.

Official sources

  1. Issues
  2. License: GPL-3.0
  3. mrwadams/attackgen on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mrwadams-attackgen.svg)](https://hysenlabs.com/projects/mrwadams-attackgen)