# The latest release is mostly about refusing to trust the model's own claims

> STRIDE GPT generates threat models, attack trees and Gherkin cases with an LLM, across a CLI, a Streamlit app and a container. What is interesting about version 0.20 is that it spends its two headline features on verification: checkpoints that refuse to resume onto changed code, and evidence checked against the file on disk.

**mrwadams/stride-gpt** — An AI-powered threat modeling tool that leverages OpenAI's GPT models to generate threat models for a given application based on the STRIDE methodology.

- Repository: https://github.com/mrwadams/stride-gpt
- Website: https://stridegpt.streamlit.app
- Stars: 1,139 · Forks: 330
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/mrwadams-stride-gpt

## Checkpoints refuse to resume onto different code

The most substantial change in version 0.20 is interruptible analysis. The analyze command now checkpoints the approved plan and each subsystem finding as it completes, written next to the report it would produce, and a resume flag reuses the subsystems that finished while rerunning the rest, then proceeds to synthesis and the data flow diagram as usual. The interesting part is the refusal logic. Resume stops if the target's git commit or the configuration hash no longer match, and it says which one changed, so that findings cannot be stitched onto a run against different code. A dirty working tree or a missing commit can be overridden with a force flag; a changed commit cannot be overridden even with it. That asymmetry is the design: uncommitted changes are a state you can acknowledge, while a different commit is a different target. A run manifest records which subsystems were reused and which were rerun, so the output is auditable after the fact.

## Every quoted snippet is checked against the file on disk

The second headline feature addresses a failure mode specific to AI code analysis. Each threat is now reported through a tool call carrying the code it came from, and every quoted snippet is checked against the file on disk before the threat is recorded. Two consequences follow, and both are about the difference between what the model said and what it did. SARIF locations point at the specific file and line range where the evidence was actually found, rather than giving every threat in a subsystem the same three files. And the `files_analyzed` field reflects what the agent actually read rather than what it claimed to have read. This matters most for SARIF, since that format is imported by code scanning tools and shown inline in a pull request, where a wrong location sends a reviewer to the wrong lines and erodes trust in every other finding in the same report.

## The diagram editor is there because XML parses and images do not

The draw.io integration is the feature that most changes the quality of the output, and the reasoning is stated in the documentation. Architecture diagrams can be created and edited directly inside the tool using the embedded diagrams.net editor, and those diagrams are parsed as XML to extract components, connections and trust boundaries. The claim is that this provides significantly richer context for threat modelling than image analysis alone, which is true for a structural reason: a trust boundary is a line with a semantic meaning, and a vision model sees a line. The existing image upload workflow is unchanged, so nothing is taken away from the multimodal path. By default the editor loads from the public hosted `embed.diagrams.net`, which is also the default threat for an air-gapped deployment, so there is a documented escape hatch: a single environment variable pointing the editor at your own draw.io instance, with the embed query parameters appended by the tool.

## The data flow diagram becomes the authoritative system model

There is a second diagram path, and it is a different idea. A data flow diagram can be generated from the application description, or an uploaded DFD image can be parsed, and the Mermaid source can then be edited live. The purpose is stated precisely: the confirmed diagram is fed back into the threat model and attack tree prompts as the authoritative system model. That word carries the weight. Everything else in the tool works from prose descriptions the user typed, and prose is where ambiguity lives; a diagram the user has confirmed pins down the components and the flows, and the prompts are conditioned on it. The CLI analysis command also emits a system-level DFD alongside its findings, so the same artefact appears in both interfaces. The loop, describe, generate, edit, confirm, feed back, is the part of this tool worth understanding.

## Four output formats and five risk vocabularies

Output goes out in four formats: Markdown, JSON, SARIF, and a self-contained HTML view for sharing with stakeholders. SARIF is the one that changes how the tool is used, because it imports into GitHub, GitLab, Azure DevOps and IDEs, which means findings can arrive as code scanning alerts rather than as a document someone has to remember to read. The risk vocabularies layered on top are worth naming individually, because a threat model that mixes them without saying so is hard to act on. STRIDE is the core method. OWASP Top 10 for Agentic Applications covers agentic systems, and OWASP LLM Top 10 covers generative AI ones, so the tool has distinct lenses for the two cases it is now most often asked about. DREAD scoring is supported for identified threats, attack trees enumerate attack paths, and Gherkin test cases are generated from the threats. MITRE ATT&CK Enterprise is used for traditional infrastructure attacks and ATLAS for machine learning and model specific ones, surfaced as markdown columns, as linked pills in HTML and as properties in SARIF.

## The published wheel is lean, and the web app is not

The dependency arrangement is the most deliberate thing in the packaging. The runtime dependencies of the CLI are short: an HTTP client, LiteLLM, Pydantic, dotenv, Typer, Rich, prompt_toolkit and defusedxml. Everything the web interface needs sits in a separate dependency group: Streamlit, the OpenAI, Anthropic and Groq clients, PyGithub, tiktoken and requests. The comment explaining why is explicit: those are kept out of the project dependencies so the published CLI wheel stays lean, and a dependency group is local or deploy-only and is never shipped in the wheel. Three consumers install from the lockfile instead, Streamlit Community Cloud, the Docker UI image and the security scan workflow, with uv.lock named as the single source of truth for every dependency. Two details there are worth a second look: the web group deliberately repeats nothing already in the runtime set, and the development group pins ruff to an exact version rather than a range.

## The container drops every capability and adds one back

The compose file for the UI image is short and defensive. New privileges are disabled, all Linux capabilities are dropped, and exactly one is added back: permission to bind a service port. That single addition is odd next to the rest, because the service listens on port 8501, which is above the privileged range and does not need the capability as a non-root user. Writable temporary filesystems are mounted with noexec and nosuid for the temporary directory, the Streamlit configuration directory and the cache directory, each size-capped. XSRF protection is switched on explicitly. The container runs as user and group 1000, matching the user in the UI Dockerfile, and has a health check against the Streamlit core health endpoint every thirty seconds with three retries. Resource limits are modest: two CPUs and two gigabytes, with half a CPU and 512 megabytes reserved, and the restart policy is unless-stopped.

## Nine keys in the example environment file, and a beta classifier

The example environment file lists eight model provider keys plus a GitHub token, because provider routing goes through LiteLLM and covers OpenAI, Anthropic, Google AI, Mistral, Groq and DeepSeek, with local hosting through an LM Studio server whose endpoint is written in as a local port. The optional diagram host variable is documented in the same file, with a note that embed query parameters are appended. Against that, the manifest classifies the project as a beta, targets Python 3.12 and newer through 3.14, and points its own homepage field at the GitHub repository rather than at the hosted Streamlit application the repository metadata advertises. The default branch is master rather than main, and the release history shows two versions in 2026, 0.19.0 in July and 0.20.0 at the end of September. The tree also carries a gitleaks configuration, a security policy, a releasing guide and an agent instructions file.

## Conclusion

Judgment: STRIDE GPT is unusual among LLM tools in treating the model's output as a claim that needs checking rather than a result to print. Two features do that work. Findings are verified against the files on disk before they are recorded, and an interrupted analysis refuses to resume onto a different commit or a changed configuration, with a force flag that covers a dirty tree but not a changed SHA. That is the right instinct for a security tool, and it is the part to look for when judging the next release. Two things to check before adopting it. Your threat model still comes from a language model, so treat the output as a draft to be reviewed by someone who knows the system. And the framework layering, STRIDE plus DREAD plus two OWASP top tens plus ATT&CK and ATLAS, is broad enough that you should decide which one your organisation actually reports in before you let the tool choose.

## FAQ

### how to use stride gpt

Provide application details such as the application type, the authentication methods in use, and whether the system is internet-facing or processes sensitive data. The tool generates a threat model from the STRIDE methodology, plus attack trees, DREAD scores, suggested mitigations and Gherkin test cases, in Markdown, JSON, SARIF or a self-contained HTML view.

### what is stride gpt

An MIT licensed Python tool that uses large language models to generate threat models and attack trees with the STRIDE methodology. It runs as a command line tool with an interactive REPL, as a Streamlit web interface, and as a Docker container image, and it also analyses a codebase or a GitHub repository directly.

### Which AI models does STRIDE GPT support?

Provider routing goes through LiteLLM, covering OpenAI, Anthropic, Google AI, Mistral, Groq and DeepSeek, plus local hosting through an LM Studio server. The listed reasoning models are the OpenAI GPT-5.4 and 5.5 series, Anthropic Claude 4.6 and 4.8 with Extended Thinking, Google Gemini 3 and the Mistral Magistral series.

### Does STRIDE GPT store my application details?

No. The feature list states there is no data storage and that application details are not saved. Configuration comes from environment variables, and the example environment file holds the provider keys plus an optional variable for pointing the embedded draw.io editor at a self-hosted instance.

### What does STRIDE GPT do with architecture diagrams?

You can create and edit them in an embedded draw.io editor, where the XML is parsed for components, connections and trust boundaries rather than treated as an image. Separately, a data flow diagram generated from your description, or parsed from an uploaded image, can be edited as Mermaid and fed back as the authoritative system model.

## Sources

- [License: MIT](https://github.com/mrwadams/stride-gpt/blob/master/LICENSE)
- [mrwadams/stride-gpt on GitHub](https://github.com/mrwadams/stride-gpt)
- [Project website](https://stridegpt.streamlit.app)
- [README](https://github.com/mrwadams/stride-gpt/blob/master/README.md)
- [Releases](https://github.com/mrwadams/stride-gpt/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mrwadams-stride-gpt
