# PurpleLlama: Meta's Toolkit for Evaluating and Defending LLM Systems

> PurpleLlama is an umbrella project from Meta that brings together evaluation benchmarks and inference-time safety tools for large language models. Its components include Llama Guard for content moderation, Prompt Guard for injection and jailbreak defense, Code Shield for filtering insecure code, and the CyberSecEval benchmark suite for measuring cybersecurity risks in LLM outputs.

**meta-llama/PurpleLlama** — Set of tools to assess and improve LLM security.

- Repository: https://github.com/meta-llama/PurpleLlama
- Stars: 4,406 · Forks: 772
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/meta-llama-purplellama

## Why Purple Teaming and What PurpleLlama Provides

In cybersecurity, a red team attacks and a blue team defends. A purple team combines both roles: it tests attacks while simultaneously building defenses. The README states that PurpleLlama borrows this concept because mitigating generative AI risks requires both attacking (red teaming) and defensive (blue team) postures, and the project intends to cover both.

The initial release focuses on two areas: cybersecurity evaluation and input/output safeguards. The safeguards are deployed at inference time to filter what goes into and comes out of an LLM. The evaluation benchmarks measure how much risk a given model presents before a safeguard is applied. The repository is organized with separate top-level directories for each component: CodeShield, CybersecurityBenchmarks, Llama-Guard, Llama-Guard2, Llama-Guard3, Llama-Guard4, Llama-Prompt-Guard-2, LlamaFirewall, and Prompt-Guard.

## Llama Guard: LLM-Based Content Moderation

Llama Guard is a series of fine-tuned models designed to classify whether LLM inputs and outputs violate content policies. The README describes Llama Guard 3 as supporting detection of the MLCommons standard hazards taxonomy and covering seven languages beyond English with a 128k context window. The models are built by fine-tuning Llama 3.1 and 3.2 model families.

Llama Guard 3 is available in several sizes: 8B, 1B, and an 11B vision variant that can reason about image content in addition to text. The vision model supports multimodal input, which is relevant for applications where users can submit images alongside text.

The README notes that Llama Guard 3 was optimized specifically to detect helpful cyberattack responses and to prevent malicious code output in environments that use code interpreters. This makes it relevant not only for general content moderation but also for the specific risk of LLM-assisted code execution environments.

## Prompt Guard: Defense Against Injection and Jailbreaks

Prompt Guard is a classifier model for detecting two categories of malicious prompt inputs. Prompt injections exploit the inclusion of untrusted third-party data into an LLM context window, causing the model to execute instructions from that data rather than from the application developer. Jailbreaks are inputs designed to override a model's safety training.

The README distinguishes these two attack types clearly: injections come from the external data environment (a retrieved document, a web page, a database record), while jailbreaks come from the user interacting directly with the model. Prompt Guard is intended to classify inputs before they reach the model, preventing either type from reaching the context window.

Llama Prompt Guard 2, a newer variant, is in the Llama-Prompt-Guard-2 directory. Both versions are listed in the license table as using the Llama 3.2 Community License.

## Code Shield: Filtering Insecure Code at Inference Time

Code Shield adds inference-time filtering for LLM-generated code. The README describes three capabilities: mitigation of insecure code suggestion risk, code interpreter abuse prevention, and secure command execution. It is designed for deployment in systems that use LLMs to write or suggest code, where the generated output could contain vulnerabilities or dangerous commands.

The README links to a Jupyter notebook (CodeShield/notebook/CodeShieldUsageDemo.ipynb) for an example of how to use Code Shield. Unlike Llama Guard and Prompt Guard, which are fine-tuned models, Code Shield is listed in the license table as MIT-licensed, which means it is the least restrictive component in the repository for commercial use.

LlamaFirewall is a newer top-level component in the repository, added after the initial release. The README does not describe it in detail in the overview section.

## CyberSecEval: Measuring LLM Cybersecurity Risk

CyberSecEval is described in the README as the first industry-wide cybersecurity safety evaluation suite for LLMs. The benchmarks are based on industry standards including CWE (Common Weakness Enumeration) and the MITRE ATT&CK framework. The suite measures two primary risk dimensions: how often a model suggests insecure code, and how willing a model is to assist in cyberattacks.

CyberSecEval 2 expanded the scope to include an LLM's propensity to abuse a code interpreter, its offensive cybersecurity capabilities, and its susceptibility to prompt injection. CyberSecEval 3 and 4 continued expanding the evaluation coverage, with results and leaderboards available on HuggingFace.

All CyberSecEval components are MIT-licensed, which means any team evaluating an LLM for cybersecurity risk can use these benchmarks regardless of what model they are testing. The benchmarks are in the CybersecurityBenchmarks directory.

## Limitations and What PurpleLlama Does Not Cover

PurpleLlama's safeguard models are fine-tuned versions of Llama models. Their accuracy depends on the training distribution of hazard examples. The README does not provide precision and recall numbers for Llama Guard on categories outside the MLCommons taxonomy, and performance on novel attack patterns not represented in training is not documented.

The repository uses git submodules (as indicated by the .gitmodules file) and contains several large subdirectories. Getting a full working environment requires understanding which components are needed: not all components require all submodules. The README does not provide a single-command setup.

The license structure is heterogeneous. Evals and benchmarks are MIT-licensed and freely usable. Safeguard models use various Llama Community License versions (Llama 2, Llama 3, or Llama 3.2 depending on the model). These Llama Community Licenses include restrictions on commercial use above certain monthly active user thresholds and require attribution. Teams must verify which license applies to each specific model component they deploy.

## Guardrails AI: A Runtime Validation Framework Approach

Guardrails AI is an open-source Python library that validates and corrects LLM outputs by defining a schema of validators that run against each response. It works as a wrapper around any LLM API call and applies a configurable set of checks after each model response, with options to retry or correct outputs that fail validation.

The difference in approach is architectural. Guardrails AI is a runtime framework that applies rules to any model's output, regardless of the model family, and integrates into existing Python code through an API wrapper. PurpleLlama is a research toolkit that includes fine-tuned classifier models (Llama Guard, Prompt Guard) and benchmark suites (CyberSecEval). Using PurpleLlama's safeguards in production requires running a separate classifier model for each input and output, while Guardrails AI applies rule-based validators without a second model call. The right choice depends on whether the threat model requires a trained model-based classifier or whether rule-based validation is sufficient.

## Licensing and Maintenance Status

The last push to the repository was on 2026-08-18. The repository is not archived. The license situation requires attention before deployment: the repository-level license file is listed as NOASSERTION in the primary language metadata, reflecting that different components use different licenses. The README provides a table mapping component types to their licenses: evals and benchmarks use MIT, while each Llama Guard and Prompt Guard variant uses the Llama Community License corresponding to the model it was built on (Llama 2 Community License, Llama 3 Community License, or Llama 3.2 Community License).

## Conclusion

PurpleLlama fits teams building products on Llama models who need ready-made input/output safety filters and standardized cybersecurity benchmarks. The CyberSecEval suite is useful to any researcher evaluating LLM cybersecurity risk regardless of which model family they use, since the benchmarks are MIT-licensed. Llama Guard and Prompt Guard are fine-tuned Llama-family models distributed under the Llama Community License, which imposes conditions on commercial use above certain user thresholds. Verify the specific Llama Community License version (2, 3, or 3.2) for each model component before deploying commercially.

## FAQ

### What is Purple Llama?

PurpleLlama is an umbrella project from Meta that provides tools and evaluations for responsible development with open generative AI models. It includes Llama Guard models for input/output content moderation, Prompt Guard for detecting prompt injections and jailbreaks, Code Shield for filtering insecure LLM-generated code, and the CyberSecEval benchmark suite for measuring cybersecurity risks.

### What is the Llama Guard model?

Llama Guard is a series of fine-tuned models in the PurpleLlama project designed to classify whether LLM inputs and outputs contain harmful or policy-violating content. Llama Guard 3 supports the MLCommons standard hazards taxonomy, seven languages, and a 128k context window, and includes an 11B vision variant for multimodal content.

### What does purple teaming mean in AI security?

The README explains that purple teaming in AI security combines red team (attacking) and blue team (defensive) approaches. The name Purple Llama reflects Meta's intent to address both attack-side evaluation (CyberSecEval benchmarks testing model risk) and defense-side tools (Llama Guard and Prompt Guard filtering harmful inputs and outputs).

### What is PurpleLlama used for in AI safety?

PurpleLlama is used to evaluate how much cybersecurity risk a language model presents (via CyberSecEval benchmarks) and to deploy runtime safeguards that filter harmful or policy-violating content in LLM inputs and outputs (via Llama Guard and Prompt Guard). The project is aimed at developers building applications on Llama models who need both measurement and mitigation tools.

## Sources

- [Issues](https://github.com/meta-llama/PurpleLlama/issues)
- [meta-llama/PurpleLlama on GitHub](https://github.com/meta-llama/PurpleLlama)
- [README](https://github.com/meta-llama/PurpleLlama/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/meta-llama-purplellama
