# NVIDIA garak: an LLM vulnerability scanner for teams who need to probe their own models

> garak is an Apache-2.0 command-line scanner that runs static, dynamic and adaptive probes against an LLM to find jailbreaks, prompt injection, hallucination and data leakage. It is built for engineers who already know which model they want to break, and it assumes you can read a JSONL log to find out what happened.

**NVIDIA/garak** — the LLM vulnerability scanner

- Repository: https://github.com/NVIDIA/garak
- Website: https://discord.gg/uVch4puUCs
- Stars: 9,368 · Forks: 1,312
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvidia-garak

## What garak actually scans for, and who ends up running it

garak checks whether an LLM can be made to fail in ways the operator does not want. The README lists the target behaviours explicitly: hallucination, data leakage, prompt injection, misinformation, toxicity generation and jailbreaks. The framing the project uses for itself is a scanner in the nmap or Metasploit tradition, applied to language models rather than hosts. That analogy is load-bearing. Like those tools, garak is a CLI you run against a target you are authorised to test, it enumerates a catalogue of techniques, and it produces a result set you have to interpret rather than a pass or fail verdict.

The audience follows from that. This is not a tool for a product manager who wants a security score. It is for red-teamers, ML platform engineers and researchers who can name a model endpoint and want evidence about how it behaves under adversarial prompts. The arXiv paper referenced in the README (arXiv:2406.11036) is the citable description of the method, and the repository ships garak-paper.pdf alongside it, which tells you the project expects academic use as well as operational use. If your question is "is our chatbot safe", garak will not answer it. If your question is "how often does this specific model emit a slur when prompted through this specific framework", garak is built for exactly that.

## Probes, detectors and generators: the three-part data flow

The architecture visible from the README and the repository layout has three roles. A generator is the target: the model family and interface garak talks to. A probe is a family of adversarial inputs. A detector judges whether a given generation counts as a failure. The README states the default behaviour plainly: garak tries all the probes it knows about on the target, using the vulnerability detectors recommended by each probe. So the default run is broad, and the pairing of probe to detector is decided by the probe author, not by you.

That default is the single most consequential design choice in the tool. It means a first run against a hosted model can be expensive and slow, because more than one generation is made per prompt. The README says the default is 10 generations per prompt, which explains the result format: figures like 840/840 at the end of a row mean the total number of generations followed by how many appeared to behave acceptably. A failure rate is derived from that ratio. Two things follow. First, the denominator matters more than the numerator when you compare runs. Second, a low failure rate is not a clean bill of health, it is a rate, and the README's own worked example shows a newer model being more susceptible to encoding-based injection than an older one, which is the opposite of what a maturity curve would suggest.

Output goes two places. Errors land in garak.log. The detailed run is written to a .jsonl file whose path is specified at the start and end of analysis. The README mentions a basic analysis script in an `anal...` path that is truncated in the text, so treat the JSONL as the ground truth and the terminal table as a summary.

## Installing garak and running a first probe

garak is a command-line tool developed on Linux and OSX. The README gives three install routes. The simplest is from PyPI, which is the one to start with:

```bash
python -m pip install -U garak
```

If you need something newer than the last PyPI release, the README offers a direct install from the main branch. This tracks unreleased code, so pin your expectations accordingly:

```bash
python -m pip install -U git+https://github.com/NVIDIA/garak.git@main
```

The third route is a source clone in its own Conda environment, which the README recommends because garak has its own dependency set. Note the Python range it specifies:

```bash
conda create --name garak "python>=3.11,<=3.13"
conda activate garak
gh repo clone NVIDIA/garak
cd garak
python -m pip install -e .
```

Before scanning anything, list the available probes. This is also the cheapest way to understand the scope of a default run:

```bash
garak --list_probes
```

A first real scan needs a target type and a target name. The README uses a Hugging Face model for its second example, checking GPT2 against the DAN 11.0 jailbreak probe. The `--spec` flag narrows the run to one probe family or, with a dot and a plugin name, one plugin:

```bash
python3 -m garak --target_type huggingface --target_name gpt2 --spec probes.dan.Dan_11_0
```

For a commercial endpoint, the target type changes and the API key is supplied as an environment variable. The README's example uses OPENAI_API_KEY with an encoding-injection probe family:

```bash
export OPENAI_API_KEY="sk-123XXXXXXXXXXXX"
python3 -m garak --target_type openai --target_name gpt-5-nano --spec probes.encoding
```

What you should see: a progress bar per loaded probe, then one row per probe evaluating that probe's results on each detector, with FAIL marking generations that showed the undesirable behaviour and a failure rate alongside the generation counts. Errors go to garak.log, and the full run is written to the JSONL file named at analysis start and end. If a generator needs a key it does not have, the README says it will tell you.

## Where garak is the wrong instrument

The default of running every known probe is convenient and also the main way to waste a budget. Against a metered API, ten generations per prompt across the full probe catalogue is a real cost, and the README gives no cost estimate, no dry-run flag and no documented way to resume an interrupted run. The practical mitigation is `--spec`, which is documented, but choosing the right spec requires knowing the probe catalogue well enough to pick, and the README does not walk through that decision.

Interpretation is the second gap. garak reports failure rates per detector. It does not tell you what rate is acceptable for your application, and the README does not propose a threshold. A detector is also a model or heuristic in its own right, so a failure rate is a measurement made through another classifier with its own error modes. The README does not document detector validation or calibration, though the requirements file includes scipy under a calibration comment, which suggests calibration work exists in the project without the README explaining how to use it.

The third limit is scope. garak probes a model or dialog system through a generator interface. It does not scan your retrieval pipeline, your tool-calling authorisation logic, your prompt template store or your deployment configuration. If the vulnerability you care about lives in the application wrapping the model, garak is aimed at the wrong layer. And the README does not document rollback, so if a scan run is interrupted or misconfigured there is no stated recovery procedure beyond starting again.

## How garak differs from general-purpose LLM evaluation suites

The closest comparison is a general LLM evaluation framework such as lm-evaluation-harness. The difference is in what counts as a result. An evaluation harness measures capability: accuracy on benchmarks, scores on tasks, a number you compare across model versions to see which is better. garak measures failure under adversarial input: whether a prompt makes the model do something it should not, expressed as a rate per probe and detector. A model can score well on benchmarks and still fail an encoding-injection probe, and the README's own example shows exactly that pattern when a newer model turns out more susceptible than an older one.

The second difference is the probe catalogue's provenance. garak's probes are implementations of external frameworks. The README names PromptInject for `probes.promptinject` and the Language Model Risk Cards paper for `probes.lmrc.SlurUsage`. That means the scanner's coverage is tied to published red-teaming literature rather than to an internal test set, which is good for defensibility and bad for coverage of anything novel or organisation-specific. If your threat model includes an attack that has not been written up, garak does not have a probe for it until someone contributes one.

The third difference is the interface. lm-evaluation-harness and similar tools are typically driven from Python or a config file. garak is a CLI first, with the README's examples all being shell invocations. That makes it easy to drop into a CI job or a shell script and awkward to embed in a larger Python analysis pipeline, though the package is importable.

## Maintenance, releases and what the Apache-2.0 licence leaves you to decide

The repository is not archived, and the last push was on 2026-09-09, which is the same day as the v0.17.0 release. The release cadence visible in the release list is roughly monthly: v0.15.1 on 2026-06-05, v0.16.0 on 2026-08-04, v0.17.0 on 2026-09-09. The pyproject.toml in the repository carries version 0.17.1.pre1, so development continues past the last tagged release, which is consistent with the README's advice to install from main if you want something fresher than PyPI. For an adopter, the practical consequence is that pinning to a tagged release is the stable path and installing from main is the way to get fixes that have not been tagged yet.

The dependency list is the real upgrade cost. requirements.txt pins or bounds a large surface: transformers, torch, datasets, langchain, litellm, openai, anthropic, cohere, boto3, plus audio and translation libraries. Several of these move quickly and several are heavy. Upgrading garak means upgrading that graph, and the README does not document a compatibility matrix or a minimum supported set beyond the Python range of 3.11 to 3.13 used in the Conda example. Budget for dependency resolution, not just for the garak version bump.

The licence is Apache-2.0, which permits commercial and internal use and requires attribution and notice retention. Beyond that, the README does not state a policy on contributed probes, on how probe results may be published, or on any trademark terms for the garak name. Those are questions for your own legal review, not something this article can settle. One concrete note the README does give: if you cloned the repository before it moved to the NVIDIA GitHub organisation, update your remote with `git remote set-url origin https://github.com/NVIDIA/garak.git`.

## Conclusion

Adopt garak if you own the model or endpoint you are scanning, you can afford the token spend that a full default probe run implies, and someone on the team will actually read the per-detector failure rates instead of the progress bar. Do not adopt it as a compliance checkbox or as a hosted service: it is a CLI you point at a target, and the README does not document a rollback path or a hosted dashboard. Before you commit, verify three things: that your target generator is listed under LLM support, that the API key environment variable the generator expects is set, and that --list_probes shows the probe families you intend to run, because the default is every probe garak knows about.

## FAQ

### What is garak used for?

garak checks whether an LLM can be made to fail in ways the operator does not want. The README lists hallucination, data leakage, prompt injection, misinformation, toxicity generation and jailbreaks among the weaknesses it probes for.

### How do I install garak?

The README's standard route is `python -m pip install -U garak`. It also documents installing from the main branch with pip and cloning the source into a Conda environment with Python between 3.11 and 3.13.

### How do I use garak?

The general syntax is `garak <options>`. You specify a target with `--target_type` and `--target_name`, and you can narrow the run with `--spec`; by default garak tries all the probes it knows on that model using the detectors recommended by each probe.

### What is garak AI?

It refers to garak, the Generative AI Red-teaming and Assessment Kit from NVIDIA, distributed on PyPI and licensed Apache-2.0. The README describes it as an LLM vulnerability scanner and compares its role to nmap or Metasploit Framework, but for LLMs.

### What does garak do?

It runs probes against a target model and uses the detectors recommended by each probe to judge the generations. Results appear as a row per probe with a failure rate, errors go to garak.log, and the detailed run is written to a .jsonl file.

## Sources

- [License: Apache-2.0](https://github.com/NVIDIA/garak/blob/main/LICENSE)
- [NVIDIA/garak on GitHub](https://github.com/NVIDIA/garak)
- [Project website](https://discord.gg/uVch4puUCs)
- [README](https://github.com/NVIDIA/garak/blob/main/README.md)
- [Releases](https://github.com/NVIDIA/garak/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvidia-garak
