# Promptfoo: LLM Evaluation and Red Teaming for AI Applications

> Promptfoo is an MIT-licensed CLI and library for testing prompts, evaluating LLM responses, and running automated red teaming and vulnerability scans against AI applications, now part of OpenAI while remaining open source.

**promptfoo/promptfoo** — Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

- Repository: https://github.com/promptfoo/promptfoo
- Website: https://promptfoo.dev
- Stars: 25,458 · Forks: 2,373
- Language: TypeScript
- License: MIT
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/promptfoo-promptfoo

## What Promptfoo Solves and Who Uses It

Teams building AI features face a problem that has no clean analog in traditional software testing: the output of an LLM call is not deterministic. The same prompt can produce different answers on different runs, different models respond differently to the same input, and a prompt that worked in development may fail in production after a model update. Promptfoo exists to replace trial-and-error prompt iteration with a systematic evaluation workflow.

The README describes two distinct use cases. The first is evals: running a set of prompts against an LLM and checking whether the outputs meet defined criteria. The second is red teaming: generating adversarial prompts automatically to find security vulnerabilities like prompt injection, jailbreaks, or data exfiltration paths in an AI application.

The README notes that Promptfoo is now part of OpenAI, but explicitly states the project remains open source and MIT licensed. It was previously an independent tool and the transition does not change the license or the open-source availability.

Supported providers listed in the README include OpenAI, Anthropic, Azure, Amazon Bedrock, Ollama, and others. The local execution model means that prompts and responses stay on the machine running the evaluation rather than going through a cloud evaluation service.

## Installing Promptfoo and Running a First Eval

Install globally via npm:

```sh
npm install -g promptfoo
```

Then initialize an example project and run an eval:

```sh
promptfoo init --example getting-started
```

Before running evals, set the API key for your provider. For OpenAI:

```sh
export OPENAI_API_KEY=sk-abc123
```

Then run the evaluation and open the results viewer:

```sh
cd examples/getting-started
promptfoo eval
promptfoo view
```

Alternatives to the npm global install include `brew install promptfoo`, `pip install promptfoo`, and `npx promptfoo@latest` to run without installing permanently. The `pf` alias is registered alongside the `promptfoo` binary.

The CLI requires Node.js 22.22.0 or higher, as specified in the package.json engines field. The Dockerfile shows an Alpine-based container build if you prefer containerized execution.

## Evaluation Configuration and Provider Comparison

Promptfoo uses declarative configuration files (YAML or JSON) to define what to test. A configuration specifies the provider or providers to call, the prompts to evaluate, and the assertions that check each response. Assertions can be string matchers, regex patterns, LLM-graded criteria, or custom JavaScript functions.

The examples directory in the repository demonstrates side-by-side model comparisons. Directory names like compare-claude-vs-gpt, compare-gpt-model-tiers, compare-deepseek-r1-vs-openai-o1, and compare-gpt-vs-claude-vs-gemini show the range of supported provider combinations. Each example is a self-contained configuration that runs the same set of prompts across multiple models and displays the results in a table.

CI/CD integration is supported, allowing evals to run as part of a pull request workflow. The README mentions code scanning as a feature that reviews pull requests specifically for LLM-related security and compliance issues. The architecture/ directory contains baseline generation scripts, suggesting the project tracks performance regressions against a baseline as part of its own development workflow.

The `promptfoo view` command opens a local web UI showing results in a comparison table. Results can be shared with teammates; the README lists sharing as a feature without specifying the sharing mechanism.

## Red Teaming: Automated Vulnerability Scanning for AI

Red teaming in Promptfoo generates adversarial test cases automatically and checks whether the AI application handles them safely. The README links to a Red Teaming guide at promptfoo.dev/docs/red-team/ for full documentation.

The red teaming workflow scans for classes of vulnerabilities in LLM-based applications: prompt injection, jailbreaks, data exfiltration, harmful content generation, and compliance violations. Promptfoo generates test inputs targeting each vulnerability class and records which ones succeed in eliciting a problematic response.

The generated security vulnerability report is separate from the evaluation results table. The README shows it as a distinct output format. This is a meaningful distinction: evals verify correctness, red teaming verifies safety. An application can pass all functional evals and still have security vulnerabilities that only adversarial inputs would expose.

The red teaming component is the portion most closely related to Promptfoo's acquisition by OpenAI. Teams building on OpenAI's models or on any other provider can use the same red teaming tooling.

## Limitations: Local Execution, Scale, and No Built-In Model Hosting

Promptfoo runs evals locally by default. The README calls this a feature: "your prompts never leave your machine." For teams with strict data handling requirements, this is an advantage. For teams that need to run thousands of parallel evals from a managed service with autoscaling, the local model places the infrastructure burden on the user.

Running evals against external LLM APIs means the eval speed is bounded by the provider's rate limits. Large eval sets against slow or rate-limited models can take minutes or hours to complete. The README does not describe a distributed execution mode within the open-source package.

Prompttoo does not host or serve models. It connects to external providers via API. Teams using local models via Ollama can evaluate those models, but Promptfoo itself does not include an inference runtime.

The project is at version 0.123.1 as of 2026-09-18. The fast version progression suggests active feature development, which also means the configuration schema and CLI flags may change between minor versions. The CHANGELOG.md tracks changes, and the package.json's renovate.json configuration suggests automated dependency updates are part of the workflow.

## Promptfoo vs DeepEval: Comparing LLM Testing Approaches

DeepEval is another LLM evaluation framework that is commonly compared to Promptfoo. DeepEval is Python-first and provides LLM-graded metrics with specific implementations for relevance, hallucination detection, toxicity, and summarization quality. It integrates with pytest for running evals as part of a standard Python test suite.

Prompttoo is TypeScript-first and CLI-driven, with configuration in YAML or JSON. It supports more provider options out of the box and has explicit red teaming functionality as a distinct feature area. DeepEval focuses on evaluation quality metrics; Promptfoo addresses both correctness evaluation and adversarial security testing.

For Python-only teams building on frameworks like LangChain or LlamaIndex, DeepEval integrates more directly into an existing Python workflow. For teams that need side-by-side model comparison across many providers with declarative configuration and CI integration, Promptfoo's YAML-based approach is easier to set up without writing test code.

## Project Status, Docker Deployment, and License

Promptfoo is MIT licensed. The README states explicitly that it remains open source and MIT licensed after the OpenAI acquisition. The package is published at version 0.123.1 with the last push on 2026-09-25 and releases shipping on a roughly weekly cadence through 2026.

A Dockerfile is included in the repository for containerized deployment. It uses Node.js 24.21.0 on Alpine Linux as the base, with Python installed for providers that require Python execution. The build args include PROMPTFOO_REMOTE_API_BASE_URL and VITE_PUBLIC_BASENAME for configuring the hosted UI in a containerized setup.

The examples directory contains over 20 ready-to-run examples demonstrating specific provider comparisons, configuration patterns, and use cases. Running `ls examples/` in a cloned repository shows the full set including examples for anthropic, azure, amazon-bedrock, and compare-agentic-sdks.

## Conclusion

Promptfoo suits teams that ship AI features and need systematic testing before and after deployment: both behavioral evals to verify outputs and red teaming to find security vulnerabilities. It is a poor fit for teams that need a hosted evaluation service with managed infrastructure, since the local execution model means they own the running environment. Before starting, verify that your LLM provider's API key is available as an environment variable and run `promptfoo init --example getting-started` to get a working example configuration. The project is at version 0.123.1 with releases shipping regularly through 2026.

## FAQ

### What is Promptfoo used for?

Promptfoo is used for evaluating LLM prompts and testing AI applications. It runs automated evals to check that model outputs meet defined criteria, and runs red teaming scans to find security vulnerabilities like prompt injection or jailbreaks.

### Is Promptfoo free?

Promptfoo is MIT licensed and free to use. It is now part of OpenAI but the README states explicitly that it remains open source and MIT licensed.

### Is Promptfoo open source?

Promptfoo is MIT licensed and publicly available on GitHub. The README confirms it remains open source following its acquisition by OpenAI.

### How do you install Promptfoo?

Install with `npm install -g promptfoo`, `brew install promptfoo`, or `pip install promptfoo`. You can also run it without installing using `npx promptfoo@latest`. Node.js 22.22.0 or higher is required.

## Sources

- [Official documentation](https://promptfoo.dev)
- [Official README](https://github.com/promptfoo/promptfoo#readme)
- [Project repository](https://github.com/promptfoo/promptfoo)
- [Release notes](https://github.com/promptfoo/promptfoo/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/promptfoo-promptfoo
