CLI tool
UKGovernmentBEIS/inspect_ai avatar
UKGovernmentBEIS/inspect_ai

Inspect: The UK AI Security Institute's Framework for LLM Evaluations

Project brief: Inspect: A framework for large language model evaluations. Inspect provides many built-in components, including facilities for prompt engineering, tool usage, multi-turn dialog, and model graded evaluations.

2,906 stars764 forksPythonMIT

At a glance

What is it?
Inspect is an open-source Python framework for running large language model evaluations, created by the UK AI Security Institute. It ships with more than 200 pre-built evaluation tasks and provides built-in components for prompt engineering, tool use, multi-turn dialog, and model-graded scoring.
Who is it for?
Inspect is the right framework for AI safety researchers and teams that need reproducible, auditable LLM evaluations with a maintained library of 200+ tasks. The MIT license imposes no commercial restrictions.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Inspect Evaluates and Who Runs It

Inspect is designed for teams that need to measure the capabilities and limitations of large language models in a systematic, reproducible way. The README introduces it as a framework for LLM evaluations with four main built-in capabilities: prompt engineering facilities for constructing structured inputs, tool usage support for evaluations where the model calls external functions, multi-turn dialog for evaluations that require back-and-forth interaction, and model-graded evaluations where a second model scores the output.

The primary audience is AI safety researchers, ML engineers, and organizations that need to run standardized evaluations as part of a model development or procurement process. The project was created by the UK AI Security Institute (AISI, at aisi.gov.uk), which is the UK government body responsible for AI safety research. This origin means the framework has been designed for the rigorous, auditable evaluation use cases that government safety research requires rather than for one-off benchmarking.

Architecture: Tasks, Solvers, Scorers, and the Sandbox

An Inspect evaluation is built around the concept of a task: a combination of a dataset of prompts, a solver that determines how the model processes each prompt, and a scorer that grades the output. The README describes components including facilities for prompt engineering and model graded evaluations as built-in, with the ability to provide extensions as separate Python packages for new elicitation and scoring techniques.

The repository tree shows a design/ directory alongside the source, which indicates that the evaluation structure is documented internally. The examples/ directory contains concrete demonstrations: examples include biology_qa.py, code_execution.py, reasoning.py, a browser-based evaluation, computer use evaluation, and multi-turn human-in-the-loop evaluation scenarios.

The sandbox mechanism appears in the examples/ directory as examples/approval/ and is also referenced in the requirements, which pin sandbox tooling versions. Sandboxes isolate code execution during evaluations, which matters for safety when evaluating models that can call tools.

Installing Inspect and Setting Up a Development Environment

For development contributions, the README gives a direct installation path. Clone the repository and install with the editable flag and the dev optional dependency group:

bash
git clone https://github.com/UKGovernmentBEIS/inspect_ai.git
cd inspect_ai
pip install -e ".[dev]"

An alternative using uv, the fast Python package manager, is also documented:

bash
uv sync --extra dev

The README notes that the uv workflow is supported but not required; the uv.lock file in the repository records a reproducible development resolution. The repository also includes a Makefile with targets for common development tasks:

bash
make check
make test

The check target runs Ruff for linting and formatting, Mypy for type checking, and a custom suppressions check. The test target runs pytest. In a uv-managed environment, these commands are prefixed with uv run.

The 200-Plus Pre-Built Evaluations

Inspect ships a collection of more than 200 evaluation tasks that are ready to run against any model without writing custom code. The README points to the evals documentation at inspect.aisi.org.uk/evals/ for the full list. These cover a range of capability and safety evaluation scenarios.

The repository structure shows an evals include-file (docs/evals/inspect-evals.mk) referenced in the Makefile, and the pyproject.toml includes a math optional dependency group (requirements-math.txt), which suggests the eval library includes mathematical reasoning tasks.

For teams that need to add their own evaluation tasks, the extension mechanism lets other Python packages contribute new tasks, solvers, and scorers to the framework. The CONTRIBUTING.md and the design/ directory serve as the starting point for understanding how to write a custom eval that fits Inspect's architecture.

The Developer Documentation and LLMs.txt Index

Inspect's documentation site publishes machine-readable documentation in three formats explicitly intended for coding agents and LLM-assisted development. The llms.txt file at inspect.aisi.org.uk/llms.txt provides a structured index of the docs. The llms-guide.txt concatenates the full user guide as Markdown. The llms-full.txt additionally bundles the API and CLI reference.

Individual documentation pages are also available as Markdown by appending .md to the .html path of any documentation URL. This is a forward-looking design choice that makes Inspect documentation consumable by AI tools without requiring web scraping.

The CLAUDE.md file at the repository root provides project-specific guidance for AI coding agents working in the Inspect codebase itself, which is consistent with the overall orientation toward AI-assisted development workflows.

Limitations: Name Collision and Extension Dependency

The name Inspect collides with Python's built-in inspect module, which provides introspection utilities for live Python objects. In search contexts, the project competes with this well-known standard library module for visibility, and in code contexts, import naming requires care when both are used in the same file.

The framework's power depends on having the right solver and scorer for your evaluation task. The 200+ pre-built evaluations cover a range of scenarios, but specialized domains (for example, domain-specific factual recall, red-teaming scenarios for niche applications, or evaluations that require specialized external tools) may require writing custom components. The README documents an extension mechanism for this purpose, but it requires Python development work.

The web UI lives in a git submodule at src/inspect_ai/_view/ts-mono/. Contributors who need to work on the TypeScript frontend must initialize this submodule separately, adding a setup step that Python-only contributors do not need.

Inspect vs. EleutherAI LM Evaluation Harness

EleutherAI's lm-evaluation-harness is the most widely used alternative open-source LLM evaluation framework. Both provide libraries of tasks, support multiple models, and can be extended with custom evaluations. The structural difference is origin and focus.

lm-evaluation-harness was designed primarily for benchmarking language model performance on academic NLP tasks (reasoning, knowledge, language understanding). Inspect was designed by a government AI safety research body and includes explicit support for multi-turn dialog evaluations, tool use evaluations, and model-graded scoring, which are evaluation patterns more relevant to assessing agentic capabilities and safety properties.

For teams running standard NLP benchmarks, lm-evaluation-harness has a larger catalog and a longer history. For teams evaluating tool-use, multi-turn, or agentic capabilities, Inspect's built-in support for those patterns reduces the custom code required.

Editorial conclusion

Inspect is the right framework for AI safety researchers and teams that need reproducible, auditable LLM evaluations with a maintained library of 200+ tasks. The MIT license imposes no commercial restrictions. The main constraint to verify before adopting it is whether the pre-built eval library covers your evaluation domain; if not, you will need to write custom tasks in Python. The uv.lock file provides a reproducible development environment, which matters for teams that need to reproduce results across machines.

Frequently asked questions

What is Inspect AI?

Inspect is a Python framework for large language model evaluations created by the UK AI Security Institute. It provides built-in support for prompt engineering, tool usage, multi-turn dialog, and model-graded scoring, and ships with more than 200 pre-built evaluation tasks.

How do you use Inspect AI?

Inspect is installed as a Python package. For development, the README gives a path using pip install -e ".[dev]" after cloning the repository, or uv sync --extra dev for uv users. Documentation is at inspect.aisi.org.uk, with a machine-readable index at inspect.aisi.org.uk/llms.txt.

What are the alternatives to Inspect AI?

EleutherAI's lm-evaluation-harness is the most widely used alternative open-source LLM evaluation framework. It covers a large library of academic NLP benchmark tasks. Inspect's advantage is its built-in support for tool use, multi-turn dialog, and model-graded evaluations.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ukgovernmentbeis-inspect-ai.svg)](https://hysenlabs.com/projects/ukgovernmentbeis-inspect-ai)