# Giskard OSS v3: evals, red teaming and test generation for agentic systems

> Giskard OSS is an Apache-2.0 Python library that splits agent testing into two stable packages: giskard-checks for evals and giskard-scan for vulnerability and RAG quality scanning. It is a rewrite, so v2 users have a migration to plan.

**Giskard-AI/giskard-oss** — 🐢 Open-Source Evaluation & Testing library for LLM Agents

- Repository: https://github.com/Giskard-AI/giskard-oss
- Website: https://docs.giskard.ai
- Stars: 5,822 · Forks: 537
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/giskard-ai-giskard-oss

## What Giskard OSS v3 is for, and who should care

Giskard OSS is an open-source Python library for testing and evaluating agentic systems. The problem it addresses is specific: an agent that returns a different but still valid answer on every run cannot be checked with assertEqual. The README frames evals as the alternative to traditional unit tests, aimed at non-deterministic outputs where the same input can produce different valid responses. That framing decides the audience. If you ship an LLM-backed feature and your regression suite is a pile of exact-match assertions, this is built for you. If your system is deterministic, ordinary pytest assertions remain cheaper and clearer.

The v3 line is a rewrite, not an increment. The README states that v3 drops heavy dependencies and ships a more powerful AI vulnerability scanner and enhanced RAG evaluation natively in giskard-scan, with no dependency on v2. It also states plainly that Giskard v2 remains available but is no longer actively maintained. That sentence matters more than any feature list, because it tells you the upgrade is a migration rather than a version bump. The one documented exception is the legacy scan for tabular and ML models, which the README says remains v2-only. Teams that came to Giskard for tabular model validation have no v3 path and should not assume one is coming.

## The modular package split: giskard-checks, giskard-scan and three foundations

The architecture is a set of focused packages, each carrying only the dependencies it needs. Two are marked stable in the README. giskard-checks handles testing and evaluation: a scenario API, built-in checks and LLM-as-judge. giskard-scan handles red teaming, prompt injection, jailbreaks, harmful content and knowledge-base quality evaluation. The README names vulnerability_scan as the successor to v2 Scan and quality_scan as the successor to v2 RAGET.

Underneath sit three libraries that are pulled in automatically and, per the README, rarely used directly: giskard-core for shared utilities and telemetry, giskard-llm for provider-agnostic LLM routing, and giskard-agents for agent and workflow orchestration. The root pyproject.toml shows how thin the top-level package is. Its only runtime dependency is giskard-checks, and everything else arrives through extras: openai, google, anthropic and azure pull giskard-llm with the matching provider SDK, litellm pulls giskard-agents, scan pulls giskard-scan, and full aggregates all-llms, all-checks, scan and litellm. The Makefile confirms the same shape from the build side, listing LIBS as giskard-core, giskard-llm, giskard-agents, giskard-checks and giskard-scan, and describing the root giskard distribution as an umbrella that ships no code of its own and only pins version ranges over those libraries.

That design has a cost worth naming. Extras mean a working install is a combination of choices, and the failure mode is an ImportError from a missing provider SDK rather than a helpful message. The pyproject classifier still reads Development Status 4 - Beta despite the stable labels on the two main packages, so the project is not claiming production maturity across the whole tree.

## Installing Giskard OSS and running a first eval

The README gives three install lines and one hard requirement: Python 3.12 or newer. The base install brings checks plus agents, llm and core; the scan extra adds the vulnerability and quality scanner; the openai extra adds a provider SDK for LLM judges and generators. Pick the extras that match what you plan to run.

```bash
pip install giskard           # checks (+ agents, llm, core)
pip install "giskard[scan]"   # + vulnerability / quality scan
pip install "giskard[openai]" # provider SDK for LLM judges / generators
```

Before the first run, decide on telemetry. The README says analytics are optional and aggregated through giskard-core, with no prompts or outputs sent. Two environment variables opt out, and the README notes they should be set before import to skip creating the identity file, though setting them later still stops further sends.

```bash
export DO_NOT_TRACK=1
# or
export GISKARD_TELEMETRY_DISABLED=1
```

The quickstart is a single scenario with one groundedness check. A plain function stands in for the system under test, the scenario wires an interaction to it, and the check asserts the answer is grounded in a supplied context. The README notes that Groundedness is an LLM judge, so the openai extra and a matching API key are required, and that the default model is openai/gpt-4o-mini.

```python
import asyncio
from giskard.checks import Scenario, Groundedness


def get_answer(inputs: str) -> str:
    return "Paris"  # replace with your model / agent


async def main() -> None:
    scenario = (
        Scenario("test_france_capital")
        .interact(inputs="What is the capital of France?", outputs=get_answer)
        .check(
            Groundedness(
                name="answer is grounded",
                context="France is in Western Europe. Its capital is Paris.",
            )
        )
    )
    result = await scenario.run()
    result.print_report()


asyncio.run(main())
```

Running that prints a report for the scenario. Two concepts carry the rest of the API. A target is any sync or async callable from inputs to outputs, optionally with a trace, which is why an existing agent function can be wrapped without restructuring it. A suite is many scenarios run together. The README also draws a distinction that trips people up: giskard.agents.Generator is an LLM client for workflows and judges, while giskard.checks input generators such as LLMGenerator synthesize user messages. Same word, different layer.

## Where Giskard OSS v3 is the wrong tool

The clearest limitation is written into the README rather than discovered later. Only the legacy scan for tabular and ML models remains v2-only, and v2 is no longer actively maintained. A team whose validation work is classical tabular ML is therefore choosing between an unmaintained branch and leaving. That is a real fork in the road, not a footnote.

The second constraint is the runtime floor. Python 3.12 or newer, stated in both the README and pyproject.toml. Repositories pinned to 3.10 or 3.11 cannot install this without an interpreter upgrade, and the pyproject classifiers list 3.12, 3.13 and 3.14, so there is no older-Python fallback to wait for.

The third is dependency weight by choice. The scan extra is separate for a reason, and the pyproject marks garak and deepteam as heavy dependencies not included by default. Installing giskard-scan without them means the scanner runs without those integrations. Installing them pulls a larger tree into your environment, which is exactly the trade the modular split is meant to let you make consciously.

Finally, the project is not a hosted service and the README does not present it as one. There is no documented rollback procedure, no result store and no dashboard described in the README. Scenario results are printed through result.print_report(). If you need historical tracking of eval outcomes across runs, that is your CI system's job, and the README does not document a built-in answer.

## Giskard Checks against generic eval frameworks

The obvious alternative is a general-purpose evaluation framework, and the difference is where the assertion lives. A typical eval library scores a dataset: you hand it a list of inputs, it produces a table of metrics, and you compare numbers between runs. Giskard Checks inverts that. The Scenario is the unit, interactions and checks are attached to it, and the result is a pass or fail report per scenario. That makes it behave like a test suite you can run in CI rather than a scoring harness you run offline.

The second difference is the target contract. The README defines a target as any sync or async callable from inputs to outputs, optionally with a trace. There is no required adapter class, no mandated dataset format, and no separate serving step. A function that returns a string is a valid target, which is what the quickstart demonstrates. Frameworks that expect a registered model endpoint or a fixed dataset schema ask for more setup before the first check runs.

The third difference is that red teaming and evaluation ship as separate packages from the same project. giskard-scan covers prompt injection, jailbreaks, harmful content and knowledge-base quality, while giskard-checks covers groundedness, conformity and LLM-as-judge. You can install one without the other. A dataset-scoring framework typically leaves adversarial testing to a different tool entirely, which means two configuration surfaces and two sets of provider credentials. Whether that consolidation is worth it depends on how much you value running both from one Python environment.

## Maintenance, upgrade cost and licence

The repository is not archived, and the last push was on 2026-09-09. Releases are recent and split by package: giskard v3.0.0, giskard-scan v1.0.0 and giskard-llm v1.0.0, all dated 2026-08-26. A three-package simultaneous 1.0 or 3.0 cut suggests the split is new, and new splits tend to move. The root pyproject pins giskard-checks at >=1.0.4,<2 and giskard-scan at >=1.0.0,<2, so the umbrella package constrains its own components within a major version. That is the practical upgrade boundary: within-major updates should be safe to take, and a major bump in any of the five libraries is the signal to read the changelog before upgrading.

Upgrade cost from v2 is the larger number. The README is explicit that v3 is a fresh rewrite with no dependency on v2, that vulnerability_scan succeeds the v2 Scan, and that quality_scan succeeds v2 RAGET. Renamed entry points mean v2 call sites will not resolve. There is no documented compatibility shim in the README, and the giskard_pypi_shim directory at the repository root is not described there, so do not assume it is one.

The licence is Apache-2.0, declared in both the README badge and pyproject.toml. That is a permissive licence with an explicit patent grant and a notice requirement, which is generally friendly to commercial use. Note the third-party angle: the repository carries a THIRD_PARTY_NOTICES.md file, and the scan extra can pull garak and deepteam, whose own licences apply to those components. Check that file before redistributing a bundled environment. This is a description of what the repository states, not legal advice.

## Conclusion

Adopt Giskard OSS v3 if you already run agents in Python and want evals and red-team scans in the same test suite; the scenario API wraps any sync or async callable, so nothing has to be rewritten to be testable. Do not adopt it if your work is still tabular or classical ML scanning, because the README says that path stays v2-only and v2 is no longer actively maintained. Before committing, verify two things yourself: that your environment is on Python 3.12 or newer, and that your provider extra is installed, since Groundedness and the other judges default to openai/gpt-4o-mini and will not run without it.

## FAQ

### Is Giskard OSS open source?

Yes. The repository is public, the primary language is Python, and the licence is Apache-2.0 as declared in both the README badge and pyproject.toml. The source is hosted at github.com/Giskard-AI/giskard-oss with documentation at docs.giskard.ai.

### What Python version does Giskard OSS v3 require?

Python 3.12 or newer. The README states this as a requirement, pyproject.toml sets requires-python to >=3.12, and the classifiers list 3.12, 3.13 and 3.14.

### What is the difference between giskard-checks and giskard-scan?

giskard-checks is the testing and evaluation package, with the scenario API, built-in checks and LLM-as-judge. giskard-scan is the red-teaming and vulnerability scanner, covering prompt injection, jailbreaks, harmful content and knowledge-base quality evaluation. Both are marked stable in the README.

### Can I use Giskard OSS v3 to scan tabular or ML models?

No. The README states that only the legacy scan for tabular and ML models remains v2-only, and that Giskard v2 remains available but is no longer actively maintained. The v3 scanner targets agentic systems.

### Does Giskard OSS send my prompts or model outputs anywhere?

The README describes telemetry as optional aggregated analytics sent through giskard-core, and states that no prompts or outputs are sent. You can opt out with DO_NOT_TRACK=1 or GISKARD_TELEMETRY_DISABLED=1, and setting either before import skips creating the local identity file.

## Sources

- [Giskard-AI/giskard-oss on GitHub](https://github.com/Giskard-AI/giskard-oss)
- [License: Apache-2.0](https://github.com/Giskard-AI/giskard-oss/blob/main/LICENSE)
- [Project website](https://docs.giskard.ai)
- [README](https://github.com/Giskard-AI/giskard-oss/blob/main/README.md)
- [Releases](https://github.com/Giskard-AI/giskard-oss/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/giskard-ai-giskard-oss
