# Evidently: an open-source framework for evaluating and monitoring ML and LLM systems

> Evidently is a Python library and optional self-hosted UI for computing data, model and LLM quality metrics, turning them into pass/fail tests, and watching them over time. The trade-off is a wide dependency surface and a monitoring service that is thinner than the hosted product.

**evidentlyai/evidently** — Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.

- Repository: https://github.com/evidentlyai/evidently
- Website: https://discord.gg/xZjKRaNp8b
- Stars: 7,923 · Forks: 918
- Language: Jupyter Notebook
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/evidentlyai-evidently

## The gap Evidently fills between a notebook and a production check

Most teams can train a model and most teams can log predictions. What is usually missing is the layer in between: something that answers whether the current batch of data still looks like the data the model was trained on, whether the model's outputs still meet a quality bar, and whether that answer is the same today as it was last week. Evidently targets exactly that layer.

The README describes it as "an open-source framework to evaluate, test and monitor ML and LLM-powered systems," and lists the intended scope: tabular and text data, predictive and generative tasks, offline evaluation and live monitoring. The audience implied by the repository is data scientists and ML engineers who already have pandas dataframes and want to attach quality checks to them without standing up a separate platform.

The project is licensed Apache-2.0 and requires Python 3.10 or newer according to pyproject.toml. It is not archived, and the last push to the default branch was on 2026-08-31.

## Reports, Test Suites and the Dataset object: how the pieces connect

The architecture is deliberately modular, and the README says so directly. There are three concepts you need to keep straight.

A Report computes and summarizes evaluations. You build one from a list of Presets or individual metrics, then call run() with data. The output object can be displayed in a notebook, serialized with .json() or .dict(), or written to a standalone HTML file with .save_html(). This is the exploratory path: run it, look at the output, decide what matters.

A Test Suite is the same Report with pass/fail conditions attached. The README describes turning any Report into a Test Suite by adding conditions, with a syntax that expresses thresholds as gt (greater than), lt (less than) and similar operators. This is the CI path: the same computation, but it returns a verdict instead of a chart.

The Dataset object is the input layer. Dataset.from_pandas() wraps a dataframe and accepts a DataDefinition plus a list of descriptors. Descriptors are row-level evaluators: the README example attaches Sentiment, TextLength and Contains to a text column, each with an alias, and then calls as_dataframe() to see the scores appended as columns. That design matters because it means the evaluation result is still a dataframe you can filter, join or export, rather than an opaque report object.

The data flow is therefore: dataframe in, descriptors and metrics applied per row or per column, results aggregated into a Report, optionally compared against thresholds, optionally persisted as HTML or JSON. The README also mentions an open architecture for exporting data and integrating with existing tools, which is consistent with the repository layout: there is an examples/grafana/ directory, suggesting a documented path for pushing results into Grafana rather than only the bundled UI.

## Installing Evidently and running a first drift report

The README gives two install paths. From PyPI:

```bash
pip install evidently
```

Or through conda-forge:

```bash
conda install -c conda-forge evidently
```

Python 3.10 or newer is required, per pyproject.toml. If you want an isolated environment first, the README also shows the standard virtualenv route with pip install virtualenv, virtualenv venv, and source venv/bin/activate.

The first real use is a data drift check. The README's tabular example loads the iris dataset from scikit-learn, splits it into a reference slice and a current slice, and runs the DataDriftPreset with the PSI method. Note the argument order in the README: the first dataframe passed to run() is the current data and the second is the reference.

```python
import pandas as pd
from sklearn import datasets

from evidently import Report
from evidently.presets import DataDriftPreset

iris_data = datasets.load_iris(as_frame=True)
iris_frame = iris_data.frame

report = Report([
    DataDriftPreset(method="psi")
],
include_tests="True")
my_eval = report.run(iris_frame.iloc[:60], iris_frame.iloc[60:])
my_eval
```

In a Jupyter notebook, evaluating my_eval renders the report inline. To keep a copy, the README shows my_eval.save_html("file.html") and notes you have to open the file from the destination folder. For programmatic use, my_eval.json() returns JSON and my_eval.dict() returns a Python dictionary.

For text and LLM evaluation, the shape is the same but the input is a Dataset with descriptors attached:

```python
import pandas as pd
from evidently import Dataset, DataDefinition
from evidently.descriptors import Sentiment, TextLength, Contains

eval_df = pd.DataFrame([
    ["What is the capital of Japan?", "The capital of Japan is Tokyo."],
    ["Who painted the Mona Lisa?", "Leonardo da Vinci."],
    ["Can you write an essay?", "I'm sorry, but I can't assist with homework."]],
                       columns=["question", "answer"])

eval_dataset = Dataset.from_pandas(pd.DataFrame(eval_df),
data_definition=DataDefinition(),
descriptors=[
    Sentiment("answer", alias="Sentiment"),
    TextLength("answer", alias="Length"),
    Contains("answer", items=['sorry', 'apologize'], mode="any", alias="Denials")
])
```

Calling eval_dataset.as_dataframe() shows the original columns plus the three descriptor scores. Running a Report over it with the TextEvals preset summarizes the distribution of those scores. The README points to the LLM evaluation tutorial for LLM-as-a-judge evaluators, which are mentioned but not shown in the quickstart.

## Launching the self-hosted monitoring UI on localhost:8000

Evidently ships a monitoring service in addition to the library. The README offers a one-command route if uv is installed:

```bash
uv run --with evidently evidently ui --demo-projects all
```

If you have already installed Evidently into an environment, the equivalent is:

```bash
evidently ui --demo-projects all
```

The README states you then visit localhost:8000 to access the UI, and that this launches a demo project. Two hosting options are described: self-host the open-source version, or sign up for Evidently Cloud, which the README marks as recommended and which it says adds dataset and user management, alerting and no-code evals. There is a live demo at demo.evidentlyai.com and a comparison page in the docs under faq/oss_vs_cloud.

That split is the most consequential design decision in the project. The library is fully open source under Apache-2.0. The operational layer around it, meaning the parts that page someone when a metric crosses a threshold, is where the README steers you toward the hosted product. If you self-host, you are running a Litestar and uvicorn service, both of which appear in the dependency list, and you own its uptime.

## Where Evidently is the wrong tool

The first limitation is dependency weight. pyproject.toml pulls in plotly, statsmodels, scikit-learn, pandas with parquet support, numpy, nltk, scipy, litestar, uvicorn, dynaconf, opentelemetry-proto and more. That is a reasonable set for a data science environment and a heavy one for a slim inference container. If your goal is a lightweight runtime check inside a serving process, this is not a small library.

The second is the beta classifier. pyproject.toml declares "Development Status :: 4 - Beta". The API surface has changed across the 0.7.x line, with v0.7.19, v0.7.20 and v0.7.21 released between January and March 2026. Anyone pinning Evidently should pin an exact version and read the release notes before moving, because the README's examples are written against a specific API shape (Dataset, DataDefinition, descriptors, presets) that earlier versions did not have.

The third is the boundary of the open-source product. The README does not document rollback, alert routing, or retention policy for the self-hosted UI; those are the areas it points at Evidently Cloud for. If your requirement is a monitored service with on-call alerting, the self-hosted path leaves that work to you.

Finally, this is a Python library. If your pipeline is JVM-based or runs entirely in a warehouse SQL dialect, Evidently does not meet you there.

## How Evidently differs from Great Expectations and WhyLabs

Great Expectations attacks a similar problem from the data quality side. Its core abstraction is an Expectation: a declarative assertion about a column, such as a value range or a null rate, attached to a data source and executed as a validation run. Evidently's core abstraction is a metric computed over two datasets, current and reference, which is what makes drift detection a first-class operation rather than something you assemble from column-level assertions. If your question is "does this column satisfy this rule," Great Expectations is the more direct fit. If your question is "has the distribution of this column moved relative to the training data," Evidently answers it without you defining the rule first.

WhyLabs takes the hosted-service approach from the start: you send profiles from a lightweight agent and the analysis and alerting happen in their platform. Evidently inverts that. The computation happens in your Python process, on your data, and the hosted layer is optional. That matters for teams who cannot ship raw data to a third party. It also means you carry the compute and the storage.

Within Evidently itself, the meaningful choice is Report versus Test Suite. Reports are for exploration and debugging; Test Suites are for regression checks in CI. The README frames them that way, and picking the wrong one is a common way to end up with a pipeline that either fails constantly on exploratory thresholds or never fails at all.

## Maintenance, versioning and what the Apache-2.0 licence means here

The repository is not archived and the last push to main was on 2026-08-31, so the codebase is being touched. The release cadence visible in the release list is three releases in the first quarter of 2026: v0.7.19 on 2026-01-05, v0.7.20 on 2026-01-09, and v0.7.21 on 2026-03-10. Nothing in the repository shows a 1.0 release, and pyproject.toml still declares Beta status.

The practical upgrade cost is API churn. Because the README's examples rely on Dataset, DataDefinition, descriptors and presets, a version bump can invalidate code you wrote against an earlier interface. Pin the version in your requirements file, and treat the release notes as required reading before any bump.

On licensing: Apache-2.0 is a permissive licence that allows commercial use, modification and redistribution, and it includes an explicit patent grant. It also requires that you preserve the licence and notice files and state significant changes. This is not legal advice; if you are embedding Evidently in a distributed product, have counsel confirm your notice obligations. Note that the licence covers the open-source library. Evidently Cloud is a separate commercial service with its own terms, and the README treats it as such.

## Conclusion

Adopt Evidently if you already work in Python and pandas and want drift detection, data validation and LLM response checks in one library, plus the option to turn any report into a pass/fail test suite for CI. Skip it if you need a managed monitoring service with alerting out of the box, or if your stack is not Python. Before committing, verify two things against a real dataset: that DataDriftPreset with method="psi" flags the shifts you actually care about rather than every column, and that the self-hosted UI on localhost:8000 covers the retention and alerting your team expects, since the README points to Evidently Cloud for those features.

## FAQ

### How to install Evidently?

The README gives two options: pip install evidently from PyPI, or conda install -c conda-forge evidently. Python 3.10 or newer is required according to pyproject.toml.

### How to use Evidently?

The README's quickstart builds a Report from a preset such as DataDriftPreset, calls report.run() with current and reference dataframes, and then displays the result or saves it with save_html(). For text data, you wrap a dataframe in a Dataset and attach descriptors like Sentiment or TextLength before running a Report with the TextEvals preset.

### What is Evidently AI used for?

The README describes it as an open-source Python library to evaluate, test and monitor ML and LLM systems, covering tabular and text data across predictive and generative tasks. It ships 100+ built-in metrics, supports custom metrics through a Python interface, and covers both offline evaluation and live monitoring.

## Sources

- [evidentlyai/evidently on GitHub](https://github.com/evidentlyai/evidently)
- [License: Apache-2.0](https://github.com/evidentlyai/evidently/blob/main/LICENSE)
- [Project website](https://discord.gg/xZjKRaNp8b)
- [README](https://github.com/evidentlyai/evidently/blob/main/README.md)
- [Releases](https://github.com/evidentlyai/evidently/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/evidentlyai-evidently
