# continuous-eval: Modular, Data-Driven Evaluation for LLM Pipelines

> continuous-eval is a Python package for scoring LLM-powered applications module by module, mixing deterministic, semantic and LLM-based metrics. It fits RAG and multi-step pipelines, but its documentation is thin on what happens when a metric fails.

**relari-ai/continuous-eval** — Data-Driven Evaluation for LLM-Powered Applications

- Repository: https://github.com/relari-ai/continuous-eval
- Website: https://continuous-eval.docs.relari.ai/
- Stars: 517 · Forks: 38
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/relari-ai-continuous-eval

## The problem continuous-eval targets: a single score hides which stage broke

Most LLM evaluation tools return one number for the whole application. That number tells you the answer was wrong, not whether the retriever missed the document or the generator ignored it. continuous-eval is built around the opposite assumption: that an LLM application is a graph of modules, and each module should be measured with metrics that fit its job. The README states the package is "created for data-driven evaluation of LLM-powered application" and lists modularized evaluation as its first differentiator: "Measure each module in the pipeline with tailored metrics." The audience is engineers running retrieval-augmented generation, code generation, agent tool use or classification pipelines who already have logged inputs and outputs and want to attribute failures to a stage. It is not a tracing product and not an observability dashboard; it consumes data you already collected and produces per-module scores and test outcomes.

## How the module graph and metric binding actually work

The mechanism visible in the README is a three-part object model: Dataset, Module, Pipeline. A Module declares a name, an input, an output type, and a list of metrics bound to specific fields. The binding is explicit. In the RAG example, the retriever module wraps PrecisionRecallF1 with `retrieved_context=ModuleOutput(page_content)`, where page_content is a plain Python function that pulls `doc["page_content"]` out of the module's raw output. That indirection is the interesting design choice: metrics do not assume a fixed schema, so a reranker returning `List[Dict[str, str]]` and a generator returning `str` can both feed the same metric family as long as you supply an extractor. Modules chain by passing one module as another's input, and `Pipeline([retriever, reranker, llm], dataset=dataset)` assembles them. The README shows `pipeline.graph_repr()` to render the graph in Mermaid format. Evaluation itself runs through EvaluationRunner, which takes the pipeline, executes the bound metrics, and returns an aggregate result object. Tests are separate: `runner.test(eval_results)` takes the results and applies threshold rules. The README's example uses `GreaterOrEqualThan(test_name="Recall", metric_name="context_recall", min_value=0.8)`, which ties a named test to a metric key and a floor. That split matters. Metrics produce numbers; tests turn numbers into pass or fail signals you can gate a deployment on. LLM-based metrics call out to a provider, and the README notes the code "requires at least one of the LLM API keys in `.env`", with `.env.example` listing OpenAI, Anthropic, Gemini, Cohere and Azure OpenAI entries plus an `EVAL_LLM` default.

## Installing continuous-eval and running your first metric

The package ships on PyPI. The README gives the install as a single pip command, and a source install through Poetry with all optional extras.

```bash
python3 -m pip install continuous-eval
```

For a source checkout, the README uses:

```bash
git clone https://github.com/relari-ai/continuous-eval.git && cd continuous-eval
poetry install --all-extras
```

The `--all-extras` flag pulls in the optional provider and semantic dependencies declared in pyproject.toml (semantic, bedrock, azure, anthropic, cohere, google). If you only need retrieval metrics, the base install is enough. For any LLM-based metric, copy `.env.example` to `.env` and fill in at least one key; the file also defines `EVAL_LLM`, which the example sets to `gpt-3.5-turbo-0125`.

The smallest useful run is a single metric on a single datum, straight from the README:

```python
from continuous_eval.metrics.retrieval import PrecisionRecallF1

datum = {
    "question": "What is the capital of France?",
    "retrieved_context": [
        "Paris is the capital of France and its largest city.",
        "Lyon is a major city in France.",
    ],
    "ground_truth_context": ["Paris is the capital of France."],
    "answer": "Paris",
    "ground_truths": ["Paris"],
}

metric = PrecisionRecallF1()
print(metric(**datum))
```

Expect a dictionary of scores rather than a single float, since precision, recall and F1 are computed together. PrecisionRecallF1 is deterministic; it compares retrieved and ground-truth context sets without calling a model.

## Running a dataset evaluation and turning scores into tests

The dataset path uses `example_data_downloader("retrieval")` to fetch a sample, then binds two retrieval metrics to a SingleModulePipeline and attaches one test. The README's script prints the aggregate results and the elapsed time, then runs the tests.

```python
from continuous_eval.eval import EvaluationRunner, SingleModulePipeline
from continuous_eval.eval.tests import GreaterOrEqualThan
from continuous_eval.metrics.retrieval import PrecisionRecallF1, RankedRetrievalMetrics

pipeline = SingleModulePipeline(
    dataset=dataset,
    eval=[
        PrecisionRecallF1().use(
            retrieved_context=dataset.retrieved_contexts,
            ground_truth_context=dataset.ground_truth_contexts,
        ),
        RankedRetrievalMetrics().use(
            retrieved_context=dataset.retrieved_contexts,
            ground_truth_context=dataset.ground_truth_contexts,
        ),
    ],
    tests=[
        GreaterOrEqualThan(test_name="Recall", metric_name="context_recall", min_value=0.8),
    ],
)
```

The README wraps the run in a `main()` function with a note that it is "important to run this script in a new process to avoid multiprocessing issues." That is a real operational constraint, not a stylistic preference: the runner uses multiprocessing, so a script that imports and evaluates at module scope can misbehave. Follow the README's structure and guard the entry point with `if __name__ == "__main__":`.

For multi-module pipelines, the same pattern repeats per stage. The README's three-step RAG example creates a retriever, a reranker and an llm module, each with its own `eval` list, then builds `Pipeline([retriever, reranker, llm], dataset=dataset)`. The reranker's `ModuleOutput(page_content)` shows that the same extractor function can be reused across stages when the output shape is similar.

## Where continuous-eval gets awkward: metric coverage and LLM dependency

Two limitations stand out from the repository itself. First, the README's own example list is narrow. The examples directory contains four files: `single_metric.py`, `evaluation_on_dataset.py`, `modular_evaluation.py` and `llm_custom_criteria.py`. The README advertises a "Comprehensive Metric Library" covering RAG, code generation, agent tool use and classification, but the walkthrough only demonstrates retrieval and answer correctness. If you need a code-generation or tool-use metric, the README does not show how to bind it; you will be reading the docs site or the source under `continuous_eval/` to find the class names and expected field names.

Second, the LLM-based metrics are not local. They require an API key and send evaluation data to the provider named in `.env`. For teams with data-residency constraints or no budget for evaluation calls, this rules out the semantic and LLM-judge metrics entirely, leaving the deterministic ones. The `semantic` extra is the escape hatch: it pulls in torch, transformers and sentence-transformers, which run locally, but pyproject.toml marks those dependencies `python = "^3.11"`, so a Python 3.10 environment cannot install them. The package's supported range is `>=3.10,<3.13`, and the semantic path is narrower than that.

A third constraint is the version cadence. The most recent release listed is v0.3.14 from 2025-01-10, while the last push to the repository was on 2026-08-10. That gap between the latest tagged release and the latest commit is worth noting before you pin a version: pyproject.toml declares `version = "0.3.14post2"`, so the installable package may carry fixes that are not in the v0.3.14 tag.

## continuous-eval compared with a single-score harness

The closest alternative approach is a single-score evaluation harness such as Ragas, which centers on end-to-end RAG quality metrics like faithfulness and answer relevancy computed over question, answer and context triples. The difference is structural rather than a matter of metric quality. A single-score harness assumes one boundary: the application in, a quality number out. continuous-eval assumes you can describe the application as a graph of modules and that you know which fields each metric needs. That buys you attribution, since a drop in the reranker's ranked retrieval metrics is visible separately from a drop in answer correctness. It costs you setup: every module needs an output type, an extractor where the raw output is nested, and a metric binding. If your application is a single prompt and a single response, the module graph is overhead with no payoff, and a flat harness will get you to a number faster. If your application is a retriever feeding a reranker feeding a generator, the flat approach will tell you the system got worse without telling you where.

## Licence and the cost of keeping up

continuous-eval is licensed Apache-2.0, declared both in the repository LICENSE file and in pyproject.toml as `license = "Apache-2.0"`. That is a permissive licence with an explicit patent grant, and it does not impose copyleft obligations on your application code. It says nothing about the terms of the LLM providers you configure in `.env`; those are separate agreements, and the evaluation data you send through LLM-based metrics is governed by them, not by this licence. This is not legal advice; check with your own counsel if the distinction matters to your organisation.

The upgrade cost is mostly dependency surface. The base install pulls in openai, tiktoken, scikit-learn, nltk, rouge, sqlglot, json-repair, jinja2 and posthog, among others. The semantic extra adds torch, transformers and sentence-transformers, which is a large download. Because the package pins `python = ">=3.10,<3.13"`, a Python 3.13 runtime is not supported, and the semantic extra additionally requires 3.11 or newer. Upgrading Python versions ahead of the package's range will block installation. The repository includes a `.pre-commit-config.yaml` and ruff configuration with `line-length = 80`, so if you plan to contribute, your formatting will be checked against that.

## Conclusion

Adopt continuous-eval if your application is a multi-stage pipeline (retriever, reranker, generator) and you need per-module scores plus threshold tests in CI. Skip it if you only need a single end-to-end score, or if you cannot send evaluation data to an external LLM provider, since the LLM-based metrics depend on API keys in .env. Before committing, verify two things: that the pinned Python range (>=3.10,<3.13) matches your runtime, and that the metric you plan to use is actually documented on the docs site, because the README stops at the retrieval and generation examples.

## FAQ

### What is continuous-eval and what is it used for?

It is an open-source Python package from Relari for data-driven evaluation of LLM-powered applications. The README describes it as measuring each module in a pipeline with tailored metrics, mixing deterministic, semantic and LLM-based metrics.

### How do I install continuous-eval?

The README gives `python3 -m pip install continuous-eval` for the PyPI package, or a git clone followed by `poetry install --all-extras` for a source install. LLM-based metrics additionally need at least one API key in a `.env` file.

### Does continuous-eval require an LLM API key?

Only for LLM-based metrics. The README states the code requires at least one LLM API key in `.env` to run them, and `.env.example` lists keys for OpenAI, Anthropic, Gemini, Cohere and Azure OpenAI plus an `EVAL_LLM` default. Deterministic metrics such as PrecisionRecallF1 do not call a model.

### What Python versions does continuous-eval support?

pyproject.toml declares `python = ">=3.10,<3.13"`. The optional semantic extra, which brings in torch, transformers and sentence-transformers, is further restricted to Python 3.11 and above.

### What licence is continuous-eval released under?

Apache-2.0, declared in both the repository LICENSE file and the `license` field of pyproject.toml.

## Sources

- [License: Apache-2.0](https://github.com/relari-ai/continuous-eval/blob/main/LICENSE)
- [Project website](https://continuous-eval.docs.relari.ai/)
- [README](https://github.com/relari-ai/continuous-eval/blob/main/README.md)
- [relari-ai/continuous-eval on GitHub](https://github.com/relari-ai/continuous-eval)
- [Releases](https://github.com/relari-ai/continuous-eval/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/relari-ai-continuous-eval
