DeepEval: An LLM Evaluation Framework Structured Around Pytest
The LLM Evaluation Framework
At a glance
- What is it?
- DeepEval is an open-source Python framework for evaluating large language model systems, built to work like Pytest for unit testing LLM applications. It provides metrics for RAG pipelines, agentic workflows, and multi-turn conversations, running LLM-as-a-judge evaluation locally.
- Who is it for?
- DeepEval is the right tool for Python developers who want to add structured evaluation to an LLM application using a test framework they already know. It handles RAG pipelines, agents, and multi-turn conversations with metrics that run locally.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 14 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
LLM Evaluation Structured Around Pytest
DeepEval positions itself as Pytest for LLM applications. The framework integrates directly with Pytest, registering a plugin via the pytest11 entry point defined in pyproject.toml. Test files use DeepEval's evaluation primitives alongside standard pytest constructs, which means existing Python test infrastructure works without modification.
The framework evaluates LLM systems in three modes: end-to-end as a black box, as complete agent trajectories across every decision and action, and at the level of individual agent steps such as LLM calls, tool use, retrieval, and sub-agent handoffs. These three modes correspond to different granularities of debugging: a failing end-to-end evaluation tells you the system output is wrong, but step-level evaluation tells you which part of the pipeline caused it.
DeepEval is described in the README as useful for choosing between models, prompts, and architectures, preventing prompt drift over time, and managing transitions between LLM providers. The .env.example file lists keys for OpenAI, Azure OpenAI, Google (Gemini), xAI Grok, Moonshot, DeepSeek, LiteLLM, Anthropic, and Openrouter, covering the providers the framework can use as evaluation judges.
How LLM-as-a-Judge Metrics Score Outputs
Most DeepEval metrics use an LLM-as-a-judge approach, where a separate LLM evaluates the output of the system under test. The framework provides several judge-based metric types, each backed by published research or a defined scoring mechanism.
G-Eval is described as a research-backed LLM-as-a-judge metric for evaluating on any custom criteria with human-like accuracy. DAG (DeepEval's graph-based deterministic LLM-as-a-judge metric builder) provides a structured way to define evaluation criteria as a directed graph.
JevEval, introduced in the v4.2.4 release, uses a model called Jev. The README describes Jev as a System One model that answers evaluation questions with calibrated probabilities instead of generated text. This is a different architecture from a generative judge: instead of producing a textual justification, it produces a probability score for each criterion.
Not all metrics require an external LLM call. The README states that some metrics use statistical methods or NLP models that run locally on your machine. This is relevant for teams that want to keep evaluation costs predictable or that cannot send evaluation data to an external API.
Configuring DeepEval and Running Your First Evaluation
DeepEval is a Python package. The pyproject.toml names it deepeval, sets the version as 4.2.6, and requires Python 3.9 to below 4.0. The package includes a CLI registered as deepeval in the poetry scripts section.
The .env.example file documents the expected environment variables. To use OpenAI as the evaluation judge, set OPENAI_API_KEY. To use Azure OpenAI, set AZURE_OPENAI_API_KEY, AZURE_OPENAI_ENDPOINT, AZURE_DEPLOYMENT_NAME, OPENAI_API_VERSION, and AZURE_MODEL_NAME. For Anthropic, set ANTHROPIC_API_KEY. For a local Ollama model, the README describes how to configure a local model endpoint.
The .env.example file covers additional providers used for voice mode evaluation, including ElevenLabs, Deepgram, Cartesia, and AssemblyAI for TTS and STT scenarios. These are optional; the core evaluation functionality uses the LLM API keys above.
The package includes a pytest plugin that registers automatically when deepeval is installed. The examples directory contains subdirectories for getting started, RAG evaluation, MCP evaluation, agent tracing, and notebook walkthroughs.
RAG, Agentic, and Multi-Turn Metric Categories
DeepEval organizes its metrics into categories that correspond to different application architectures.
For RAG pipelines, the framework provides answer relevancy, faithfulness, contextual recall, contextual precision, and contextual relevancy. Faithfulness measures whether the output factually aligns with the retrieval context, which is distinct from answer relevancy, which measures whether the output addresses the input. The RAGAS metric computes an average of answer relevancy, faithfulness, contextual precision, and contextual recall as a single composite score.
For agentic workflows, the metrics include task completion, tool correctness, goal accuracy, step efficiency, plan adherence, plan quality, tool use quality, and argument correctness. Tool correctness checks whether the right tools were called with the right arguments. Step efficiency measures whether the agent took unnecessary steps to complete a task.
For multi-turn conversational systems, the framework offers knowledge retention, conversation completeness, turn relevancy, and turn faithfulness. Knowledge retention evaluates whether a chatbot retains factual information across the turns of a conversation, which is a distinct failure mode from single-turn hallucination.
The README notes that the framework covers LangChain and OpenAI-based implementations, and the dev dependencies in pyproject.toml include langchain, langchain_core, langchain_community, langchain_text_splitters, and anthropic as development dependencies.
Limitations and Cases Where DeepEval Is Not Enough
DeepEval's LLM-as-a-judge approach introduces an evaluation cost. Each metric call that uses a judge LLM sends a request to an external API (or to a local model). For large test suites with many test cases and multiple metrics per case, the evaluation cost in time and API spend can be substantial.
The local package does not store evaluation results across runs in a way that supports team-wide comparison. The README explicitly points users toward the Confident AI platform for comparing iterations, sharing evaluation reports, and monitoring AI in production. Confident AI is described as an enterprise product and is separate from the open-source DeepEval package.
The framework requires that the evaluated system return structured outputs in the format each metric expects. For example, faithfulness requires both the generated output and the retrieval context. Plugging in a system that does not surface its retrieval context requires wrapping or instrumenting the pipeline to capture that data.
The pyproject.toml shows the project uses poetry for builds and locks dependencies. Teams using pip-only environments or strict dependency management need to verify that the package installs correctly in their specific environment.
Comparing DeepEval to Static Benchmarks
The alternative to a framework like DeepEval is evaluating LLM systems using static benchmarks: fixed datasets with pre-scored answers where the system is evaluated on how often it matches the reference output. Static benchmarks are reproducible and cheap to run once the dataset is fixed, but they measure performance on a specific distribution and do not generalize to the system's actual inputs.
DeepEval takes a different approach by running evaluation on the data your system actually processes and using metrics that apply to your specific evaluation criteria. G-Eval, for example, accepts custom criteria in natural language, which means the metric can be tuned to what matters for a specific application without building a new dataset.
The trade-off is that LLM-as-a-judge evaluation introduces its own source of variance: the judge LLM can produce inconsistent scores across runs, and its calibration depends on the model and prompt used for evaluation. JevEval's probability-based scoring is one approach to address this, but it is a newer addition and the README does not document its calibration methodology in detail.
Editorial conclusion
DeepEval is the right tool for Python developers who want to add structured evaluation to an LLM application using a test framework they already know. It handles RAG pipelines, agents, and multi-turn conversations with metrics that run locally. Teams that need full production observability and evaluation sharing beyond what the open-source package provides are directed by the README to the Confident AI platform, which is a separate paid product. The package is version 4.2.6, requires Python 3.9 to 3.x, and is Apache-2.0 licensed. The last push was on 2026-09-16.
Frequently asked questions
What is DeepEval used for?
DeepEval is used to evaluate LLM applications including RAG pipelines, AI agents, and chatbots. It provides metrics for answer relevancy, faithfulness, task completion, tool correctness, and multi-turn conversation quality, structured to work like Pytest unit tests.
How do I install DeepEval?
DeepEval is a Python package named deepeval in pyproject.toml, requiring Python 3.9 or later. The README points to deepeval.com/docs/getting-started for installation steps. After installation, set the relevant API key in a .env file using the OPENAI_API_KEY or equivalent variable for your chosen LLM judge.
What is G-Eval in DeepEval?
G-Eval is a research-backed LLM-as-a-judge metric in DeepEval for evaluating outputs against any custom criteria. You define the evaluation criteria in natural language, and G-Eval uses a judge LLM to score the output with human-like accuracy.
Is DeepEval free and open source?
The DeepEval Python package is open source under the Apache-2.0 license. Using it requires your own API key for the judge LLM, which incurs API costs. The Confident AI platform for team-level evaluation sharing and production monitoring is a separate paid product.
What is the difference between DeepEval and RAGAS?
RAGAS is a specific metric in DeepEval that computes a composite score from answer relevancy, faithfulness, contextual precision, and contextual recall. DeepEval is the broader framework that includes RAGAS as one of many metrics, alongside agentic metrics, multi-turn metrics, and custom G-Eval criteria.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/confident-ai-deepeval)