All concepts
Concept

What is LLM evaluation?

LLM evaluation (evals) is the practice of measuring how a language model or an LLM-based system performs on defined tasks, using datasets, metrics and scoring rules. It turns vague impressions about output quality into numbers and pass/fail results you can compare over time.

Published September 28, 2026

How LLM evaluation works

An eval has three moving parts: a dataset, a way to run the model over it, and a scoring rule. The dataset is a set of inputs, usually paired with expected outputs, reference answers or grading criteria. The runner sends each input to a model or a full pipeline (retrieval, tools, prompt templates) and records what comes back. The scorer then assigns a value to each response. Scoring methods fall into two broad families. Deterministic checks compare output to a target string, a JSON schema, a regular expression or a unit test, and they return an exact pass or fail. Model-based scoring asks another model to judge the response against a rubric, which handles open-ended answers but introduces its own variance and cost. Some frameworks support both, plus human review for cases where neither is trustworthy. The unit of measurement matters as much as the metric. A single accuracy number hides which inputs failed, so most tooling stores per-example results alongside aggregate scores. That is what makes an eval useful for regression testing: you rerun the same dataset after changing a prompt or swapping a model, then compare per-example outcomes, not just the average. Latency, token cost and refusal rates are often recorded in the same pass, because a model that answers correctly but slowly or expensively is a different engineering decision than one that is simply wrong.

When you need evals and when you do not

You need evals when a change to your system could silently degrade output. Prompt edits, model version bumps, retrieval changes and tool-calling logic all shift behaviour in ways that a few manual spot checks will miss. If you ship an LLM feature to other people, a small regression suite pays for itself the first time a provider updates a model under you. You also need them when you must justify a choice: picking between two models, two chunk sizes or two prompt strategies is guesswork without a fixed dataset to compare against. You do not need a full evaluation platform for a throwaway prototype, a single internal script, or a task where correctness is trivially checkable in production (for example, a classifier whose labels you already log). Building a dataset is the expensive part, and it is wasted effort if the feature may be deleted next week. A middle path is worth naming: start with twenty to fifty hand-written examples that cover the failure modes you actually fear, run them by hand, and only invest in automation once the feature survives. The trap is treating eval infrastructure as a prerequisite for any LLM work at all. It is not. It is a maintenance cost you take on when the system is expected to live.

Common pitfalls and limits

The first limit is contamination. Public benchmarks are scraped into training data sooner or later, so a high score on a well-known test set may reflect memorisation rather than capability. pathwaycom/arc-task-gen exists precisely because of this: according to its description, it generates original ARC-AGI-1-style tasks distribution-matched to the public eval set, so a model's score can be checked against problems it is unlikely to have seen. The repository ships the generator scripts and a short instruction file, not a hosted service. The second limit is the judge. Model-based scoring is convenient but noisy; a judge model has its own biases, and its verdicts drift when the judge is updated. Any eval that relies on an LLM judge should be spot-checked against human labels, and the judge version should be pinned. The third limit is dataset drift. A dataset written for one product version goes stale as the product changes, and a stale dataset rewards behaviour you no longer want. The fourth limit is overfitting to the eval itself. Once a team iterates against a fixed set of examples, improvements on that set stop predicting improvements in the wild. Hold-out sets and periodic dataset refreshes are the usual mitigation, and neither is free. Finally, aggregate scores compress information. A model that scores 80 percent on a benchmark may fail every example in the category you care about most. Per-slice reporting is more work and more honest.

How evals show up in open-source projects

Several projects in this space treat evaluation as their core product. openai/evals pairs a benchmark registry with a YAML-driven framework for running your own tests. According to our analysis, it is a strong fit if you already call OpenAI models and want reproducible comparisons, and it is not a general-purpose evaluation platform. open-compass/opencompass is a platform for running large-scale evaluations across open-weight checkpoints and commercial APIs, with support for a long list of models and datasets. confident-ai/deepeval takes a different angle: it turns evaluation into something closer to a test suite than a dashboard, running research-backed metrics locally against any model. Its trade-off, per our analysis, is a wide metric surface, documentation that functions as the real product, and a hosted platform one click away. langfuse/langfuse and langwatch/langwatch bundle evaluation with observability. Langfuse covers tracing, evals, metrics and prompt management, and our analysis notes it runs from a single docker compose file, is still being pushed to, and carries a licence file that is not plain MIT text. LangWatch bundles tracing, datasets, offline evaluation and prompt optimisation into a single Apache-2.0 codebase, aimed at teams that want regression testing and production observability without stitching five tools together. Kiln-AI/Kiln pairs a desktop app with an MIT-licensed Python library so one dataset can flow through evals, prompt optimisation, RAG and fine-tuning; our analysis flags it as a young project with an unusual licence file and documentation that leans on the app. ConardLi/easy-dataset converts PDFs, Markdown and other documents into structured datasets for fine-tuning, RAG and evaluation, with an AGPL license and reliance on external LLM APIs shaping where it fits. jeinlee1991/chinese-llm-benchmark is a Chinese-language capability benchmark scoring hundreds of models across seven domains, paired with a defect library; our analysis describes it as a reference table, not a test harness you run yourself. huggingface/pytorch-image-models (timm) is not an LLM tool, but it ships train, eval, inference and export scripts for image models, which is the same split between a model and the harness that measures it.

In practice

LLM evaluation is less a product category than a discipline: define a dataset, run the system over it, score the results, and keep the whole thing versioned so comparisons mean something. Start with a small set of examples you wrote yourself and a scoring rule you can defend. If you want a framework, openai/evals and open-compass/opencompass are the two most direct starting points, while langfuse/langfuse and langwatch/langwatch are the option when you also need tracing in production. Read the licence file and the documentation before committing; several of these projects have terms or maintenance patterns that do not match their reputation.

huggingface/pytorch-image-modelsThe largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more37,172 stars · Pythonlangfuse/langfuseGitHub describes it as 🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23. The repository metadata lists TypeScript as its primary language. The metadata lists the NOASSERTION license. This article stays within the project description and details documented in the GitHub repository README.35,105 stars · TypeScriptopenai/evalsEvals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.19,513 stars · Pythonconfident-ai/deepevalThe LLM Evaluation Framework18,313 stars · PythonConardLi/easy-datasetA powerful tool for creating datasets for LLM fine-tuning 、RAG and Eval14,961 stars · JavaScriptpathwaycom/arc-task-genGenerates original ARC-AGI-1-style tasks distribution-matched to the public eval set.11,108 stars · Pythonopen-compass/opencompassOpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.7,450 stars · Pythonjeinlee1991/chinese-llm-benchmark非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。6,457 starsKiln-AI/KilnBuild, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.5,071 stars · Pythonlangwatch/langwatchThe platform for LLM evaluations and AI agent testing4,886 stars · TypeScriptPrimeIntellect-ai/verifiersOur library for RL environments + evals4,655 stars · PythonEvolvingLMMs-Lab/lmms-evalOne-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks4,437 stars · Python

Sources

  1. huggingface/pytorch-image-models repository
  2. langfuse/langfuse repository
  3. openai/evals repository
  4. confident-ai/deepeval repository
  5. ConardLi/easy-dataset repository