Model or dataset
modelscope/evalscope avatar
modelscope/evalscope

EvalScope: One-Command LLM Evaluation and Performance Benchmarking

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

3,470 stars498 forksPythonApache-2.0

At a glance

What is it?
EvalScope is a Python framework from the ModelScope Community for evaluating LLMs, VLMs, and AI agents against standard benchmarks and custom datasets, with a single CLI entry point, a web dashboard for result comparison, and an agent loop that records full per-sample traces.
Who is it for?
EvalScope suits teams that need a single tool to run capability benchmarks, stress-test inference performance, and evaluate agents end-to-end, particularly if they are already in the ModelScope ecosystem. It is less suited to teams that need a simple, minimal evaluation harness with few dependencies, since EvalScope's optional extras bring in OpenCompass, VLMEvalKit, and RAGEval as backends.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 9 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What EvalScope Does and Who It Is For

Evaluating a large language model requires more than running it on a benchmark dataset: different benchmarks use different scoring methods, different model APIs have different request formats, and comparing models reliably requires fixing the evaluation setup across runs. EvalScope addresses this by providing a unified evaluation framework with a single CLI command that handles benchmark data, model querying, scoring, and result storage.

The project targets ML engineers and researchers who evaluate LLMs, vision-language models, embedding models, rerankers, and AIGC models. The README positions it as a one-stop framework: capability evaluation, inference performance stress testing, and visualization in one package.

The last push was on 2026-09-21. The most recent release at the time of writing was v1.12.0, published on 2026-09-16. The license is Apache-2.0.

Quick Start: Running a Benchmark in One Command

Install EvalScope with pip:

bash
pip install evalscope

The README shows a one-line evaluation against the GSM8K benchmark:

bash
evalscope eval --model your-model-name --api-url $OPENAI_API_BASE_URL --api-key $OPENAI_API_KEY --eval-type openai_api --datasets gsm8k --limit 5

This command queries the model through an OpenAI-compatible API endpoint, runs the first 5 examples from GSM8K, scores the responses, and writes a report. The --eval-type openai_api flag tells EvalScope to format requests as OpenAI Chat completions, which works with any API server that implements that interface, including locally hosted models via vLLM or other inference engines.

The evalscope CLI entry point is registered in pyproject.toml at evalscope.cli.cli:run_cmd. Python 3.10 or later is required.

Benchmark Coverage and Multi-Backend Integration

EvalScope ships with built-in support for MMLU, C-Eval, GSM8K, and many more industry-recognized benchmarks. The What's New section of the README shows the pace at which new benchmarks are added: the v1.12.0 cycle added AutomationBench, JobBench, MiniWoB, OmniDocBench-v1.6, PerceptionBench, ScreenSpot-Pro, PLawBench, PMC-VQA, HiPhO, LogicVista, and CC-OCR-V2, among others.

For evaluations that require specialized backends, EvalScope integrates OpenCompass, VLMEvalKit for vision-language models, and RAGEval for retrieval-augmented generation evaluation. These backends are optional extras; installing the base package with pip install evalscope does not pull them in. The pyproject.toml defines separate optional dependency groups for each backend (opencompass, vlmeval, rag, perf, agentx, aigc, sandbox).

The v1.11.0 release introduced published evaluation versions for reproducible benchmark results. Pinning an evaluation to a published version ensures that the dataset, scoring method, and configuration are fixed, so results can be compared across model versions.

Agent Evaluation and the External Agent Bridge

EvalScope includes an agent evaluation mode that drives benchmarks inside a controlled multi-turn AgentLoop. Each benchmark run records a full per-sample Agent Trace with tool calls, responses, and step timing. The trace data is visualizable through the web dashboard.

A notable addition described in the What's New section is the External Agent Bridge mode. This allows evaluating off-the-shelf agent CLIs such as Claude Code and OpenAI's Codex directly through EvalScope. The bridge transparently forwards each CLI's LLM traffic (covering Anthropic Messages, OpenAI Chat, and OpenAI Responses APIs including SSE streaming) to the evaluation model and records the full trajectory as an agent trace. Custom runners can be registered with the @register_runner decorator.

Agent benchmarks supported through EvalScope include SWE-bench, SWE-bench Pro, SWE-bench Multilingual, GAIA (with multi-turn ReAct and bash in a Docker sandbox), BigCodeBench, BrowseComp, and MCP-Atlas. The GAIA integration uses the official rule-based scorer.

Inference Performance Testing

Beyond capability evaluation, EvalScope includes a performance stress testing module that measures inference server behavior under load. Key metrics tracked include TTFT (time to first token) and TPOT (time per output token), both of which matter for interactive applications.

The performance module supports a --data-source flag for unified dataset loading and parallelized request generation. A --duration flag sets a wall-clock budget for benchmark modes. The Trie agentic trace replay feature allows replaying recorded multi-turn agent traces as performance tests, with three bundled dataset plugins (trie_agentic_coding, trie_code_qa, trie_office_work) that replay real traces with per-turn token caps and tool-call latency simulation.

Vendor Verifier benchmarks (k2_verifier, kimi_verifier, minimax_verifier) validate whether third-party API deployments faithfully reproduce official model behavior, using a shared FunctionCallAdapter base class.

Visualization and the Web Dashboard

EvalScope provides an interactive Web Dashboard for comparing evaluation results across models. The dashboard supports multi-dimensional model comparison, a report overview, and per-sample prediction inspection. The README shows screenshots of a dashboard overview view, a model comparison view, a report overview, and a prediction details view.

Dashboard functionality is part of the app optional dependency group in pyproject.toml. The evalscope web command serves the dashboard. For programmatic access to results without the web interface, evaluation outputs are stored as JSON reports that the framework reads back for subsequent comparisons.

Arena Mode provides pairwise battle evaluation: multiple models receive the same prompts and their responses are ranked against each other. This is different from benchmark scoring, which measures absolute performance; Arena Mode measures relative preference.

Limitations: Dependency Surface and Configuration Complexity

EvalScope's breadth is also its main cost. The full set of optional extras (OpenCompass, VLMEvalKit, RAGEval, sandbox, agentx) brings in a large dependency surface. Teams that only need to run a few text benchmarks against an OpenAI-compatible API do not need most of this, but the configuration options and documentation surface area still scale with the full framework.

The RAG evaluation module was refactored in v1.11.0 to upgrade to MTEB 2.x and RAGAS 0.4.x with Pydantic-based configs. Projects that pinned to older EvalScope versions and used the RAG module will need to migrate their configurations to the new schema.

EvalScope is primarily designed for evaluation, not for training or fine-tuning. It does not replace a model serving framework; it requires a running model endpoint to evaluate against. For teams using non-OpenAI-compatible model APIs, adapter work is required to match the evalscope eval --eval-type parameter.

Alternative: lm-evaluation-harness and the Trade-off in Focus

The EleutherAI lm-evaluation-harness is the most widely used alternative for LLM benchmark evaluation. It emphasizes reproducibility and a large benchmark library, with direct model loading via HuggingFace Transformers rather than requiring an API server. The trade-off is that lm-evaluation-harness focuses on text benchmark evaluation and does not include inference performance testing, agent evaluation modes, or a web dashboard.

EvalScope's differentiation is the combination of capability evaluation, performance stress testing, agent trace recording, and a visual dashboard in one package. For teams that need all of these in a single tool and are comfortable with ModelScope's ecosystem, EvalScope covers more ground. Teams that only need benchmark scores and prefer a leaner, model-direct evaluation path may find lm-evaluation-harness simpler to operate.

Editorial conclusion

EvalScope suits teams that need a single tool to run capability benchmarks, stress-test inference performance, and evaluate agents end-to-end, particularly if they are already in the ModelScope ecosystem. It is less suited to teams that need a simple, minimal evaluation harness with few dependencies, since EvalScope's optional extras bring in OpenCompass, VLMEvalKit, and RAGEval as backends. Before running evaluations, review the v1.11.0 release notes on published evaluation versions: pinning to a published version is the reliable way to produce reproducible benchmark results.

Frequently asked questions

How do I install and run EvalScope for the first time?

Run pip install evalscope, then use the evalscope eval command with --model, --api-url, --api-key, --eval-type openai_api, and --datasets to specify what to evaluate. The README shows gsm8k as an example dataset and --limit 5 to run a quick test against the first five examples.

Can EvalScope evaluate vision-language models and agents?

Yes. EvalScope supports LLMs, VLMs via the VLMEvalKit backend, embedding and reranker models, AIGC models, and agents through a multi-turn AgentLoop that records per-sample Agent Traces. The External Agent Bridge mode also supports evaluating off-the-shelf agent CLIs like Claude Code.

What is the difference between evalscope eval and evalscope perf?

evalscope eval runs capability benchmarks and scores model responses against reference answers using datasets like GSM8K or MMLU. evalscope perf measures inference server performance under load, tracking metrics like TTFT (time to first token) and TPOT (time per output token).

Official sources

  1. License: Apache-2.0
  2. modelscope/evalscope on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/modelscope-evalscope.svg)](https://hysenlabs.com/projects/modelscope-evalscope)