# OpenCompass: The LLM Evaluation Platform Behind the CompassRank Leaderboard

> OpenCompass is a Python framework for evaluating large language models across more than 100 datasets, supporting API backends for GPT-4, Claude, Llama, and dozens of other models. It is the evaluation engine behind the CompassRank public leaderboard and is used by teams that need reproducible benchmarks across multiple models and datasets in a single run.

**open-compass/opencompass** — OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.

- Repository: https://github.com/open-compass/opencompass
- Website: https://opencompass.org.cn/
- Stars: 7,450 · Forks: 867
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/open-compass-opencompass

## What OpenCompass Evaluates and Who Runs It

Language model evaluation has a structural problem: picking a model based on a single benchmark is unreliable because models can be tuned to score well on specific tests while performing poorly elsewhere. OpenCompass addresses this by providing a single framework that runs evaluations across a wide range of datasets in one configured job, so teams get a cross-dataset view rather than a single number.

The primary users are ML researchers who are developing or selecting LLMs and need consistent evaluation methodology. Teams adopting models for production use the platform to compare candidates on task-specific benchmarks before committing. The project is also the engine behind CompassRank at rank.opencompass.org.cn, a public leaderboard showing results for models including Llama, Mistral, InternLM2, GPT-4, Qwen, GLM, and Claude.

The Apache-2.0 licence covers the framework code. The latest tagged release at the time of writing is version 0.5.4, published on 2026-08-26. The repository received a push on 2026-09-17.

## Architecture: Configuration Files, Datasets, and the Run Loop

OpenCompass is driven by Python configuration files. Each evaluation job specifies a model configuration, a list of dataset configurations, and a summariser. From version 0.4.0 onwards, all dataset configs, model configs, and summariser configs were consolidated into the opencompass Python package itself rather than living in separate directories on disk. The README marks this as a breaking change and instructs users to update any config references that pointed to the old ./configs/ paths.

The entry point for running evaluations is run.py in the repository root. The examples/ directory contains dozens of pre-built evaluation scripts named after specific benchmarks and model types, from eval_api_demo.py for API-based models to eval_OlympiadBench.py and eval_TheoremQA.py for specialist reasoning benchmarks.

For large evaluation jobs, OpenCompass supports concurrent inference across tasks together with evaluation watching. Completed inference tasks can be monitored and subsequent evaluation steps triggered without waiting for the full batch to finish. Heartbeat mechanisms are included to detect and recover from stalled tasks. The README describes the concurrent inference implementation at opencompass/tasks/openicl_infer_concurrent.py and the evaluation watcher at opencompass/tasks/openicl_eval_watch.py.

## Installing OpenCompass and Running a First Evaluation

The repository uses a standard Python package structure with setup.py and a requirements.txt that delegates to requirements/runtime.txt. The setup.py hooks into the install process to download NLTK tokenisation data (punkt) automatically, which means the first install requires internet access and adds an NLTK dependency that some environments restrict.

Full installation steps are documented at opencompass.readthedocs.io/en/latest/get_started/installation.html rather than in the README itself. The repository's top-level structure includes a requirements/ directory with separate requirement files for different use cases.

Once installed, running a pre-built example follows the pattern of pointing the runner at an example configuration file. The examples/ directory provides dozens of starting points. For evaluating an API model, eval_api_demo.py is listed as the entry point. For multimodal evaluation, the README points to eval_mmbench_vlmevalkit.py and eval_mmmu_pro_vlmevalkit.py after the VLMEvalKit integration added in 2026-08-25.

For the VLMEvalKit integration, the README notes that it enables native multimodal dataset loading, inference through OpenAI-compatible APIs, and evaluation with official VLMEvalKit metrics. This means teams evaluating vision-language models can use OpenCompass as a unified runner instead of switching between frameworks.

## API Model Ecosystem and Backend Integration

OpenCompass supports a wide range of API backends. The July 2026 update added support for the OpenAI Responses API and LiteLLM AI Gateway alongside updated SDK integrations for Gemini and Anthropic. The corresponding implementation files are named in the README: opencompass/models/openai_response.py, opencompass/models/litellm_api.py, opencompass/models/gemini_sdk_api.py, and opencompass/models/claude_sdk_api.py.

LiteLLM AI Gateway support is significant because LiteLLM acts as a proxy layer that normalises calls to dozens of model providers behind a single interface. A team running OpenCompass through LiteLLM can switch between providers without changing their OpenCompass configuration.

For multi-turn instruction following evaluation, the README notes that multi-round inference was added to GenInferencer alongside support for the Multi-IF dataset. This allows evaluating how models perform on tasks that require maintaining context across multiple exchanges, not just single-turn completions.

The RawPromptTemplate addition from March 2026 lets evaluators pass original benchmark prompts and structured conversations to models without formatting transformations. The README describes this as serving API models, ChatML datasets, and cases where additional prompt content needs to be appended on the model side.

## LLM-as-Judge, Math Verification, and CascadeEvaluator

Not all evaluation tasks can be scored by string matching against reference answers. For open-ended tasks, OpenCompass provides GenericLLMEvaluator, which uses an LLM to score model outputs. The README describes this as LLM-as-judge evaluation and points to docs/en/advanced_guides/llm_judge.md for configuration details.

For mathematical reasoning tasks, MATHVerifyEvaluator provides specialised scoring that understands mathematical notation and equivalence rather than requiring exact string matches. The README points to docs/en/advanced_guides/math_verify.md.

The CascadeEvaluator, added in April 2025, allows multiple evaluators to run in sequence on the same output. The README describes this as allowing custom evaluation pipelines for complex assessment scenarios, where a first evaluator might filter responses before a second applies a more expensive LLM-based scoring step.

OpenCompass also includes a repeat analysis tool at tools/analyze_repeat.py for detecting repetitive content and looping model outputs. The README describes this as working on current evaluation tasks or existing evaluation results, which makes it useful for diagnosing a specific model's failure mode without rerunning the full evaluation.

## Real Limitations of the OpenCompass Framework

OpenCompass is not a lightweight tool. Evaluating a model across a large dataset suite requires significant compute, especially for locally hosted models. The framework is designed for batch runs, not interactive single-prompt testing.

The configuration system is Python-based and requires reading the documentation to set up correctly. The README does not provide a copy-paste quickstart command for a new user. Installation documentation lives at an external readthedocs site rather than inline in the README, which creates a gap between arriving at the repository and running anything.

Version 0.4.0 introduced a breaking change that moved all dataset and model configuration files from the ./configs/ directory into the Python package. The README explicitly warns about this. Any automation or scripts that referenced the old file paths need updating after upgrading through that version boundary.

NLTK downloads a punkt data file on install via the custom DownloadNLTK hook in setup.py. This adds external network access as a hard requirement during installation, which is a problem in air-gapped environments.

The framework targets model evaluation, not dataset creation or training. Teams that want to build custom datasets for their domain need to author OpenCompass-compatible configuration files, which requires reading the dataset format documentation.

## How OpenCompass Compares to lm-evaluation-harness

EleutherAI's lm-evaluation-harness is the other widely adopted Python evaluation framework for language models. Both run batched evaluations across benchmark datasets. The qualitative difference is in integration scope and orientation.

lm-evaluation-harness focuses primarily on autoregressive language models run locally or via Hugging Face's transformers library. Its strength is evaluating open-weight models on a curated set of academic benchmarks. OpenCompass has a broader API model ecosystem out of the box, including the OpenAI Responses API, LiteLLM, Gemini, and Anthropic SDK integrations. OpenCompass also operates the CompassRank public leaderboard at rank.opencompass.org.cn, which gives it a community infrastructure beyond just the evaluation code.

OpenCompass has more built-in infrastructure for large-scale parallel evaluation jobs, including concurrent inference with evaluation watching and heartbeat monitoring. lm-evaluation-harness is simpler to get started with for a team that wants to evaluate a local model against a single benchmark set.

For teams evaluating both API-hosted models and locally hosted open-weight models in the same pipeline, OpenCompass's unified configuration system covering both categories is a practical advantage. For teams that only need local open-weight model evaluation against standard English benchmarks, lm-evaluation-harness's lower setup friction is a real consideration.

## Conclusion

OpenCompass is a practical choice for ML teams that need to benchmark multiple LLMs against a standardised dataset suite and want a framework that already handles API model integration, evaluation watching, and LLM-as-judge scoring. It is not the right choice for teams that need a quick single-command evaluation without reading documentation; the configuration system is Python-based and the full installation guide lives at opencompass.readthedocs.io. Teams migrating from an older OpenCompass configuration should check the v0.4.0 breaking change notice in the README, which consolidated all config files into the package. Confirm which datasets are included before use by browsing CompassHub at hub.opencompass.org.cn.

## FAQ

### What is OpenCompass?

OpenCompass is a Python-based LLM evaluation platform that runs model benchmarks across more than 100 datasets. It supports both locally hosted models and API-based models including GPT-4, Claude, Llama, and others, and powers the CompassRank public leaderboard.

### What is OpenCompass used for in AI development?

OpenCompass is used to benchmark large language models against standardised datasets, compare multiple models in a single evaluation run, and detect performance regressions after fine-tuning. It includes LLM-as-judge scoring and math verification evaluators for open-ended tasks.

### How do I find installation instructions for OpenCompass?

The README links to the full installation guide at opencompass.readthedocs.io/en/latest/get_started/installation.html. The repository uses setup.py and a requirements.txt pointing to requirements/runtime.txt. Note that the setup process downloads NLTK data automatically, requiring internet access.

## Sources

- [License: Apache-2.0](https://github.com/open-compass/opencompass/blob/main/LICENSE)
- [open-compass/opencompass on GitHub](https://github.com/open-compass/opencompass)
- [Project website](https://opencompass.org.cn/)
- [README](https://github.com/open-compass/opencompass/blob/main/README.md)
- [Releases](https://github.com/open-compass/opencompass/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/open-compass-opencompass
